Here’s How Much Data Gets Used By Generative AI Tools For Each Request

Generative AI Tools

While the world is going wild over the potential benefits of generative AI, there’s little attention paid to the data deployed to build and operate these tools.

Let’s look at a few examples to explore what’s involved in determining data use, and why this matters for end users as well as operators.

Text-based generative AI usage

Text-based generative AI tools are a marvel of modern technology, but their usage isn’t without cost. The amount of data these tools consume depends significantly on the complexity and length of the request, as well as the given tool’s sophistication level.

For instance, GPT-3 by OpenAI consumes nearly 175 billion parameters to perform its tasks. It utilizes vast amounts of text from books, websites, and other resources to generate human-like responses. Every character inputted in your request can account for around 4 bytes (assuming unicode encoding), with a typical request averaging at about 2000 characters or less. It adds up when we consider millions of requests being made per day.

Meanwhile, GPT-4 is even more intensive, with 1.8 trillion parameters and a dataset that exceeds a petabyte at its disposal. So while providing an exact figure is not always easy, the data usage collectively across this class-leading platform is undeniably vast, and growing with each iteration.

<strong>Here’s How Much Data Gets Used By Generative AI Tools For Each Request</strong>

Image-oriented AI data consumption

Image-based generative AI tools are fascinating beasts, but they are also voracious consumers of data, easily eclipsing their text-focused counterparts.

For example, Generative Adversarial Networks (GANs) are often used to create realistic images from random noise. In doing so, they gobble up substantial amounts of data. To train just one style-GAN model requires a dataset containing thousands or even millions of high-res images.

Let’s break it down in more simple terms. The average size for an HD image is around 2 MB. If we assume the AI is trained with about 1 million such images, this would mean several terabytes of storage is necessary initially.

This highlights that when employing GANs or similar heavy-duty image-generating AIs, resource usage can scale substantially, which is something key to keep in mind as you plan your projects.

Of course, for image tools that aren’t specifically generative, such as AI editing solutions, the data usage is substantially smaller. For instance, being able to automatically alter photo backdrops is subtractive rather than generative. If you’re interested in the ins and outs of background changer software, it’s a good idea to learn more about this technique to appreciate its advantages.

Speech synthesis tool data needs

Now, let’s focus on speech synthesis tools, which form another category of generative AI that demonstrates the hefty data consumption involved in this tech.

Tools like Google’s Text-to-Speech or Amazon Polly translate textual information into spoken voice output, a process that involves significant amounts of data. Using deep learning technologies, these AIs average about 2MB per minute for standard-quality audio.

Delving deeper, if the tool needs to generate an hour’s worth of audio content, such as for audiobooks, it may require up to roughly 120 MB of processed data. That figure doesn’t include the initial training sets involving hundreds upon thousands of recorded human speech hours.

So while the final product is often just a few minutes long and not excessively large, it’s critical to remember that producing this requires vast reserves of underlying processed and unprocessed data.

Data use in music AI

It’s a good move to contextualize the functionality of music-related generative AI tools by understanding their data use.

These AIs, like OpenAI’s MuseNet or Sony’s Flow Machines, creatively compose music. However, their ingenuity is of course founded on a deluge of data. Thousands of MIDI files that could be anywhere from 10 KB to several megabytes each are needed for training these models.

For instance, a one-minute piece generated approximately requires about 1 MB when converted to MP3 format. But this lean output once again belies the immense resources used during the model-training phase, with copious musical samples processed.

So while these ingenious AIs churn out beautiful symphonies seemingly from thin air, the volume of data that’s been poured into their production and ongoing operation is certain to be staggering.

Chatbots and their data demands

Chatbots are a common use-case of generative AI technology, but these virtual assistants also demand substantial data.

To begin with, generating conversational responses necessitates a considerable backend infrastructure. On average, answering one chat request typically uses around 1-2 KB if only the final output text is weighed up.

However, that doesn’t tell the whole story. Behind-the-scenes there’s extensive data usage during training phases, which is the same story across every generative AI solution we’ve discussed. Consider the hundreds of thousands or more conversational scripts used to train big-league players like Microsoft’s Azure Bot or Google Dialogflow.

Add this to ongoing fine-tuning and you start grasping the true scale of requisite raw information. So even though your Alexa’s response to tomorrow’s weather may seem simplistic, it represents the tip of the iceberg in terms of the data used.

Final thoughts

It’s worth concluding on the point that while generative AI is heavily data-reliant, tools are becoming more efficient even if the datasets on which they are trained keep growing. So while it’s worth being cautious about their impact, this is certainly not a deal-breaker for anyone thinking about using them.

Amazon Announces “Amazon Q,” the Company’s Generative AI Assistant

In a striking move within the competitive landscape of productivity software and generative AI chatbots, Amazon recently unveiled its latest innovation: “Amazon Q.” This announcement, made at the AWS Reinvent conference in Las Vegas, marks Amazon's assertive stride into a domain where tech giants like Microsoft and Google have already established a significant presence.

The introduction of Amazon Q is not just a new product launch; it's a statement of Amazon's ambition and technological prowess in the rapidly evolving field of AI-driven software solutions.

The emergence of Amazon Q can be viewed in the context of the recent success of Microsoft-backed OpenAI's ChatGPT. Since its launch, ChatGPT has revolutionized the way generative AI is perceived, offering human-like text generation based on brief inputs.

Amazon's Q steps into this arena, signaling not only a challenge to Microsoft's growing influence in AI but also an effort to redefine the capabilities of chatbots in the professional world. This move by Amazon could be seen as a strategic endeavor to capture a share of the market that has been intrigued and captivated by the possibilities of AI, as demonstrated by ChatGPT's popularity.

Image: Amazon

Features and Accessibility of Amazon Q

Amazon Q emerges as a sophisticated chatbot designed to seamlessly integrate with Amazon Web Services (AWS), offering an array of functionalities tailored for the modern workplace. Central to its appeal is the ability to assist users in navigating the expansive ecosystem of AWS, offering real-time troubleshooting and guidance. This integration signifies a leap in making AWS's complex array of services more accessible and user-friendly.

In terms of availability, Amazon has launched Amazon Q with an initial free preview, allowing users to experience its capabilities without immediate cost. This approach mirrors the common strategy in software services, where early adopters are given a chance to test and provide feedback.

Following this period, it will transition to a tiered pricing model. The standard business user tier is set at $20 per person per month, while a more feature-rich tier for developers and IT professionals will cost $25 per person per month. This pricing strategy positions the chatbot competitively against similar offerings.

Image: Amazon

Integration and Functionality

The integration of Amazon Q extends beyond AWS, encompassing popular communication applications such as Salesforce’s Slack and various text-editing tools used by software developers. This versatility underscores Amazon's ambition to make the chatbot an indispensable part of the professional toolkit, facilitating smoother workflows and enhanced communication.

A standout feature of Amazon Q is its ability to automate modifications to source code, significantly reducing the workload for developers. This functionality is not just a time-saver; it represents a shift in how AI can be leveraged to enhance coding efficiency and accuracy.

Furthermore, it boasts the capability to connect with over 40 enterprise systems. This extensive connectivity means users can access and interact with information across various platforms like Microsoft 365, Dropbox, Salesforce, Zendesk, and AWS's own S3 data-storage service. The ability to upload and query documents within these interactions further enhances the chatbot’s utility, making it a comprehensive tool for managing a wide array of business functions.

Amazon’s Expansion with Amazon Q

The launch of Amazon Q marks a significant milestone in Amazon's strategic expansion into AI-assisted productivity software. Its integration with AWS, competitive pricing, and innovative features position it as a formidable player in a market dominated by tech giants like Microsoft and Google. Amazon Q's potential to streamline AWS service navigation and enhance the efficiency of various professional tasks is notable.

Looking ahead, Amazon Q could have far-reaching implications for Amazon and the broader AI-assisted productivity software sector. It has the potential to not only diversify Amazon's product offerings but also to redefine how businesses interact with AI tools. The success of Amazon Q could pave the way for more advanced AI applications in the workplace, further blurring the lines between human and machine-led tasks. As the AI landscape continues to evolve, Amazon Q could emerge as a key driver in shaping the future of AI integration in professional environments.

Maximizing business value with ETL for Big Data

image-6

In today’s digital era, big data has emerged as a pivotal asset for organizations across various industries. The vast volumes of data generated every moment offer unparalleled insights into customer behaviors, market trends, and operational efficiencies. However, the true value of this data lies not just in its quantity but in the ability to process and analyze it effectively. This is where Extract, Transform, Load (ETL) processes become indispensable.

ETL is the backbone of data engineering, serving as a critical pipeline for transforming raw data into meaningful information. The extraction phase involves retrieving data from multiple sources, which could range from traditional databases to real-time streaming platforms.

In the transformation step, we cleanse, normalize, and reformat this diverse data, ensuring it adheres to a consistent structure and format, crucial for accurate analysis. Finally, the loading phase involves transferring this processed data into a data warehouse or analytical tool for further exploration and decision-making. This systematic approach not only streamlines data management but also enhances the reliability and usability of the information, turning big data into an invaluable resource for business intelligence and strategic planning.

Understanding Big Data and its business implications

Big data, characterized by its immense volume, varied formats, rapid generation, and the need for veracity, has become a cornerstone in contemporary data engineering. Volume refers to the massive quantities of data generated every second, from terabytes to petabytes, posing significant storage and processing challenges. Variety encompasses the diverse types of data, ranging from structured data in databases to unstructured data like emails, videos, and social media interactions. This diversity demands sophisticated parsing and integration techniques to ensure cohesive data analysis.

Velocity highlights the speed at which data is created and needs to be processed, often in real-time, to maintain its relevance. This necessitates robust and agile processing systems capable of handling streaming data and quick turnarounds. Lastly, veracity underscores the importance of accuracy and reliability in data. In the realm of big data, ensuring the quality and authenticity of information is paramount, as even minor inaccuracies can lead to significant misinterpretations in large-scale analyses.

In the business context, big data plays a critical role in driving decisions and spurring innovation. By analyzing vast datasets, organizations can uncover patterns and insights that were previously inaccessible, leading to more informed strategic decisions, personalized customer experiences, and efficient operations. However, managing big data comes with its own set of challenges, including ensuring data privacy, securing data against breaches, and maintaining data integrity. Overcoming these challenges is essential for businesses to truly leverage the potential of big data in driving growth and competitive advantage.

The role of ETL in Big Data management

The Extract, Transform, Load (ETL) process is fundamental in big data management, serving as the critical pathway for turning vast amounts of raw data into valuable insights. At its core, ETL encompasses three key stages:

Extract: This initial phase involves gathering data from various sources. In big data environments, these sources are diverse and voluminous, ranging from on-premises databases to cloud-based storage and real-time data streams. The challenge lies in efficiently extracting data without compromising its integrity, necessitating robust and versatile extraction mechanisms.

Transform: Once extracted, the data often exists in a raw, unstructured format that is not immediately suitable for analysis. The transformation stage is where this data is converted into a more usable format. This involves cleaning (removing inaccuracies or duplicates), normalizing (ensuring data is consistent and in standard formats) and enriching the data. For big data, transformation is a complex task due to the variety and volume, requiring advanced algorithms and processing power to handle the data effectively.

Load: The final stage involves transferring the processed data into a data warehouse or analytical tool for further analysis and reporting. In big data scenarios, this often means dealing with large volumes of data being loaded into systems that support high-level analytics and real-time querying.

The significance of ETL in big data management cannot be overstated. It enables efficient handling of large datasets, ensuring data quality and consistency, which are crucial for accurate analysis. Moreover, ETL acts as a bridge between the raw, often unstructured data, and the actionable insights businesses need. By effectively extracting, transforming, and loading data, ETL processes unlock the value hidden within big data, facilitating informed decision-making and strategic business planning. This transformation from data to insights is not just a technical process but a business imperative in the data-driven world.

Deep dive into ETL processes for Big Data

In the realm of big data, the ETL process is tailored to handle the complexity and volume inherent in diverse data sources. During extraction, data engineers employ techniques like parallel processing and incremental loading to efficiently pull data from various sources, be it traditional databases, cloud storage, or real-time streams. These techniques are crucial in managing the high volume and velocity of big data, ensuring a seamless extraction process without system overload.

Transformation is arguably the most intricate phase in big data ETL. It involves cleaning the data by removing inconsistencies and errors, which is vital to maintain data quality. Normalization is another key aspect, where data from different sources is brought to a common format, facilitating unified analysis. Advanced strategies like data deduplication and complex transformations are often employed, using tools capable of handling big data’s scale and complexity.

When it comes to loading, efficiency is paramount, especially with large datasets. Common techniques like bulk loading involve loading large volumes of data in batches, significantly reducing the time and resources required. Partitioning data before loading it into a data warehouse or lake is another strategy that enhances query performance and data management.

The retail industry provides a real-world example. A large retailer might use ETL to integrate customer data from various sources – online transactions, in-store purchases, and customer feedback. In this scenario, ETL processes not only handle the sheer volume of data but also transform it into a structured format suitable for analysis. ETL in big data scenarios showcases its transformative power by using this transformed data for personalized marketing, inventory optimization, and enhancing customer experiences.

Optimizing ETL for business intelligence

Integrating ETL Processes with business intelligence

Integrating ETL processes with business intelligence (BI) tools is crucial for deriving actionable insights from big data. This integration allows for the seamless flow of cleansed and transformed data into BI platforms, enabling real-time analytics and reporting. ETL plays a pivotal role in ensuring that the data feeding into these tools is accurate, consistent, and formatted appropriately for complex analyses.

In the context of data warehousing and data lakes, ETL is essential for structuring and organizing data. Data warehouses tailor ETL processes to transform and load data into a format optimized for query performance and data integrity. For data lakes, especially those handling unstructured or semi-structured data, ETL is critical in tagging and cataloging data, making it searchable and usable for analytics purposes.

ETL’s contribution extends to predictive analytics and informed decision-making. By efficiently processing and preparing large datasets, ETL enables businesses to leverage advanced predictive models and machine learning algorithms. These models can analyze historical and real-time data, providing businesses with foresight into market trends, customer behavior, and operational efficiencies. Thus, optimized ETL processes are not just about data management; they are a cornerstone in a company’s strategic decision-making toolkit.

Best practices in ETL for maximizing business value

Data quality management in ETL processes

Effective data quality management in ETL involves implementing validation rules and consistency checks to ensure accuracy and reliability. Additionally, regularly auditing and cleansing data helps maintain its integrity throughout the ETL process. In terms of data security and compliance, it’s crucial to employ encryption and access controls during data extraction and loading. Furthermore, adhering to regulatory standards like GDPR and HIPAA is essential for compliance and maintaining consumer trust.

Continuous monitoring and optimization of ETL processes are vital. Additionally, this includes performance tuning of ETL jobs, regularly updating the data transformation logic to align with evolving business needs, and proactively identifying bottlenecks. Such practices not only improve efficiency but also ensure that the ETL process continually adds value to the business by providing high-quality, secure, and relevant data for decision-making.

Conclusion

ETL processes are indispensable in the realm of big data, serving as a critical conduit for transforming vast, unstructured datasets into actionable business insights. The efficiency and effectiveness of ETL directly influences a business’s ability to harness the full potential of big data, driving informed decision-making and strategic planning. In an era where data is a key competitive differentiator, investing in robust ETL tools and processes is not just advisable but essential. Businesses that prioritize and refine their ETL practices unlock the true value of their data, fostering growth and innovation in today’s data-driven marketplace.

This AI-powered Eufy robot vacuum and mop is still only $650 after Cyber Monday

Eufy Clean X9 Pro CleanerBot

What's the Cyber Monday deal?

Amazon dropped the price for the AI-powered Eufy X9 Pro robot vacuum and mop to only $650 as part of its Cyber Monday deals and the discount is still available today.

Why this deal is ZDNET-recommended

The Eufy Clean X9 Pro CleanerBot, a new 2-in-1 robot vacuum, boasts a deep cleaning, hands-free mopping experience, coupled with 5,500pa of suction power. It also uses some AI navigation features to maneuver throughout your house.

Also: Best robot vacuum deals: Get a Roomba or Shark on sale now

Initially, I was less than enthusiastic about trying out yet another robot mop vacuum (I'd tested a similar one recently), but once I watched the Eufy X9 Pro work its way across my home floors, my mind was changed.

ZDNET RECOMMENDS

Eufy Clean X9 Pro CleanerBot

This is the perfect robot vacuum and mop for homes with hard floors, even if there are carpets and rugs in between.

View at Amazon

The CleanerBot truly lives up to the name, outperforming my old Roborock and the Yeedi MopStation Pro in vacuum and mop functions. The suction power, 5,500pa at maximum capacity, is outstanding. And the main brush is bristle-less, made of silicone wedges instead that are just as effective at cleaning floors.

In my limited experience (as I've only tested this model for about a week), the primary silicone brush makes it less likely for the X9 Pro to get tangled, as it's easier to scoop debris up than sweep it.

The mopping function on the Eufy X9 Pro CleanerBot is one of the two features that impressed me the most. The X9 Pro has two rotating mop pads — which I love in a robot vac/mop combo — which put 2.2 lbs of downward pressure to break down tough stains, a particularly useful feat for my home of children and pets.

Review: Roborock S8 Pro Ultra: This 2-in-1 vacuum can do just about everything

The other outstanding feature, and probably my favorite, is the use of AI for navigation, obstacle avoidance, and mapping. The CleanerBot has time-of-flight sensors and an AI camera system, called AI See, that helps detect and avoid objects so the vacuum doesn't suck up your kids' socks or stuffed animals.

It also uses iPath Laser Navigation to create maps of your home, which separates the rooms by color in the Eufy Clean app and even shows you the obstacles that the robot has found in each room. When you review the map after cleaning, you'll find things like power cords, shoes, and trash cans marked on the map.

Eufy isn't the first to use this technology for obstacle avoidance and mapping, but it is a great feature. I hate having to pick up every last bit of paper my kids dropped before I can start cleaning — only to have the robot vacuum get stuck anyway on a power cord somewhere.

Also: This robot vacuum has a brilliant self-cleaning feature I didn't know I needed

The Eufy Clean app lets you customize settings for charging, cleaning intensity, voice, and more. And it also enables you to choose from the rooms that the robot automatically created on the map so you can send it to clean just that area, like a muddy entryway. You can choose to clean zones as small as 1.6 ft by 1.6 ft on the map in case of spills.

The Eufy Clean X9 Pro CleanerBot easily adjusts to uneven surfaces to cross up to 2 cm barriers.

Beyond the AI See camera set, the CleanerBot has a sensor to detect floor types in case you're running the X9 Pro in vacuum and mopping mode and it reaches a carpet or a rug. Once the robot detects a rug or carpet, it raises the mop pads to keep them off the mat and only vacuums on the soft surface.

Also: Best robot vacuums you can buy right now

Here's another thing I was glad to see: The X9 returns dutifully to its station to wash the mop pads rather than wait until they're overdue for a cleaning. I don't want to see my robot mop dragging dry, dirty mopping pads minutes after it should've returned for a refresh, but I haven't found this to be a problem with the X9.

ZDNET's buying advice

The Eufy Clean X9 Pro CleanerBot is available for sale at $650 and is the perfect option for someone looking for a robot vacuum and mop combination for a home with a lot of hard floors, whether that's tile or hardwood, with some carpet or rugs mixed in.

It doesn't have a self-emptying dustbin, and the dustbin itself has to be emptied after each cleaning as it's pretty tiny. Still, the mopping feature and the suction power are impressive, especially as the mop can pick up stains and dirt that my Yeedi MopStation Pro left behind.

Featured reviews

AWS Unveils Graviton4, Trainium2 for Faster, Affordable AI Model Building

AWS Re:invent Adam VP

At re:Invent in Las Vegas, Amazon Web Services (AWS) announced two new AI chips –AWS Graviton4 , AWS Trainium2. The new chips aim to provide advancements in price performance and energy efficiency for a wide range of customer workloads, including machine learning training and generative AI applications.

Graviton4 offers up to 30% better compute performance, 50% more cores, and 75% more memory bandwidth than Graviton3. Trainium2 delivers up to 4x faster training than its first generation, with deployment capability in EC2 UltraClusters of up to 100,000 chips.

(Source: Business Wire)

David Brown, VP of Compute and Networking at AWS said that Graviton4 marks the fourth generation they have delivered in just five years, and is the most powerful and energy-efficient chip ever built. “Silicon underpins every customer workload, making it a critical area of innovation for AWS,” he added.

He said that it has more than 50K customers for Graviton, and its other cloud providers are still just talking about making them, and are yet to deliver first server processors. At Ignire 2023, Microsoft recently launched Azure Maia 100 AI Accelerator, its first in-house custom AI system on a chip.

Some of its customers leveraging AWS chips include Anthropic, Databricks, Datadog, Epic, Honeycomb, SAP and others. Naveen Rao, VP of generative AI at Databricks said that AWS Trainium gave them the scale and high performance needed to train our Mosaic MPT models, and at a low cost.

“AWS Graviton4 instances are the fastest EC2 instances we’ve ever tested, and they are delivering outstanding performance across our most competitive and latency-sensitive workloads,” said Roman Visintine, lead cloud engineer at Epic Games.

Juergen Mueller, CTO of SAP SE said that as part of the migration process of SAP HANA Cloud to AWS Graviton-based Amazon EC2 instances, we have already seen up to 35% better price performance for analytical workloads.

Graviron4-powered R8g instances are available today in preview, with general availability planned in the coming months. Check out here. Trainium2 is said to be available in Amazon Ec2 Trn2 instances Check it out here.

The post AWS Unveils Graviton4, Trainium2 for Faster, Affordable AI Model Building appeared first on Analytics India Magazine.

Data-driven supply chain part 2: The theory of constraints & the concept of the information supply chain.

The viability of the ‘Viable Vision’.

Data-driven supply chain part 2: The theory of constraints & the concept of the information supply chain.

I did hear about the Theory of Constraints (TOC) off and on through the late 90s, but I didn’t pay much attention until late 2001. One of the i2 consultants I met at their annual meet in Malaysia had one too many- and ended up lecturing me on how TOC was going to change the world. Ken Sharma, one of the two partners who set up i2 worked for Goldratt Institute before he joined hands with Sanjiv Sidhu. I bought Eli Goldratt’s book ‘The Goal’ on the way back to the Airport, and I was immediately hooked. The simplicity and the logic in the model were real, and the book was a great read…more like a fast-paced thriller than a dull business book.

The market for TOC started picking up in India, especially after the introduction of “viable vision”. The concept caught the imagination of every aspirational CEO; after all who can resist the chance to turn the company’s topline into its bottom line, JUST in four years?

In simple terms, a company implementing TOC can aim to convert its current top-line $ number into its bottom-line $ number in four years. Seems incredible, but Goldratt believed it is a perfectly ‘viable vision’ for a company to have. He used to conduct one-day workshops exclusively for the CEOs / CxOs to explain why such a vision is viable. Goldratt was so convincing, that many CEOs wanted to try it out. Besides, Goldratt was willing to share the risk; he would take the majority of his fees if, and only if, the company managed to achieve its stated goal.

While there were a few big successes, there were a few disappointments too. I would assume a ‘lack of commitment from the top management’ must be one of the key underlying reasons for failures. Over a period of time, TOC consultants stopped talking about ‘viable vision’, but they continue to get engaged especially by growth-oriented companies of medium size. The goals being set for the TOC projects still continue to be challenging, but definitely not as ambitious as in the case of ‘viable vision’.

Are unresolved Bottlenecks in the information supply chain the reasons TOC projects fail?

Strictly speaking, TOC projects do not ‘completely’ fail. However, they can fail to deliver the intended value. They do produce results, but the scale and value delivered may depend on how exactly you have designed and implemented your project. In an overwhelming majority of cases, the ‘vision’ for converting your topline number into your bottom line, in four years or less, may turn out to be ‘not viable’.

There is a fair amount of academic criticism of TOC. Some claim there is nothing new in TOC (Steyn H., 2000), it has heavily borrowed from pre-existing concepts, some claim it is the same as the theory proposed by Wolfgang Mewes (Mewes. W., 1963), it does not work for product mix decisions, etc… The majority of the academics seem to suggest Prof. Goldratt’s work does not have the necessary ‘rigor’ to be called a serious academic theory. Goldratt published a paper titled “Standing on the shoulders of Giants” (Goldratt, 2009) acknowledging different pre-existing concepts and people who inspired TOC.

While there is no consensus on why implementations fail, many consultants attribute the failure to the following reasons:

  1. Limited Scope definition & Lack of commitment from Top Management… Typically cases where the management is extra-cautious, end up choosing a small sub-process for a pilot which is inherently unsuitable for TOC implementation or to demonstrate its value conclusively.
  2. The key constraint identified is NOT the ‘real key’ constraint.
  3. The customer was not able to handle the business disruption through the ‘period of pain’ and ended up going back to his old process model.
DDSC-Figure-3-The-Traditional-TOC-Implementations-–-Root-causes-for-Failures-1
Figure 3: The Traditional TOC Implementations – Root-causes for Failures

In my view: TOC projects fail because:

  1. They fail to recognize that the enterprise value chain is NOT a unitary chain, but a complex network of multiple supply chains, each dependent on the other.
  2. They fail to resolve the bottlenecks in the all-important Information supply chain.

Let me explain this further.

My experience implementing TOC: The importance of the information supply chain

I spent a good part of the last three decades in consulting and working for services companies. I tried the concept of TOC as the COO of a large digital content services KPO for one of its divisions.

The whole experience was fascinating.., The division was badly run for years and had just lost the confidence (and 3/4th of business) of its largest client – A Global Publishing company… The morale was low, and my boss (and the MD) privately confessed to me that he ran out of ideas. I was asked to drive the business personally and do what I could to bring it back on track. There was already enough confusion in the system and people, so I did not mention anywhere that I was driving a TOC implementation so as not to confuse them further. Since bringing the business back on track was a bigger priority for me, I wanted to simultaneously try TOC and Service-chain Optimization (called SCO, as a concept less popular than TOC) together…Further, I always believed creating a ‘seamless, granular, and drill-down visibility into the entire value chain is the most important fix. So building visibility into the supply chain (service-chain in this case) was my biggest priority.

In lieu of a standard due diligence, I asked all the teams in the division to create a ‘daily status report’ (DSR) — a report that clearly lists out “status of every job” as it progresses through different stages of the workflow. I mandated a ‘standard reporting format’ that evolved and became more and more granular over a period of time. I ensured that all the granular elements (pertaining to different clusters) on different pages, added up to the overall division’s reporting numbers on a front page (that we called At-a-glance) as an ‘Executive-Summary’ for the top management.

In my mind, the daily status report was meant to be a digital twinthat reflects the true status of each of the jobs in each of the workflow stages, across the complete value chain. While the digital twin was not enabling the ‘collecting and monitoring the information (data + insights)’ in real-time, the 6–12 hours lag was acceptable for us to start with.

DDSC-Figure-4-Applying-TOC-for-Resolving-Bottlenecks-in-the-Information-Supply-chain
Figure 4: Applying TOC for Resolving Bottlenecks in the Information Supply-chain

And every day, I used to have a daily status meeting with all the cluster-heads and the team-leads, where I would run through every job that is delayed or likely to be delayed. It was easy enough to discover which particular stage in the workflow was a bottleneck, or likely to become a bottleneck. A quick ramping-up of the capacity in critical stages of the workflow fixed the bottlenecks to a large extent. Besides the bottlenecks, the other teething problems involved ‘communication’ or the lack of it. Information and the instructions, between the clients and the delivery teams, were being lost in the translation, leading to rework in critical workflow stages, constraining the capacity further. Then to my surprise, I discovered a similar communication drop was happening between different stages in the internal workflow. The missing-data, and the back and forth for the clarifications, were the root cause of bottlenecks in most cases.

We started creating a space within the daily status report for all the critical pieces of information that need to be exchanged between different stages of the internal workflow and with the clients’ teams. We created a color code to alert the client on the ‘missing pieces of information’ so as to ensure they clarify on a priority.

Soon with the workflow bottlenecks fixed, and with smoother information flow, the teething delivery issues that crippled the division a few months back started disappearing. Simultaneously, customers began to recognize the significant and welcome improvements in delivery and started diverting more and more business from different vendors back to our company. We started creating a separate section within the daily status report for the client to list out the new business that is likely to come our way. The advanced information helped the clusters to plan and ramp up their capacities in anticipation of the business.

Within nine months, the division had become the most efficient (and most profitable) business unit in the company. The top-line run-rate ($ sales per month) of the division more than tripled by the end of 9th month. The concept of daily status reports was subsequently rolled-out to other divisions with similar quantum of success.

Here are my key learnings from the experiment:

  1. The enterprise value-chain is not one unitary chain, but a complex network, or an interlinked web of multiple supply-chains.
DDSC-Figure-5-Supply-chain-network-the-Underlying-Layer-of-Information-Supply-chain
Figure 5: Supply-chain network – the Underlying Layer of Information Supply-chain
  • Every physical goods supply chain has a Services supply chain supporting it, a set of people providing services to ensure the physical goods supply chain functions well.
  • For example: You need people to make documents like way bills, receiving reports, invoices, etc., at every stage in the supply chain… the efficiency and the speed of the Services supply chain can affect the speed and efficiency of the Physical-goods-supply-chain. For e.g.: any delay in creating the GST-Invoices can hold the truck from leaving the factory premises.
  • Every Physical-goods supply -chain has a Cash-flow- supply -chain supporting the value exchange each time the title to the goods changes from one entity to the other or each time a legal bailment is created…. A Cash-flow supply chain that pays for every service that is delivered to enable the Physical goods supply chain. For e.g.: Any delay in releasing the payment for one or more of the services such as transport, can hold the physical goods supply-chain to a grinding halt.
  • Every Physical goods supply -chain is supported by a Data supply -chain / Information supply -chain… Information that gets passed on from one workflow stage to the next either in the form of a physical document (like a GST Invoice, or a Waybill), or as a ‘digital document’ (For e.g.: SAP uses a format called IDoc (Internal Document) to send out data pertaining to a unique transaction for exchanging information with external applications)
  • For a Supply – chain to function well at full capacity, the highest possible throughput, the multiple layers of supply – chains need to work in tandem, perfectly in equilibrium. Any bottlenecks in the Cash-flow Supply -chain, the Information- Supply -chain, or the Internal Services Supply -chain will definitely create bottlenecks in the Physical goods supply chain.

2. Creating ‘seamless, drill-down visibility into supply-chain’ is the most important first step for any supply-chain engagement — irrespective of the industry segment. The rule is as important for the Physical goods supply – chain in the manufacturing sector, as it is for the service – chain in the services industry.

The ideal mode of supply-chain visibility should be in the form of a ‘Digital Supply-chain Twin (e.g.: Similar to a SCADA system) that reflects the ‘instant truth’ of every job at every stage of the supply-chain. Instant truth means instant exchange of information between the Physical-goods Supply – chain and the Digital Twin.

3. TOC implementations in the manufacturing sector, that focus purely on the ‘constraints’ within the Physical-goods Supply-chain, and ignore the constraints in the ‘Services Supply -chain, Cash-flow Supply -chain, and Information Supply -chain,… are scripted for failure.

In the case of Services-companies, the efficiency and effectiveness of the Service-chain would depend on the efficiency & effectiveness of the layers of the Cash-flow Supply-chain and the Information Supply-chain supporting it. The Information Supply-chain is far more important in services companies.

DDSC-Figure-6-The-Importance-of-Resolving-Bottlenecks-in-the-Information-Supplychain
Figure 6: The Importance of Resolving Bottlenecks in the Information Supply-chain – A Necessary condition for Supply-chain Optimization

4. The series of ‘in-process buffers’ created to ease the constraints (in the workflow) in a typical TOC implementation, would mean incremental work-in-process inventory, which in turn means incremental working-capital, and can lead toincremental stress in the Cash-flow supply-chain.

There are numerous articles on how ‘improving the throughput’ in one workflow stage could create incremental stress and a bottleneck in the subsequent workflow stages downstream.

DDSC-Figure-7-Throughput-Maximization-A-Function-of-Harmony-between-Interdependent-Supply-chains
Figure 7: Throughput Maximization – A Function of Harmony between Interdependent Supply-chains

5. Maximizing the organizational ‘throughput’ in a TOC implementation, needs to be done without upsetting the delicate balance, or the ‘equilibrium between different interdependent layers of supply chains (of physical goods, internal services, and information).

Digital (data-driven) supply chain

And that brings me to the next big question. What if every decision at every stage in the workflow is driven by ‘data’… 100% relevant and sufficient data that can support the decisions that drive the supply chain?

How exactly do you recognize and list what supply-chain decisions are being made at each stage of the supply -chain? And to add to the complexity, each of such decisions in the physical-goods supply-chain may influence the vitals of the cash-flow supply-chain, or the demand on the services supply-chain.

In the case study I mentioned above, we did go through real granular data on every work-flow-stage, for every job, and more importantly, institutionalized a method to generate and utilize such data to make service-chain decisions’. In our case, the data was a ‘real’ priority; hence we had a deliberate strategy to identify and source the data that supports decisions at every workflow stage.

But I am not sure if a typical TOC implementation ‘prioritizes the data that supports supply-chain decisions’. I am also not sure if TOC implementations take into account the concept of interdependent multiple layers of supply-chains, and the bottlenecks that may occur in Cash-flow and Information (Data + Insights) supply-chains that in turn affect the Physical goods supply-chain.

So how exactly does one go about building a data-driven supply-chain? Most articles and books I came across have only dealt with generic directions. …Not useful for someone looking for specific directions for building a ‘digital supply-chain’.

The following response (picture on the right) has been generated by Open-AI Chat. Surprisingly, this generative AI does a better job than most articles. It does mention “identifying process bottlenecks” in the very first step.

DDSC-Figure-8-How-to-Build-a-Digital-Supply-chain-–-Chat-GPT4-Answers
Figure 8: How to Build a Digital Supply-chain – Chat GPT4 Answers!

Read further: Part-3 of this Series covers the immediate and pressing need for digitizing your supply chain; the key features of the digital supply chain, and how is going to be different from the traditional supply chain & a broad roadmap for building a Data-driven, AI-powered supply chain.

Read earlier: Part-1 of this series explains how data-driven decision-making in supply -chain has evolved over the years, from the earliest rudimentary decision-support systems, such as i2 Technologies and SAP-APO, to the supply chains of the future, which will heavily rely on Big data, AI-ML and Generative AI).

AWS Subtly Calls Out OpenAI ChatGPT Security Flaws, Introduces Bedrock Guardrails 

At AWS’s re:Invent, AWS’s chief Adam Selipsky, subtly called out OpenAI’s security flaws, while introducing its security and safety features in Amazon Bedrock.

Citing CNBC’s report, ‘Microsoft briefly restricted employee access to OpenAI’s ChatGPT, citing security concerns,‘ as part of his slide, Selipsky introduced Guardrails for Amazon Bedrock.

He stressed on the importance of responsible AI, and how AWS has been integrated this into its platform from day one, “An important component of responsible AI is promoting the interaction between consumers and the applications to avoid harmful outcomes, and the easiest way to do this is actually placing limits on what information models can and can’t do,” shared Selipsky.

“With Guardrails for Amazon Bedrock, you can consistently implement safeguards to deliver relevant and safe user experiences aligned with your company policies and principles.” the company said in its blog post.

Guardrails enable users to set restrictions on topics and apply content filters, eliminating undesirable and harmful content from interactions within applications. This provides an additional layer of control beyond the safeguards inherent in foundation models (FMs).

Guardrails can be applied to all LLMs in Amazon Bedrock, encompassing fine-tuned models and Agents for Amazon Bedrock.

“OpenAI has a unique perspective on safety, driven by scientific measurement and lessons from iterative deployment”, OpenAI’s former board member Greg Brockman posted on X minutes after AWS’s Guardrails were announced.

OpenAI has consistently emphasised that it refrains from using API data for training its models. In an effort to build trust with enterprises, the AI startup introduced ChatGPT Enterprise earlier this year.

The post AWS Subtly Calls Out OpenAI ChatGPT Security Flaws, Introduces Bedrock Guardrails appeared first on Analytics India Magazine.

DSC Weekly 28 November 2023

Announcements

  • The data security landscape is shifting as organizations embrace SaaS-based applications, public cloud services, data analytics platforms, and AI/ML workloads. Unfortunately, such progress has been met by a resurgence of ransomware attacks and a record-breaking number of data breach disclosures. In light of these revelations, new consumer privacy protection acts have imposed heavier financial penalties, leading to changes in the cybersecurity insurance marketplace and how organizations measure potential risk. Register for the free Data Security Risks and Challenges summit to learn about the latest changes in data security and explore new technologies designed to protect data and customers.
  • Ransomware attacks show no signs of slowing down. This year marked a record-breaking year for ransomware attacks, as they surged 74% by the first three months of 2023. Organizations require not only a solid prevention plan, but they need established recovery solutions to ensure they bounce back from attacks that can cause irreparable economic and reputational damage. The newer and more treacherous modern threat landscape forces organizations to take a second look at cyber insurance and the security it can ensure against fallout from an attack. Join the upcoming Ransomware Preparedness: Strategies for a Secure Future summit to hear leading experts discuss actionable strategies to prevent ransomware attacks, mitigate damage, and select the best cyber insurance option for your organization.

Top Stories

  • Data-driven, AI-powered supply chain part 3: Imagining the Future – Supply chain 5.0
    November 28, 2023
    by Krishna Pera
    The immediate and pressing need for ‘digitizing’ your supply-chain One may conclude: ‘Digitizing’ the supply-chain has become a survival necessity for companies to stay competitive. Apart from a substantial jump in the efficiency-effectiveness, the customer-experience, and upside to revenues, companies can expect a huge-huge cost-saving.
  • Trusted, automated data sharing across spreadsheets and other documents
    November 28, 2023
    by Alan Morrison
    Earlier in the fall, Charles Hoffman joined our non-profit Dataworthy Collective (DC) that focuses on best practices in trusted knowledge graph development. Hoffman is a CPA, consultant and former PwC auditor who works with clients who use the Extensible Business Reporting Language (XBRL).
  • Generative AI: Precursor to Autonomous Analytics
    November 25, 2023
    by Bill Schmarzo
    We are living in a time of unprecedented change and innovation. Generative AI (GenAI) has created new horizons for us to explore the possibilities of AI in creating novel and diverse content. But the real revolution lies in the next step of the journey – autonomous analytics, an emerging category of analytics that can learn, adapt, and act with minimal human intervention as they interact with their environment.
Education_DSC_160x600-2

In-Depth

  • Data-driven supply chain part 2: The theory of constraints & the concept of the information supply chain.
    November 28, 2023
    by Krishna Pera
    The viability of the ‘Viable Vision’. I did hear about the Theory of Constraints (TOC) off and on through the late 90s, but I didn’t pay much attention until late 2001.
  • Maximizing business value with ETL for Big Data
    November 28, 2023
    by Ovais Naseem
    In today’s digital era, big data has emerged as a pivotal asset for organizations across various industries. The vast volumes of data generated every moment offer unparalleled insights into customer behaviors, market trends, and operational efficiencies.
  • Here’s How Much Data Gets Used By Generative AI Tools For Each Request
    November 28, 2023
    by Erika Balla
    While the world is going wild over the potential benefits of generative AI, there’s little attention paid to the data deployed to build and operate these tools. Let’s look at a few examples to explore what’s involved in determining data use, and why this matters for end users as well as operators.
  • The role of generative AI in shaping the e-commerce landscape
    November 28, 2023
    by Pritesh Patel
    Generative AI is rapidly altering the landscape for e-commerce professionals with applications ranging from effective supply chain management to tailored client experiences. The application of generative AI has transformed e-commerce and offered cutting-edge fixes to improve almost all facets of online enterprises.
  • DSC Weekly 21 November 2023
    November 21, 2023
    by Scott Thompson
    Read more of the top articles from the Data Science Central community.

AWS unveils new Trainium AI chip and Graviton 4, extends Nvidia partnership

aws-graviton4-and-aws-trainium2-prototype

The Graviton 4 chip, left, is a general-purpose microprocessor chip being used by SAP and others for large workloads, while Trainium 2 is a special-purpose accelerator chip for very large neural network programs such as generative AI.

At its annual AWS re:Invent developer conference in Las Vegas, Amazon on Tuesday announced a new version of Trainium 2, its dedicated chip for training neural networks. Trainium 2 is tuned specifically for training so-called large language models (LLMs) and foundation models — the kinds of generative AI programs such as OpenAI's GPT-4.

The company also unveiled a new version of its custom microprocessor, Graviton 4, and said it is extending its partnership with Nvidia to run Nvidia's most advanced chips in its cloud computing service.

Also: The future of cloud computing, from hybrid to edge to AI-powered

The Trainium 2 is designed to handle neural networks with trillions of parameters, or neural weights, which are the functions of the program's algorithm that give it scale and power, generally speaking. Scaling to larger and larger parameters is a focus of the entire AI industry.

The trillion-parameter count has become something of an industry obsession because of the fact that the human brain is believed to contain 100 trillion neuronal connections — making a trillion-parameter neural network program seem related to the human brain, whether or not it in fact is.

The chips are "designed to deliver up to four times faster training performance and three times more memory capacity" than their predecessor, "while improving energy efficiency (performance/watt) up to two times," said Amazon.

Amazon is making the chips available in instances of its EC2 cloud computing service known as "Trn2" instances. The instance offers 16 of the Trainium 2 chips operating in concert, which can be extended to 100,000 instances, Amazon said. Those larger instances are interconnected using the company's networking system, called the Elastic Fabric Adapter, which can provide for a total of 65 exaFLOPs of computing power. (One exaFLOP is a billion, billion floating point operations per second.)

Also: AWS unveils local cloud zones for exclusive customer use

At that scale of compute, said Amazon, "Customers can train a 300-billion parameter LLM in weeks versus months."

Besides serving customers, Amazon has additional incentives to continue to push the envelope on AI silicon. The company has invested $4 billion in privately held generative AI startup Anthropic, a group that broke off from OpenAI. That investment puts the company in a position to compete with Microsoft's exclusive deal with OpenAI.

The Graviton 4 chip, which is built on the microprocessor intellectual property of ARM Holdings, competes with processors from Intel and Advanced Micro Devices based on the older x86 chip standard. The Graviton 4 has "30% better compute performance," Amazon said.

Also: Why Nvidia is teaching robots to twirl pens and how generative AI is helping

Unlike the Trainium chips for AI, Graviton processors are meant to run more conventional workloads. Amazon AWS said customers — including Datadog, DirecTV, Discovery, Formula 1, Nielsen, Pinterest, SAP, Snowflake, Sprinklr, Stripe, and Zendesk — use the Graviton chips "to run a broad range of workloads, such as databases, analytics, web servers, batch processing, ad serving, application servers, and microservices."

SAP said in prepared remarks that it has been able to achieve "35% better price performance for analytical workloads" running its HANA in-memory database on the Graviton chips, and that "we look forward to evaluating Graviton4, and the benefits it can bring to our joint customers."

The new chips follow by two years the introduction in 2021 of Graviton 3 and the original Trainium.

Amazon's news follows the introduction by Microsoft last week of its first chips for AI. Alphabet's Google, the other cloud titan alongside Amazon and Microsoft, preceded both in 2016 with the first cloud chip for AI, the TPU, or Tensor Processing Unit, of which it has since offered multiple generations.

Also: Amazon turns Fire TV Cube into a thin client for enterprises

In addition to the two new chips, Amazon said it extended its strategic partnership with AI chip giant Nvidia. AWS will be the first cloud service to run the forthcoming GH200 Grace Hopper multi-chip product from Nvidia, which combines the Grace ARM-based CPU and the Hopper H100 GPU chip.

The GH200 chip, which is supposed to start shipping next year, is the next version of the Grace Hopper combo chip, announced earlier this year, which is already shipping in its initial version in computers from Dell and others.

The GH200 chips will be hosted on AWS via Nvidia's purpose-built AI computers, the DGX, which the two companies said will speed up the training of neural networks with more than a trillion parameters.

Nvidia said it will make AWS its "primary cloud provider for its ML research and development."

Artificial Intelligence

Salmonn: Towards Generic Hearing Abilities For Large Language Models

Hearing, which involves the perception and understanding of generic auditory information, is crucial for AI agents in real-world environments. This auditory information encompasses three primary sound types: music, audio events, and speech. Recently, text-based Large Language Model (LLM) frameworks have shown remarkable abilities, achieving human-level performance in a wide range of Natural Language Processing (NLP) tasks. Additionally, instruction tuning, a training method using pairs of reference responses and user prompts, has become popular. This approach trains large language models to more effectively follow open-ended user instructions. However, current research is increasingly focused on enhancing large language models with the capability to perceive multimodal content.

Focusing on the same, in this article, we will be talking about SALMONN or Speech Audio Language Music Open Neural Network, a state of the art open speech audio language music neural network built by incorporating speech and audio encoders with a pre-trained text-based large language model into a singular audio-text multimodal model. The SALMONN model enables Large Language Models to understand and process generic audio inputs directly, and deliver competitive performance on a wide array of audio & speech tasks used in training including auditory information-based question answering, speech recognition and translation, speaker verification, emotion recognition, audio & music captioning, and much more. We will be taking a deeper dive into the SALMONN framework, and explore its working, architecture, and results across a wide array of NLP tasks. So let’s get started.

SALMONN : An Introduction to Single Audio-Text Multimodal Large Language Models

SALMONN stands for Speech Audio Language Music Open Neural Network, and it is a single audio-text multimodal large language model framework capable of perceiving and understanding three basic audio or sound types including speech, audio events, and music. The SALMONN model enables Large Language Models to understand and process generic audio inputs directly, and deliver competitive performance on a wide array of audio & speech tasks.

To boost its performance on both speech, and non-speech audio tasks, the SALMONN framework employs a dual encoder structure consisting of a BEATs audio encoder, and a speech encoder sourced from the Whisper speech model. Additionally, the SALMONN framework also uses a window-level Q-Former or query Transformer as a connection module to effectively convert an output sequence of variable-length encoder to augmented audio tokens of a variable number, and ultimately achieve high temporal resolution for audio-text alignment. The LoRA or Low Rank Adaptation approach is used as a cross-modal adaptor to the Vicuna framework to align its output space with its augmented input space in an attempt to further boost its performance. In the SALMONN framework, the ability to perform cross-modal tasks unseen during the training phase lost during training of instructions as cross-modal emergent abilities which is the primary reason why the SALMONN framework implements an additional few-shot activation stage to regain the LLM framework’s general emergent abilities.

Furthermore, the framework makes use of a wide array of audio events, music benchmarks, and speech benchmarks to evaluate its cognitive hearing abilities, and divides the benchmarks in three levels. At the first benchmark level, the framework trains eight tasks in instruction training including translation, audio captioning, and speech recognition. The other two benchmark levels are untrained tasks with the second level benchmark consisting of 5 speech-based Natural Language Processing tasks like slot filling and translation to untrained languages relying on high-quality multilingual alignments between text and speech tokens. The final level benchmark tasks attempt to understand speech and non-speech auditory information for speech-audio co-reasoning and audio-based storytelling.

To sum it up, the SALMONN framework is

  1. The first multimodal large language model capable of understanding and perceiving general audio inputs including audio events, speech, and music to the maximum of its ability.
  2. An attempt to analyze cross-modal emergent abilities offered by implementing the LoRA scaling factor, and using an extra budget-friendly activation stage during training to activate cross-modal emergent abilities of the framework.

SALMONN : Architecture and Methodology

In this section, we will be having a look at the architecture, training method, and experimental setup for the SALMONN framework.

Model Architecture

At the core of its architecture, the SALMONN framework synchronizes and combines the outputs from two auditory encoders following which the framework implements a Q-Former at the frame level as a connection module. The output sequence generated by the Q-Former is merged with text instruction prompts and it is then provided as an input to the LoRA adaptation approach to generate the required response.

Auditory Encoders

The SALMONN framework makes use of two auditory encoders: a non-speech BEATs audio encoder, and a speech encoder sourced from OpenAI’s Whisper framework. The BEATs audio encoder is trained to use the self-supervised iterative learning approach in an attempt extract non-speech high-level audio semantics whereas the speech encoder is trained on a high amount of weakly supervised data for speech recognition and speech translation tasks with the output features of the encoder suitable to include background noise and speech information. The model first tokenizes the input audio, and follows it up by masking and predicting it in training. The resulting auditory features of these two encoders complement each other, and are suitable for both speech, and non-speech information.

Window Level Q-Former

Implementing the Q-Former structure is a common approach used in the LLM frameworks to convert the output of an image encoder into textual input tokens, and some modification is needed when dealing with audio tokens of varying lengths. To be more specific, the framework regards the encoder output of the input image as a concatenated encoder output sequence, and the Q-Former deploys a fixed number of trainable queries to transform the encoder output sequence into textual tokens using stacked blocks of Q-Former. A stacked Q-Former block resembles a Transformer decoder block with the exceptions being removing casual masks in the self-attention layers, and the use of a fixed number of trainable static queries in the initial blocks.

LoRA and LLM

The SALMONN framework also deploys a Vicuna LLM which is a LLaMA large language model framework fine-tuned to follow instructions more accurately, and effectively. The LoRA framework is a common method used for parameter-efficient fine-tuning, and its inclusion in the SALMONN framework to value weight matrices and adapt the query in the self-attention layers.

Training Method

The SALMONN framework makes use of a three-stage cross-modal training approach. The training stage comprises a pre-training stage, and an instruction tuning stage that are included in most visual LLM frameworks, and an additional activation tuning stage is implemented to resolve over-fitting issues encountered during audio captioning and speech recognition tasks.

Pre-Training Stage

To limit the gap observed between pre-trained parameters including encoders & LLM, and randomly initialized parameters including adaptor & connection modules, the SALMONN framework uses a large amount of audio captioning and speech recognition data to pre-train the LoRA and Q-Former components. These tasks contain vital auditory information about the key contents of audio events both speech and non-speech, and neither of them require complex understanding or reasoning to learn alignment between textual and auditory information.

Instruction Fine-Tuning Stage

The instruction fine-tuning stage implemented in the SALMONN framework resembles the one implemented in NLP and visual LLM frameworks by using a list of audio events, music tasks and speech events to fine-tune audi-text instructions. The tasks are prioritized on the basis of their importance across different tests including phone recognition, overlapping speech recognition, and music captions. Furthermore, textual information paired with audio data forms the basis for generating instruction prompts.

Task Over-Fitting

Even when implementing only the first two training stages, the SALMONN framework delivers competitive results on instruction tuning tasks, although the performance is not up to the mark when performing cross-modal tasks, especially on tasks that require cross-modal co-reasoning abilities. Specifically, the model occasionally violates instruction prompts that result in the generation of irrelevant or incorrect responses, and this phenomenon is referred to as task overfitting in the SALMONN framework, and the Activation Tuning stage is implemented to resolve these overfitting issues.

Activation Tuning Stage

An effective approach to resolve overfitting issues is to regularize intrinsic conditional language models using longer and more diverse responses like storytelling or auditory-information based question answering. The framework then generates the pair training data for such tasks using text paired with audio or speech or music captions.

Task Specifications

To evaluate SALMONN’s zero-shot cross-modal emergent abilities, developers have included 15 speech, audio and music tasks divided across three levels.

Level 1

In the first level, tasks are used for instruction tuning, and therefore, they are the easiest set of tasks that the SALMONN framework has to perform.

Level 2

The second level consists of untrained tasks, and the complexity level is higher when compared to level 1 tasks. In level 2, tasks are Natural Language Processing based tasks including speech keyword extraction that is used to evaluate the framework's accuracy when extracting certain keywords using speech. Other tasks include SQQA or Spoken Query-based Question Answering that evaluates the common sense knowledge the framework extracts using speech questions, a SF or Speech-based Slot Filling task to evaluate the accuracy of slot values, and finally, there are two AST tasks for English to German, and English to Japanese conversions.

Level 3

The complexity of tasks in Level 3 is the maximum when compared to other two levels, and it includes SAC or Speech Audio Co-Reasoning, and Audio-based Storytelling tasks. The SAC task requires the SALMONN framework to understand a question included in the audio clip fed to the model, find supportive evidence using audio events or music in the background, and finally generate an appropriate reason to answer the question. The Audio-based storytelling tasks require the model to generate a meaningful story based on the auditory information sourced from general audio inputs.

Results

Level 1 Tasks

The following table demonstrates the results on Level 1 tasks, and as it can be observed, the SALMONN framework returns competitive results on Level 1 tasks with or without activation-tuning.

Level 2 and 3 Tasks

Although the SALMONN framework returns competitive results on Level 1 tasks even without fine-tuning, the same cannot be said for Level 2 and Level 3 tasks as without activation, the SALMONN framework suffers heavily from over-fitting on tasks. The performance dips even further on SQQA, SAC, and Storytelling tasks with emphasis on multimodal interactions, and the SALMONN framework struggles to follow instructions without activation tuning. However, with activation tuning, the results improve considerably, and the results are included in the following image.

Discounting LoRA Scaling Factor

Discounting LoRA Scaling Factor evaluates the influence of using time-test discounting of the LoRA scaling factor to minimize overfitting issues on tasks. As it can be observed in the following figure, a decrease in the LoRA scaling factor to 2.0 elevates the cross-modal reasoning ability of the SALMONN framework on ASR & PR tasks, SQQA tasks, Storytelling tasks, and SAC tasks respectively.

Evaluating Task-Overfitting

To emphasize on activation tuning, the SALMONN framework analyzes the changes in perplexity during the three training stages, and as it can be seen in the following image, perplexity changes for AAC and ASR tasks have small final values post the first training stage, indicating the model’s learning of cross-modal alignments.

Furthermore, the perplexity of the PR task also drops post instruction tuning owing to its reliance on the LoRA component to learn the output tokens. It is also observed that although instruction tuning helps in reducing the perplexity on Storytelling and SAC tasks, the gap is still large enough to perform the tasks successfully unless an additional activation stage is added or the LoRA component is removed.

Activation Tuning

The SALMONN framework dives into different activation methods including training the model on text-based QA task pairs with long answers, or using audio-based long written stories, whereas using long speech transcriptions for ASR tasks. Both the Q-Former and LoRA components are fine-tuned using these three methods. Furthermore, the framework ignores the audio and Q-Former inputs in an attempt to fine-tune the LoRA and Vicuna components as an adaptive text-based large language model, and the results are demonstrated in the following image, and as it can be seen, the model cannot be activated by ASR ( training ASR with long labels), nor Story or Text-based by training LoRA component using text prompt inputs.

Final Thoughts

In this article, we have talked about SALMONN or Speech Audio Language Music Open Neural Network, a single audio-text multimodal large language model framework capable of perceiving and understanding three basic audio or sound types including speech, audio events, and music. The SALMONN model enables Large Language Models to understand and process generic audio inputs directly, and deliver competitive performance on a wide array of audio & speech tasks.

The SALMONN framework delivers competitive performance on a wide array of trained tasks including audio captioning, speech translation & recognition, and more while generalizing to a host of untrained understanding tasks including speech translation for keyword extracting and untrained languages. Owing to its abilities, the SALMONN framework can be regarded as the next step towards enhancing the generic hearing abilities of large language models.