This week in AI: AI-powered personalities are all the rage

This week in AI: AI-powered personalities are all the rage Devin Coldewey Kyle Wiggers 9 hours

Keeping up with an industry as fast-moving as AI is a tall order. So until an AI can do it for you, here’s a handy roundup of recent stories in the world of machine learning, along with notable research and experiments we didn’t cover on their own.

Last week during its annual Connect conference, Meta launched a host of new AI-powered chatbots across its messaging apps — WhatsApp, Messenger and Instagram DMs. Available for select users in the U.S., the bots are tuned to channel certain personalities and mimic celebrities including Kendall Jenner, Dwyane Wade, MrBeast, Paris Hilton, Charli D’Amelio and Snoop Dogg.

The bots are Meta’s latest bid to boost engagement across its family of platforms, particularly among a younger demographic. (According to a 2022 Pew Research Center survey, only about 32% of internet users aged 13 to 17 say that they ever use Facebook, an over-50% decline from the year prior.) But the AI-powered personalities are also a reflection of a broader trend: the growing popularity of “character-driven” AI.

Consider Character.AI, which offers customizable AI companions with distinct personalities, like Charli D’Amelio as a dance enthusiast or Chris Paul as a pro golfer. This summer, Character.AI’s mobile app pulled in over 1.7 million new installs in less than a week while its web app was topping 200 million visits per month. Character.AI claimed that, moreover, as of May, users were spending on average 29 minutes per visit — a figure that the company said eclipsed OpenAI’s ChatGPT by 300% as ChatGPT usage declined.

That virality attracted backers including Andreessen Horowitz, who poured well over $100 million in venture capital into Character.AI, which was last valued at $1 billion.

Elsewhere, there’s Replika, the controversial AI chatbot platform, which in March had around 2 million users — 250,000 of whom were paying subscribers.

That’s not to mention Inworld, another AI-driven character success story, which is developing a platform for creating more dynamic NPCs in video games and other interactive experiences. To date, Inworld hasn’t shared much in the way of usage metrics. But the promise of more expressive, organic characters, driven by AI, has landed Inworld investments from Disney and grants from Fortnite and Unreal Engine developer Epic Games.

So clearly, there’s something to AI-powered chatbots with personalities. But what is it?

I’d wager to say that chatbots like ChatGPT and Claude, while undeniably useful in decidedly professional contexts, don’t hold the same allure as “characters.” They’re not as interesting, frankly — and it’s no surprise. General-purpose chatbots were designed to complete specific tasks, not hold an elivening conversation.

But the question is, will AI-powered characters have staying power? Meta’s certainly hoping so, considering the resources it’s pouring into its new bot collection. I’m not sure myself — as with most tech, there’s a decent chance the novelty will wear off eventually. And, then it’ll be onto the next big thing — whatever that ends up being.

Here are some other AI stories of note from the past few days:

  • Spotify tests AI-generated playlists: References discovered in the Spotify app’s code indicate the company may be developing generative AI playlists users could create using prompts, Sarah reports.
  • How much are artists making from generative AI? Who knows? Some generative AI vendors, like Adobe, have established funds and revenue sharing agreements to pay artists and other contributors to the data sets used to train their generative AI models. But it’s not clear how much these artists can actually earn, TC learned.
  • Google expands AI-powered search: Google opened up its generative AI search experience to teenagers and introduced a new feature to add context to the content that users see, along with an update to help train the search experience’s AI model to better detect false or offensive queries.
  • Amazon launches Bedrock in GA, brings CodeWhisperer to the enterprise: Amazon announced the general availability of Bedrock, its managed platform that offers a choice of generative AI models from Amazon itself and third-party partners through an API. The company also launched an enterprise tier for CodeWhisperer, Amazon’s AI-powered service to generate and suggest code.
  • OpenAI entertains hardware: The Information reports that storied former Apple product designer Jony Ive is in talks with OpenAI CEO Sam Altman about a mysterious AI hardware project. In the meantime, OpenAI — which is planning to soon release a more powerful version of its GPT-4 model with image analysis capabilities — could see its secondary-market valuation soar to $90 billion.
  • ChatGPT gains a voice: In other OpenAI news, ChatGPT evolved into much more than a text-based search engine, with OpenAI announcing recently that it’s adding new voice and image-based smarts to the mix.
  • The writers’ strike and AI: After almost five months, the Writers Guild of America reached an agreement with Hollywood studios to end the writers’ strike. During the historic strike, AI emerged as a key point of contention between the writers and studios. Amanda breaks down the relevant new contract provisions.
  • Getty Images launches an image generator: Getty Images, one of the largest suppliers of stock images, editorial photos, videos and music, launched a generative AI art tool that it claims is “commercially safer” than other, rival solutions on the market. Prior to the launch of its own tool, Getty had been a vocal critic of generative AI products like Stable Diffusion, which was trained on a subset of its image content library.
  • Adobe brings gen AI to the web: Adobe officially launched Photoshop for the web for all users with paid plans. The web version, which was in beta for almost two years, is now available with Firefly-powered AI tools such as generative fill and generative expand.
  • Amazon to invest billions in Anthropic: Amazon has agreed to invest up to $4 billion in the AI startup Anthropic, the two firms said, as the e-commerce group steps up its rivalry against Microsoft, Meta, Google and Nvidia in the fast-growing AI sector.

More machine learnings

When I was talking with Anthropic CEO Dario Amodei about the capabilities of AI, he seemed to think there were no hard limits that we know of — not that there are none whatsoever, but that he had yet to encounter a (reasonable) problem that LLMs were unable to at least make a respectable effort at. Is it optimism or does he know of what he speaks? Only time will tell.

In the meantime, there’s still plenty of research going on. This project from the University of Edinburgh takes neural networks back to their roots: neurons. Not the complex, subtle neural complexes of humans, but the simpler (yet highly effective) ones of insects.

From the paper, a diagram showing views of the robot and some of its vision system data.

Ants and other small bugs are remarkably good at navigating complex environments, despite their more rudimentary vision and memory capabilities. The team built a digital network based on observed insect neural networks, and found that it was able to successfully navigate a small robot visually with very little in the way of resources. Systems in which power and size are particularly limited may be able to use the method in time. There’s always something to learn from nature!

Color science is another space where humans lead machines, more or less by definition: we are constantly striving to replicate what we see with better fidelity, but sometimes that fails in ways that in retrospect seem predictable. Skin tone, for example, is imperfectly captured by systems designed around light skin — especially when ML systems with biased training sets come into play. If an imaging system doesn’t understand skin color, it can’t expose and adjust the exposure and color properly.

Images from Sony research on more inclusive skin color estimation.

Sony is aiming to improve these systems with a new metric for skin color that more comprehensively but efficiently defines it using a color scale as well as perceived light/dark levels. In the process of doing this they showed that bias in existing systems extends not just to lightness but to skin hue as well.

Speaking of fixing photos, Google has a new technique almost certainly destined (in some refined form) for its Pixel devices, which are heavy on the computational photography. RealFill is a generative plug-in that can fill in an image with “what should have been there.” For instance, if your best shot of a birthday party happens to crop out the balloons, you give the system the good shot plus some others from the same scene. It figures out that there “should” be some balloons at the top of the strings and adds them in using information from the other pictures.

Image Credits: Google/

It’s far from perfect (they’re still hallucinations, just well informed hallucinations) but used judiciously it could be a really helpful tool. Is it still a “real” photo though? Well, let’s not get into that just now.

Lastly, machine learning models may prove more accurate than humans in predicting the number of aftershocks following a big earthquake. To be clear (as the researchers emphasize), this isn’t about “predicting” earthquakes, but characterizing them accurately when they happen so that you can tell whether that 5.8 is the type that leads to three more minor quakes within an hour, or only one more after 20 minutes. And the latest models are still only decent at it, under specific circumstances — but they are not wrong, and they can work thorough large amounts of data quickly. In time these models may help seismologists better predict quakes and aftershocks, but as the scientists note, it’s far more important to be prepared; after all, even knowing one is coming doesn’t stop it from happening.

KDnuggets Top Posts for August 2023: Forget ChatGPT, This New AI Assistant Will Change the Way You Work

KDnuggets Top Posts for August 2023: Forget ChatGPT, This New AI Assistant Will Change the Way You Work

It's time for the top posts of August 2023! Check out what people were checking out of what was published in August.

As a reminder: top posts are defined as the posts with the highest number of views normalized over the first 14 days of post publication.

Check them out below!

  1. Forget ChatGPT, This New AI Assistant Is Leagues Ahead and Will Change the Way You Work Forever by Abid Ali Awan
  2. 7 Projects Built with Generative AI by Eugenia Anello
  3. Best Python Tools for Building Generative AI Applications Cheat Sheet by KDnuggets
  4. Harnessing ChatGPT for Automated Data Cleaning and Preprocessing by Bala Priya C
  5. Data Scientists Need to Specialize to Survive the Tech Winter by Nate Rosidi
  6. 7 Steps to Mastering Data Cleaning and Preprocessing Techniques by Eugenia Anello
  7. 5 Ways You Can Use ChatGPT's Code Interpreter For Data Science by Abid Ali Awan
  8. The Best Courses for AI from Universities with YouTube Playlists by Nisha Arya

Thanks to everyone who contributed, and looking forward to the top posts of next month!

More On This Topic

  • Forget ChatGPT, This New AI Assistant Is Leagues Ahead and Will Change the…
  • A New Way of Managing Deep Learning Datasets
  • A new book that will revolutionize the way your organization approaches…
  • I Used ChatGPT (Every Day) for 5 Months. Here Are Some Hidden Gems That…
  • What Is ChatGPT Doing and Why Does It Work?
  • 5 ChatGPT Features to Boost your Daily Work

ChatGPT’s new web browsing feature is a big disappointment. Use this plugin instead

aisearch2-gettyimages-1631437508

Last week, OpenAI (the folks who brought you the wildly disruptive ChatGPT) announced that the AI chatbot will no longer be limited to data from before September 2021 (or in the case of GPT-4, from before January 2022). In fact, it will be able to browse the web and provide insights into current data. The only limitation? This feature is available exclusively to paying Plus and Enterprise customers.

In this article, we'll explore exactly what the new ChatGPT Plus update does, plus a lot of what it fails to do.

Also: How does ChatGPT actually work?

TL;DR: It's an odd beast, and it's quite disappointing.

How to enable Browse with Bing

The new browsing capability is provided via a beta option called Browse with Bing. To enable it, you must be using ChatGPT Plus. Then go to the Settings menu, choose Beta features, and turn on Browse with Bing.

When you start a new session, you'll have to decide if you want to run Browse with Bing, Advanced Data Analysis, or plugins.

As I discussed a few weeks ago, you can only run one option at a time, and you have to change to an entirely new session every time you want to change to a different option.

Also: My two favorite ChatGPT Plus plugins and the remarkable things I can do with them

What exactly is Browse with Bing doing?

Neither Microsoft nor OpenAI are saying much about what Browse with Bing actually does. You may recall that ChatGPT's browsing feature was available for a short time last summer, only to be disabled because it went a bit rogue, including bypassing paywalls. Now, it's back, but OpenAI's entire corpus of information on it is limited to a single paragraph.

For ChatGPT to have web access, it needs to do two things: search and retrieve. It needs to construct some sort of search string or strings from the prompt given and then pass that search string to a crawled index of the Internet. Clearly, this is where Bing comes in. Bing, like Google, has a representation of the entire Internet in its indexes and can return search results.

Next, ChatGPT must be able to retrieve the contents of web pages based on their URLs, extract the content from the ads, process the content in context, and provide answers to the prompt giver.

All of this usually works in the background, via APIs and calls. I assume that Browse with Bing in ChatGPT does the same thing. But, oddly enough, that's not how it's represented.

Also: I asked ChatGPT, Bing, and Bard what worries them. Google's AI went Terminator on me

I issued a query that caused Browse with Bing to look on ZDNET for some information. Here's a rundown of the notification screens I got during that interaction:

Sorry about the rough quality. I captured this in video and had to convert it to show you the individual frame-by-frame notifications.

What's weird about this is that Browse with Bing says "Clicking on www.zdnet.com," as if it were actually clicking a mouse. Then it says "scrolling page" as if it were actually scrolling the page.

It's highly unlikely that this feature is implemented with a bunch of robot screens and mice actually clicking on pages. So why indicate progress that way? Is it some attempt to consumerize or dumb down notifications for the Bing audience?

It doesn't hurt anything, but it's strange.

Browse with Bing vs WebPilot

I've been using ChatGPT Plus to access current Web information for months now. At first, I used the MixerBox WebSearchG plugin, but found that to be unreliable. For the past few months, I've been using the WebPilot plugin and have been quite pleased.

Does Browse with Bing bring anything to the table that WebPilot doesn't? Nope. In fact, it's more limited. A lot more limited.

Also: Extending ChatGPT: Can AI chatbot plugins really change the game?

To demonstrate this, I'll share three tests. Each test compares the results of Browse with Bing inside ChatGPT Plus to WebPilot running as a ChatGPT Plus plugin. Because you can't run Browse with Bing and WebPilot in the same session, all of these tests were run in separate sessions.

The first two tests used ZDNET as a test environment. For the third, I tasked ChatGPT with analyzing a breaking story in the news.

Finding an article reference

In July, I wrote a story about Google storage where I used the term "infraquake." I then referenced that story in another article in September.

My test was to see which of the two add-ons could find that information, given the prompt "what does gewirtz mean by infraquake." Note that I didn't tell ChatGPT that this was a ZDNET article, nor did I tell it which Gewirtz I was referring to.

Here's how Browse with Bing responded (the first time):

As you can see, Browse with Bing had no idea what I was talking about. However, WebPilot found it quickly:

Some time later in my testing, I once again asked the Browse with Bing add-on and this time, it found both articles.

Interestingly, Browse with Bing does provide a footnote reference for source information. However, note from the screenshot that while it discusses two articles, it only provides one source citation.

I should mention that one of the benefits of Browse with Bing is supposed to be source citations, but I found that a number of my tests ran, provided information, and did not provide a footnote link for source data. So it's not consistent.

Comparing writing contexts

This next assignment was more complex. Let me give you a bit of background. Both Sabrina Ortiz and I write about AI for ZDNET. We both cover it from different angles and I wanted to see if ChatGPT could ascertain our respective approaches. We both have author pages, which list our articles. So I fed ChatGPT this prompt:

Here's the result. The difference couldn't have been more stark. Not only does WebPilot provide a profile and some recent articles, it actually compares the types of articles we write. On the other hand, Browse with Bing fails to find any content to compare.

This is presented small, so click the zoom button to see the full text.

The failure of Browse with Bing in comparison to the excellent job of WebPilot is rather astounding.

Briefing on current news

The WGA writers' strike was settled just last week. I asked both tools to "Give me a complete briefing on the status of the writers' strike."

This is presented small, so click the zoom button to see the full text.

As you can see, Browse with Bing did provide some value here. But it's not nearly as comprehensive as the reply from WebPilot.

Next, I wanted to see if the AI could describe the details of the final agreement. I asked, "Describe the deal."

This is presented small, so click the zoom button to see the full text.

As you can see, Browse with Bing provided a truncated but still relevant answer. However, the more detailed information provided by WebPilot is substantially more useful.

What does it all mean?

The bottom line is that Browse with Bing is disappointing. Fortunately, WebPilot does everything that you would have expected from Browse with Bing, so good web search functionality is available to you already. Browse with Bing is just a weird little extension that seems mired by… something… maybe too many Zoom or Teams meetings and a ton of intra-organizational compromise?

ChatGPT is exciting and disruptive. ChatGPT with Advanced Data Analysis is game-changing. ChatGPT with the WebPilot plugin lets you do remarkable things.

Also: Generative AI will far surpass what ChatGPT can do. Here's everything on how the tech advances

ChatGPT with Browse with Bing is meh. It doesn't make anything new possible. It doesn't do anything better than any other solution. My best description is that it's a web summarizing tool that just phones it in.

It's still in beta, so maybe it will get better. Until then, give it a pass.

You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter on Substack, and follow me on Twitter at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, and on YouTube at YouTube.com/DavidGewirtzTV.

Artificial Intelligence

Humata AI summarizes and answers questions about your PDFs

Humata AI summarizes and answers questions about your PDFs Kyle Wiggers 10 hours

Cyrus Khajvandi, a Stanford biology graduate and two-time entrepreneur, often found it challenging to stay on top of scientific research while managing his daily workload. Recognizing that he wasn’t the only one — and that AI technology was becoming more accessible — Khajvandi began developing an AI platform to summarize and answer questions about documents, particularly scientific studies.

The platform, Humata AI, launched in February, with former Labelbox founder Dan Rasmuson joining as CTO. And it quickly gained traction — processing tens of millions of pages of files, growing to a user base of millions and securing $3.5 million in funding from investors, including Google’s Gradient Ventures, ARK invest and M13.

“Our mission at Humata is to empower people and organizations to make smarter and faster decisions by being able to ask questions across all their files,” Khajvandi told TechCrunch in an email interview. “Humata is like [OpenAI’s] ChatGPT for all your files.”

Humata is exceptionally simple in its execution. True to the premise, the platform simply lets users ask questions about their files — namely PDF files — and get answers. Users can upload one or more PDFs and ask questions across them; Khajvandi says that customers include not only academics but professionals in law, the oil and gas industry and customer support.

Now, chatbots like the aforementioned ChatGPT and Anthropic’s Claude offer similar file-analyzing features. But Khajvandi makes the case that Humata — in part because of its limited functionality and focus — is more robust.

“People can ask AI any question and get the answer from their own data instantly with highlighted references,” he said. “This is possible because of the recent advancements in AI enabling every worker to get instant answers to their questions.”

Humata AI

Image Credits: Humata AI

Now, AI isn’t necessarily the best at summarizing. Fast Company tested ChatGPT’s ability to sum up articles, and found that the model had a tendency to get content wrong, leave pieces out or outright invent facts not contained in the documents it summarized.

There’s also the obvious privacy question. Companies — and individual users, for that matter — might not feel comfortable uploading their documents to Humata’s platform for processing — particularly if the documents contain sensitive info.

Khajvandi stands by Humata’s summarization skills, claiming that the company trained its models on “diverse datasets” and “rigorously tested” them for bias. He also says that Humata only collects “necessary data,” and has implemented “strong safeguards” to prevent unauthorized access.

“We ensure informed consent, helping users understand what they’re agreeing to,” Khajvandi added. “As our AI systems advance, we’re careful not to infer sensitive information without explicit permission. We adhere to legal and ethical standards across different regions and cultures, making Humata enterprise-ready.”

Humata, which now has thousands of customers on its paid plan (or so Khajvandi claims), plans to put the capital it has raised so far ($3.58 million, inclusive of a pre-seed round) toward enhancing its AI capabilities, improving the user experience and expanding its market reach.

“We chose to raise now because we’ve seen a growing demand for efficient, AI-driven solutions in synthesizing insights from vast volumes of enterprise files,” Khajvandi said. “The funds will help us develop new features, refine our existing products and expand into new markets, ultimately by empowering businesses to make better and faster decisions with their private data using Humata.”

Automate Graphic Design Activity with ChatGPT Canva Plugin

Most of us must already have heard about the Canva company before. For you who did not know, Canva is a platform that provides free and paid platform services for any graphic design activity. It’s mostly aimed at social media and presentation design, but the platform can be used for any graphical design.

Recently, Canva made a move to include their plugin into the ChatGPT. We all know that ChatGPT is a platform that can help us in our daily work, such as planning, answering questions, and code debugging. The addition of the Canva plugin in ChatGPT opens various graphical design activities we can perform on the platform.

How does it work? Let’s get into it.

Setup the Canva Plugin

To access the ChatGPT Canva plugin, we need to subscribe to ChatGPT Plus, which costs around 20 USD per month when this article was written. After you have subscribed as the plus user, go to the settings and turn on the Plugins option within the Beta features.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

Access the Plugin store under the GPT-4 model selection with everything set.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

In the Plugin store, type Canva and install the plugin. It would only take a second, and now we can use the plugin for our activity.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author Trying Out Canva Plugin

Canva were perfect for any social media or presentation activity, so let’s try to use the plugin for any of these activities. First, we want to create a Twitter post about Introduction to Python. In this case, I would use the prompt: “Help me create a twitter post about Introduction to Python and guide me to make a nice post design that would attract attention.”

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

The result is the selection of various Canva templates to design our social media post template.

Currently, the plugin cannot access this template directly, so we need to design them ourselves. However, we can still ask for guidance from the ChatGPT. ChatGPT asks for our template selection in the following prompt, and I select the third one. In the process, ChatGPT then brings the content suggestion and asks if we need help to add more design.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

We certainly need more help in creating our social media post design. In this case, we can ask ChatGPT to guide us on which element is the best to add to our design.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

With the suggestion coming from the ChatGPT, we can further ask them for guidance to add these elements into Canva.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

You can keep asking for further guidance from the ChatGPT to improve your design until you are satisfied.

You can still ask the ChatGPT plugin to do various design activities. For example, the below sample is the prompt for “Help me create a logo for my data science business that targets the data enthusiast.”.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

The ChatGPT would provide a selection of templates once more. This time, it’s all the template for creating a logo. Like the previous example, we can ask further for customization guidance.

Automate Graphic Design Activity with ChatGPT Canva Plugin
Image by Author

The guidance would help you create that perfect logo for your business.

While it’s not a total automation process, the ChatGPT Canva plugin helps us by automating the template selection process and providing guidance to achieve the best graphic design. You can use the plugin to help you create a presentation or combine another plugin to provide you with better guidance.

Conclusion

ChatGPT Canva plugin is a plugin that helps the user to automate certain graphic design activities, including template selection and how to improve the design template. The plugin is only available for ChatGPT Plus users so don’t forget to subscribe if you want to use the Canva plugin.
Cornellius Yudha Wijaya is a data science assistant manager and data writer. While working full-time at Allianz Indonesia, he loves to share Python and Data tips via social media and writing media.

More On This Topic

  • Noteable Plugin: The ChatGPT Plugin That Automates Data Analysis
  • 5 Tasks To Automate With Python
  • Automate Microsoft Excel and Word Using Python
  • The Prefect Way to Automate & Orchestrate Data Pipelines
  • Automate the Boring Stuff with GPT-4 and Python
  • Automate Your Codebase with Promptr and GPT

Generative AI will far surpass what ChatGPT can do. Here’s everything on how the tech advances

A woman's hand is asking an AI chatbot pre-typed questions & the AI website is answering.

More than any of the many headline achievements of artificial intelligence — winning at chess, predicting the folding of proteins, labeling cats and dogs — the form of AI known as generative AI has captivated the global imagination.

ChatGPT became the fastest-growing software program in history in January, reaching a hundred million users in less than two months from its public debut. It spawned numerous rivals, both proprietary programs such as Google's Bard, and open-source alternatives such as the University of California at Berkeley's Koala. The rush of excitement has prompted an arms race between tech giants Microsoft and Google and their peers, and a surge in the business of AI chip maker Nvidia.

Excitement over large language models has led to a flowering of numerous proprietary and open-source programs of increasing scale just for text. The diagram is from the 2023 paper "Emotional Intelligence of Large Language Models" by Xuena Wang and colleagues at Tsinghua University.

All of this fervent activity has its roots in the simple fact that unlike past AI programs, which mostly produced a numeric score — a "1" for a cat picture, a "0" for a dog picture — ChatGPT, and image programs such as Stability.ai's Stable Diffusion, and OpenAI's DALL-E, reproduce something of the world.

By outputting a paragraph, a picture, or even the skeleton of a computer program, such programs are mirroring society's creations.

The mirroring aspect is going to increase dramatically in a very short span of time.

Today's generative programs will seem primitive in comparison to the powers of programs that will be prevalent at the end of this year as they output many more kinds of things.

Moving to multiple modalities

What computer scientists call mixed modalities, or "multi-modality," will take center stage, as programs fuse text, images, "point clouds" of physical space, sounds, video, and entire computer functions as smart applications.

The mixed modality will make possible far more capable programs and will contribute to a long-held goal of continuous learning. It may even advance the goal of "embodied AI" by giving a lift to robotics.

"ChatGPT was made for entertainment, and it does a lot of things really well, but it's, sort-of, a demo," said Naveen Rao, founder of AI startup MosaicML, in an interview with ZDNET. "Now we have to start thinking about, well, if I'm using this for a purpose, how do I make that better?"

Rao, whose company was acquired by Databricks for its expertise in running AI programs, now serves as vice president of generative AI at Databricks.

Also: Meta's AI image generator says language may be all you need

Part of that improvement will be making generative AI more than just a personal "Co-pilot," like Microsoft's GitHub Co-pilot, which assists a single individual typing in a chat prompt. The programs will instead become collaborative, for teams, said Emad Mostaque, founder and CEO of Stability.ai, in an interview with ZDNET.

"A lot of AI is just used as a one-to-one thing, or it's an autonomous agent," said Mostaque. "It's at iPhone 2G phase now, where it's just a single mode and you cut and paste, whereas I think the most exciting thing is how we can collaborate better and tell better stories with it, and that's not a solitary endeavor."

One of the things that is "fundamentally missing," said Databricks's Rao, "is the multi-modal-ness of the world," given that "large language models are very one-dimensional in that they only see the world through text."

Modalities refer to the nature of the input and the output, such as text, image, or video. A variety of modalities are possible and have been explored with increasing diversity, because the same basic concepts that drive ChatGPT can be applied to any type of input.

"Multi-modality is the way, definitely," said Mostaque. "You'll need models of every type, and maybe if you bring them together, it'll be amazing."

"The language-only stuff got a lot of traction and excitement, and so the media focuses on that, but people are working seriously on other things," said Jim Keller, a renowned computer-chip designer who is CEO of AI chip startup Tenstorrent, in an interview with ZDNET. Keller is betting his company on the prospect that handling mixed modalities will be one of the big AI demands going forward.

A machine for any kind of data

In a large language model, which is the heart of ChatGPT's technology, text is turned into a token, a quantified mathematical representation. The machine then has to find what is missing from either masked parts of an entire phrase, or the latter part of a phrase. It is the act of recreation that brings about the paragraphs that ChatGPT spits out.

Likewise, in the case of images, the widely used diffusion process — popularized by Stability.ai's Stable Diffusion version — corrupts images with noise, and the act of recreating the original image trains a neural network to generate high-fidelity images.

The same processes of recovering what's missing or corrupted are spreading rapidly to numerous modalities, or, types of data. For example, in a recent issue of Nature magazine, University of Washington biologist David Baker and team corrupted the amino acid sequences of proteins via a process they call RFdiffusion. That process will train a neural network to produce a protein, in simulation, a novel synthetic protein, that has desired properties.

Such a synthesis can dramatically cut down on the number of proteins that need to be invented and tested to come up with novel antibodies for diseases. (The Nature article is behind a paywall, but a free version is posted on the bioRxiv file server. More information can be found at the Baker Lab website.)

The RFdiffusion process developed by the Baker Lab of the Insitute for Protein Design at the University of Washington corrupts amino acid sequences to then synthesize a novel protein structure much in the way image diffusion creates pictures.

"We have labs for every modality," said Stability.ai's Mostaque, who claims his company and OpenAI are "the only two independent multi-modal companies," outside of the tech giants such as Google. That multiple modality includes a lab at Stability.ai just for audio, he said, a lab just for code generation, even a lab for biology that works on things such as re-creating fMRI images using the Stable Diffusion technology.

The magic, however, happens when more modalities are combined. The "breakthrough," said Mostaque, came in work last year by Katherine Crowson and several other researchers who trained an image-generating neural network to keep refining their output until the output satisfied a text-based prompt. They found that re-working images to match the "semantic" content of the text improved image quality. Crowson is now at Stability.ai, noted Mostaque.

That image-text work has been proceeding swiftly at numerous institutions. The AI researchers at Meta have proposed a combination of text and image machines called CM3Leon that excels at not merely outputting text or outputting images, but carrying out tasks that involve both at the same time such as identifying objects in a given image or generating captions from a given image.

Meta's CM3Leon neural network mixes images and text to perform multiple tasks, like describing in detail a given image, or altering a given image with precision. It's detailed in the 2023 paper, "Scaling Autoregressive Multi-Modal Models: Pre-training and Instruction Tuning," by Lilu Yu and colleagues at Meta AI.

A richer picture of the world

The combination of multiple modalities starts to build a richer picture of the world for the neural network. Databricks's Rao cites the neuroscience concept of "stereognosis," which means to know the world by sense of touch. If someone asks how much change you have in your pocket, you can feel the coins and tell by size and weight without seeing them. "I have a representation of the world and objects that are actually represented in multiple modalities," he said. "If I can learn concepts that span modalities, then we've done something interesting."

The idea that different senses flesh out understanding is echoed in the multi-modal experiments being carried out. Research is active into how to make so-called "backbone" neural networks that can mix and match a dizzying array of modalities, and they show intriguing performance benefits.

Scholars at Carnegie Mellon University recently offered what they call a "High-Modality Multimodal Transformer," which combines not just text, image, video, and speech but also database table information and time series data. Lead author Paul Pu Liang and colleagues report that they observed "a crucial scaling behavior" of the 10-mode neural network. "Performance continues to improve with each modality added, and it transfers to entirely new modalities and tasks."

Carnegie Mellon's 2023 paper "High-Modality Multimodal Transformer" by Paul Liang and colleagues combines not just text, image, video, and speech but also database table information and time series data.

Scholars Yiyuan Zhang and colleagues at the Multimedia Lab of The Chinese University of Hong Kong boosted the number of modalities to a dozen in their Meta-Transformer. Its point clouds model 3D vision, while its hyper-spectral sensing data represents electromagnetic energy reflected back from the ground to fly-over images of landscapes.

The Meta-Transformer is the future of generative AI, with tons of data of all different kinds being fused to have a richer sense of what is being produced as output. It's explored in the 2023 paper, "Meta-Transformer: A Unified Framework for Multimodal Learning," by Yiyuan Zhang and colleagues at the Multimedia Lab of the Chinese University of Hong Kong and OpenGVLab at the Shanghai AI Laboratory.

Making a story book from multiple modes

The immediate payoff of multi-modality will simply be to enrich the output of a thing such as ChatGPT in ways that go far beyond the "demo" mode. A children's story book, a book with text passages combined with pictures illustrating the text, is one immediate example. By combining the language and image attributes, the kinds of pictures created by the diffusion process can be more subtly controlled from picture to picture.

As explained by scientists at Google and lead author Wan-Duo Kurt Ma of Victoria University of Wellington in New Zealand, a process known as directed diffusion can move the cat — or a castle, or a bird — through various scenes, creating a series of images that afford not only greater control but transitions as in a narrative.

A technique called directed diffusion can move an entity — a cat, a castle, a bird — through various scenes, creating a series of images that afford not only greater control, but also transitions as in a narrative. It's detailed in the 2023 paper "Directed Diffusion: Direct Control of Object Placement Through Attention Guidance" by Wan-Duo Kurt Ma and colleagues at Victoria University of Wellington and Google Research.

Similarly, Hyeonho Jeong of Korea's Sungkyunkwan University, along with scholars at the Korea Advanced Institute of Science & Technology, came up with yet another twist on diffusion — latent diffusion — which they detailed in a recent paper. They claim it gives access to many more details in an image at a low level of granularity.

The result is the ability to generate story books where a character moves through different scenarios image by image, like adding knobs to the text prompt to dial in different scenarios. The consistency of the object across images is what they call "Iterative Coherent Identity Injection."

A technique called latent diffusion extends image-making with what its inventors call "Identity Injection" in order to script a character's movement through story book images.

Just as with the protein synthesis at the Baker Lab, the applications of mixed modality can become pretty wild. Another recent paper by Chenyu Tang and colleagues at Cambridge University's Department of Engineering proposes constructing a "digital twin," a computer simulation of the human body, with all the organs and tissues rendered, and the flows of blood and such depicted, by combining data from multiple medical instruments in the same process as stable diffusion.

"Both movement sensors (such as accelerometers, EMG sensors, etc.) and biochemical sensors (for detecting disease-corresponding biomarkers, such as saliva sensors, sweat sensors, etc.) can produce specific outputs for the patient," the authors wrote. "Although these outputs have distinct patterns, they all correspond to the same disease."

The "digital twin" of the human body could be enabled by combining data from multiple medical instruments in the same process as stable diffusion. The diagram represents the "five-level roadmap for body DT [digital twin]," as seen in the 2023 paper "Human Body Digital Twin: A Master Plan" by Chenyu Tang and colleagues at the University of Cambridge.

Special modal masters

How the modalities get put together will be as important as which ones, said Stability.ai's Mostaque. "The final bit will be composition, as these building blocks that we build are put into proper software that is AI-first, that reimagines all of this creation, consumption, and these process flows with these cool new tools," he said.

While some massive models such as Google's PaLM LLM or GPT-4 may be called in, a lot of mixed modality will happen as an orchestration of components, he said. "How do you bring together models in really interesting ways, and have many different models working together to achieve the outcomes that you want to really augment that?"

While PaLM and GPT-4 can be powerful, he said, there's ample evidence that "a lot more specialized models can outperform" the biggest programs. As a result, "We're gonna have a lot of specialist models, I think, across the modalities," he said, a process of "de-constructing" the technology into its appropriate roles, "and then some multi-modal models that can do everything, and they're called at the appropriate time for the appropriate thing."

Robotics is the next AI frontier

The mixing of modalities is noteworthy for the realm of embodied AI — in the form of robotics.

Sergey Levine, associate professor in the electrical engineering department at the University of California at Berkeley, told ZDNET that as it relates to generative AI, systems in robotics have a significant role.

"The multi-modal stuff is quite exciting," added Levine, a member of the University's Berkeley Artificial Intelligence Research facility who also works with teams at Google.

By processing images and text, a multi-modal neural network is already able to produce "high-level robot commands," he said. The code that a roboticist would ordinarily write to instruct a robot can be "fully automated, essentially," said Levine.

"What we want is the ability to quickly and easily command the robots to do stuff," said Levine. "Bridging that gap is something that language models are gonna be great at."

Also: DeepMind's RT-2 makes robot control a matter of AI chat

Levine helped oversee an early demonstration at Google that was published recently, called PaLM-E, which the Google researchers call "An Embodied Multimodal Language Model." The robot is able to follow a series of instructions such as "bring me the rice chips from drawer," which the language model breaks down into atomic instructions, such as "go to the drawer," "open the drawer," "pick the green rice chip bag," etc.

A subsequent work, by Google's DeepMind unit, called RT-2, builds upon PaLM-E by adding the ability to generate spatial coordinates for the robot. Levine calls that work "a significant advance."

As with the concept of stereognosis, Levine argues that increasing modalities may bring an enriched model of the world and thereby bring some basic reasoning abilities.

Also: DeepMind's RT-2 makes robot control a matter of AI chat

If large language models and diffusion models can integrate the process of "taking previous images and predicting [text] descriptions, and taking previous descriptions and predicting images," said Levine, "now they might start, kind-of, drilling further down in terms of how they understand the world."

A primitive example of world knowledge is a robot bartender that Levine has worked on, which checks people's I.D. "You can actually tell the language model, write me some code for a robot bartender, and it generates some logic to do that, and if someone orders a cup of water, that's not an alcoholic beverage," and therefore doesn't require an I.D. check.

We're going to need a lot more memory

The combination of robotics and multi-modality has more profound implications because it expands the appetite for data dramatically. Today's generative AI such as ChatGPT has no explicit memory. It only works on the last bunch of stuff you typed at the prompt, and after a while, it forgets things from long ago.

Using mixed modality that includes many more data samples will force generative AI to develop something like a real memory of data. "When we start moving to multi-modal models, now that starts being much more demanding on context," said Levine, "because the current prototype of that model takes in one image, but maybe you want to give it a thousand images.

"Maybe you want to show it a tour of your house so that it knows where everything in your house is, so that when you ask it to bring you the car keys, it can sort of examine its memory and figure out where the car keys are — now that requires a much longer context."

Also: Microsoft, TikTok give generative AI a sort of memory

Video data can be equally if not more critical for letting a robot build a portrait of the world. Those videos, coupled with text and point clouds and other modalities, become a simulator by which a robot can build a model of the world, said Levine. "If these models essentially provide a way to learn very high fidelity simulators, that could have a very, significant impact in the future."

Expanding to thousands of images and possibly hours of video, perhaps gigabytes of point-cloud, 3D data, to train multi-modal programs, means ChatGPT and the rest will have to dramatically expand their access to data via a so-called memory bank.

Many efforts are underway to "augment" language models with what's called retrieval from a database. That can be seen in Meta's CM3Leon program, which lets the software dip into a database and find relevant images.

Efforts such as the Hyena technology at Stanford University and Canada's MILA institute attempt to dramatically expand what can be fed into a program's prompt so that any amount of data can be input, of any modality.

Also: This new technology could blow away GPT-4 and everything like it

That means that along with mixed modality, the successors to ChatGPT will be able to juggle far greater context — whole books, series of articles, movies, and records of physical structures in three dimensions. It also means that the context for any task can become much more tailored to an individual or a group's acquired knowledge. Mostaque said such models will not only bring the generalized knowledge of GPT-4, but also specific knowledge, as well as the knowledge of your team, your company, and beyond.

"I think that's the big unlock, when it goes enterprise next year," said Mostaque, referring to the imminent popular adoption of generative AI in corporate settings.

TikTok owner ByteDance's "Self-Controlled Memory system" can reach into a data bank of hundreds of turns of dialogue, and thousands of characters, to give any language model capabilities superior to that of ChatGPT in answering questions about past events. It's shown in the 2023 paper, "Unleashing Infinite-Length Input Capacity for Large-scale Language Models with Self-Controlled Memory System," by Xinnian Liang and colleagues at ByteDance AI Lab.

Continuous learning attainable

As multi-modality expands to video and audio and point clouds and all the rest, Keller, the CEO of AI chip company Tenstorrent, believes that more advanced generative models, especially those coming from the open-source software community, will lead to a profound change in the field's distinction between training and inference.

Training is when a neural net is first developed. It is an extremely costly scientific process, with hundreds or even thousands of GPUs used. Inference is when the finished network is used to make predictions for end users, a much less demanding process that is widely deployed as a cloud service.

But "the generative models actually use quite a few features from training in inference," said Keller. A program such as Stability.ai's Stable Diffusion, for generating images, updates its neural network during inference, he said. "It is multi-pass: it has a back pass" as well as the typical forward process of predictions, so that "it looks like it's in training mode."

For that reason, "I think the AI engine of the future … will have a fairly diverse set of capabilities that won't look like inference versus training," but more like a fusion of the two.

If Keller is right, the future generative models could be the start of a long-held goal of continuous learning for machine learning, also sometimes called online learning, whereby a generative neural network is not fixed once trained but evolves continually as people use it more.

"I think this is going to be the case" agreed Stability.ai's Mostaque. "Continuous learning will be key, because the way we do it now, teaching [the model] the same thing over and over, is not appropriate."

Already, said Mostaque, things such as Stability.ai's "Dream Booth," which lets one build a customized version of an image, are moving beyond the rigid notion of re-training a language-image model to something more fluid. He said these become personal avatars — and over the next few months — a kind of hyper-Dream Booth that allows for the personalization of all your images in real time.

"That's why continuous learning will be so important: to enable that continuous process so that it evolves."

Artificial Intelligence

Spotify spotted developing AI-generated playlists created with prompts

Spotify spotted developing AI-generated playlists created with prompts Sarah Perez @sarahintampa / 8 hours

Following the successful launch of Spotify’s AI-powered DJ feature and, more recently, added support for AI-translated podcasts, Spotify now appears to be developing another means of using AI in its app: AI-powered playlists. References discovered in the app’s code indicate the company may be developing generative AI playlists users could create using prompts.

The new additions were uncovered by tech veteran-turned-investor Chris Messina, who posted screenshots of code in Spotify’s app that refer to “AI playlists” and “playlists based on your prompts.” He theorized that creating these may be an option within the Blend genre, where the tastes of different users are mixed to create a playlist with songs that everyone likes.

Image Credits: Spotify code via Chris Messina (opens in a new window)

Reached for comment, Spotify declined to confirm its plans around AI playlists.

“At Spotify, we are constantly iterating and ideating to improve our product offering and offer value to users. But we don’t comment on speculation around possible new features and do not have anything new to share at this time,” a spokesperson told TechCrunch.

That said, Spotify may have already been laying the groundwork for AI playlists created with prompts with it developed a feature called Niche mixes, which today allows users to build unique playlists based on a description alone. These Niche mixes currently let you specify almost anything to create a playlist — from a genre to a vibe or an aesthetic, like Cottagecore Indie Mix, Bubblegum Pop Mix, Discofox Mix, Feel Good Driving Mix, Fun Road Trip Mix, Travel Mashup Mix, and others.

However, Spotify told us when the Niche Mixes were launched back in March that they were not AI-powered, despite their initial appearance. Instead, Spotify said the mixes were driven by the company’s personalization tech and algorithms.

Image Credits: Spotify code via Chris Messina (opens in a new window)

Messina tells TechCrunch his new findings indicate the new AI playlists would be built using prompts, as well, but he hasn’t yet uncovered the feature in the public app, only in the code.

He suspects the feature may be tied to Blend because there are also code references that indicate users could invite others to create AI playlists together.

Image Credits: Spotify code via Chris Messina

All the lines of code were discovered in the latest build of the Spotify app, so this is clearly a new, in-development feature.

Of course, not all features that a company builds internally to test make it to the public, but it is an indication of how Spotify is thinking about the role AI could play when it comes to music personalization.

The company had suggested previously that features like the AI DJ wouldn’t be the limit to Spotify’s adoption of AI technologies. There’s a team at Spotify that’s working on the latest language models, Spotify’s head of Personalization, Ziad Sultan, told us in a conversation at the company’s event earlier this year. In fact, Spotify has a few hundred people working on personalization and machine learning techniques, including a large research team that’s working to understand “all the possibilities across Large Language Models, across generative voice, across personalization,” he had said.

Around the same time, TechCrunch also heard that Spotify had been toying around with an ChatGPT-like chatbot that would allow users to request music, but nothing was settled in terms of a public launch on that front.

Go From AI Novice to Advanced User for Just $30

Man and woman on a background of a robot hand holding the OpenAI icon.
Image: StackCommerce

We all know that artificial intelligence (AI) is firmly entrenched in our future, but many of us don’t know how to use it to our professional advantage today. Whether you’re trying to turbocharge your career or have a company of your own, you’d do well to learn how AI can help you with the AI-Powered Productivity & Learning Bundle. It’s on sale now for just $29.99 – a huge saving on its regular price of $436.

What’s in the bundle?

This bundle consists of four courses with over a hundred hours of lectures on ChatGPT, the metaverse and other valuable AI tools. The information you’ll learn can help you to improve efficiency, increase your productivity, advance your career or transform your business.

The Boost Your Productivity with AI course will teach you how to use AI to strengthen your time management skills by exploring how to implement various AI tools. You’ll learn tips and best practices for using them to maximum effect.

You’ve surely heard of ChatGPT by now, but Introduction to ChatGPT will take you from complete novice to power user. The course starts with the fundamentals before teaching you how to use ChatGPT for writing, business analysis, research, and much more.

Metaverse Essentials for Beginners also starts with the basics. First, the course covers metaverse technologies, then analyzes the metaverse’s potential impact on society. You’ll then learn how to identify opportunities for profitable metaverse investments by minimizing risk. You’ll even learn about the role of non-fungible tokents (NFTs) in the metaverse.

Educators will be thrilled with the AI Resources for Teaching course. It first provides an introduction to various AI tools and explains how they work. The course then demonstrates how you can assess your teaching skills and improve them using these AI tools in the most effective way. The courses are taught by instructors from International Open Academy, which has a 4.4/5 star instructor rating.

The best thing about this bundle, aside from its affordability, is that you can learn at your own pace in your own time. You don’t need to pay tuition or work another commute into your schedule. You certainly don’t need any specialized equipment – you can even access these courses on a phone or tablet.

Grab this deal while you can: Get the AI-Powered Productivity & Learning Bundle today while it’s on sale for just $29.99.

Prices and availability are subject to change.

Subscribe to the Executive Briefing Newsletter

Discover the secrets to IT leadership success with these tips on project management, budgets, and dealing with day-to-day challenges.

Delivered Tuesdays and Thursdays Sign up today

Want to Become a Data Scientist? Part 2: 10 Soft Skills You Need

Want to Become a Data Scientist? Part 2: 10 Soft Skills You Need
Image by Author

This is Part 2 of the skills required to become a data scientist. A lot of people speak about hard skills when it comes to being a data scientist. Companies will list out different tools and software that they would like you to know, but when you’re at your interview it is how you perceive yourself that matters the most.

This comes from your soft skills and personality.

So rather than blabbering on, let’s just get right into it.

Communication

Communication is key. You’ve probably heard that so many times and it can get very annoying — but it matters. Especially when you’re working in a technical field, it is very important to be able to communicate these technical concepts to non-technical stakeholders. Reminding yourself that not everybody is technically inclined and you will need to ensure you have effective communication to explain valuable insights, findings from your analysis and data-driven decisions.

Problem-Solving

Dealing with complex and unstructured problems every day requires you to be able to solve problems. You will need to comb through the task, break it down and figure out the issues with proposed solutions.

You may not be able to instantly look at a piece of data and find the issue straight away, this is why problem-solving skills are important.

Critical Thinking

As part of your problem-solving skills, when you are trying to find solutions to your problem or task-at-hand, you need to be a critical thinker. You need to understand the problem you’re facing and how you will choose the appropriate methods towards your solution.

This includes evaluating the quality of the data, and how you interpret the results to make data-driven decisions as well as avoid biases.

Business Understanding

You will need to have a good understanding of the business model and implement business skills. You will always have to keep at the back of your mind: ‘How is this company going to use this analytics?’. When you have a well-rounded understanding of this, you will be able to figure out what to do with the analytics, such as creating an application, a report, etc.

Time Management

As a data scientist, you will manage multiple tasks throughout your day. Juggling these tasks can take a toll, and get you frustrated very easily. Managing your time will relieve you from stress.

Once you've had a few trial runs of what a data science project lifecycle looks like, you will be able to understand how much time each phase requires. You can then use this experience to manage your tasks such as data cleaning, analysis, and more more effectively.

Teamwork

Going hand in hand with time management, you will see that having an effective method and process for the data science project lifecycle in place requires teamwork. As a student data scientist, you will be the sole person working on the project. Once you start with a company, these tasks can be split up between the data science team. Not only does it effectively take workload off your shoulders, but it gives everybody on the team to experience the tasks included.

Teamwork is only effective when communication is in place — remember this! Always communicate with your team members about what you are doing, if you’re blocked on something, or the outcome of your task.

Data science projects consist of cross-functional teams, therefore you will have to collaborate with other experts such as business analysts, product managers, and more.

Storytelling and Presentation

As I mentioned before, a part of your communication skills is to understand that every stakeholder may or may not be technically inclined. Therefore, you will need to take this into consideration when narrating and presenting your analytical findings.

You can practice your data storytelling skills via blogs, as it is a good way to explain technical concepts in a simpler format. Presenting your findings can be done through powerpoint presentations, data visualizations and more.

Practicing these will make your life easier as stakeholders will have fewer questions due to the way the findings were presented.

Domain Expertise

Working with a company and dealing with day-to-day tasks will help build your skills and make you more proficient. However, you will need to go above and beyond when working in a field that is very innovative.

Whatever it is that you’re interested in, I would highly advise you to be an expert in that field. This allows your skills and knowledge to be transferable and you can apply this in your day-to-day tasks.

Self-Development

In a field that is constantly evolving, keeping on top of things is very very important. Your learning will not end once you land your first data science job. You will be constantly learning new things, and you will need to dedicate time out of your working day to learn about these things.

I’m not saying you have to go full blown crazy back into education, but you will need to read articles, news and learn how new tools and softwares work. This will increase your skill-set and make your daily tasks more efficient.

Governance and Security

As a data scientist, you will be working with sensitive information. There are ethical guidelines that you will need to follow when collecting data, using it, as well as sharing it. You need to remember that some data is private information, therefore what you do with it is very important.

You want to look into the ethics, bias, and security around your company's processes and policies.

Wrapping it up

I hope this was a quick and easy guideline on the soft skills that you need as a data scientist. A lot of these skills you will naturally build and progress in a work setting, but it is always good to know what you’re up against.

Happy learning!
Nisha Arya is a Data Scientist, Freelance Technical Writer and Community Manager at KDnuggets. She is particularly interested in providing Data Science career advice or tutorials and theory based knowledge around Data Science. She also wishes to explore the different ways Artificial Intelligence is/can benefit the longevity of human life. A keen learner, seeking to broaden her tech knowledge and writing skills, whilst helping guide others.

More On This Topic

  • Want to Become a Data Scientist? Part 1: 10 Hard Skills You Need
  • Want to Use Your Data Skills to Solve Global Problems? Here’s What You Need…
  • 9 Skills You Need to Become a Data Engineer
  • These Soft Skills Can Make or Break Your Data Science Career
  • 6 Soft Skills for Data Scientists Working Remotely
  • Everything you Need to Become a SAS Certified Data Scientist

A Deep Dive into Retrieval-Augmented Generation in LLM

Retrieval Augmented Generation Illustration using Midjourney

Imagine you're an Analyst, and you've got access to a Large Language Model. You're excited about the prospects it brings to your workflow. But then, you ask it about the latest stock prices or the current inflation rate, and it hits you with:

“I'm sorry, but I cannot provide real-time or post-cutoff data. My last training data only goes up to January 2022.”

Large Language Model, for all their linguistic power, lack the ability to grasp the ‘now‘. And in the fast-paced world, ‘now‘ is everything.

Research has shown that large pre-trained language models (LLMs) are also repositories of factual knowledge.

They've been trained on so much data that they've absorbed a lot of facts and figures. When fine-tuned, they can achieve remarkable results on a variety of NLP tasks.

But here's the catch: their ability to access and manipulate this stored knowledge is, at times not perfect. Especially when the task at hand is knowledge-intensive, these models can lag behind more specialized architectures. It's like having a library with all the books in the world, but no catalog to find what you need.

OpenAI's ChatGPT Gets a Browsing Upgrade

OpenAI's recent announcement about ChatGPT's browsing capability is a significant leap in the direction of Retrieval-Augmented Generation (RAG). With ChatGPT now able to scour the internet for current and authoritative information, it mirrors the RAG approach of dynamically pulling data from external sources to provide enriched responses.

ChatGPT can now browse the internet to provide you with current and authoritative information, complete with direct links to sources. It is no longer limited to data before September 2021. pic.twitter.com/pyj8a9HWkB

— OpenAI (@OpenAI) September 27, 2023

Currently available for Plus and Enterprise users, OpenAI plans to roll out this feature to all users soon. Users can activate this by selecting ‘Browse with Bing' under the GPT-4 option.

Chatgpt New Browsing Feature

Chatgpt New ‘Bing' Browsing Feature

Prompt engineering is effective but insufficient

Prompts serve as the gateway to LLM's knowledge. They guide the model, providing a direction for the response. However, crafting an effective prompt is not the full-fledged solution to get what you want from an LLM. Still, let us go through some good practice to consider when writing a prompt:

  1. Clarity: A well-defined prompt eliminates ambiguity. It should be straightforward, ensuring that the model understands the user's intent. This clarity often translates to more coherent and relevant responses.
  2. Context: Especially for extensive inputs, the placement of the instruction can influence the output. For instance, moving the instruction to the end of a long prompt can often yield better results.
  3. Precision in Instruction: The force of the question, often conveyed through the “who, what, where, when, why, how” framework, can guide the model towards a more focused response. Additionally, specifying the desired output format or size can further refine the model's output.
  4. Handling Uncertainty: It's essential to guide the model on how to respond when it's unsure. For instance, instructing the model to reply with “I don’t know” when uncertain can prevent it from generating inaccurate or “hallucinated” responses.
  5. Step-by-Step Thinking: For complex instructions, guiding the model to think systematically or breaking the task into subtasks can lead to more comprehensive and accurate outputs.

In relation to the importance of prompts in guiding ChatGPT, a comprehensive article can be found in an article at Unite.ai.

Challenges in Generative AI Models

Prompt engineering involves fine-tuning the directives given to your model to enhance its performance. It's a very cost-effective way to boost your Generative AI application accuracy, requiring only minor code adjustments. While prompt engineering can significantly enhance outputs, it's crucial to understand the inherent limitations of large language models (LLM). Two primary challenges are hallucinations and knowledge cut-offs.

  • Hallucinations: This refers to instances where the model confidently returns an incorrect or fabricated response. Although advanced LLM has built-in mechanisms to recognize and avoid such outputs.

Hallucinations in LLMs

Hallucinations in LLM

  • Knowledge Cut-offs: Every LLM model has a training end date, post which it is unaware of events or developments. This limitation means that the model's knowledge is frozen at the point of its last training date. For instance, a model trained up to 2022 would not know the events of 2023.

Knowledge cut-off in LLMS

Knowledge cut-off in LLM

Retrieval-augmented generation (RAG) offers a solution to these challenges. It allows models to access external information, mitigating issues of hallucinations by providing access to proprietary or domain-specific data. For knowledge cut-offs, RAG can access current information beyond the model's training date, ensuring the output is up-to-date.

It also allows the LLM to pull in data from various external sources in real time. This could be knowledge bases, databases, or even the vast expanse of the internet.

Introduction to Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) is a framework, rather than a specific technology, enabling Large Language Models to tap into data they weren't trained on. There are multiple ways to implement RAG, and the best fit depends on your specific task and the nature of your data.

The RAG framework operates in a structured manner:

Prompt Input

The process begins with a user's input or prompt. This could be a question or a statement seeking specific information.

Retrieval from External Sources

Instead of directly generating a response based on its training, the model, with the help of a retriever component, searches through external data sources. These sources can range from knowledge bases, databases, and document stores to internet-accessible data.

Understanding Retrieval

At its essence, retrieval mirrors a search operation. It's about extracting the most pertinent information in response to a user's input. This process can be broken down into two stages:

  1. Indexing: Arguably, the most challenging part of the entire RAG journey is indexing your knowledge base. The indexing process can be broadly divided into two phases: Loading and Splitting.In tools like LangChain, these processes are termed “loaders” and “splitters“. Loaders fetch content from various sources, be it web pages or PDFs. Once fetched, splitters then segment this content into bite-sized chunks, optimizing them for embedding and search.
  2. Querying: This is the act of extracting the most relevant knowledge fragments based on a search term.

While there are many ways to approach retrieval, from simple text matching to using search engines like Google, modern Retrieval-Augmented Generation (RAG) systems rely on semantic search. At the heart of semantic search lies the concept of embeddings.

Embeddings are central to how Large Language Models (LLM) understand language. When humans try to articulate how they derive meaning from words, the explanation often circles back to inherent understanding. Deep within our cognitive structures, we recognize that “child” and “kid” are synonymous, or that “red” and “green” both denote colors.

Augmenting the Prompt

The retrieved information is then combined with the original prompt, creating an augmented or expanded prompt. This augmented prompt provides the model with additional context, which is especially valuable if the data is domain-specific or not part of the model's original training corpus.

Generating the Completion

With the augmented prompt in hand, the model then generates a completion or response. This response is not just based on the model's training but is also informed by the real-time data retrieved.

Retrieval-Augmented Generation

Retrieval-Augmented Generation

Architecture of the First RAG LLM

The research paper by Meta published in 2020 “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” provides an in-depth look into this technique. The Retrieval-Augmented Generation model augments the traditional generation process with an external retrieval or search mechanism. This allows the model to pull relevant information from vast corpora of data, enhancing its ability to generate contextually accurate responses.

Here's how it works:

  1. Parametric Memory: This is your traditional language model, like a seq2seq model. It's been trained on vast amounts of data and knows a lot.
  2. Non-Parametric Memory: Think of this as a search engine. It's a dense vector index of, say, Wikipedia, which can be accessed using a neural retriever.

When combined, these two create an accurate model. The RAG model first retrieves relevant information from its non-parametric memory and then uses its parametric knowledge to give out a coherent response.

RAG ORIGNAL MODEL BY META

Original RAG Model By Meta

1. Two-Step Process:

The RAG LLM operates in a two-step process:

  • Retrieval: The model first searches for relevant documents or passages from a large dataset. This is done using a dense retrieval mechanism, which employs embeddings to represent both the query and the documents. The embeddings are then used to compute similarity scores, and the top-ranked documents are retrieved.
  • Generation: With the top-k relevant documents in hand, they're then channeled into a sequence-to-sequence generator alongside the initial query. This generator then crafts the final output, drawing context from both the query and the fetched documents.

2. Dense Retrieval:

Traditional retrieval systems often rely on sparse representations like TF-IDF. However, RAG LLM employs dense representations, where both the query and documents are embedded into continuous vector spaces. This allows for more nuanced similarity comparisons, capturing semantic relationships beyond mere keyword matching.

3. Sequence-to-Sequence Generation:

The retrieved documents act as an extended context for the generation model. This model, often based on architectures like Transformers, then generates the final output, ensuring it's coherent and contextually relevant.

Document Search

Document Indexing and Retrieval

For efficient information retrieval, especially from large documents, the data is often stored in a vector database. Each piece of data or document is indexed based on an embedding vector, which captures the semantic essence of the content. Efficient indexing ensures quick retrieval of relevant information based on the input prompt.

Vector Databases

Vector Database

Source: Redis

Vector databases, sometimes termed vector storage, are tailored databases adept at storing and fetching vector data. In the realm of AI and computer science, vectors are essentially lists of numbers symbolizing points in a multi-dimensional space. Unlike traditional databases, which are more attuned to tabular data, vector databases shine in managing data that naturally fit a vector format, such as embeddings from AI models.

Some notable vector databases include Annoy, Faiss by Meta, Milvus, and Pinecone. These databases are pivotal in AI applications, aiding in tasks ranging from recommendation systems to image searches. Platforms like AWS also offer services tailored for vector database needs, such as Amazon OpenSearch Service and Amazon RDS for PostgreSQL. These services are optimized for specific use cases, ensuring efficient indexing and querying.

Chunking for Relevance

Given that many documents can be extensive, a technique known as “chunking” is often used. This involves breaking down large documents into smaller, semantically coherent chunks. These chunks are then indexed and retrieved as needed, ensuring that the most relevant portions of a document are used for prompt augmentation.

Context Window Considerations

Every LLM operates within a context window, which is essentially the maximum amount of information it can consider at once. If external data sources provide information that exceeds this window, it needs to be broken down into smaller chunks that fit within the model's context window.

Benefits of Utilizing Retrieval-Augmented Generation

  1. Enhanced Accuracy: By leveraging external data sources, the RAG LLM can generate responses that are not just based on its training data but are also informed by the most relevant and up-to-date information available in the retrieval corpus.
  2. Overcoming Knowledge Gaps: RAG effectively addresses the inherent knowledge limitations of LLM, whether it's due to the model's training cut-off or the absence of domain-specific data in its training corpus.
  3. Versatility: RAG can be integrated with various external data sources, from proprietary databases within an organization to publicly accessible internet data. This makes it adaptable to a wide range of applications and industries.
  4. Reducing Hallucinations: One of the challenges with LLM is the potential for “hallucinations” or the generation of factually incorrect or fabricated information. By providing real-time data context, RAG can significantly reduce the chances of such outputs.
  5. Scalability: One of the primary benefits of RAG LLM is its ability to scale. By separating the retrieval and generation processes, the model can efficiently handle vast datasets, making it suitable for real-world applications where data is abundant.

Challenges and Considerations

  • Computational Overhead: The two-step process can be computationally intensive, especially when dealing with large datasets.
  • Data Dependency: The quality of the retrieved documents directly impacts the generation quality. Hence, having a comprehensive and well-curated retrieval corpus is crucial.

Conclusion

By integrating retrieval and generation processes, Retrieval-Augmented Generation offers a robust solution to knowledge-intensive tasks, ensuring outputs that are both informed and contextually relevant.

The real promise of RAG lies in its potential real-world applications. For sectors like healthcare, where timely and accurate information can be pivotal, RAG offers the capability to extract and generate insights from vast medical literature seamlessly. In the realm of finance, where markets evolve by the minute, RAG can provide real-time data-driven insights, aiding in informed decision-making. Furthermore, in academia and research, scholars can harness RAG to scan vast repositories of information, making literature reviews and data analysis more efficient.