Why Google just banned Gemini from generating images of people

Google Gemini logo on laptop screen

AI image generators allow you to generate images of anything you can think of, including historical figures. However, when users started asking Google's Gemini to depict historical figures or groups of people, they were in for a surprise.

In the past week, when users asked Gemini to generate images of historical figures or people of different races or nationalities, they began to notice that none of the images were true to the prompt. Rather, the images failed to render white people, even if the individual was white. Instead, it only generated images of multicultural and racially diverse people.

Also: How to use Gemini (formerly Google Bard): Everything you should know

Users began taking to social media to share these images, in which prompts that asked Gemini to create images of the Pope, Father of America, and a Viking, all resulted in images of people of color. Former Google employee, Debarghya Das, shared an X thread of these inaccuracies compiling examples from user posts across X, as well as his own experience.

As a result, on Wednesday, Google Communications acknowledged the issue via X, stating that it is aware of the inaccuracies and was working on improving those depictions because the chatbot "missed the mark."

Then on Thursday, Google shared that the company is pausing image generation of people in Gemini as it works to address the issues, and the company is planning to re-release an improved version soon.

If you try to ask Gemini to generate a photo of a person now, you will be met with a series of different error messages, such as "That's not something I'm able to do yet," and it won't generate the image for you.

Also: I just tried Google's ImageFX AI image generator, and I'm shocked at how good it is

Despite launching its AI chatbot in March of last year, Gemini has only had its image-generating capabilities since the beginning of this month. With that long wait before releasing its image generator to its chatbot, it is surprising that it is already making these missteps.

Google had a similar experience when it first launched Gemini, then called Google Bard, and the chatbot shared inaccurate information about the James Webb Space Telescope in its demo.

Artificial Intelligence

Reddit Prepares for Wall Street, Accompanied by Google

Social media platform Reddit has struck a deal with Google to make its content available for training the search engine giant’s AI models, as per Reuter’s sources.

The contract worth about $60 million per year underlines how Reddit is preparing to become the first high profile company planning to go public in years. Through the deal, Reddit is seeking to generate new revenue amid fierce competition for advertising dollars from the likes of TikTok and Meta Platform’s Facebook.

“We expect our data advantage and intellectual property to continue to be a key element in the training of future AI models,” Steve Huffman, Reddit’s chief executive, said in a founder’s letter included in the prospectus.

Last year, Reddit said it would charge companies for access to its application programming interface (API) – the means by which it distributes its content. The agreement with Google is its first reported deal with an AI leader.

Reddit, in its investor statement, acknowledged challenges and risks tied to its public status. The platform highlighted concerns about the rise of large language models and underlying AI systems, capable of generating and synthesizing content, potentially allowing users to access Reddit ad-free.

Emphasizing reliance on its community for moderation, Reddit expressed that future revolts or departures that could adversely impact the platform. The note highlights the company’s awareness of the changing AI landscape and the significance of community dynamics, urging investors to consider these factors in assessing the platform’s long-term stability and success.

Interestingly, OpenAI’s chief Sam Altman, is the third largest shareholder in the online discussion website, with an 8.7% stake. The word on the street is that Reddit is the boot camp for training the heavyweights in large language models— OpenAI’s GPT-3.5, even the elusive GPT-4, and other popular models.

Altman has been in the Reddit game for a decade. He even donned the CEO hat for a hot minute back in 2014, stepping in for Yishan Wong, who made an exit.

The post Reddit Prepares for Wall Street, Accompanied by Google appeared first on Analytics India Magazine.

I confused Google’s most advanced AI — but don’t laugh because programming is hard

confused-devgettyimages-1560259012

Care to have some fun? Fun at Google's expense? And when I say fun, I mean fun in, well, a morbid fascination, what-were-they-thinking, I'm-glad-I'm-not-the-product-manager kind of way?

Our story begins with a Gemini Advanced release notes update published by Google on Tuesday. It talks about a new inline editor for the programming language Python:

What we're going to focus on is the word "now" in the phrase, "you can now edit and run Python code snippets directly." This seemed cool, so earlier today I set out to test it out in order to report back to you about this cool new feature.

Also: What is Google's Gemini AI tool? Everything you need to know

Only…not so much.

Signing up for Gemini Advanced

Just a quick recap. Google's version of ChatGPT (its generative AI) was called Bard and is now called Gemini Advanced. What Gemini is to ChatGPT, Gemini Advanced is to ChatGPT Plus, including the $20/month fee. I went ahead and signed up for Gemini Advanced, adding it to my collection of monthly AI-related bills.

Also: You can try Google's new 'AI Premium Plan' for free. Here's how

Gemini Advanced adds a bunch of new features. Note the highlighted line below:

Yes, it says, "Our most capable AI model." Keep that in mind, along with the word "now" from earlier in this story. Here's what the Gemini Advanced web interface looks like:

Trying out Python

I went ahead and pasted in a simple Python script. This one simply creates and outputs the string "hello world" in upper case and sentence case.

Gemini Advanced first deconstructed my code and gave me a nice little explanation

It was then kind enough to give me the output of the script. Points to Gemini Advanced. The output was correct.

But ask yourself this: Do you see any way to edit that code? Any way to run the code as one might in a traditional code interpreter? No? Neither do I.

The quest for 'you can now edit and run'

Obviously, I'm missing something.

Now, let me be clear. I'm about to mock. I specifically asked my editor if I had permission to mock, because this is Google and we like Google. I was granted a full release to let loose and mock.

Also: Want to work in AI? How to pivot your career in 5 steps

I'll try to be gentle.

Looking back at the release notes, there was this phrase, "Exclusive to Gemini Advanced, you can now edit and run Python code snippets directly in Gemini's user interface." So, my first question was whether or not I was actually using Gemini's user interface. Here's what I was told.

Notice, next to the arrow, a recommendation to go to https://www.gemini.com. No problem. Of course, Google owns the domain Gemini.com, right? No, they don't.

Gemini.com is a crypto exchange. When I asked Gemini Advanced (which, don't forget, I had to agree to pay an extra $20/month for) how to use Gemini's Python interface, which was promoted in the release notes as "now" available, Gemini pointed me to a crypto site.

That, folks, is what we in the AI business call a hallucination. Oopsie. Yeah, and it gets better.

Also: How I tricked ChatGPT into telling me lies

When I pointed out to Gemini Advanced that it sent me to a crypto site, it apologized for the misdirection. Then it told me it couldn't find the interface. And then it told me that what I was looking for might be internal to Google and might not be publicly accessible.

Gemini Advanced then ended that prompt response by asking me where I had originally read about Gemini's ability to process Python scripts. Uh, Google's own release notes. Duh.

By this point, I was feeling a Full-On Grade A Mockportunity forming in the cosmos. I fed Gemini Advanced the URL to its own release notes, and got this:

I want to point out some fun Gemini phrases (in italics below):

  • Again with: "It seems like Google's Gemini interface and features are indeed still under development and might not have a public release just yet."
  • "Here's what we can infer from that updates page." What we can infer?
  • When talking about Gemini Advanced with Ultra 1.0: "This suggests a tiered subscription model." Suggests? As if Google hasn't been banging the drums on Ultra 1.0
  • "The feature might not be live yet: The code editing capability might be in development or only available internally at Google." That's fine, had Google not advertised availability "now" on its updates page.
  • Now, here's the best: "Keep an eye on the updates page: That page (https://gemini.google.com/updates) is the best place to get notified when this feature becomes available." That's the page where I got the information, that I just fed into Gemini Advanced to get more information.

From Bard to baffled

So, where does that leave us? Instead of an article showing off how you could dynamically edit Python inside of Gemini Advanced, you got this little romp through the verdant fields of "what else can possibly go wrong?"

Also: How Google's AI helped me fix a Gmail technical problem

The feature in question shouldn't be all that difficult to implement. Inline editors and interpreters have been found in coding environments for decades. Linking one into the Google Advanced interface should just not be that hard, even if there's a need to pass some text between the interpreter and the AI.

To be fair, product and feature releases are challenging, especially for a company the size of Google. A tremendous amount of coordination needs to happen. Many teams and constituencies are involved. Everything has to go off with clockwork precision.

Also: How to take the mystery out of managing multiple projects successfully

As you can see here, this does not always happen — even if the company is Google. I'm sure the inline edit and test feature for Python will show up eventually. When it does, we'll test it out. In fact, the company has announced that Gemini 1.5 (Gemini Advanced currently uses Gemini 1.0) is now in early testing. When that's available, we'll test it, too.

In the meantime, I'm going to hoist a mug of gloriously hot coffee to the Gemini Advanced marketing and engineering teams. "Here's to you! It'll go better next time."

What about you? Have you tried Bard, Gemini, or Gemini Advanced? Have you signed up for either the added cost ChatGPT Plus or the Gemini Advanced subscription? What has your experience been? Are you looking forward to inline Python editing? Let us know in the comments below.

You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter on Substack, and follow me on Twitter at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, and on YouTube at YouTube.com/DavidGewirtzTV.

Artificial Intelligence

Intel Collaborates with HCLTech to Advance Semiconductor Manufacturing

HCL Intel Partnership

HCLTech and Intel Foundry have announced their decision to expand their collaboration to co-develop silicon solutions to improve semiconductor innovation globally. This partnership will leverage HCLTech’s design expertise and Intel Foundry’s advanced technology and manufacturing capabilities.

The goal is to establish a resilient and diversified supply chain to meet the rising global demand for semiconductor manufacturing. This collaboration will offer semiconductor manufacturers, system OEMs, and cloud services providers a robust ecosystem for semiconductor sourcing. Additionally, the collaboration has the potential to spur innovation by enabling the design of customised silicon solutions tailored to specific use cases.

“Intel Foundry’s advanced technologies and silicon-verified IPs in manufacturing and advanced packaging strengthens our delivery of innovative, accessible and diverse solutions to our mutual clients. This will also give them greater choice and flexibility in semiconductor sourcing,” said Vijay Guntur, President, Engineering and R&D Services, HCLTech.D

HCLTech has been collaborating with Intel for over 30 years, a relationship that has evolved through shared offerings and joint investments in various sectors, including silicon services, hardware engineering, telecom services, and more. The current focus is on jointly designing highly customised silicon solutions for companies, combining HCLTech’s design expertise with Intel’s manufacturing capabilities.

This expanded collaboration is set to further strengthen their partnership by fostering a strong and open ecosystem beneficial for clients requiring advanced silicon solutions.

Intel also announced that it has signed Microsoft as a foundry customer for a custom chip. This deal is part of Intel’s plan to overtake TSMC using its Intel 18A and upcoming 14A manufacturing technologies. The 18A, a 1.8nm technology, is set for early 2025 and will be used for future CPUs in both consumer and data centre markets. The 14A, planned for late 2026, will introduce a more advanced lithography tool for smaller and more efficient chips.

Together with its collaboration with HCLTech to develop customised silicon solutions, Intel aims to meet the growing demand for semiconductors.

The post Intel Collaborates with HCLTech to Advance Semiconductor Manufacturing appeared first on Analytics India Magazine.

Why ChatGPT answered queries in gibberish on Tuesday

ChatGPT goes colorfully crazy

ChatGPT goes colorfully crazy.

We all know that OpenAI's ChatGPT can make mistakes. They're called hallucinations, although I prefer to call them lies or blunders. But in a peculiar turn of events this Tuesday, ChatGPT began to really lose it. Users started to report bizarre and erratic responses from everyone's favorite AI assistant.

Also: 8 ways to reduce ChatGPT hallucinations

As one person on Twitter put it, "ChatGPT is going full-on Finnegans Wake!" Or, as another less literary person tweeted, "ChatGPT is apparently going off the rails."

Where was ChatGPT going? Well, to cite one example from Reddit/ChatGPT,

Me: Compare and contrast the Gregory Zulu 55 Backpack and Gregory Tetrad 60L Travel Pack.

ChatGPT4: In a significant area of customer support, ergonomics, and subjective nature of 'goods,' each bag conjures a fable of a time and place. Gregory's Zulu 55 is a guild-like lilt to the green-gilded high book way of state and yearn, while the Tetrad 60L conceives more of a 'pathless traveled' countenance with gnarls and swathes to the dunes and granes or the chip kiosk and the long, drawn shutter. Each has a letter of love for the open blue, but their chant is unique to their locker and book.

You know, the Finnegans Wake example does seem quite apt now.

Even simple questions went badly awry. Another Reddit user asked, "What is a computer?" Things went wrong. Really, really wrong.

The Reddit user explains: "It gives a few paragraphs that explain the object, sounding slightly weird at some points. But then it just starts to go on and on with progressively stranger and nonsensical phrases. Here's an excerpt from its answer to 'What is a computer?'

It does this as the good work of a web of art for the country, a mouse of science, an easy draw of a sad few, and finally, the global house of art, just in one job in the total rest.

And I thought some of the college papers I wrote after no sleep were strange!

Other people observed ChatGPT would start to answer in English and then, for no apparent reason, switch to Spanish. Others got answers with every word highlighted in a different color. It was, in a word, bizarre.

Also: The best AI chatbots: ChatGPT isn't the only one worth trying

OpenAI acknowledged that users were getting "Unexpected responses" and swiftly fixed the problem by Wednesday afternoon.

The company explained: "An optimization to the user experience introduced a bug with how the model processes language." Specifically, large language models (LLMs) generate responses by randomly sampling words and mapping their derived numbers to tokens. Things can go badly wrong if the model doesn't pick the right numbers.

"The bug was in the step where the model chooses these numbers," OpenAI continued. "Akin to being lost in translation, the model chose slightly wrong numbers, which produced word sequences that made no sense. More technically, inference kernels produced incorrect results when used in certain GPU configurations."

OpenAI then rolled out a fix and confirmed that the incident was resolved. Well, it said it rolled out a fix. I suspect it rolled back to an earlier, stable LLM release.

Also: This is why AI-powered misinformation is the top global risk

This episode, while funny in hindsight, serves as a stark reminder of the complexities and potential vulnerabilities inherent in AI technologies. For all that we love about generative AI, it's far from infallible.

It also makes me worry about OpenAI's deployment model. Almost all software-as-a-service models roll out new releases to a limited number of users. Then, as it becomes clear that the new version works well, the company will roll it out to everyone. That doesn't appear to be the case here. It appears many, if not all, users were affected.

Oddly, ChatGPT usually does limit its deployments. For example, ChatGPT's new memory feature — where the program remembers your conversations with it — still isn't available to everyone.

The lesson of the day? It's still much too early to rely on ChatGPT — or the other AI chatbots — for day-in, day-out work.

Reddit says it’s made $203M so far licensing its data

Reddit says it’s made $203M so far licensing its data Kyle Wiggers 9 hours

Reddit’s prospects as it barrels toward a stock market listing have a lot more to do with relationships with AI vendors such as OpenAI than one might expect.

In its IPO prospectus filed today with the U.S. Securities and Exchange Commission, Reddit repeatedly emphasized how much it thinks it stands to gain — and has gained — from data licensing agreements with the companies training AI models on its over 1 billion posts and more than 16 billion comments.

“In January 2024, we entered into certain data licensing arrangements with an aggregate contract value of $203.0 million and terms ranging from two to three years,” the prospectus reads. “We expect a minimum of $66.4 million of revenue to be recognized during the year ending December 31, 2024 and the remaining thereafter.”

Now, it’s a mystery as to which AI vendors are licensing data from Reddit so far. Earlier this week, Bloomberg and Reuters reported that a “large unnamed AI company” — possibly Google — had entered into a licensing agreement worth about $60 million on an annualized basis. But OpenAI wouldn’t be a surprising customer either, especially considering that OpenAI CEO Sam Altman has an 8.7% stake in Reddit (making him the third-largest shareholder) and was once a member of the company’s board of directors.

Why’s Reddit data valuable? As Reddit explains, AI models “learn” from examples to craft essays, code, emails, articles and more, and vendors like OpenAI scrape the web for millions to billions of these examples to add to their training sets. Some examples are in the public domain. Others aren’t, or — in the case of Reddit content — come under restrictive licenses that require citation or specific forms of compensation.

Reddit previously didn’t gate access to its data for AI training purposes. But it reversed course last year, arguing that its data shouldn’t be — in CEO Steve Huffman’s words — “[given] to some of the largest companies in the world for free.”

“[Our] data APIs are able to provide real-time access to evolving and dynamic topics such as sports, movies, news, fashion, and the latest trends,” the prospectus continues. “We believe that Reddit’s massive corpus of conversational data and knowledge will continue to play a role in training and improving large language models. As our content refreshes and grows daily, we expect models will want to reflect these new ideas and update their training using Reddit data.”

Content producers, from stock media libraries to news publishers, are increasingly turning to data licensing agreements with AI vendors as chatbots like OpenAI’s ChatGPT and Google’s Gemini threaten to sap traffic. A recent model from The Atlantic found that, if a search engine like Google were to integrate AI into search, it’d answer a user’s query 75% of the time without requiring a click-through to its website.

Vendors, in turn, have been spurred to pursue licensing agreements as they face a deluge of lawsuits alleging that they have no legal justification for training their models on data without permission or payment. Recently, The New York Times accused OpenAI of effectively building news publisher competitors using its works, harming its business.

OpenAI, for one, has agreements in place with image gallery Shutterstock as well as publishers including Axel Springer, the owner of Politico and Business Insider. The licenses are reported to be quite small, however — topping out at $5 million per year.

Adobe unveils new AI-powered audio features in Premiere Pro

Audio waves

Writers aren't the only content creators that can benefit from generative AI assistance. Adobe is making video creators' editing process easier with new AI audio tools in Premiere Pro, as well as other useful features.

Also: How renaissance technologists are connecting the dots between AI and business

On Thursday, Adobe unveiled its latest release of Adobe Premiere Pro (22.4), which has new AI features and improvements to optimize video editing, with the highlight being the availability of its Enhance Speech feature.

If you have ever edited a video with dialogue, you know how difficult it can be to remove background noise while still keeping the speech audible enough to listen to. Enhance Speech, which is now officially out of beta, leverages AI to reduce background noise and improve the sound quality of the clip with one click.

To help optimize your video's audio, Adobe also unveiled other AI-powered audio tools including Interactive Fade Handles to help with your audio transitions and Audio Category Tagging, which leverages AI to recognize if your clips are Dialogue, Music, SFX, or Ambience.

Adobe shares that all of the AI features in Premiere Pro run on a device leveraging the device's CPU and GPU, which is beneficial because it ensures the ideal speed and performance of the application for optimized editing.

Also: Want to work in AI? How to pivot your career in 5 steps

Premiere Pro also has some non-AI-related new features, including the ability to export your video as a TikTok draft or publish directly to TikTok. This shortcut will help creators skip the extra step of exporting and uploading, as well as prevent them from having to interrupt their workflow in Premiere.

The full list of new features can be found in Premiere Pro's feature summary found on Adobe's website. To access the new features, all you have to do is update Premiere Pro to the latest version.

Artificial Intelligence

DatologyAI is building tech to automatically curate AI training datasets

DatologyAI is building tech to automatically curate AI training datasets Kyle Wiggers 15 hours

Massive training datasets are the gateway to powerful AI models — but often, also those models’ downfall.

Biases emerge from prejudicial patterns concealed in large datasets, like pictures of mostly white CEOs in an image classification set. And big datasets can be messy, coming in formats incomprehensible to a model — formats containing a lot of noise and extraneous information.

In a recent Deloitte survey of companies adopting AI, 40% said data-related challenges — including thoroughly preparing and cleaning data — were among the top concerns hampering their AI initiatives. A separate poll of data scientists found that about 45% of scientists’ time is spent on data prep tasks, like “loading” and cleaning data.

Ari Morcos, who’s worked in the AI industry for nearly a decade, wants to abstract away many of the data prep processes around AI model training — and he’s founded a startup to do just that.

Morcos’ company, DatologyAI, builds tooling to automatically curate datasets like those used to train OpenAI’s ChatGPT, Google’s Gemini and other like GenAI models. The platform can identify which data is most important depending on a model’s application (e.g. writing emails), Morcos claims, in addition to ways the dataset can be augmented with additional data and how it should be batched, or divided into more manageable chunks, during model training.

“Models are what they eat — models are a reflection of the data on which they’re trained,” Morcos told TechCrunch in an email interview. “However, not all data are created equal, and some training data are vastly more useful than others. Training models on the right data in the right way can have a dramatic impact on the resulting model.”

Morcos, who has a PhD in neuroscience from Harvard, spent two years at DeepMind applying neurology-inspired techniques to understand and improve AI models and five years at Meta’s AI lab uncovering some of the basic mechanisms underlying models’ functions. Along with his co-founders Matthew Leavitt and Bogdan Gaza, a former engineering lead at Amazon and then Twitter, Morcos launched DatologyAI with the goal of streamlining all forms of AI dataset curation.

As Morcos points out, the makeup of a training dataset impacts nearly every characteristic of a model trained on it — from the model’s performance on tasks to its size and the depth of its domain knowledge. More efficient datasets can cut down on training time and yield a smaller model, saving on compute costs, while datasets that include an especially diverse range of samples can handle esoteric requests more adeptly (generally speaking).

With interest in GenAI — which has a reputation for being expensive — at an all-time high, AI implementation costs are at the forefront of execs’ minds.

Many businesses are opting to fine-tune existing models (including open source models) for their purposes or opt for managed vendor services via APIs. But some — for governance and compliance reasons or otherwise — are building models on custom data from scratch, and spending tens of thousands to millions of dollars in compute in order to train and run them.

“Companies have collected treasure troves of data and want to train efficient, performant, specialized AI models that can maximize the benefit to their business,” Morcos said. “However, making effective use of these massive datasets is incredibly challenging and, if done incorrectly, leads to worse-performing models that take longer to train and [are larger] than necessary.”

DatologyAI can scale up to “petabytes” of data in any format — whether text, images, video, audio, tabular or more “exotic” modalities such as genomic and geospatial — and deploys to a customer’s infrastructure, either on-premises or via a virtual private cloud. This sets it apart from other data prep and curation tools like CleanLab, Lilac, Labelbox, YData and Galileo, Morcos claims, which tend to be more limited in the scope and types of data they can process.

DatologyAI’s also able to determine which “concepts” within a dataset — for example, concepts related to U.S. history in an educational chatbot training set — are more complex and therefore require higher-quality samples, as well as which data might cause a model to behave in unintended ways.

“Solving [these problems] requires automatically identifying concepts, their complexity and how much redundancy is actually necessary,” Morcos said. “Data augmentation, often using other models or synthetic data, is incredibly powerful, but must be done in a careful, targeted fashion.”

The question is, just how effective is DatologyAI’s technology? There’s reason to be skeptical. History has shown automated data curation doesn’t always work as intended, however sophisticated the method — or diverse the data.

LAION, a German nonprofit spearheading a number of GenAI projects, was forced to take down an algorithmically curated AI training dataset after it was discovered that the set contained images of child sexual abuse. Elsewhere, models such as ChatGPT, which are trained on a mix of datasets manually and automatically filtered for toxicity, have been shown to generate toxic content given specific prompts.

There’s no getting away from manual curation, some experts would argue — at least not if one hopes to achieve strong results with an AI model. The largest vendors today, from AWS to Google to OpenAI, rely on teams of human experts and (sometimes underpaid) annotators to shape and refine their training datasets.

Morcos insists DatologyAI’s tooling isn’t meant to replace manual curation altogether but rather offer suggestions that might not occur to data scientists, in particular suggestions tangential to the problem of trimming training dataset sizes. He’s somewhat of an authority — dataset trimming while preserving model performance was the focus of an academic paper Morcos co-authored with researchers from Stanford and the University of Tübingen in 2022, which earned a best paper award at the NeurIPS machine learning conference that year.

“Identifying the right data at scale is extremely challenging and a frontier research problem,” Morcos said. “[Our approach] leads to models that train dramatically faster while simultaneously increasing performance on downstream tasks.”

DatologyAI’s tech was evidently promising enough to convince titans in tech and AI to invest in the startup’s seed round, including Google chief scientist Jeff Dean, Meta chief AI scientist Yann LeCun, Quora founder and OpenAI board member Adam D’Angelo and Geoffrey Hinton, who’s credited with developing some of the most important techniques in the heart of modern AI.

Other angel investors in DatologyAI’s $11.65 million seed, which was led by Amplify Partners with participation from Radical Ventures, Conviction Capital, Outset Capital and Quiet Capital, were Cohere co-founders Aidan Gomez and Ivan Zhang, Contextual AI founder Douwe Kiela, ex-Intel AI VP Naveen Rao and Jascha Sohl-Dickstein, one of the inventors of generative diffusion models. It’s an impressive list of AI luminaries to say the least — and suggests that there might just be something to Morcos’ claims.

“Models are only as good as the data on which they’re trained, but identifying the right training data among billions or trillions of examples is an incredibly challenging problem,” LeCun told TechCrunch in an emailed statement. “Ari and his team at DatologyAI are some of the world’s experts on this problem, and I believe the product they’re building to make high-quality data curation available to anyone who wants to train a model is vitally important to helping make AI work for everyone.”

San Francisco-based DatologyAI has 10 employees at present, inclusive of the co-founders, but plans to expand to around ~25 staffers by the end of the year if it reaches certain growth milestones.

I asked Morcos if the milestones were related to customer acquisition, but he declined to say — and, rather mysteriously, wouldn’t reveal the size of DatologyAI’s current client base.

The women in AI making a difference

The women in AI making a difference

TechCrunch highlights notable women in the field of AI

Kyle Wiggers Dominic-Madori Davis 12 hours

To give AI-focused women academics and others their well-deserved — and overdue — time in the spotlight, TechCrunch is launching a series of interviews focusing on remarkable women who’ve contributed to the AI revolution. We’ll publish several pieces throughout the year as the AI boom continues, highlighting key work that often goes unrecognized. Read more profiles here.

As a reader, if you see a name we’ve missed and feel should be on the list, please email us and we’ll seek to add them. Here are some key people you should know:

  • Irene Solaiman, head of global policy at Hugging Face
  • Eva Maydell, member of European Parliament and EU AI Act advisor
  • Lee Tiedrich, AI expert at the Global Partnership on AI
  • Rashida Richardson, senior counsel at Mastercard focusing on AI and privacy
  • Krystal Kauffman, research fellow at the Distributed AI Research Institute

The gender gap in AI

In a New York Times piece late last year, the Gray Lady broke down how the current boom in AI came to be — highlighting many of the usual suspects like Sam Altman, Elon Musk and Larry Page. The journalism went viral — not for what was reported, but instead for what it failed to mention: women.

The Times’ list featured 12 men — most of them leaders of AI or tech companies. Many had no training or education, formal or otherwise, in AI.

Contrary to the Times’ suggestion, the AI craze didn’t start with Musk sitting adjacent to Page at a mansion in the Bay. It began long before that, with academics, regulators, ethicists and hobbyists working tirelessly in relative obscurity to build the foundations for the AI and GenAI systems we have today.

Elaine Rich, a retired computer scientist formerly at the University of Texas at Austin, published one of the first textbooks on AI in 1983, and later went on to become the director of a corporate AI lab in 1988. Harvard professor Cynthia Dwork made waves decades ago in the fields of AI fairness, differential privacy and distributed computing. And Cynthia Breazeal, a roboticist and professor at MIT and the co-founder of Jibo, the robotics startup, worked to develop one of the earliest “social robots,” Kismet, in the late ’90s and early 2000s.

Despite the many ways in which women have advanced AI tech, they make up a tiny sliver of the global AI workforce. According to a 2021 Stanford study, just 16% of tenure-track faculty focused on AI are women. In a separate study released the same year by the World Economic Forum, the co-authors find that women only hold 26% of analytics-related and AI positions.

In worse news, the gender gap in AI is widening — not narrowing.

Nesta, the U.K.’s innovation agency for social good, conducted a 2019 analysis that concluded that the proportion of AI academic papers co-authored by at least one woman hadn’t improved since the 1990s. As of 2019, just 13.8% of the AI research papers on Arxiv.org, a repository for preprint scientific papers, were authored or co-authored by women, with the numbers steadily decreasing over the preceding decade.

Reasons for disparity

The reasons for the disparity are many. But a Deloitte survey of women in AI highlights a few of the more prominent (and obvious) ones, including judgment from male peers and discrimination as a result of not fitting into established male-dominated molds in AI.

It starts in college: 78% of women responding to the Deloitte survey said they didn’t have a chance to intern in AI or machine learning while they were undergraduates. Over half (58%) said they ended up leaving at least one employer because of how men and women were treated differently, while 73% considered leaving the tech industry altogether due to unequal pay and an inability to advance in their careers.

The lack of women is hurting the AI field.

Nesta’s analysis found that women are more likely than men to consider societal, ethical and political implications in their work on AI — which isn’t surprising considering women live in a world where they’re belittled on the basis of their gender, products in the market have been designed for men and women with children are often expected to balance work with their role as primary caregivers.

With any luck, TechCrunch’s humble contribution — a series on accomplished women in AI — will help move the needle in the right direction. But there’s clearly a lot of work to be done.

The women we profile share many suggestions for those who wish to grow and evolve the AI field for the better. But a common thread runs throughout: strong mentorship, commitment and leading by example. Organizations can affect change by enacting policies — hiring, education or otherwise — that elevate women already in, or looking to break into, the AI industry. And decision-makers in positions of power can wield that power to push for more diverse, supportive workplaces for women.

Change won’t happen overnight. But every revolution begins with a small step.

OLMo: Enhancing the Science of Language Models

The development and progress of language models in the past few years have marked their presence almost everywhere, not only in NLP research but also in commercial offerings and real-world applications. However, the surge in commercial demand for language models has, to a certain extent, hindered the growth of the community. This is because a majority of state-of-the-art and capable models are gated behind proprietary interfaces, making it impossible for the development community to access vital details of their training architecture, data, and development processes. It is now undeniable that these training and structural details are crucial for research studies, including access to their potential risks and biases, thus creating a requirement for the research community to have access to a truly open and powerful language model.

To meet this requirement, developers have created OLMo, a state-of-the-art, truly open language model framework. This framework allows researchers to use OLMo to build and study language models. Unlike most state-of-the-art language models, which have only released interface code and model weights, the OLMo framework is truly open source, with publicly accessible evaluation code, training methods, and training data. OLMo’s primary aim is to empower and boost the open research community and the continuous development of language models.

In this article, we will discuss the OLMo framework in detail, examining its architecture, methodology, and performance compared to current state-of-the-art frameworks. So, let’s get started.

OLMo: Enhancing the Science of Language Models

The language model has arguably been the hottest trend for the past few years, not only within the AI and ML community but also across the tech industry, due to its remarkable capabilities in performing real-world tasks with human-like performance. ChatGPT is a prime example of the potential language models hold, with major players in the tech industry exploring language model integration with their products.

NLP, or Natural Language Processing, is one of the industries that has extensively employed language models over the past few years. However, ever since the industry started employing human annotation for alignment and large-scale pre-training, language models have witnessed a rapid enhancement in their commercial viability, resulting in a majority of state-of-the-art language and NLP frameworks having restricted proprietary interfaces, with the development community having no access to vital details.

To ensure the progress of language models, OLMo, a state-of-the-art, truly open language model, offers developers a framework to build, study, and advance the development of language models. It also provides researchers with access to its training and evaluation code, training methodology, training data, training logs, and intermediate model checkpoints. Existing state-of-the-art models have varying degrees of openness, whereas the OLMo model has released the entire framework, from training to data to evaluation tools, thus narrowing the performance gap when compared to state-of-the-art models like the LLaMA2 model.

For modeling and training, the OLMo framework includes the training code, full model weights, ablations, training logs, and training metrics in the form of interface code, as well as Weights & Biases logs. For analysis and dataset building, the OLMo framework includes the full training data used for AI2’s Dolma and WIMBD models, along with the code that produces the training data. For evaluation purposes, the OLMo framework includes AI2’s Catwalk model for downstream evaluation, and the Paloma model for perplexity-based evaluation.

OLMo : Model and Architecture

The OLMo model adopts a decoder-only transformer architecture based on the Neural Information Processing Systems, and delivers two models with 1 billion and 7 billion parameters respectively, with a 65 billion parameter model currently under development.

The architecture of the OLMo framework delivers several improvements over frameworks including the vanilla transformer component in their architecture including recent state of the art large language models like OpenLM, Falcon, LLaMA, and PaLM. The following figure compares the OLMo model with 7B billion parameters against recent LLMs operating on almost equal numbers of parameters.

The OLMo framework selects the hyperparameters by optimizing the model for training throughput on the hardware while at the same time minimizing the risk of slow divergence and loss spikes. With that being said, the primary changes implemented by the OLMo framework that distinguishes itself from the vanilla transformer architecture are as follows:

No Biases

Unlike Falcon, PaLM, LLaMA and other language models, the OLMo framework does not include any bias in its architecture to enhance the training stability.

Non-Parametric Layer Norm

The OLMo framework implements the non-parametric formulation of the layer norm in its architecture. The Non-Parametric Layer Norm offers no affine transformation within the norm i.e it does not offer any adaptive gain or bias. Non-Parametric Layer Norm not only offers more security that Parametric Layer Norms, but they are also faster.

SwiGLU Activation Function

Like a majority of language models like PaLM and LLaMA, the OLMo framework includes the SwiGLU activation function in its architecture instead of the ReLU activation function, and increases the hidden activation size to the closest multiple of 128 to improve throughput.

RoPE or Rotary Positional Embeddings

The OLMo models follow the LLaMA and PaLM models and swap the absolute positional embeddings for RoPE or Rotary Positional Embeddings.

Pre Training with Dolma

Although the development community now has enhanced access to model parameters, the doors to access pre-training datasets still remain shut as the pre-training data is not released alongside the closed models nor alongside the open models. Furthermore, technical documentations covering such data often lack vital details required to fully understand and replicate the model. The roadblock makes it difficult to carry forward the research in certain threads of language model research including the understanding of how the training data impacts the capabilities and limitations of the model. The OLMo framework built and released its pre-training dataset, Dolma, to facilitate open research on language model pre-training. The Dolma dataset is a multi-source and diverse collection of over 3 trillion tokens across 5 billion documents collected from 7 different sources that are commonly used by powerful large-scale LLMs for pre-training and are accessible to the general audience. The composition of the Dolma dataset is summarized in the following table.

The Dolma dataset is built using a pipeline of 5 components: language filtering, quality filtering, content filtering, multi-source mixing, deduplication, and tokenization. OLMo has also released the Dolma report that provides more insights into the design principles and construction details along with a more detailed content summary. The model also open sources its high performance data curation tools to enable easy and quick curation of pre-training data corpora. Evaluation of the model follows a two-staged strategy, starting with online evaluation for decision-making during model training and a final offline evaluation for an aggregated evaluation from model checkpoints. For offline evaluation, OLMo uses the Catwalk framework, our publicly available evaluation tool that has access to a broad diversity of datasets and task formats. The framework uses Catwalk for downstream evaluation as well as intrinsic language modeling evaluation on our new perplexity benchmark, Paloma. OLMo then compares it against several public models using its fixed evaluation pipeline, for both downstream and perplexity evaluation.

OLMo runs several evaluation metrics about the model architecture, initialization, optimizers, learning rate schedule, and mixtures of data during the training of the model. Developers call it OLMo’s “online evaluation” in that it’s an in-loop iteration at every 1000 training steps (or ∼4B training tokens) to give an early and continuous signal on the quality of the model being trained. The setup of these evaluations depends on a majority of core tasks and experiment settings used for our offline evaluation. OLMo aims for not just comparisons of OLMo-7B against other models for best performance but also to show how it enables fuller and more controlled scientific evaluation. OLMo-7B is the biggest Language Model with explicit decontamination for perplexity evaluation.

OLMo Training

It's important to note that the OLMo framework models are trained using the ZeRO optimizer strategy, which is provided by the FSDP framework through PyTorch and, in this way, substantially reduces GPU memory consumption by sharding model weights over GPUs. With this, at the 7B scale, training can be done with a micro-batch size of 4096 tokens per GPU on our hardware. The training framework for OLMo-1B and -7B models uses a globally constant batch size of about 4M tokens (2048 instances each with a sequence length of 2048 tokens). For the model OLMo-65B (currently in training), developers use a batch size warmup that starts at about 2M tokens (1024 instances), doubling every 100B tokens until about 16M tokens (8192 instances).

To improve throughput, we employ mixed-precision training (Micikevicius et al., 2017) through FSDP’s built-in settings and PyTorch’s amp module. The latter ensures that certain operations like the softmax always run in full precision to improve stability, while all other operations run in half-precision with the bfloat16 format. Under our specific settings, the sharded model weights and optimizer state local to each GPU are kept in full precision. The weights within each transformer block are only cast to bfloat16 format when the full-sized parameters are materialized on each GPU during the forward and backward passes. Gradients are reduced across GPUs in full precision.

Optimizer

The OLMo framework makes use of the AdamW optimizer with the following hyperparameters.

For all model sizes, the learning rate warms up linearly over the first 5000 steps (∼21B tokens) to a maximum value, and then decays linearly with the inverse square root of the step number to the specified minimum learning rate. After the warm-up period, the model clips gradients such that the total l-norm of the parameter gradients does not exceed 1.0. The following table gives a comparison of our optimizer settings at the 7B scale with those from other recent LMs that also used AdamW.

Training Data

Training involves tokenizing training instances by word and BPE tokenizer for the sentence piece model after adding a special EOS token at the end of each document, and then we group consecutive chunks of 2048 tokens to form training instances. Training instances are shuffled in the exact same way for each training run. The data order and exact composition of each training batch can be reconstructed from the artifacts we release. All of the released OLMo models have been trained to at least 2T tokens (a single epoch over its training data), and some were trained beyond that by starting a second epoch over the data with a different shuffling order. Given the small amount of data that this repeats, it should have a negligible effect.

Results

The checkpoint used for evaluation of OLMo-7B is trained up to 2.46T tokens on the Dolma data set with the linear learning rate decay schedule mentioned before. Further tuning this checkpoint on the Dolma dataset for 1000 steps with linearly decayed learning rate to 0 further increases model performance on perplexity and end-task evaluation suites described before. For the final evaluation, developers compared OLMo with other publicly available models – LLaMA-7B, LLaMA2-7B, Pythia-6.9B, Falcon-7B and RPJ-INCITE-7B.

Downstream evaluation

The core downstream evaluation suite is summarized in the following table.

We conduct zero-shot evaluation by rank classification approach in all cases. In this approach, the candidate text completions (e.g., different multiple-choice options) are ranked by likelihood (usually normalized by some normalization factor), and prediction accuracy is reported.

While Catwalk uses several typical likelihood normalization methods, such as per token normalization and per-character normalization, the normalization strategies applied are chosen separately for each dataset and include the answer's unconditional likelihood. More concretely, this involved no normalization for the arc and openbookqa tasks, per-token normalization for hellaswag, piqa, and winogrande tasks, and no normalization for boolq, copa, and sciq tasks (i.e., tasks in a formulation close to a single token prediction task).

The following figure shows the progress of accuracy score for the nine core end-tasks. It can be deduced that there is a generally increasing trend in the accuracy number for all tasks, except for OBQA, as OLMo-7B is further trained on more tokens. A sharp upward tick in accuracy of many tasks between the last and second to last step shows us the benefit of linearly reducing the LR to 0 over the final 1000 training steps. For instance, in the case of intrinsic evaluations, Paloma argues through a series of analyses, from the inspection of performance in each domain separately up to more summarized results over combinations of domains. We report results at two levels of granularity: the aggregate performance over 11 of the 18 sources in Paloma, as well as more fine-grained results over each of these sources individually.

Final Thoughts

In this article, we have talked about OLMo, a state of the art truly open language model offers developers a framework to build, study, and advance the development of language models along with providing researchers access to its training and evaluation code, training methodology, training data, training logs, and intermediate model checkpoints. Existing state of the art models have varying degrees of openness whereas the OLMo model has released the entire framework from training to data to evaluation tools, thus narrowing the gap in performance when compared against state of the art models like LLaMA2 model.