Does your business need a chief AI officer?

People walking in an illustration of a brain

From automation to machine learning and generative artificial intelligence (AI), emerging technologies are changing the companies we work for and the roles we fulfill.

While it's already clear these technologies will have a big impact on repetitive manual tasks and mid-tier managerial responsibilities, what do AI and automation mean for the upper echelons of the organization? CIOs are traditionally the executives who oversee technology within a business. These custodians are a steady hand when it comes to the procurement and deployment of IT.

Also: If AI is the future of your business, should the CIO be the one in control?

However, the use of AI has risen significantly during the past 12 months with the launch of OpenAI's ChatGPT and a slew of similar technologies.

The rapid pace of development means it would be tough for anyone, let alone a CIO who's already responsible for managing day-to-day IT operations, to oversee the rollout of AI.

So, given the rapid rise of emerging technologies, do businesses need a chief AI officer?

Avivah Litan, distinguished VP analyst at Gartner, says chief AI officers are being appointed as a best-practice technique in some organizations — and she gives an example.

Also: The ethics of generative AI: How we can harness this powerful technology

"Global financial institutions do have chief AI risk officers reporting into CFOs or the bank's chief risk officer," she said to ZDNET. However, not every business faces the risks of a bank. And Litan says many organizations won't be able to afford a specialist AI office.

For most companies, the best way to deal with the nascent risks and opportunities of AI is to create a collective approach that brings experts from across the business together.

"That's why you see task forces," she says. "And, generally, it's much better to have line risk. It's better to have clear responsibilities and task forces." Implementing and overseeing AI successfully is a complicated conundrum that touches on all parts of the enterprise.

Litan says all organizations must recognize AI is not only a novel phenomenon, but one that can't be treated like just any other IT application.

Also: 56% of professionals are unsure if their companies have ethical guidelines for AI use

"It's definitely different. From the security point of view, that's been our message all along: It's a new vector," she says. "And you can't use the old controls to manage risk. The same is true with opportunities. You can't just use your old business processes to manage AI opportunities. It's a different animal."

Jarrod Phipps, executive vice president and CIO at auto specialist Holman, is another expert who says any decision about whether a business needs a chief AI officer is likely to depend on size and scale. Yet, he believes having someone senior who's responsible for AI should be seen as a good thing for most organizations.

"I don't think it can hurt — and the reason is that there is now a completely new security paradigm in place," says Phipps in an interview with ZDNET, echoing the sentiments of Gartner's Litan.

He envisages a chief AI leader as sitting somewhere between the chief information security officer and a strategic data leader. The AI leader must ensure opportunities are grabbed without taking undue risks.

Also: A thorny question: Who owns code, images, and narratives generated by AI?

"There's a security and data privacy element to what we're going to be doing, but there's also this dramatic experience component to it. And you must constantly keep the balance between the two," he says. "You'll have to make some hard trade-offs. And any time there's hard trade-offs to be made, having consistency around how you execute those trade-offs is really important."

Phipps says Holman has created an AI council. Much like the AI task forces that Litan envisages, the council at Holman analyzes potential use cases and consists of a range of senior-level people from across the company, including Phipps.

In the longer term, he believes a chief AI leader who is singularly responsible for emerging technology could help businesses balance the risks and rewards of AI.

That's a sentiment that resonates with Lily Haake, head of technology and digital executive search at recruiter Harvey Nash, who told ZDNET that it's "entirely possible" companies will need a chief AI officer in the future.

Also: What technology analysts are saying about the future of generative AI

However, we're not at that point yet.

"Most organizations aren't even piloting this stuff and the data quality isn't there in many companies to need that sort of investment, but potentially, in the longer term, those requirements will exist," she says. When that situation arises, Haake expects chief AI officers will have strengths in some technical areas, such as algorithms, natural language processing, and machine learning, that an average employee wouldn't be expected to understand.

"And that person will need to have the compliance and the regulatory oversight as well. Where they will sit, I don't know, because I think there's an argument they would still sit under the CIO or the CDO," she says. "But everything depends on how quickly AI proliferates. Maybe in 10 or 15 years, there'll be a chief AI officer who sits on the executive board and has a helicopter view across the organization because, by then, AI is likely going to be imperative to every facet of the business."

For now, though, the idea of appointing another tech chief to sit alongside a CIO, CTO, CDO, or CISO is probably a step too, says Omer Grossman, global CIO at CyberArk.

Also: A new role emerges for software leaders: Overseeing generative AI

In the case of the growing range of enterprises that like to think of themselves as technology businesses, AI will simply be part of everyday working practices. In companies where AI is a key element of the business's products or services, Grossman expects another senior executive to take on AI leadership responsibilities.

"This is where you would think the head of R&D or the chief product officer, who is the person who's responsible for the company delivering its products or services, should basically oversee the use and embedding of AI capabilities into products," he says.

Like Gartner's Litan, Grossman said AI leaders are likely to play an important role in heavily regulated sectors, particularly financial services. "In a more traditional enterprise, such as a bank, there might be a chief AI officer who has an AI center of excellence or some kind of architecture or framework that elevates all the company's deliverables using AI." Grossman said.

Andy Moore, chief data officer (CDO) at Bentley Motors, is another executive who is unsure that most businesses need a single executive who's responsible for AI.

Also: Everyone wants responsible AI, but few people are doing anything about it

Like other experts, Moore refers to size and scale as being deciding factors. He says Bentley isn't a big enough company to start building its own costly large language models (LLMs) from scratch, which is the kind of project you'd expect a C-level chief AI officer to oversee.

At present, AI leadership success for most companies will be about looking for best-of-breed LLMs — and that's probably a skill set that's within the compass of a CIO and a dedicated task force that helps guide the business on use cases, risks, and rewards.

"It's about having that mindset that says, 'OK, what LLM will give me what I need, and how do I make sure I don't leave a footprint of my data on there, while still being able to draw in the benefits of emerging technology," says Moore.

Artificial Intelligence

Hugging Face has a two-person team developing ChatGPT-like AI models

Hugging Face has a two-person team developing ChatGPT-like AI models Kyle Wiggers 12 hours

AI startup Hugging Face offers a wide range of data science hosting and development tools, including a GitHub-like portal for AI code repositories, models and datasets, as well as web dashboards to demo AI-powered applications.

But some of Hugging Face’s most impressive — and capable — tools these days come from a two-person team that was formed just in January.

H4, as it’s called — “H4” being short for “helpful, honest, harmless and huggy” — aims to develop tools and “recipes” to enable the AI community to build AI-powered chatbots along the lines of ChatGPT. ChatGPT’s release was the catalyst for H4’s formation, in fact, according to Lewis Tunstall, a machine learning engineer at Hugging Face and one of H4’s two members.

“When ChatGPT was released by OpenAI in late 2022, we started brainstorming on what it might take to replicate its capabilities with open source libraries and models,” Tunstall told TechCrunch in an email interview. “H4’s primary research focus is around alignment, which broadly involves teaching LLMs how to behave according to feedback from humans (or even other AIs).”

H4 is behind a growing number of open source large language models, including Zephyr-7B-α, a fine-tuned, chat-centric version of the eponymous Mistral 7B model recently released by French AI startup Mistral. H4 also forked Falcon-40B, a model from the Technology Innovation Institute in Abu Dhabi — modifying the model to respond more helpfully to requests in natural language.

To train its models, H4 — like other research teams at Hugging Face — relies on a dedicated cluster of more than 1,000 Nvidia A100 GPUs. Tunstall and his other H4 co-worker, Ed Beeching, are based remotely in Europe, but receive support from several internal Hugging Face teams, among them the model testing and evaluation team.

“The small size of H4 is a deliberate choice, as it allows us to be more nimble and adapt to an ever-changing research landscape,” Beeching told TechCrunch via email. “We also have several external collaborations with groups such as LMSYS and LlamaIndex, who we collaborate with on joint releases.”

Lately, H4 has been investigating different alignment techniques and building tools to test how well techniques proposed by the community and industry really work. The team this month released a handbook containing all the source code and datasets they used to build Zephyr, and H4 plans to update the handbook with code from its future AI models as they’re released.

I asked whether H4 had any pressure from Hugging Face higher-ups to commercialize their work. The company, after all, has raised hundreds of millions of dollars from a pedigreed cohort of investors that includes Salesforce, IBM, AMD, Google, Amazon Intel and Nvidia. Hugging Face’s last funding round valued it at $4.5 billion — reportedly more than 100 times the company’s annualized revenue.

Tunstall said that H4 doesn’t directly monetize its tools. But he acknowledged that the tools do feed into Hugging Face’s Expert Acceleration Program, Hugging Face’s enterprise-focused offering that provides guidance from Hugging Face teams to build custom AI solutions.

Asked if he sees H4 in competition with other open source AI initiatives, like EleutherAI and LAION, Beeching said that it isn’t H4’s objective. Rather, he said, the intention is to “empower” the open AI community by releasing the training code and datasets associated with H4’s chat models.

“Our work would not be possible without the many contributions from the community,” Beeching said.

How to Use Duet AI to Create Images for Slides & Backgrounds

Google Workspace customers with the Duet AI add-on may generate images in Google Slides and Google Meet in a web browser. Duet AI in Google Slides or Meet offers an alternative to laboriously drawing custom images yourself or selecting from sterile stock photos; instead, you can type text to describe your desired image.

As always, make sure that your use of generated AI images complies with your organization’s guidelines for use and attribution.

Visit Duet AI

Jump to:

  • Using Duet AI in Google Slides and Meet
  • What types of images can Duet AI create?

Using Duet AI in Google Slides and Meet

When using Duet AI, Figure A shows how to create an image in Google Slides (i.e., Create image with Duet AI), and Figure B shows how to access the background image creation option in Google Meet (i.e., Generate a background). Activate the feature, enter text that describes an image, optionally select a style from the drop-down menu and then wait a few seconds for the system to generate images.

Figure A

Google Slides image creation.
Select the Duet AI icon in Google Slides on the web, then enter a text prompt and select Create to generate images. Image: Andy Wolber/TechRepublic

Figure B

Google meet background generator.
When using Google Meet on the web in Chrome to use Duet AI to create a background image, select the three-dot More menu | Apply Visual Effects | Generate A Background. Then enter a text prompt and select Create Samples to generate images. Select a sample to apply it as your virtual background. Image: Andy Wolber/TechRepublic

Drop-down style menu options differ between Google Slides and Google Meet. The Google Slides style drop-down defaults to No Style, but you also might select Photography, Background, Vector Art, Sketch, Watercolor, Cyberpunk and I’m Feeling Lucky options. While the I’m Feeling Lucky option is a nod to an early Google search feature that automatically took you to a first result, in this case it lets the system select a style. Similarly, the Google Meet background generator style defaults to No Style with the available options of Photography, Sci-Fi, Fantasy, 3D Animation, Illustration and Film Noir.

How to review the Duet AI generated images

You may review the generated images. When you select a generated image either by clicking or tapping it, the system adds it either as a background in Google Meet or an image in Google Slides.

If you’re not happy with any of the generated images, select View More to try again. You might also edit the text prompt to describe your desired image differently. Google’s Duet AI support page suggests that you might obtain better results when your text describes the subject, setting, distance, materials and background.

What types of images can Duet AI create?

The variety of images that Duet AI can create in Google Slides and Google Meet is vast. To give you a sense of the range and quality available, I generated five distinct types of images in different styles: an object, a scene, people, an idea and a sign. The images on the left below were the first four images the system generated, which I inserted on a Google Slide and then captured as a screenshot. The image to the right is a similar prompt used in Google Meet. Since the style options differ, the choice is noted in each case below.

Generate an object

With no style selected, the prompt “Laptop on a desk in an office” produced images that suggest a straightforward photograph of a common office scene in Google Slides and Google Meet (Figure C).

Figure C

Duet AI generated laptop images for Google Slide and Google Meet.
Laptop images generated by Duet AI in Google Slides (left) and Google Meet (right). Image: Andy Wolber/TechRepublic

Generate a scene

A prompt of “Beautiful nature scene of bird flying over the Rio Grande” resulted in an image in both Google Slides and Google Meet (Figure D) that depicted a river with varying quantities of birds in flight. The watercolor style in Google Slides and the illustration style in Google Meet evoked the quality of hand-created works.

Figure D

Duet AI generated scene images for Google Slide and Google Meet.
Scene images generated by Duet AI in Google Slides (left) and Google Meet (right). Image: Andy Wolber/TechRepublic

Generate an image of an abstract idea

The prompt “Abstract illustration of a neural network” explored how the system might show a concept. The results differed, with Google Slides set to vector art style showing roughly brain-shaped images, some nodes, lines and patterns, while Google Meet set to sci-fi style produced glowing lines and nodes of light in roughly circular patterns set against a dark horizon (Figure E).

Figure E

Duet AI generated abstract illustrations for Google Slide and Google Meet.
Abstract illustration images generated by Duet AI in Google Slides (left) and Google Meet (right). Image: Andy Wolber/TechRepublic

Generate an image with people

In my testing, the system sometimes declined to generate images with people. The prompt “Two people shaking hands, photorealistic” set to photography style in both Google Slides and Meet produced results (Figure F). However, close examination of the handshake reveals this aspect of Duet AI can still be refined to depict hands accurately.

Figure F

Duet AI generated handshake images for Google Slide and Google Meet.
Handshake images generated by Duet AI in Google Slides (left) and Google Meet (right). Image: Andy Wolber/TechRepublic

Generate a sign with text

Next, I tried a request to generate a “Sign that says ‘encourage experimentation,'” with the style option set to sketch in Slides and fantasy in Meet, respectively (Figure G). Both results adhered to the selected style, with the sketches seemingly in pencil and the fantasy image signs glowing against a dark background. The sign text in every case consists of random marks — unless it is somehow written in a language known only to Duet AI.

Figure G

Duet AI generated signs for Google Slide and Google Meet.
Signs generated by Duet AI in Google Slides (left) and Google Meet (right). Image: Andy Wolber/TechRepublic

Generate an image from literature

When prompted with the wonderfully descriptive first paragraph of James Joyce’s short story Two Gallants from his book “Dubliners,” Google Slides generated the following image (Figure H, left). Repeated attempts often similarly produced just one image in response, unlike nearly all of the above prompts that resulted in four sample images. The complexity of the text prompt likely affected the number of images the system could generate within a system-defined response time. In Google Meet, the system declined to produce an image in response to this prompt.

Figure H

Duet AI generated a text prompt of the first paragraph of James Joyce’s Two Gallants'produced images for Google Slide and Google Meet.
A text prompt of the first paragraph of James Joyce’s Two Gallants produced the image generated by Google Slides (left) and a “…can’t help with that” response in Google Meet (right). Image: Andy Wolber/TechRepublic

Mention or message me on Mastodon (@awolber) to let me know how you use Duet AI to generate images in Google Slides or backgrounds in Google Meet. What prompts and style settings produce images you prefer?

In the race to innovate with generative AI, your competitive advantage isn’t just speed—it’s also your data

gettyimages-1152447168

Generative AI, a type of AI that can create new content and ideas including conversations, stories, images, videos, and music has taken the world by storm. Like all AI, generative AI is powered by ML models—very large models that are pre-trained on vast amounts of data and commonly referred to as foundation models (FMs). And consumer-facing applications like ChatGPT have demonstrated how powerful the latest machine learning models have become.

Organizations can apply generative AI across all lines of business including engineering, marketing, customer service, finance, and sales to transform nearly every aspect of how they work. They're creating virtual assistants and new call center features to enhance the customer experience, boosting employee productivity with conversational search and code generation, and improving operations with document processing and data-enriched cybersecurity. These changes are happening across industries from healthcare to financial services as companies use generative AI to produce better results, faster.

This doesn't mean that generative AI alone will transform your business, though. To fully realize the benefits of generative AI, you need to differentiate the applications you build with it, which requires going deeper. Generative AI relies on data not only to generate content, but also to learn and evolve. Every great generative AI application is supported by a solid data strategy that helps you customize your models and build competitive advantage.

If you want to build GenAI applications that are unique to your business needs, your organization's data will be what provides the differentiator.

Check your data foundation

Generative AI, with its ability to create content, relies on data. In a simple context, the better the quality and relevance of the input data, the more refined and applicable the outputs. Data doesn't just feed AI; it shapes it, offering a foundation upon which the AI learns and evolves.

There are several ways that organizations might leverage data in their generative AI applications. While some companies will build and train their own large language models (LLMs) with vast amounts of data, many more will use their organizational data to fine-tune existing foundation models for their unique business needs or add context to prompts through Retrieval Augmented Generation (RAG), a framework for feeding LLMs accurate, up-to-date information from external sources to improve LLM responses.

For example, if you are an online travel agency that wants to generate personalized travel itineraries, you'd want to use customer profile data in your databases to tailor recommendations based on things like past trips, web history, and travel preference. You could then marry that data with other company data like flight and hotel inventory, promotions, and similar travel details.

The key to making all of these use cases work is quality data. In fact, according to the Amazon Web Services CDO survey, the number one challenge for organizations in realizing the potential of generative AI is data quality.

"The biggest disservice companies can do is to only develop a generative AI strategy," says Archana Vemulapalli, head of product and global strategy for data and AI at AWS. "You need to have a data and AI strategy." Vemulapalli suggests building a data strategy that begins with data collection and ends with data governance, with each step ensuring your data is accessible, reliable, and secure.

First, it's essential to lay the foundation of your data strategy with a scalable infrastructure for data storage. Generative AI relies on vast amounts of data, including text, images, and videos, so your infrastructure will need to be capable of handling the volume, variety, and velocity of data you'll collect. That includes breaking down data silos and bringing together data controlled by different groups across your organization. You'll also want to choose data storage that's optimized for your particular use cases. For instance, generative AI applications often use vector data, so you'll need data stores that are capable of searching and storing that type of data.

Next, the quality of data used to train generative AI models significantly impacts their performance, since models learn from the data they're trained on. Along those lines, it's important to ensure that your data is representative of your datasets and you've taken steps with those datasets to identify and mitigate bias.

It's also important to have tools to easily connect your different data sources. These tools can include data integration platforms, APIs, and connectors to software-as-a-service (SaaS) applications, on-premises data stores, and other clouds.

Finally, you need to ensure that your builders have easy but governed access to data. Establishing data governance practices is crucial to promote the integrity, security, and compliance of your data. This encompasses defining data standards, access controls, data lineage, and data lifecycle management. It also involves implementing security measures to protect sensitive data and consideration of relevant data protection regulations. Data governance encourages data use that's reliable, traceable, and compliant with privacy and legal requirements.

Once you have a solid end-to-end data foundation in place, you're ready to innovate with generative AI.

Start with a small but mighty problem

In laying out a plan for using generative AI, start by focusing on business goals. "Think about what levers you want to exercise with generative AI," Vemulapalli says. "Is the goal to drive customer experience, find new revenue streams, or build a new product out and see how it scales? Get your strategy aligned."

The next step is to find a use case that can show meaningful impact quickly. "What time-consuming, difficult, or impossible problems could generative AI help solve? Where do you have data to help in this process?" Vemulapalli says. "Think big about the opportunities, but start small. Start with a problem that causes day-to-day irritations—one that your organization will see real value in fixing."

"Just pick one known pain point and solve for that. Don't wait around for a silver bullet use case. Your use cases will evolve," she says. "Just experiment, just get going." And once you have your use case identified, you can workback backwards to identify the relevant data needed.

Choose and customize a Foundation Model

Generative AI is powered by ML models—very large models that are pre-trained on vast amounts of data and commonly referred to as Foundation Models (FMs). FMs learn to apply their knowledge within a wide range of contexts through pre-training exposure to internet-scale data in all its various forms and myriad patterns, and these "general FMs" can be used out of the box for some use cases. But many organizations are looking for FMs that can be customized to perform domain-specific functions unique to the organization. In this case, the FM must be "fine-tuned" to the organization's proprietary data.

AWS developed Amazon Bedrock for exactly this purpose. With the comprehensive capabilities of Amazon Bedrock, you can easily experiment with a variety of top FMs, privately customize them with your data using techniques such as fine-tuning and retrieval augmented generation (RAG), and create managed agents that execute complex business tasks—from booking travel and processing insurance claims to creating ad campaigns and managing inventory—all without writing any code.

Imagine a content marketing manager who works at a leading fashion retailer and needs to develop fresh, targeted ad and campaign copy for an upcoming new line of handbags. To do this, they provide a few labeled examples of their best performing taglines from past campaigns, along with the associated product descriptions. Bedrock makes a separate copy of the base foundational model that is accessible only to the customer and trains this private copy of the model that will then automatically start generating effective social media, display ad, and web copy for the new handbags. Now, the marketing manager has a new ad campaign informed by their historical data, without having to invest in a new model or incremental training, all while keeping the organization's data private and secure.

Whichever FM an organization chooses, it's critical to keep data private and secure and to retain control over who can access the models. "You want to ensure that the right guardrails are in place to protect your organization's data and IP," Vemulapalli says. "Your data is your differentiator and your ultimate competitive advantage."

Train and develop responsibly

New challenges in handling data responsibly stem from the vast size of generative AI's open-ended foundation models trained by billions of parameters, and raise new issues in defining, measuring, and mitigating responsible AI concerns across the development cycle. Accuracy, fairness, intellectual property considerations, toxicity, and privacy must all be considered on a new level.

"Consider your stance on responsible AI, transparency, data collection, security, and privacy with AI," Vemulapalli says. "How can you ensure the technology is used accurately, fairly, and appropriately?" Organizations should train on these considerations, bake them into governance and compliance frameworks, and factor them into vendor selection processes to select partners who share the same values.

"Everyone is committed to being responsible," Vemulapalli says. "What's important is how it is executed and enforced."

There's also the matter of training and upskilling your people. Consider the technical skills required to use this new technology and how to infuse them into your organization. You might look at building technical skills alongside skills like critical thinking and problem-solving. We ultimately want people, assisted by AI, to solve real business challenges and critically assess and question inferences from ML models. This is particularly important with generative AI models that distill data rather than provide considered answers.

Be ready for the next thing

Establishing an end-to-end data foundation is imperative to a successful generative AI strategy, and treating your data as your greatest asset will guide your steps along that journey. That solid data foundation will, in turn, set you up to innovate faster. "Customers are being thoughtful and fast," Vemulapalli says. "We are currently in a period of intense experimentation and quickly transitioning to at-scale implementations. Everyone recognizes the need to move quickly."

Learn more about innovating with generative AI at AWS for Data.

Fakespot Chat, Mozilla’s first LLM, lets online shoppers research products via an AI chatbot

Fakespot Chat, Mozilla’s first LLM, lets online shoppers research products via an AI chatbot Sarah Perez @sarahintampa / 9 hours

Earlier this year, Mozilla acquired Fakespot, a startup that leverages AI and machine learning to identify fake and deceptive product reviews. Now, Mozilla is launching its first LLM (large language model) with the arrival of Fakespot Chat, an AI agent that will help consumers as they shop online by answering questions about the product or even suggesting questions that could be useful in your product research.

There’s some irony in using AI to combat the scourge of fake reviews, which are today also often crafted using AI technology, like GPT. As CBNC reported in April, a number of Amazon product reviews were transparently created via ChatGPT as they began with the phrase “As an AI language model,” which tends to be part of ChatGPT’s responses. In July, TripAdvisor told The Guardian it had already removed over 20,000 reviews it believed contained AI-generated text from across over 15,000 properties in its system. The U.S. Federal Trade Commission has also proposed a rule that would attempt to ban fake product reviews, warning that AI is making the problem even worse.

But Fakespot has been using AI, including generative AI technologies, to make the online shopping process more trustworthy, not less. For instance, it launched a generative AI feature called Pros and Cons last year, that could replace the need for reading reviews by writing up its own summaries of a product’s positives and negatives. The feature was trained on billions of data points, with the model itself using five different models under its hood, the company said.

Image Credits: Mozilla Fakespot

This week, Fakespot Chat launched into testing, allowing shoppers to ask an AI chatbot about a product they’re considering, similar to how you could ask a salesperson for help if you were shopping in a physical store in the real world. The technology uses AI and machine learning to sort through the product reviews, sorting real from fake, to answer the user’s questions. The information from your chat session is saved to improve the experience for others, Mozilla notes, but users don’t have to create an account or divulge personal information for the experience to work.

The feature is available via the Fakespot Analyzer or it can be used on an Amazon.com product from Fakespot’s browser extension. For the former, you’d copy and paste the URL of the product into the analyzer to ask your questions, but if using the browser add-on, the analysis starts automatically. When the analysis is complete, Fakespot Chat appears on the right-hand side of the analysis page alongside other features, like Pros and Cons, as well as Fakespot’s Review Grades and Highlights. You can then interrogate the AI agent about the product as you weigh your purchase decisions.

The product joins other Mozilla-led AI initiatives, including a $30 million commitment to build a startup and community called Mozilla.ai focused on creating an independent open-source AI ecosystem. It also hosted its first Responsible AI Challenge which encouraged builders to compete for prize money by creating trustworthy AI solutions.

Mozilla admits its new AI chatbot may not always get things right, so it invites users to submit feedback if they think the model could be improved.

“Ultimately, our goal with Fakespot Chat is to reduce your product research time and lead you to better purchasing decisions,” said Saoud Khalifah, Founder and Director of Fakespot at Mozilla, in an announcement.

Mozilla buys Fakespot, a startup that identifies fake reviews, to bring shopping tools to Firefox

10 Major AI Updates at GitHub Universe 2023

GitHub’s parent company Microsoft is seeing big growth in the generative AI business, as the company’s CEO, Satya Nadella, told Wall Street that the company’s paying customers for its GitHub Copilot software rose by 40% in the September quarter from the prior quarter.

“We have over 1 million paid copilot users in more than 37,000 organizations that subscribe to copilot for business,” said Nadella, “with significant traction outside the United States.” Building on the existing global user base, the platform has made new major AI announcements at ongoing annual GitHub conference — Universe 2023.

In the official statement announcing the launch, Thomas Dohmke, CEO, GitHub, said: “In March, we shared our vision of a new future of software development with Copilot X, where AI infuses every step of the developer lifecycle. Our vision has manifested itself into a new reality for the world’s developers.” He further stated, “Just as GitHub was founded on Git, today we are re-founded on Copilot.”

Here are 10 AI update made at GitHub Universe 2023:

Copilot Chat

With GitHub’s latest Copilot Chat, the platform is making natural language the go-to programming language for developers. Now you can debug and find errors with ease, just by chatting.

The chatbot powered by OpenAI’s GPT-4 will be generally available in December 2023. Apart from users who have a GitHub Copilot subscription it will also be available to verified teachers, students, and maintainers of popular open source projects for free.

Slash Commands and Context Variables

Fixing or improving code has never been easier! GitHub introduces slash commands and context variables, making tasks like code fixes and test generation a breeze with simple commands like /fix and /tests.

Inline Chat Integration

Say hello to the new inline Copilot Chat, enabling developers to discuss specific lines of code seamlessly within their coding flow and editor.

One-Click Actions

GitHub’s Copilot now offers powerful shortcuts with just a click! Speed up your development process with smart actions that streamline tasks like fixing suggestions, reviewing pull requests, and generating responses.

JetBrains Suite Integration

Copilot Chat will be soon available in the JetBrains suite of IDEs, making it easier for users to access AI-powered assistance directly within their preferred coding environment.

The feature is available to preview starting today.

Chat on Mobile App and GitHub.com

GitHub is bringing Copilot Chat to their mobile app, ensuring that developers can access its powerful features even while on the move, enhancing their coding experience anytime, anywhere.

The bot will also be available on Github.com combined with the power of GitHub’s advanced code search, GitHub is enabling Copilot Chat to understand and help with the latest changes to popular open source projects.

Copilot Enterprise

GitHub Copilot initially boosted developers’ speed by 55% as an IDE autocomplete function. Now, GitHub is introducing Copilot for Enterprise which will help teams in codebase orientation, documentation creation, personalized suggestions, and swift pull request reviews.

The feature will be generally available from February 2024 at $39 USD per user monthly.

GitHub Copilot Partner Program

GitHub is teaming up with over 25 leading partners, including Datastax, LaunchDarkly, Postman, Hashicorp, and Datadog, to expand Copilot’s capabilities and create an ecosystem of AI-driven coding solutions.

AI-Powered Security

GitHub Copilot employs an LLM-based security system, detecting and hampering insecure code patterns, like hardcoded credentials and SQL injections. Henceforth, GitHub’s Advanced Security will feature AI-powered tools to uncover and mitigate vulnerabilities and sensitive data in code, for upgraded application security.

GitHub Copilot Workspace

GitHub Next’s research team has introduced the AI-driven GitHub Copilot Workspace, to facilitate idea translation into code for developers. This upcoming platform signifies GitHub’s exploration of software development.

Set for a 2024 launch, Copilot Workspace will enable seamless code creation through natural language and AI.

The post 10 Major AI Updates at GitHub Universe 2023 appeared first on Analytics India Magazine.

Learn How to Design & Deploy Responsible AI Systems

Sponsored Content

Learn How to Design & Deploy Responsible AI Systems

How do you build AI solutions that are trusted, responsible, and effective, too? Get the practical guidance you need in this white paper from Teradata. Download now.

DOWNLOAD WHITEPAPER NOW

More On This Topic

  • Learn How to Design & Deploy Responsible AI Systems
  • Coding Ethics for AI & AIOps: Designing Responsible AI Systems
  • Design effective & reliable machine learning systems!
  • Machine Learning Systems Design: A Free Stanford Course
  • Towards a Responsible and Ethical AI
  • Put Responsible AI into Practice—attend the digital event on…

Microsoft reveals plans to protect elections from deepfakes and other misinformation

Voting booths

There's been no shortage of misinformation related to elections in recent years, but the emergence of generative AI could take those dangers to a whole new level. And that's why Microsoft is taking steps to enhance election cybersecurity, including fighting deepfakes and changing search results to promote reliable information — not just in the US but also around the world.

In a blog post detailing its efforts, Microsoft president Brad Smith and corporate vice president, Technology for Fundamental Rights Teresa Hutson outlined a five-part plan to protect electoral processes.

Also: We're not ready for the impact of generative AI on elections

First, Microsoft said, it will introduce a "Content Credentials as a Service" tool that lets users digitally "sign" media to authenticate it with C2PA (or Coalition for Content Provenance and Authenticity) watermarking. In short, creators will be able to attach a signature to an image that shows when, how, and by whom it was made. Anywhere that content goes, the digital signature follows, and clicking on an embedded pin shows its history.

Additionally, a "Campaign Success Team" within Microsoft will advise and support political campaigns as they navigate AI and cybersecurity challenges, including protecting the authenticity of their own content. An "Election Communications Hub" is also being created to help democratic governments secure their election processes with access to Microsoft security and support teams in the weeks leading up to their election.

Also: How to use Bing Chat (and how it's different from ChatGPT)

Microsoft will also be using its voice to support legal changes that protect campaigns from deepfakes, the company said, as well as other misuse of new technologies — primarily the "Protect Elections from Deceptive AI Act" that bans the use of artificial intelligence to generate deceptive content about federal candidates in political ads (with important exceptions for parody, satire, and newsroom use).

Lastly, Microsoft will be making sure voters get the best information by offering "authoritative" election information when people search with Bing. The search engine will pair with the National Association of State Election Directors to proactively promote trusted sources of news around the world, making sure trusted information appears higher in search results than unreliable information.

Four main principles guide Microsoft's efforts, including the beliefs that:

  • Voters have a right to transparent and authoritative information regarding elections.

  • Candidatesshould be able to make it clear when content originates from their campaign and have recourse when their likeness is distorted by AI for the purpose of deceiving the public.

  • Political campaigns should protect themselves from cyber threats and be able to navigate AI with affordable and easily deployed tools.

  • Election authorities should have tools to ensure a secure and resilient election process.

Also: 3 ways Microsoft's new Secure Future Initiative aims to tackle growing cyber threats

Microsoft isn't the only company taking steps to fight election misinformation, as Facebook recently banned political campaigns from using its generative AI tools and started requiring campaigns to disclose any other content that was AI-generated.

Google’s AI-powered search experience expands globally to 120+ countries and territories

Google’s AI-powered search experience expands globally to 120+ countries and territories Sarah Perez @sarahintampa / 8 hours

Google’s AI-powered search experience is rolling out worldwide, after initial launches in select markets including the U.S., India, and Japan. Starting today, the AI-based conversational experience known as SGE, or Search Generative Experience, will be available in over 120 new countries and territories, globally. It will also support four new languages: Spanish, Portuguese, Korean, and Indonesian.

These join other supported languages, including English, Hindi, and Japanese. In addition, SGE will see other minor improvements, starting in the U.S., in terms of asking follow-up questions and using features like translations and definitions.

Launched earlier this year, SGE is Google’s answer to Bing Chat, the OpenAI-powered AI chatbot experience available through Bing search and Microsoft’s Edge browser. Similar to Bing Chat, SGE lets web users interact with an AI using natural language. Users can ask questions and receive responses that aren’t just a list of links, as Google has historically offered, but are fully-formed answers delivered in complete sentences, with references cited.

The experience has been steadily updated with new features following its arrival, including AI-powered summaries of paywalled articles, definitions of terms you may not be familiar with in certain subjects (like STEM, economics, history, and others), improvements to its coding-related answers, as well as the ability to generate images and write drafts, among other things. It also recently opened up to U.S. teens, ages 13-17.

Today, in addition to the global expansion, Google will begin testing a new way for users to ask follow-up questions directly on the search results page. Now, as you explore a topic, you’ll be able to see your prior questions and search results, including Search ads in dedicated slots throughout the page, Google says. The company is positioning this as an easier way to dive deeper into a topic, but it’s also about making sure its ads business stays relevant in the AI-powered search era.

Image Credits: Google

This update will arrive first in the U.S. in English in the weeks ahead.

Another improvement is coming to SGE’s translation feature. When you ask Search to translate a phrase where some words could have more than one possible meaning, you can tap on those words and pick the meaning that relates to what it is you want to say. This option may also appear when you need to specify the gender for a particular word.

Image Credits: Google

This feature will initially come to U.S. users for English-to-Spanish translations in the weeks ahead, and more countries will be added in the future.

Another small tweak involves the newly added definitions feature that allows users to ask for definitions of unfamiliar words found in answers about select educational topics in their AI-powered overviews. Now, in addition to science, economics, and history, you can also ask for definitions in areas like coding and health information. When available, these words will be highlighted, so you can hover over them to preview the definition and related images.

This option will arrive over the next month in English in the U.S. with more countries to follow.

“We’re at the beginning of a long arc of innovation, and we’re excited by the progress so far,” Hema Budaraju, Google’s Senior Director of Product Management for Search, tells TechCrunch. “Now, even more people around the world can use generative AI in Search for everyday help and we look forward to expanding to even more countries in the future.”

For reference, the full list of countries and territories that now have access to SGE includes the following:

  • American Samoa
  • Angola
  • Antigua and Barbuda
  • Bahamas
  • Bangladesh
  • Barbados
  • Belize
  • Benin
  • Bhutan
  • Bolivia
  • Botswana
  • Brazil
  • Brunei
  • Burkina Faso
  • Burundi
  • Cambodia
  • Cameroon
  • Cape Verde
  • Central African Republic
  • Chad
  • Chile
  • Christmas Island
  • Cocos (Keeling) Islands
  • Colombia
  • Comoros
  • Congo [DRC]
  • Congo [Republic]
  • Cook Islands
  • Costa Rica
  • Côte d’Ivoire
  • Dominica
  • Dominican Republic
  • Ecuador
  • El Salvador
  • Equatorial Guinea
  • Eritrea
  • Eswatini
  • Ethiopia
  • Fiji
  • French Guiana
  • Gabon
  • Gambia
  • Ghana
  • Grenada
  • Guadeloupe
  • Guam
  • Guatemala
  • Guinea
  • Guinea-Bissau
  • Guyana
  • Haiti
  • Honduras
  • Indonesia
  • Jamaica
  • Kenya
  • Kiribati
  • Kyrgyzstan
  • Laos
  • Lesotho
  • Liberia
  • Madagascar
  • Malawi
  • Malaysia
  • Maldives
  • Mali
  • Marshall Islands
  • Mauritius
  • Mexico
  • Micronesia
  • Mongolia
  • Mozambique
  • Myanmar
  • Namibia
  • Nauru
  • Nepal
  • New Zealand
  • Nicaragua
  • Niger
  • Nigeria
  • Niue
  • Northern Mariana Islands
  • Pakistan
  • Palau
  • Panama
  • Papua New Guinea
  • Paraguay
  • Peru
  • Philippines
  • Puerto Rico
  • Rwanda
  • Saint Kitts and Nevis
  • Saint Lucia
  • Saint Vincent and the Grenadines
  • Samoa
  • São Tomé and Príncipe
  • Senegal
  • Seychelles
  • Sierra Leone
  • Singapore
  • Solomon Islands
  • Somalia
  • South Africa
  • South Korea
  • South Sudan
  • Sri Lanka
  • Suriname
  • Taiwan
  • Tajikistan
  • Tanzania
  • Thailand
  • Timor-Leste
  • Togo
  • Tokelau
  • Tonga
  • Trinidad and Tobago
  • Turkmenistan
  • Turks and Caicos Islands
  • Tuvalu
  • U.S. Virgin Islands
  • Uganda
  • United States Minor Outlying Islands
  • Uruguay
  • Uzbekistan
  • Vanuatu
  • Venezuela
  • Vietnam
  • Western Sahara
  • Zambia
  • Zimbabwe

Google’s AI search experience adds AI-powered summaries, definitions and coding improvements

Google’s AI-powered search experience can now generate images, write drafts

NLP Rise with Transformer Models | A Comprehensive Analysis of T5, BERT, and GPT

Guide on NLP

Natural Language Processing (NLP) has experienced some of the most impactful breakthroughs in recent years, primarily due to the the transformer architecture. These breakthroughs have not only enhanced the capabilities of machines to understand and generate human language but have also redefined the landscape of numerous applications, from search engines to conversational AI.

To fully appreciate the significance of transformers, we must first look back at the predecessors and building blocks that laid the foundation for this revolutionary architecture.

Early NLP Techniques: The Foundations Before Transformers

Word Embeddings: From One-Hot to Word2Vec

In traditional NLP approaches, the representation of words was often literal and lacked any form of semantic or syntactic understanding. One-hot encoding is a prime example of this limitation.

One-hot encoding is a process by which categorical variables are converted into a binary vector representation where only one bit is “hot” (set to 1) while all others are “cold” (set to 0). In the context of NLP, each word in a vocabulary is represented by one-hot vectors where each vector is the size of the vocabulary, and each word is represented by a vector with all 0s and one 1 at the index corresponding to that word in the vocabulary list.

Example of One-Hot Encoding

Suppose we have a tiny vocabulary with only five words: [“king”, “queen”, “man”, “woman”, “child”]. The one-hot encoding vectors for each word would look like this:

  • “king” -> [1, 0, 0, 0, 0]
  • “queen” -> [0, 1, 0, 0, 0]
  • “man” -> [0, 0, 1, 0, 0]
  • “woman” -> [0, 0, 0, 1, 0]
  • “child” -> [0, 0, 0, 0, 1]

Mathematical Representation

If we denote V as the size of our vocabulary and wi​ as the one-hot vector representation of the i-th word in the vocabulary, the mathematical representation of wi​ would be:

wi​=[0,0,…,1,…,0,0] where the i-th position is 1 and all other positions are 0.where the i-th position is 1 and all other positions are 0.

The major downside of one-hot encoding is that it treats each word as an isolated entity, with no relation to other words. It results in sparse and high-dimensional vectors that do not capture any semantic or syntactic information about the words.

The introduction of word embeddings, most notably Word2Vec, was a pivotal moment in NLP. Developed by a team at Google led by Tomas Mikolov in 2013, Word2Vec represented words in a dense vector space, capturing syntactic and semantic word relationships based on their context within large corpora of text.

Unlike one-hot encoding, Word2Vec produces dense vectors, typically with hundreds of dimensions. Words that appear in similar contexts, such as “king” and “queen”, will have vector representations that are closer to each other in the vector space.

For illustration, let's assume we have trained a Word2Vec model and now represent words in a hypothetical 3-dimensional space. The embeddings (which are usually more than 3D but reduced here for simplicity) might look something like this:

  • “king” -> [0.2, 0.1, 0.9]
  • “queen” -> [0.21, 0.13, 0.85]
  • “man” -> [0.4, 0.3, 0.2]
  • “woman” -> [0.41, 0.33, 0.27]
  • “child” -> [0.5, 0.5, 0.1]

While these numbers are fictitious, they illustrate how similar words have similar vectors.

Mathematical Representation

If we represent the Word2Vec embedding of a word as vw​, and our embedding space has d dimensions, then vw​ can be represented as:

vw​=[v1​,v2​,…,vd​] where each vi​ is a floating-point number representing a feature of the word in the embedding space.

Semantic Relationships

Word2Vec can even capture complex relationships, such as analogies. For example, the famous relationship captured by Word2Vec embeddings is:

vector(“king”) – vector(“man”) + vector(“woman”)≈vector(“queen”)vector(“king”) – vector(“man”) + vector(“woman”)≈vector(“queen”)

This is possible because Word2Vec adjusts the word vectors during training so that words that share common contexts in the corpus are positioned closely in the vector space.

Word2Vec uses two main architectures to produce a distributed representation of words: Continuous Bag-of-Words (CBOW) and Skip-Gram. CBOW predicts a target word from its surrounding context words, whereas Skip-Gram does the reverse, predicting context words from a target word. This allowed machines to begin understanding word usage and meaning in a more nuanced way.

Sequence Modeling: RNNs and LSTMs

As the field progressed, the focus shifted toward understanding sequences of text, which was crucial for tasks like machine translation, text summarization, and sentiment analysis. Recurrent Neural Networks (RNNs) became the cornerstone for these applications due to their ability to handle sequential data by maintaining a form of memory.

However, RNNs were not without limitations. They struggled with long-term dependencies due to the vanishing gradient problem, where information gets lost over long sequences, making it challenging to learn correlations between distant events.

Long Short-Term Memory networks (LSTMs), introduced by Sepp Hochreiter and Jürgen Schmidhuber in 1997, addressed this issue with a more sophisticated architecture. LSTMs have gates that control the flow of information: the input gate, the forget gate, and the output gate. These gates determine what information is stored, updated, or discarded, allowing the network to preserve long-term dependencies and significantly improving the performance on a wide array of NLP tasks.

The Transformer Architecture

The landscape of NLP underwent a dramatic transformation with the introduction of the transformer model in the landmark paper “Attention is All You Need” by Vaswani et al. in 2017. The transformer architecture departs from the sequential processing of RNNs and LSTMs and instead utilizes a mechanism called ‘self-attention' to weigh the influence of different parts of the input data.

The core idea of the transformer is that it can process the entire input data at once, rather than sequentially. This allows for much more parallelization and, as a result, significant increases in training speed. The self-attention mechanism enables the model to focus on different parts of the text as it processes it, which is crucial for understanding the context and the relationships between words, no matter their position in the text.

Encoder and Decoder in Transformers:

In the original Transformer model, as described in the paper “Attention is All You Need” by Vaswani et al., the architecture is divided into two main parts: the encoder and the decoder. Both parts are composed of layers that have the same general structure but serve different purposes.

Encoder:

  • Role: The encoder's role is to process the input data and create a representation that captures the relationships between the elements (like words in a sentence). This part of the transformer does not generate any new content; it simply transforms the input into a state that the decoder can use.
  • Functionality: Each encoder layer has self-attention mechanisms and feed-forward neural networks. The self-attention mechanism allows each position in the encoder to attend to all positions in the previous layer of the encoder—thus, it can learn the context around each word.
  • Contextual Embeddings: The output of the encoder is a series of vectors which represent the input sequence in a high-dimensional space. These vectors are often referred to as contextual embeddings because they encode not just the individual words but also their context within the sentence.

Decoder:

  • Role: The decoder's role is to generate output data sequentially, one part at a time, based on the input it receives from the encoder and what it has generated so far. It is designed for tasks like text generation, where the order of generation is crucial.
  • Functionality: Decoder layers also contain self-attention mechanisms, but they are masked to prevent positions from attending to subsequent positions. This ensures that the prediction for a particular position can only depend on known outputs at positions before it. Additionally, the decoder layers include a second attention mechanism that attends to the output of the encoder, integrating the context from the input into the generation process.
  • Sequential Generation Capabilities: This refers to the ability of the decoder to generate a sequence one element at a time, building on what it has already produced. For example, when generating text, the decoder predicts the next word based on the context provided by the encoder and the sequence of words it has already generated.

Each of these sub-layers within the encoder and decoder is crucial for the model's ability to handle complex NLP tasks. The multi-head attention mechanism, in particular, allows the model to selectively focus on different parts of the sequence, providing a rich understanding of context.

Popular Models Leveraging Transformers

Following the initial success of the transformer model, there was an explosion of new models built on its architecture, each with its own innovations and optimizations for different tasks:

BERT (Bidirectional Encoder Representations from Transformers): Introduced by Google in 2018, BERT revolutionized the way contextual information is integrated into language representations. By pre-training on a large corpus of text with a masked language model and next-sentence prediction, BERT captures rich bidirectional contexts and has achieved state-of-the-art results on a wide array of NLP tasks.

BERT

BERT

T5 (Text-to-Text Transfer Transformer): Introduced by Google in 2020, T5 reframes all NLP tasks as a text-to-text problem, using a unified text-based format. This approach simplifies the process of applying the model to a variety of tasks, including translation, summarization, and question answering.

t5 Architecture

T5 Architecture

GPT (Generative Pre-trained Transformer): Developed by OpenAI, the GPT line of models started with GPT-1 and reached GPT-4 by 2023. These models are pre-trained using unsupervised learning on vast amounts of text data and fine-tuned for various tasks. Their ability to generate coherent and contextually relevant text has made them highly influential in both academic and commercial AI applications.

GPT

GPT Architecture

Here's a more in-depth comparison of the T5, BERT, and GPT models across various dimensions:

1. Tokenization and Vocabulary

  • BERT: Uses WordPiece tokenization with a vocabulary size of around 30,000 tokens.
  • GPT: Employs Byte Pair Encoding (BPE) with a large vocabulary size (e.g., GPT-3 has a vocabulary size of 175,000).
  • T5: Utilizes SentencePiece tokenization which treats the text as raw and does not require pre-segmented words.

2. Pre-training Objectives

  • BERT: Masked Language Modeling (MLM) and Next Sentence Prediction (NSP).
  • GPT: Causal Language Modeling (CLM), where each token predicts the next token in the sequence.
  • T5: Uses a denoising objective where random spans of text are replaced with a sentinel token and the model learns to reconstruct the original text.

3. Input Representation

  • BERT: Token, Segment, and Positional Embeddings are combined to represent the input.
  • GPT: Token and Positional Embeddings are combined (no segment embeddings as it is not designed for sentence-pair tasks).
  • T5: Only Token Embeddings with added Relative Positional Encodings during the attention operations.

4. Attention Mechanism

  • BERT: Uses absolute positional encodings and allows each token to attend to all tokens to the left and right (bidirectional attention).
  • GPT: Also uses absolute positional encodings but restricts attention to previous tokens only (unidirectional attention).
  • T5: Implements a variant of the transformer that uses relative position biases instead of positional embeddings.

5. Model Architecture

  • BERT: Encoder-only architecture with multiple layers of transformer blocks.
  • GPT: Decoder-only architecture, also with multiple layers but designed for generative tasks.
  • T5: Encoder-decoder architecture, where both the encoder and decoder are composed of transformer layers.

6. Fine-tuning Approach

  • BERT: Adapts the final hidden states of the pre-trained model for downstream tasks with additional output layers as needed.
  • GPT: Adds a linear layer on top of the transformer and fine-tunes on the downstream task using the same causal language modeling objective.
  • T5: Converts all tasks into a text-to-text format, where the model is fine-tuned to generate the target sequence from the input sequence.

7. Training Data and Scale

  • BERT: Trained on BooksCorpus and English Wikipedia.
  • GPT: GPT-2 and GPT-3 have been trained on diverse datasets extracted from the internet, with GPT-3 being trained on an even larger corpus called the Common Crawl.
  • T5: Trained on the “Colossal Clean Crawled Corpus”, which is a large and clean version of the Common Crawl.

8. Handling of Context and Bidirectionality

  • BERT: Designed to understand context in both directions simultaneously.
  • GPT: Trained to understand context in a forward direction (left-to-right).
  • T5: Can model bidirectional context in the encoder and unidirectional in the decoder, appropriate for sequence-to-sequence tasks.

9. Adaptability to Downstream Tasks

  • BERT: Requires task-specific head layers and fine-tuning for each downstream task.
  • GPT: Is generative in nature and can be prompted to perform tasks with minimal changes to its structure.
  • T5: Treats every task as a “text-to-text” problem, making it inherently flexible and adaptable to new tasks.

10. Interpretability and Explainability

  • BERT: The bidirectional nature provides rich contextual embeddings but can be harder to interpret.
  • GPT: The unidirectional context may be more straightforward to follow but lacks the depth of bidirectional context.
  • T5: The encoder-decoder framework provides a clear separation of processing steps but can be complex to analyze due to its generative nature.

The Impact of Transformers on NLP

Transformers have revolutionized the field of NLP by enabling models to process sequences of data in parallel, which dramatically increased the speed and efficiency of training large neural networks. They introduced the self-attention mechanism, allowing models to weigh the significance of each part of the input data, regardless of distance within the sequence. This led to unprecedented improvements in a wide array of NLP tasks, including but not limited to translation, question answering, and text summarization.

Research continues to push the boundaries of what transformer-based models can achieve. GPT-4 and its contemporaries are not just larger in scale but also more efficient and capable due to advances in architecture and training methods. Techniques like few-shot learning, where models perform tasks with minimal examples, and methods for more effective transfer learning are at the forefront of current research.

The language models like those based on transformers learn from data which can contain biases. Researchers and practitioners are actively working to identify, understand, and mitigate these biases. Techniques range from curated training datasets to post-training adjustments aimed at fairness and neutrality.