Photoshop for music? Adobe unveils new AI tools for audio

Sound/light waves

Adobe's generative AI tools, such as generative fill, have transformed the photo editing game. Now, Adobe is attempting to do the same but for audio with a new AI audio editing platform.

On Wednesday, the Adobe Research Team unveiled Project Music GenAI Control, an early-stage generative AI music generation and editing tool that allows users to generate music from text prompts, with AI acting as a music co-creator.

Also: The best AI image generators

Like with other text-to-music generators, Project Music GenAI Control, will enable users to enter a simple text prompt such as "powerful rock," "happy dance," "intense hip hop" or a reference melody to have a new music track output within seconds.

However, the biggest highlight of Adobe's take is that users can adjust the track for elements such as intensity, tempo, structure, repeating patterns, and more, directly within their workflow, as seen in the video below.

Additionally, users can ask the platform to loop the audio, fade it out, or extend a track, which is especially ideal for video creators who often struggle to find a track that meets the length and needs of their video.

"One of the exciting things about these new tools is that they aren't just about generating audio — they're taking it to the level of Photoshop by giving creatives the same kind of deep control to shape, tweak, and edit their audio," says Nicholas Bryan, senior research scientist at Adobe Research. "It's a kind of pixel-level control for music,"

Also: I tried Copilot Notebook: Microsoft's new AI tool offers two handy prompt features

Adobe Research is still developing Project Music GenAI Control in collaboration with colleagues at the University of California, San Diego, and as a result, the tool is not generally available to the public yet. However, if you want to try a music generation tool for yourself there are some tools you can test now, such as MusicFX by Google and Stable Audio by Stability AI.

Artificial Intelligence

Windows 11’s big new update is full of AI and rolling out today — here’s what’s in it

windows-11-gettyimages-1237969343

Microsoft's AI steamroller continues to barrel on. The latest development, rolling out via Windows Update starting today, is a set of changes to Copilot in Windows 11, the ChatGPT-based tool that gets its own dedicated key on the keyboards of new Windows PCs shipping this year. (As the company notes, Copilot for Windows is still in preview and is available only in select markets.)

The biggest introduction is the addition of plugins that can tie third-party apps and services into the feature. According to a blog post announcing this capability, the first wave of plugins includes OpenTable (so you can ask Copilot to make dinner reservations) and Instacart (to help you put together a grocery list). Other plugins due to arrive "over the next month" include those from Shopify, Klarna, and Kayak.

Also: 11 Windows touchpad tricks to help you work faster and smarter

The common thread for those plugins highlights their real purpose: Each one allows Microsoft to monetize the activity of Windows customers using the new AI tools by taking a cut of the transactions through the partner apps. Those multi-billion-dollar investments aren't going to pay for themselves, after all.

Another update that's due to arrive in about a month — "late March" is the vague timeframe in the announcement — is the addition of new Windows skills for Copilot. On a Windows 11 PC, you will be able to type a prompt and have Copilot perform the following actions in Settings: turn battery saver on or off; open the Device Information, System Information, or Battery information pages; or open the Storage page.

For Windows 11 users who need accessibility assistance, Copilot will be able to launch an array of options, including Live Captions, Narrator, and Screen Magnifier. You'll also be able to turn on Voice Access and Voice Input, adjust text size, and open the Contrast Themes page.

If you just want some information about the device running Windows, you can ask Copilot to show available Wi-Fi networks, display the current IP address, show available storage space, or empty the Recycle Bin.

Of course, typing a prompt to do any of those things arguably saves little time compared to simply opening the app and doing those things with the graphical interface. Consider this, for now, mostly a proof of concept. The real payoff comes down the line when Windows can respond reliably to voice input and display any system information or adjust any setting.

Other AI features are showing up in apps that are included with Windows 11. In the Photos app, a new Generative Erase feature matches the capability that Google Photos has been doing for a while on Android devices: erasing distracting elements (like people walking in the background of a beach sunset) from photos.

Also: Here's the best way to upgrade to Windows 11 Pro, and why you should

Meanwhile, the Clipchamp video editing app gets a Silence Removal feature that zaps awkward pauses from the soundtrack of a recorded video.

Both of those features are scheduled to roll out starting today.

Not all the new features in this update are AI-related. Probably the most welcome change, at least in this reviewer's opinion, is the long-awaited ability to turn off the clickbait news headlines from the Widgets board and use it exclusively for useful, at-a-glance information, such as weather, calendar items, and to-do lists.

Also: I turned my laptop into a desktop PC and I've never been more productive

Once this update is installed, you'll also find improvements to accessibility features, including the ability to create custom Voice Shortcuts that can automate repetitive tasks, along with the ability to use voice commands to work with multiple displays.

On PCs that are connected to Android phones, you'll be able to use the phone's front-facing camera as a webcam on video conferencing apps.

For enterprise customers, the biggest news is a major change in update management. The Windows Update for Business deployment service will merge with Windows Autopatch into a single update management solution for Windows PCs, Microsoft 365 applications, Microsoft Edge, and Teams. Naturally, it will offer AI support.

Also: Microsoft Copilot vs. Copilot Pro: Is the subscription fee worth it?

According to Microsoft, these "new experiences" will begin rolling out today via Windows Update and the Microsoft Store, with the features themselves being turned on in a controlled release for consumers. The company says it anticipates "broad availability" by April's Patch Tuesday release.

Windows 10

Telstra Among APAC Telcos Seeking Innovation, Resilience and Growth in 2024

Telecommunications companies in the Asia-Pacific region have entered 2024 facing cost pressures amid limited returns from their 5G investments to date. S&P Global Ratings forecasts that, while 5G adoption in many APAC markets remains too low to boost revenues, price rises such as labour and electricity are squeezing margins.

Cost pressures a prime driver for AI and product innovation

Profile photo of Jayanth Nagarajan.
Jayanth Nagarajan, Head of Telecom Industry for Asia Pacific & Japan, Amazon Web Services

Telecommunications firm Telstra in Australia and other telcos across the region are responding with efforts to innovate and grow their revenue. According to Amazon Web Services’ Head of Telecommunications in Asia-Pacific and Japan Jayanth Nagarajan, key themes that will dominate agendas for regional telco executives in 2024 include:

  • Leveraging data, generative AI and machine learning at scale.
  • Building resilient, cost-effective and robust core networks.
  • Growing business revenue through new products and services.

“Telcos are nationally mandated critical infrastructure in many countries, so deploying resilient and cost effective 5G networks is a given,” Nagarajan said. “But the corollary to that is cost pressures are real across the industry, and there’s a focus on looking for ways to build more resilient solutions while reducing total cost of ownership.” (Figure A)

Line graph showing telecommunications companies in APAC are battling to grow revenues amid rising cost pressures.
Figure A: Telecommunications companies in APAC are battling to grow revenues amid rising cost pressures. Image: S&P Global Ratings

AI top of mind for Asia-Pacific telecommunications companies

Artificial intelligence featured heavily at the Mobile World Congress 2024. While early use cases for telcos have focused on enhancing customer care and employee productivity, Nagarajan observes cost pressures mean telcos are now looking at AI as more than a search engine or chatbot, with use cases including network optimisation.

Customer service, employee productivity and networks are key use cases

A large number of telcos identified immediate use cases in the areas of customer service and employee productivity. Just some of those, according to AWS, have included forecasting the future needs of customers, suggesting retention offers and personalising customer recommendations to create a connected customer journey.

“We are also seeing a lot of employee productivity use cases like automating RFP responses and providing training and support,” Nagarajan explained, including giving engineers access to enterprise knowledge bases while in the field.

However Nagarajan said telcos are now also increasingly using AI to help optimise and automate networks, including smart planning, deployment and troubleshooting of their networks.

SEE: How telcos in Australia and New Zealand are embracing use cases for AI.

“We are starting to see the use of AI used in areas like the radio network and the SMO layer to drive finer grained control over radio towers and how they radiate,” said Nagarajan. “This makes them far more energy efficient. Telcos are extremely energy hungry, so any reduction to energy utilization is something they can take back to the CFO and the business.”

Telstra sees AI becoming an integrated part of network software

Leading Australian telco Telstra has been using AI “for a heck of a long time,” according to Telstra’s Network Technology Development & Innovation Executive Channa Seneviratne. Telstra’s CEO Vicki Brady recently revealed that AI was already being used to improve half of the telco’s key processes. This included automatic detection and resolution of fixed service faults, in addition to solving customer issues faster.

Seneviratne said that, while the deployment of AI was important from a customer service point of view, it would also be increasingly important for networks.

“What you’ll see as we go to 5G, to 5G-Advanced and then 6G is that AI will be more embedded in the network platform,” Seneviratne said. “So offline AI is there, but what you will see in the future is more integration of AI and machine learning models into software.”

Telcos building resilient, cost effective and robust core networks

The resilience and cost effectiveness of telecommunications core networks has become critical. For example, Telstra is using AI in its effort to make its network more automated, efficient, resilient and secure, but this means the need for GPUs is rising. Telstra’s Network Applications and Cloud Executive Shailin Sehgal said Telstra has needed to make enough room in its power budget for this to be efficient and sustainable.

SEE: IBM Consulting seeing evidence of innovation in Australia AI deployments.

In February 2024, Telstra announced a deployment of Ericsson’s bare metal cloud environment, also known as Cloud Native Infrastructure Solution, which can boost 5G core efficiency with high-density microprocessors from Advanced Micro Devices. With closed-loop liquid cooling, energy consumption is reduced by 49%, described as “a major technological leap” for its private cloud infrastructure.

Telecommunications firms making greater use of cloud services

Telecommunications companies have been escalating their use of the cloud to support network resilience and reduce cost, according to Nagarajan. AWS has witnessed more BSS and OSS applications moving at scale, as well as more network workloads. There was also an expansion in new use cases, including using the elasticity of the cloud for resilience enhancement and disaster recovery among telcos.

Landmark Telstra trial uses AWS to boost network resilience in a crisis

Telstra recently announced a hybrid network architecture resiliency trial that would ensure continuity of voice calls during unplanned network interruptions, where the primary solution is not working correctly. Using Nokia’s IP Multimedia Subsystem software in a collaboration with cloud provider AWS, Telstra said the trial showed the potential for on-demand capacity in an unplanned network disruption scenario.

“When deployed at scale this will provide enhanced network resiliency by allowing call control to be supported by the cloud, rather than experiencing a network disruption,” Telstra said in a statement.

Telstra is working with AWS and Nokia on expanding the trial beyond Voice over LTE on IMS, with hopes it could include the full scale of network provided services, including data services like Eftpos and text messaging.

Growing business revenue through new products and services

Telecommunications companies are looking for new ways to boost revenue.

“They’re looking to innovate quicker, and they’re looking to build in-house capabilities to develop and introduce new services for their customers,” Nagarajan explained. “But we all know the challenges that the industry is facing; legacy applications, traditional long-standing workforce development processes, being steeped in legacy regulatory environments.”

Monetising 5G a priority in the Asia-Pacific region

Much of the focus is on monetising investments in 5G networks.

“For the past few years I think we’ve heard the promise of added revenue from 5G, but we’ve yet to see that windfall which is the reality check,” Nagarajan commented.

Telstra, for one, has seen its 5G coverage increase to 87% of the Australian population as of the beginning of 2024, with 48% of mobile traffic now on 5G, though not all of this is stand-alone 5G.

Nagarajan argues 2023 could have marked a tipping point for 5G, with global and Asia Pacific 5G coverage growing, devices starting to catch up, and use cases for 5G technologies growing in sophistication.

“The last barrier to realising 5G potential is the community across industry and the cross-functional partnerships that are needed to create 5G services and to reduce the barriers to accelerating innovation,” Nagarajan said.

Telecommunications companies look to APIs to boost revenue

Nagarajan said one avenue for growth has been through operators exposing network capabilities through APIs. This is enabling telcos to leverage the developer community to bring services together from multiple parties to deliver an end customer outcome.

Telstra to offer 5G slicing product for enterprise customers

Telstra announced in February 2024 it would bring a 5G slicing product to market in 2024, which could offer guaranteed 5G performance to paying enterprise customers. Following the development of a new proof-of-value engine — which is able to analyse 5G traffic every second to check if a customer has received the committed performance — Telstra has said it will be able to create multiple virtual networks with different performance characteristics, so enterprises can access customised 5G solutions.

Morph Studio turns your Stability AI-generated video clips into full films — here’s how to try it

Morph Studio

The next era of AI involves video generation, with giants in the space such as Google and OpenAI recently unveiling new text-to-video models. Now Stability AI and Morph are joining the party, with a unique take on a video-generating platform.

On Wednesday, Morph announced a collaboration with Stability AI in developing Morph Studio, an all-in-one video generation solution meant to help take users from an idea into an entire film by leveraging generative AI.

Also: The best AI image generators of 2024: Tested and reviewed

The platform combines text-to-video, text-to-image, and video-to-video generation in one place, which users can access via a storyboard format to create and edit shots together. Then they can connect ideas, try different versions, and change the styles, as seen in the video below.

Once the video is completed to the user's liking, they can export it to any post-production software such as Adobe Premiere Pro, Davinci Resolve, and Avid Media Composer, to complete additional edits.

As with DALL-E or Midjourney, Morph Studio users will be able to upload their workflows to the gallery to share their work within the community and even have others use the work in their own projects.

Also: Photoshop for music? Adobe unveils new AI tools for audio

What makes Morph Studio noteworthy? The platform takes the video generation process one step further by helping users create stories from the generated clips and even directly exporting them into video editing platforms to enable complete film projects. Morph Studio has not been released to the public yet; however, if interested, you can join the waitlist here simply by adding your email address.

Artificial Intelligence

Vimeo’s new AI hub vows to organize your team’s videos every which way

Vimeo Central

Vimeo was quick to adopt generative AI into its platform. In June 2023, the video company added a suite of generative AI tools for optimizing video creation. Now, Vimeo is turning its AI efforts toward working professionals with a new video hub.

On Thursday, Vimeo unveiled Vimeo Central, an AI-powered hub designed for business teams to tap into the full potential of video communications leveraging generative AI tools.

Also: YouTube Create expands to more countries — here's why you should try it

"In today's distributed work era, effective team collaboration and employee engagement carries new urgency," said Vimeo's interim CEO Adam Gross. "With Vimeo Central, companies of all sizes can become video-first organizations."

According to Vimeo, the video hub works in tandem with a customer's various video platforms — including Zoom, Webex, Google Drive, Dropbox, and Box — to support the uploading of all content into one video library, a secure centralized hub.

Not only does this make it easier to locate any video needed, but it also allows teams to interact more productively with all video-related information by leveraging advanced tools that can search for not only video titles but also video transcripts.

Vimeo Central also features AI tools to help users get up-to-date with the information they miss. For example, users can condense hours of video into a short highlight reel, bypassing the need to watch the entire video themselves.

The AI tools can also generate text summaries, add chapters to videos, generate titles and FAQs, and even answer questions with the exact moment in the recording that users should watch to find their answers.

Also: Beyond programming: AI spawns a new generation of job roles

Vimeo Central also features a recording studio with a built-in teleprompter, analytics to track team-watching data, including information about who on the team watched what, and a new editor for message polishing.

Even though many video conferencing platforms such as Zoom and Otter.ai have implemented similar features, what makes Vimeo's offering stand out is that it incorporates video from all different platforms in one place, making it a streamlined hub. Vimeo Central is an enterprise-only package, but some of its capabilities, including AI, are offered as self-serve consumer plans.

Artificial Intelligence

Windows 11 Update Brings New Tricks to Microsoft Copilot

Microsoft has announced a slew of new features and capabilities for Windows 11, including new skills for Microsoft Copilot and an AI-powered “generative erase” tool for photo editing, as well as various features for users targeting patching, security and accessibility.

Many of the new features will begin rolling out today via Windows Update and as new apps available on the Microsoft Store. Most features will be enabled by default in the March 2024 optional non-security preview release for all editions of Windows 11 versions 23H2 and 22H2, according to Microsoft.

New features announced for Copilot Preview for Windows 11

As part of the Copilot Preview for Windows 11, Microsoft’s generative AI tool is receiving new integrations with popular apps as well as accessibility features to help users navigate Windows 11 via prompts and voice commands.

For example, starting in late March 2024, Windows 11 users will be able to change system settings through prompts typed directly into Copilot in Windows, currently accessible in the Copilot Preview via an icon on the taskbar, or by pressing Windows + C.

Microsoft Copilot will be able to perform the following actions:

  • Turn on/off battery saver.
  • Show device information.
  • Show system information.
  • Show battery information.
  • Open storage page.
  • Launch Live Captions.
  • Launch Narrator.
  • Launch Screen Magnifier.
  • Open Voice Access page.
  • Open Text size page.
  • Open contrast themes page.
  • Launch Voice input.
  • Show available Wi-Fi network.
  • Display IP Address.
  • Show Available Storage.

The new third-party app integrations for Copilot will give Windows 11 users new ways to interact with various applications. For example, making business lunch reservations through OpenTable.

In a blog post, Yusuf Mehdi, executive vice president and consumer chief marketing officer at Microsoft, said the company would be “adding new ways to connect and get things done” via integrations with apps including Shopify, Klarna and Kayak over the next month.

AI-powered Generative Erase features for photos and videos

Other new AI features for Windows 11 rolling out today include a new, AI-powered Generative Erase tool, which sounds reminiscent of Google’s Magic Eraser tool for Google Photos. Generative Erase allows users to remove unwanted objects or artifacts from their photos in the Photos app.

Likewise, Microsoft’s video editing tool Clipchamp is receiving a Silence Removal tool, which functions much as the name implies ­— it allows users to remove gaps in conversation or audio from a video clip.

Voice controls get accessibility enhancements in this new Windows 11 update

Voice access is another focal point of Microsoft’s latest Windows 11 update, detailed in a separate blog post by Windows Commercial Product Marketing Manager Harjit Dhaliwal.

Users can now use voice controls to navigate between multiple displays, aided by number and grid overlays that provide easy switching between screens. Voice access now supports multiple languages, including French, German and Spanish, and users can create custom voice shortcuts for frequently-used commands.

Also, improvements to Microsoft’s built-in screen reader tool, Narrator, will make it better at detecting handwriting and text inside of images, said Microsoft. It will also be capable of detecting and announcing bookmarks in documents and comments in Microsoft Word.

Users will be able to preview different narrator voices before choosing to download them, with a new keyboard command enabling users to navigate between images on their screen more easily.

Windows Autopatch to offer “unified enterprise update management”

For business users, Microsoft has overhauled update management for Windows 11 to offer a more streamlined experience that gives users more granular control over apps and devices.

Significantly, Windows Autopatch — Microsoft’s update service — is being integrated with the Windows Update for Business deployment service, which Microsoft said was being done in response to customer feedback and to “make the update ecosystem easier to understand.”

DOWNLOAD: This Power Checklist for Managing and Troubleshooting Windows User Accounts from TechRepublic Premium

According to Mehdi, Windows Autopatch will now “leverage AI to program the necessary updates and reduce the impact on team productivity.” Other Autopatch updates include:

  • The ability to import Update rings for Windows 10 and later (preview).
  • Customer defined service outcomes (preview).
  • Improved data refresh speed and reporting accuracy.

“Autopatch will now become the unifying Windows update management solution, providing a single way in which organizations can manage updates while maintaining the highest level of control,” Medhi added.

More Windows 11 updates: Sharing, Casting and Cloud PC

Rounding out the updates for Windows 11 include:

  • More options for sharing content to nearby Windows devices, as well as faster transfer speeds.
  • Improved detection and suggestions for Casting to nearby devices.
  • Passwordless sign in for Cloud PC via a dedicated mode for Windows 365 Boot and fast account switching via Windows 365 Switch.
  • A hover function for Snap layouts that makes it easier to view and switch between layout options.

How and when can I access the new Windows 11 features?

The new features and capabilities for Windows 11 will be rolling out from today until April 2024, at which point all eligible users should have access to them. Most features will be available via Windows Update, and new apps will be available via Microsoft Store updates.

“Windows 11 devices will get new functionality at different times, as we will be gradually rolling out some of these new features over the coming weeks initially via controlled feature rollout (CFR) to consumers,” said Medhi.

He continued: “We anticipate broad availability for most new features by the April 2024 security update release for all eligible devices. Most of these new Windows 11 features will be enabled by default in the March 2024 optional non-security preview release for all editions of Windows 11, versions 23H2 and 22H2. IT admins who want to get the new Windows 11 features can use the enable and control optional updates policy to enable optional updates for their managed devices.”

To ensure they get the updates when they become available, users with eligible devices running Windows 11 versions 22H2 and 23H2 can go to Settings | Windows Update and turn on Get The Latest Updates As Soon As They’re Available and then select Check For Updates.

A summary of all the new features is available in the Windows Update configuration documentation. Users can also stay up to date on rollout plans via the Windows release health dashboard.

Meet Copilot for Finance, Microsoft’s latest AI chatbot — here’s how to preview it

financeabstractgettyimages-1295799789

As part of its ongoing campaign to integrate generative artificial intelligence into everything, Microsoft on Thursday announced the public preview of Copilot for Finance, a tool meant to help those in the finance office perform tasks such as analyzing variance in sales, speeding up collections, and tracking down missing invoices.

"Sixty-two percent of finance professionals say they are stuck in the drudgery of data entry and review cycles," noted Charles Lamanna, Microsoft chief vice president of business applications and platforms, in a blog post announcing the software. "Copilot for Finance can help free up time for finance to play more of a strategic role in delivering counsel and insights to the business by streamlining financial tasks, automating workflows, and providing insights in the flow of work."

Also: Microsoft launches two new Copilots, adding AI-guidance for Service and Sales

Copilot for Finance, noted the company, handles things such as variance analysis in sales. The program pops up a natural-language summary of where actual sales may fall short of projected sales, with an indication of the reason for the variance.

The program will "quickly conduct a variance analysis in Excel using natural language prompts to review data sets for anomalies, risks, and unmatched values," according to the blog post. "This type of analysis helps finance provide strategic insights to business leaders about where it is meeting, exceeding, or falling short of planned financial outcomes and why."

Using suggested prompts — such as "help me understand forecast to actuals variance data" — the program can be made to pull data directly from across ERP and financial system records and analyze the sources.

In a companion blog entry, Microsoft head of business applications marketing Emily He describes the program's capabilities: "[Copilot for Finance can] streamline audits by pulling and reconciling data with a simple prompt, revolutionize collections by automating communication and payment plans, and accelerate financial reporting by detecting variances with ease."

For example, a company's accounts receivable specialist doesn't need to first pull data from an ERP system and then put it into Excel, writes He, they can simply ask at the prompt and the Copilot pulls that data. That can speed up an audit of accounts receivable to find missing amounts.

The "potential time and cost savings are substantial" from such operations, writes He.

Also: Microsoft's GitHub Copilot pursues the absolute 'time to value' of AI in programming

The Finance program is the latest installment in Microsoft's proliferation of generative AI in the past year, including its Microsoft 365 Copilot for the productivity suite and Dynamics 365 Copilot for enterprise resource planning functions. The Finance Copilot is embedded in Microsoft 365 so that it can work with the productivity apps in the suite.

Microsoft is offering a Copilot for Finance demo here, and you can sign up for the public preview here.

In November, Microsoft unveiled two more Copilots, Copilot for Service and Copilot for Sales.

Microsoft has not disclosed the pricing of Copilot for Finance. "We will share the pricing for Copilot for Finance at GA later this year," Microsoft told ZDNET in an email.

Artificial Intelligence

Gemma: Google Bringing Advanced AI Capabilities through Open Source

Google Open Source LLM Gemma

The field of artificial intelligence (AI) has seen immense progress in recent years, largely driven by advances in deep learning and natural language processing (NLP). At the forefront of these advances are large language models (LLMs) – AI systems trained on massive amounts of text data that can generate human-like text and engage in conversational tasks.

LLMs like Google's PaLM, Anthropic's Claude, and DeepMind's Gopher have demonstrated remarkable capabilities, from coding to common sense reasoning. However, most of these models have not been openly released, limiting their access for research, development, and beneficial applications.

This changed with the recent open sourcing of Gemma – a family of LLMs from Google's DeepMind based on their powerful proprietary Gemini models. In this blog post, we'll dive into Gemma, analyzing its architecture, training process, performance, and responsible release.

Overview of Gemma

In February 2023, DeepMind open sourced two sizes of Gemma models – a 2 billion parameter version optimized for on-device deployment, and a larger 7 billion parameter version designed for GPU/TPU usage.

Gemma leverages a similar transformer-based architecture and training methodology to DeepMind's leading Gemini models. It was trained on up to 6 trillion tokens of text from web documents, math, and code.

DeepMind released both raw pretrained checkpoints of Gemma, as well as versions fine-tuned with supervised learning and human feedback for enhanced capabilities in areas like dialogue, instruction following, and coding.

Getting Started with Gemma

Gemma's open release makes its advanced AI capabilities accessible to developers, researchers, and enthusiasts. Here's a quick guide to getting started:

Platform Agnostic Deployment

A key strength of Gemma is its flexibility – you can run it on CPUs, GPUs, or TPUs. For CPU, leverage TensorFlow Lite or HuggingFace Transformers. For accelerated performance on GPU/TPU, use TensorFlow. Cloud services like Google Cloud's Vertex AI also provide seamless scaling.

Access Pre-trained Models

Gemma comes in different pre-trained variants depending on your needs. The 2B and 7B models offer strong generative abilities out-of-the-box. For custom fine-tuning, the 2B-FT and 7B-FT models are ideal starting points.

Build Exciting Applications

You can build a diverse range of applications with Gemma, like story generation, language translation, question answering, and creative content production. The key is leveraging Gemma's strengths through fine-tuning on your own datasets.

Architecture

Gemma utilizes a decoder-only transformer architecture, building on advances like multi-query attention and rotary positional embeddings:

  • Transformers: Introduced in 2017, the transformer architecture based solely on attention mechanisms has become ubiquitous in NLP. Gemma inherits the transformer's ability to model long-range dependencies in text.
  • Decoder-only: Gemma only uses a transformer decoder stack, unlike encoder-decoder models like BART or T5. This provides strong generative capabilities for tasks like text generation.
  • Multi-query attention: Gemma employs multi-query attention in its larger model, allowing each attention head to process multiple queries in parallel for faster inference.
  • Rotary positional embeddings: Gemma represents positional information using rotary embeddings instead of absolute position encodings. This technique reduces model size while retaining position information.

The use of techniques like multi-query attention and rotary positional embeddings enable Gemma models to reach an optimal tradeoff between performance, inference speed, and model size.

Data and Training Process

Gemma was trained on up to 6 trillion tokens of text data, primarily in English. This included web documents, mathematical text, and source code. DeepMind invested significant efforts into data filtering, removing toxic or harmful content using classifiers and heuristics.

Training was performed using Google's TPUv5 infrastructure, with up to 4096 TPUs used to train Gemma-7B. Efficient model and data parallelism techniques enabled training the massive models with commodity hardware.

Staged training was utilized, continuously adjusting the data distribution to focus on high-quality, relevant text. The final fine-tuning stages used a mixture of human-generated and synthetic instruction-following examples to enhance capabilities.

Model Performance

DeepMind rigorously evaluated Gemma models on a broad set of over 25 benchmarks spanning question answering, reasoning, mathematics, coding, common sense, and dialogue capabilities.

Gemma achieves state-of-the-art results compared to similarly sized open source models across the majority of benchmarks. Some highlights:

  • Mathematics: Gemma excels on mathematical reasoning tests like GSM8K and MATH, outperforming models like Codex and Anthropic's Claude by over 10 points.
  • Coding: Gemma matches or exceeds the performance of Codex on programming benchmarks like MBPP, despite not being specifically trained on code.
  • Dialogue: Gemma demonstrates strong conversational ability with 51.7% win rate over Anthropic's Mistral-7B on human preference tests.
  • Reasoning: On tasks requiring inference like ARC and Winogrande, Gemma outperforms other 7B models by 5-10 points.

Gemma's versatility across disciplines demonstrates its strong general intelligence capabilities. While gaps to human-level performance remain, Gemma represents a leap forward in open source NLP.

Safety and Responsibility

Releasing open source weights of large models introduces challenges around intentional misuse and inherent model biases. DeepMind took steps to mitigate risks:

  • Data filtering: Potentially toxic, illegal, or biased text was removed from the training data using classifiers and heuristics.
  • Evaluations: Gemma was tested on 30+ benchmarks curated to assess safety, fairness, and robustness. It matched or exceeded other models.
  • Fine-tuning: Model fine-tuning focused on improving safety capabilities like information filtering and appropriate hedging/refusal behaviors.
  • Terms of use: Usage terms prohibit offensive, illegal, or unethical applications of Gemma models. However, enforcement remains challenging.
  • Model cards: Cards detailing model capabilities, limitations, and biases were released to promote transparency.

While risks from open sourcing exist, DeepMind determined Gemma's release provides net societal benefits based on its safety profile and enablement of research. However, vigilant monitoring of potential harms will remain critical.

Enabling the Next Wave of AI Innovation

Releasing Gemma as an open source model family stands to unlock progress across the AI community:

  • Accessibility: Gemma reduces barriers for organizations to build with cutting-edge NLP, who previously faced high compute/data costs for training their own LLMs.
  • New applications: By open sourcing pretrained and tuned checkpoints, DeepMind enables easier development of beneficial apps in areas like education, science, and accessibility.
  • Customization: Developers can further customize Gemma for industry or domain-specific applications through continued training on proprietary data.
  • Research: Open models like Gemma foster greater transparency and auditing of current NLP systems, illuminating future research directions.
  • Innovation: Availability of strong baseline models like Gemma will accelerate progress on areas like bias mitigation, factuality, and AI safety.

By providing Gemma's capabilities to all through open sourcing, DeepMind hopes to spur responsible development of AI for social good.

The Road Ahead

With each leap in AI, we inch closer towards models that rival or exceed human intelligence across all domains. Systems like Gemma underscore how rapid advances in self-supervised models are unlocking increasingly advanced cognitive capabilities.

However, work remains to improve reliability, interpretability, and controllability of AI – areas where human intelligence still reigns supreme. Domains like mathematics highlight these persistent gaps, with Gemma scoring 64% on MMLU compared to estimated 89% human performance.

Closing these gaps while ensuring the safety and ethics of ever-more-capable AI systems will be the central challenges in the years ahead. Striking the right balance between openness and caution will be critical, as DeepMind aims to democratize access to benefits of AI while managing emerging risks.

Initiatives to promote AI safety – like Dario Amodei's ANC, DeepMind's Ethics & Society team, and Anthropic's Constitutional AI – signal growing recognition of this need for nuance. Meaningful progress will require open, evidence-based dialogue between researchers, developers, policymakers and the public.

If navigated responsibly, Gemma represents not the summit of AI, but a basecamp for the next generation of AI researchers following in DeepMind's footsteps towards fair, beneficial artificial general intelligence.

Conclusion

DeepMind's release of Gemma models signifies a new era for open source AI – one that transcends narrow benchmarks into generalized intelligence capabilities. Tested extensively for safety and broadly accessible, Gemma sets a new standard for responsible open sourcing in AI.

Driven by a competitive spirit tempered with cooperative values, sharing breakthroughs like Gemma raises all boats in the AI ecosystem. The entire community now has access to a versatile LLM family to drive or support their initiatives.

While risks remain, DeepMind's technical and ethical diligence provides confidence that Gemma's benefits outweigh its potential harms. As AI capabilities grow ever more advanced, maintaining this nuance between openness and caution will be critical.

Gemma takes us one step closer to AI that benefits all of humanity. But many grand challenges still await along the path to benevolent artificial general intelligence. If AI researchers, developers and society at large can maintain collaborative progress, Gemma may one day be seen as a historic basecamp, rather than the final summit.

New Book: Statistical Optimization for GenAI and Machine Learning

This book is for participants in my AI and machine learning certification program. However, it is now free and available to everyone. With tutorials, enterprise-grade projects and solutions, it covers state-of-the-art material on topics such as generative adversarial networks (GAN), specialized LLM, data synthetization, as well as classical machine learning. It is a work in progress, as I regularly add new projects and new techniques.

The focus is on fast, simple, and better methods, by comparison to vendor solutions. For instance: NoGAN, better evaluation metrics, xLLM (specialized multi-LLM with taxonomies and self-tuning), variable-length embeddings, generating observations outside the training set range, or fast probabilistic vector search. In classical ML, I discuss simple techniques, for instance to synthesize geo-spatial data. Also, I show how to increase speed, size and quality without neural networks.

This textbook is an invaluable resource to instructors and professors teaching AI, or for corporate training. Also, it is useful to prepare for job interviews or to build a robust portfolio. And for hiring managers, there are plenty of original interview questions. The amount of Python code accompanying the solutions is considerable, using a vast array of libraries as well as home-made implementations showing the inner workings and improving existing black-box algorithms. By itself, this book constitutes a solid introduction to Python and scientific programming. The code is also on my GitHub repository.

snap
The first three steps in the main xLLM project

Contents

  • Machine learning optimization (including NoGAN, automated SQL)
  • Time series and spatial processes
  • Scientific computing (including synthetic music)
  • Generative AI (with explainable AI)
  • Data visualizations and animations
  • NLP and large language models
  • Glossary: GAN and synthetic data
  • Glossary: GenAI and LLMs
  • Introduction to extreme LLM (xLLM) and GPT
  • Bibliography
  • Index

How to Get Your Copy?

The 158-page textbook is free. You can read or download it on GitHub, here. The Python code is also on GitHub, in the same repository. To not miss future updates about the book (I have more projects to add or finish) sign-up to my newsletter, here. Upon signing-up, you will get a code to access member-only content. There is no cost. The same code gives you a 20% discount on all my eBooks in my eStore, here.

Author

Towards Better GenAI: 5 Major Issues, and How to Fix Them

Vincent Granville is a pioneering GenAI scientist and machine learning expert, co-founder of Data Science Central (acquired by a publicly traded company in 2020), Chief AI Scientist at MLTechniques.com and GenAItechLab.com, former VC-funded executive, author (Elsevier) and patent owner — one related to LLM. Vincent’s past corporate experience includes Visa, Wells Fargo, eBay, NBC, Microsoft, and CNET. Follow Vincent on LinkedIn.

Free Site Reliability Engineering Course From Google + Uplimit

Sponsored Content

Free Site Reliability Engineering Course From Google + Uplimit

Google is offering a free two-week course on Site Reliability Engineering in partnership with Uplimit! The course is designed to prepare engineers of all backgrounds for SRE-focused industry roles, and is taught by a veteran Site Reliability Engineer at Google.

The cohort starts on March 11. Spots are limited — claim yours here!

SIGN UP NOW!

More On This Topic

  • Supercharge Your AI Journey! Join Uplimit's Free Building AI…
  • Elevate Your Search Engine Skills with Uplimit's Search with ML Course!
  • Accelerate Your Machine Learning Journey with Uplimit's Metaflow…
  • SAS Analytics Pro – now available for on-site or containerized…
  • Free Data Engineering Course for Beginners
  • What Google Recommends You do Before Taking Their Machine Learning…