Data Science Hiring Process at Zoho

For Indian SaaS unicorn Zoho, AI has grown to be an integral part of its operations. In 2011, it began its journey with data science and now has developed its complete technology infrastructure internally, spanning from data centres to AI software tools. This in-house ingenuity empowers Zoho to effectively address various business obstacles encountered by its clientele and users.

“Our goal is to democratise access to AI for businesses of all sizes, significantly lowering the entry barrier in terms of cost and the volume of data needed for training,” said Ramprakash Ramamoorthy, Director – AI Research at Zoho, in an exclusive interview.

With its headquarters situated in Chennai, the company recently crossed 100 million users across its expansive suite of over 55 business applications, without relying on external investments, becoming the first bootstrapped SaaS firm to achieve this milestone. This achievement comes on the heels of Zoho’s remarkable accomplishment of reaching $1 billion in annual revenue the previous year.

Founded back in 1996 by Sridhar Vembu and Tony Thomas, Zoho has evolved into a worldwide frontrunner in delivering cloud-centric software solutions tailored for both enterprises and individuals. The firm boasts an extensive portfolio of applications and services that encompass diverse facets of corporate functions, such as managing customer relationships (CRM), facilitating email and collaborative tools, enhancing office productivity software, streamlining financial and accounting processes, automating marketing endeavors, and much more.

Currently Hiring

Zoho has over ten job openings in AI and ML across various locations. These roles require a strong knowledge of algorithms, machine learning, and experience with handling large datasets. Responsibilities include collaborating with teams to identify AI and ML applications, developing models from concept to production, data collection and analysis, using ML techniques to solve real-world problems, optimizing models for performance, and staying updated in the AI field.

Both seasoned professionals with a profound understanding of data analytics, data visualisation techniques, and coding skills, as well as recent graduates eager to learn and contribute to challenging projects, are suitable for these positions. The desired qualities in potential candidates include a track record of past achievements, a forward-looking attitude towards accomplishing tasks, and a penchant for collaborative teamwork and ownership.

Inside Zoho’s Data Science Team

At the heart of Zoho’s AI efforts is its adaptable AI research team, which operates within the company’s hybrid AI infrastructure, allowing it to swiftly respond to market demands and promote innovation. The AI research team is a vital part of Zoho’s central research hub, ZLabs, which influences strategic AI, analytics, and other critical areas. Zoho collaborates closely with its central AI team and product-specific AI experts to implement tailored AI solutions.

Initially, AI was used for tasks like finding unusual patterns and understanding sentiments. Today, AI is deeply integrated into Zoho’s technology, powering various features in its products. These AI-driven features include statistical ML, computer vision, natural language processing, and predictive analytics for business users. Importantly, Zoho emphasizes user privacy when developing and using AI models.

AI, ML, and data science are not just for product development at Zoho; they also play a significant role in the company’s IT operations. In IT monitoring, AI steps in to find anomalies, anticipate potential problems, and pinpoint the origins of issues. Zoho harnesses the power of AI not only to scrutinize operational data for more accurate forecasts but also to trim costs. On another front, conversational AI interfaces grant instant access to vital organizational data, making the process of handling tickets a breeze. And Zoho doesn’t stop there; they put AI to work in bolstering security by actively detecting and preventing malware, keeping an eye on user behavior, and spotting phishing threats before they cause havoc.

Riding the Generative AI Wave

Simultaneously, the company is also experimenting with generative AI. Earlier this year, Zoho stated that it’s increasing its investment in R&D and plans to introduce additional generative AI solutions gradually between 2023 and 2024. They explore various avenues for the implementation of generative AI. Within its AI strategy, three fundamental principles—Privacy, Experience, and Value—guide its approach to integrating generative AI technologies.

Zoho has integrated its AI engine, Zia, which operates securely on the Zoho cloud, with OpenAI’s ChatGPT. As a result, a total of 13 generative AI applications and integrations, all powered by ChatGPT, are now available within the Zoho ecosystem, solidifying ChatGPT’s growing influence in the Zoho landscape.

One notable feature, “Ask Zia,” empowers users to inquire about its data, with AI providing responses based on its comprehension of the data. Moreover, “Zia Presentations” can generate statistics and reports using user data from CRM. Its use of generative AI extends to tasks such as automated translation, grammar error detection, and more.

“The collaboration with OpenAI has been pivotal in enabling Zoho customers to embrace generative capabilities by simply adding the OpenAI key to its accounts. This integration ensures that sensitive organizational context remains securely within the confines of its accounts, a crucial aspect from a privacy perspective,”

In June, Vembu said that Zoho is also preparing to build its proprietary LLM, with the company’s devoted R&D unit, where data will always be retained within the Zoho ecosystem.

Tech Stack

Zoho relies on its own suite of products and internal tools to enhance customer performance. Initially, Java was the primary programming language, favored for its compatibility with operational procedures and Zoho’s unique infrastructure, which uses co-located datacenters rather than public cloud providers like AWS or Azure.

Over time, Java remained the choice for statistical machine learning models, while Python gained popularity for neural networks due to open-source frameworks. Zoho also maintains native models in C for on-premise and mobile integration. They use their proprietary language, Deluge, across over 55 applications, enabling tailored solutions for diverse customer needs.

Hiring Process

The hiring process for data science roles at the company focuses on technical skills, problem-solving abilities, and alignment with company values. For entry-level positions, it includes a programming and aptitude test, multiple rounds of programming interviews, and personal interviews to evaluate communication skills and cultural fit. Lateral hires undergo a customized process based on experience, involving programming assessments, discussions on data science concepts, and presentations on past challenges.

Key skills valued by the company include a strong programming foundation in Java, Python, and C.

Upon joining the data science team, candidates can expect a positive and open work environment that encourages ownership, personal development, and creativity. Zoho expects team members to bring their best skills, commitment, and passion, fostering growth, innovation, and collaboration through experimentation and learning from mistakes.

Tips To Crack The Interview

Prospective data science candidates should approach interviews with honesty and preparation, emphasizing a strong foundation in fundamental concepts and programming skills. The company values technical proficiency, personality, values, and potential contributions during the interview process.

Key messages for candidates applying for data science jobs at the company:

  • Deep understanding of fundamental concepts is essential for success.
  • Proficiency in programming languages is crucial.
  • Hands-on practice with real-world data is encouraged.
  • Utilize available digital resources for learning and growth.

The company prioritises privacy as a fundamental right and ethics in innovation. It values qualities like curiosity, creativity, and a commitment to problem-solving in prospective candidates. Those who share these values are encouraged to apply.

Work Culture

Zoho’s work culture is distinct in its rejection of the hustle culture in favor of a supportive and enriching environment. It centers around two core pillars: people and research and development (R&D). The primary focus is on fostering individual growth and contributions, with transparency and an open-door policy ensuring all team members feel valued. This approach encourages teamwork and community spirit.

Furthermore, Zoho’s work culture places a strong emphasis on R&D and technical knowledge. Employees enjoy various perks like 24/7 food availability, comprehensive family health insurance, transportation benefits, and paid time-off. While ESOPs aren’t offered due to private ownership, the company has a profit-sharing plan to reward employees based on company success.

In terms of gender diversity, Zoho takes a non-traditional approach, prioritizing talent alignment with values rather than traditional qualifications. Notably, the Zoho Labs division achieves a balanced 50:50 gender ratio, showcasing the company’s dedication to diversity and inclusion.

What sets Zoho apart from competitors is its emphasis on autonomy, long-term growth, values over metrics, community commitment, privacy, innovation, R&D focus, and openness.

Check out the opportunities now.

Read more: Data Science Hiring Process at Digit Insurance

The post Data Science Hiring Process at Zoho appeared first on Analytics India Magazine.

Singapore arms firefighters with AI glasses

punggol-fire-station-facade

The Punggol Fire Station is Singapore's first smart fire station.

Singapore is arming its firefighters with smart glasses capable of inspecting and identifying defects in their equipment, so the frontliners can be better prepared to respond to emergencies.

The two-year pilot project encompasses the use of 5G and artificial intelligence (AI), including augmented reality (AR) technologies, to automate processes and facilitate real-time remote assistance.

Also: Apple 'smart glasses' could be an iPhone accessory and cheaper than Vision Pro, suggests new patent

The initiative involves several public sector agencies — Infocomm Media Development Authority (IMDA), Home Team Science and Technology Agency (HTX), and Singapore Civil Defence Force (SCDF) — as well as local telco StarHub and tech vendor IBM.

The pilot will roll out at SCDF's smart fire station located in Punggol, on the latest generation of fire emergency vehicles, and test AI-powered visual inspection and AR-enabled remote assistance.

These will involve the use of voice-activated commands and dynamic content displayed on 5G-connected smart glasses to enable firefighters to inspect their equipment and monitor inventory. Data collected from the inspections then will be sent to a centralized dashboard to identify defects and provide status updates of equipment in real-time, so SCDF can be better prepared to respond to emergencies.

With 5G enabling ultra-low latency and high-bandwidth connectivity, the agency also will be able to run AR-enabled remote assistance. This will enable emergency frontliners on the ground to have live access to fire investigation specialists through video interactions and receive help in scene processing and analysis.

IBM's AI-powered computer vision application, Maximo Visual Inspection, automates the inspection process and generates aggregated maintenance data. This is pushed to a central dashboard to equip firefighters with data insights they then can use to make operational decisions.

The project subsequently will be integrated with backend services and ongoing field testing, with IMDA and IBM helping SCDF to fold data governance and cybersecurity processes into the operations.

Also: Meta has a secret VR headset that may have a key advantage over Apple's Vision Pro

"SCDF has been piloting smart technologies to reshape the way frontliners are trained and operations are executed," said Ling Young Ern, SCDF's deputy commissioner for future technology and public safety. "The 5G-enabled fire engine makes secure, real-time transmission of high-quality videos and images achievable."

Noting that Singapore is among the first global markets to have nationwide standalone 5G network coverage, IMDA's assistant chief executive Ong Chen Hui said the government agency has been focused on efforts to commercialize 5G projects. These have included rollouts in verticals such as maritime.

Another 5G pilot involving the use of smart glasses was launched in August 2022 at Singapore's Keppel Offshore & Marine (Keppel O&M), where AR/VR technologies were tapped to facilitate site inspection and remote monitoring.

Apple Launches iCringe with a Sustainability Twist

At Apple’s Wonderlust event last day, besides unveiling a suite of products including the iPhone 15, iPhone 15 Plus, new Apple Watch Ultra 2, and Series 9, among others, the company also introduced a new member – Mother Nature.

Well, standing by their commitment to reduce their entire carbon footprint to zero by 2030, Apple’s watch series 9, released last day, is the first completely carbon-neutral product by the company. Alongside, all the items were linked by one common goal – making the environment greener and reducing carbon footprint.

‘Mother Nature’ Makes Apple Leather Free

In a rather cringe skit featuring actor-comedian Octavia Spencer and also marking the acting debut of chief Tim Cook, the company shares its efforts to minimise its impact on the environment by addressing climate change through renewable energy and energy efficiency, using greener materials, and inventing new ways to conserve precious resources. Their goal is to achieve a net-zero climate impact on their devices by 2030.

The new iPhone cases and Apple Watch bands are not going to be of leather anymore as it has several challenges due to their connection to cattle farming, which can have high environmental costs. This might also be Apple’s aim to challenge the status quo in the luxury market by introducing alternative materials, like FineWoven, and following the lead of other luxury car brands that have embraced non-leather fabrics and sustainability.

The Apple Watch Series Nine, for instance, incorporates 100% recycled aluminium, gold, tin, copper, tungsten, and even recycled cobalt in the battery. They’ve also revamped the sport loop watchband to contain 82% recycled yarn.

Apple’s influence in the market is significant, and its decision to abandon leather is likely to prompt other brands to follow suit. While leather may not disappear overnight, Apple’s move is expected to reduce overall demand for leather, potentially reshaping the premium image of the material it is often associated with, in the long term.

The team, is also eliminating plastic from their packaging and using clean electricity for their operations and has many suppliers committed to using renewable energy. They are also reducing transportation emissions by shipping more products via ocean and investing in global projects to protect the environment, like tree planting in Paraguay and Brazil. These efforts have resulted in a remarkable 78% decrease in the carbon footprint of the Apple Watch Series Nine.

USB C Port: Boon or Bane?

It is true that Apple is making impactful strives towards a greener environment. However, there is a slight hiccup.

Last day, the company also announced that they are transitioning to USB-C ports which has garnered mixed reactions. Even though some perceive it as a positive and more economical step, it’s causing a significant e-waste problem due to people’s lack of awareness about proper charger disposal. A recent survey in the US revealed that 75% of respondents don’t dispose of their chargers correctly, and 55% discard them as general waste, worsening electronic waste issues. This problem is magnified when you consider that 54,000 metric tons of chargers are discarded globally each year.

These statistics emphasise the urgent need for increased awareness and responsible disposal practices. While new regulations mandate USB-C ports on mobile phones, the immediate concern is whether people know how to dispose of old charging cables properly. E-waste is a rapidly growing global concern, projected to increase by 21% between 2019 and 2030, with staggering amounts of chargers discarded annually.

What Others Are Doing

Similar to Apple, both Google and Microsoft share the goal of becoming carbon-negative by 2030. Microsoft has already made progress, reducing its own emissions by 22% and those of its suppliers and customers by 0.5%. They also plan to rely entirely on renewable energy sources by 2025 and have introduced initiatives like the Microsoft Cloud for Sustainability to reduce emissions, improve water access, and reduce waste. Microsoft is actively addressing the environmental impact of generative AI models and data centres by implementing “green software” principles.

Google is also committed to sustainability, aiming for full reliance on carbon-free energy and a net-zero emissions footprint by 2030. They match their electricity usage with renewable energy, expand clean energy efforts to third-party data centres, and use AI for sustainability, such as flood prediction with Flood Hub. Google’s core products like Maps promote eco-friendly routing to reduce carbon emissions. They also focus on sustainable campuses and watershed projects to conserve water and operate their business more sustainably.

Yet, it’s important to highlight that, as of now, Apple stands alone among major tech giants in achieving the milestone of a completely carbon-neutral product.

Read more: Apple Hates AI So Much That It…

The post Apple Launches iCringe with a Sustainability Twist appeared first on Analytics India Magazine.

The ethics of generative AI: How we can harness this powerful technology

Person holding abstract AI interface

With new discoveries about generative AI's capabilities being announced every day, people in various industries are seeking to explore the extent to which AI can propel not only our daily tasks but also bigger, more complex projects.

However, with these discoveries are concerns and questions about how to regulate the use of generative AI. Simultaneously, lawsuits against OpenAI are emerging, and the ethical use of generative AI is a glaring concern.

As updated AI models evolve newer capabilities, legal regulations still lie in a gray area. What we can do now is educate ourselves on the challenges that come with using powerful technology and learn what guardrails are being put in place against the misuse of technology that holds enormous potential.

Use AI to combat AI manipulation

From situations such as lawyers citing false cases that ChatGPT created, to college students using AI chatbots to write their papers, and even AI-generated pictures of Donald Trump being arrested, it is becoming increasingly difficult to distinguish between what is real content, what was created by generative AI, and where the boundary is for using these AI assistants. How can we hold ourselves accountable while we test AI?

Also: A thorny question: Who owns code, images, and narratives generated by AI?

Researchers are studying ways to prevent the abuse of generative AI by developing methods of using it against itself to detect instances of AI manipulation. "The same neural networks that generated the outputs can also identify those signatures, almost the markers of a neural network," said Dr. Sarah Kreps, director and founder of the Cornell Tech Policy Institute.

Also: 7 advanced ChatGPT prompt-writing tips you need to know

One method of identifying such signatures is called "watermarking," by which a kind of "stamp" is placed on outputs that have been created by generative AI such as ChatGPT. This helps to distinguish what content has and hasn't been subjected to AI. Though studies are still underway, this could potentially be a solution to distinguishing between content that has been altered with generative AI and content that is truly one's own.

Dr. Kreps gave the analogy of researchers using this stamping method to that of teachers and professors scanning students' submitted work for plagiarism, where one can "scan a document for these kinds of technical signatures of ChatGPT or GPT model."

Also: Who owns the code? If ChatGPT's AI helps write your app, does it still belong to you?

"OpenAI [is] doing more to think about what kinds of values it encodes into [its] algorithms so that it's not including misinformation or contrary, contentious outputs," Dr. Kreps told ZDNET. This has especially been a concern as OpenAI's first lawsuit was due to ChatGPT's hallucination that created false information about Mark Walters, a radio host.

Digital literacy education

Back when computers were gaining momentum for school use, it was common to take classes like computer lab to learn how to find reliable sources on the internet, make citations, and properly do research for school assignments. Consumers of generative AI can do the same as they did when they were first starting to learn how to use a piece of technology: Educate yourself.

Also: The best AI chatbots

Today, with AI assistants such as Google Smart Compose and Grammarly, using such tools is common if not universal. "I do think that this is going to become so ubiquitous, so 'Grammarly-ed' that people will look back in five years and think, why did we even have these debates?" Dr. Kreps said.

However, until further regulations are put in place, Dr. Kreps says, "Teaching people what to look for I think is part of that digital literacy that would go along with thinking through being a more critical consumer of content."

For instance, it is common that even the most recent AI models create errors or factually incorrect information. "I think these models are better now at not doing these repetitive loops that they used to, but they'll do little factual errors, and they'll do it sort of very credibly," Dr. Kreps said. "They'll make up citations and attribute an article incorrectly to someone, those kinds of things, and I think being aware of that is really helpful. So scrutinizing outputs to think through, 'does this sound right?'"

Also: These are my 5 favorite AI tools for work

AI instruction should start at the most basic level. According to the Artificial Intelligence Index Report 2023, K-12 AI and computer science education grew in both the US and the rest of the world since 2021, reporting that since then, "11 countries, including Belgium, China, and South Korea have officially endorsed and implemented a K-12 AI curriculum."

Time allocated to AI topics in classrooms included algorithms and programming (18%), data literacy (12%), AI technologies (14%), ethics of AI (7%), and more. In a sample curriculum in Austria, UNESCO reported that "they also gain an understanding of ethical dilemmas that are associated with the use of such technologies, and become active participants on these issues."

Beware of biases

Generative AI is capable of creating images based on the text that a user inputs. This has become problematic for AI art generators such as Stable Diffusion, Midjourney, and DALL-E, not only because they're images that artists did not give permission to use, but also because these images are created with clear gender and racial biases.

Also: The best AI art generators

According to the Artificial Intelligence Index Report, The Diffusion Bias Explorer by Hugging Face inputted adjectives with occupations to see what kinds of images Stable Diffusion would output. The stereotypical images that were generated revealed how an occupation is coded with certain adjective descriptors. For example, "CEO" still significantly generated images of men in suits when a variety of adjectives such as "pleasant" or "aggressive" were inputted. DALL-E, also had similar results with "CEO," producing images of older, serious men in suits.

Images from Stable Diffusion of "CEO" with different adjectives.

Images from DALL-E of a "CEO."

Midjourney was shown to have a similar bias. When asked to produce an "influential person," it generated four older white males. However, when given the same prompt by AI Index later on, Midjourney did produce an image of one woman out of its four images. Images of "someone who is intelligent" revealed four generated images of white, older, males who were wearing glasses.

Images from Midjourney of an "influential person."

According to Bloomberg's report on generative AI bias, these text-to-image generators also show a clear racial bias. Over 80% of the images generated by Stable Diffusion with the keyword "inmate," contained people with darker skin. However, according to the Federal Bureau of Prisons, less than half of the US prison population is made up of people of color.

Images from Stable Diffusion of "inmate."

Furthermore, the keyword "fast-food worker" attributed images of people with darker skin tones 70% of the time. In reality, 70% of fast-food workers in the US are white. For the keyword "social worker," 68% of images generated were of people with darker skin tones. In the US, 65% of social workers are white.

Images from Stable Diffusion of "fast-food worker."

What are the ethical questions that experts are posing?

Currently, researchers are exploring hypothetical questions to unmoderated models to test how AI models such as ChatGPT would respond. "What topics should be off-limits for ChatGPT? Should people be able to learn the most effective assassination tactics?" Dr. Kreps posed the types of questions that researchers are examining.

"That's just one sort of fringe example or question but it's one where if [it's] an unmoderated version of the model, you could put that question in or 'how to build an atomic bomb' or these things which maybe you could have done on the internet but now you're getting in one place, a more definitive answer. So they're thinking through those questions and trying to come with a set of values that they would encode into those algorithms" Dr. Kreps said.

Also: 6 harmful ways ChatGPT can be used by bad actors, according to a study

According to the Artificial Intelligence Index Report, between 2012 and 2021, the number of AI incidents and controversies increased 26 times. With more controversies arising due to new AI capabilities, the need to carefully consider what we're inputting into these models is pressing.

More importantly, if these generative AI models are drawing from data that is already available on the internet, such as the statistics from occupations, should they allow the risk of continuing to create misinformation and depict stereotypical images? If the answer is yes, AI could play a detrimental role in reinforcing humans' implicit and explicit biases.

Questions also remain about who owns code and the risk of liability exposure using AI-generated code, as well as the legal implications of using images that AI generates. Dr. Kreps gave an example of the controversy surrounding copyright violation when it comes to asking an art generator to create an image in the style of a specific artist.

"I think that some of those questions are ones that would have been hard to anticipate because it was hard to anticipate just how quickly these technologies would diffuse," Dr. Kreps said.

Whether these questions will finally be answered when the use of AI tools like ChatGPT starts to plateau is still a mystery, but data shows that we may be past ChatGPT's peak, as it experienced its first drop in traffic in June.

The ethics of AI moving forward

Many experts believe the use of AI isn't a new concept and this is evident in our use of AI to perform the simplest tasks. Dr. Kreps also gave examples of using Google Smart Compose when sending an email and checking essays for errors through Grammarly. With the growing presence of generative AI, how can we move forward so that we can coexist with it without becoming consumed by it?

"The people have been working with these models for years and then they come out with ChatGPT, and you have 100 million downloads in a short amount of time," Dr. Kreps said. "With that power [comes] a responsibility to study more systematically some of these questions now that are coming up."

Also: ChatGPT and the new AI are wreaking havoc on cybersecurity in exciting and frightening ways

According to the Artificial Intelligence Index Report, the number of bills with "artificial intelligence" that were passed increased from only one in 2016 to 37 in 2022 among 127 countries. Furthermore, the report also shows that in 81 countries, parliamentary records about AI show that global legislative proceedings having to do with AI have increased almost 6.5 times since 2016.

Though we are witnessing a push for stronger legal regulations, much is still unclear, according to experts and researchers. Dr. Kreps suggests that the "most effective" way of using AI tools is "as an assistant rather than a replacement for humans."

While we await further updates from lawmakers, companies, and teams are taking their own precautions when using AI. For instance, ZDNET has begun to include disclaimers at the end of its explainer pieces that use AI-generated images to show how to use a specific AI tool. OpenAI even has its Bug Bounty program in which the company will pay people to look for ChatGPT bugs.

Regardless of what regulations are eventually implemented and when they are solidified, the responsibility comes down to the human using AI. Rather than fearing generative AI's growing capabilities, it is important to focus on the consequences of our inputs into these models so that we can recognize when AI is being used unethically and act accordingly to combat these attempts.

Artificial Intelligence

4 ways generative AI can stimilate the creator economy

Photo AI generates an image of ZDNET's Maria Diaz.

The Photo AI tool generates an image of ZDNET's Maria Diaz.

Artificial intelligence is here to stay. It's not a fad or a craze — it's a movement. Sure, the buzzwords will become less popular with time so ChatGPT and AI won't dominate the news. But AI tools will become the basis of how we do things in life and work, much like the industrial revolution changed the foundation of our modern existence and the internet did subsequently.

The form of AI you've likely heard the most about is generative AI. It's the one that ChatGPT uses to create a letter, that MidJourney uses to create striking images, and Bard uses to translate documents and summarize bodies of text.

Content creators are one professional group that seems to be embracing the power of AI; from a survey of 1,000 content creators conducted by Lightricks, a video and image editing apps developer, 84% reported they were likely to use AI if it would save them time or money, and 86% said they would use it if positively impact their creative process.

But on the flip side, generative AI is also the same technology that can create deep fakes, which are images and videos that closely resemble the likeness of others to the point of proving hard to determine whether they're real.

Also: Why generative AI so popular: Everything you need to know

Generative AI can also decrease the authenticity of shared content if someone uses it instead of originally-created content. Because it trains on massive amounts of data that multiple creators and authors have already created, it can raise red flags for copyright infringement.

What is the creator economy?

The vast success of social media has resulted in its growth into a full-blown business model. A long way from your Myspace Top 8 and glitter GIFs, we've found a way to monetize and create an economic model from our social media habits.

The creator economy is the socioeconomic system where independent creators monetize their content, directly or indirectly. Content creators, also called influencers, produce and share the material with their audience.

The material could be anything from written content in blogs, emails, newsletters, videos, photography, or a combination of several in one or more social media networks. A content creator can monetize content on YouTube and TikTok and sell it on Instagram, for example.

The audience consumes the creator's content, and, in many cases, the creator can benefit from that consumption — in other cases, the audience member has to take further action to benefit the creator, like purchasing subscriptions or merchandise.

Content creators have many ways to monetize their content, from advertising revenue offered through views, as is common on YouTube, to brand sponsorships, affiliate marketing, merchandising, offering paid subscriptions to exclusive content, and more.

Though the term "influencer" is known for its negative connotation and is probably enough to elicit images of an excessively privileged world traveler in beige lace dresses and flower crowns, it is as accurate as can be to define the role of a content creator: "One who exerts influence: a person who inspires or guides the actions of others," according to Merriam-Webster.

With the understanding of what the creator economy is, let's get into how generative AI can change the game for influencers.

1. AI image creation

If you've consumed any media in the past few years, you've likely seen some AI-generated images, even if you've been unaware of them.

There are many widely available AI art generators that you can go and sign up for as quickly as you can sign up for ChatGPT. Bing, Microsoft's search engine, even has its AI-powered Image Creator that you can use with the same account you use to check Outlook or sign into Xbox, and it's not half bad.

OpenAI has its own DALL-E 2, but many others range in quality, ease of use, and price. The content creator and ZDNET's own David Gewirtz, who you may know from the YouTube channel, Advanced Geekery, detailed how he created images with MidJourney for an Etsy shop.

According to the Lightricks survey, 53% of creators use generative AI to create photo and video backgrounds, making this the most common use of AI among content creators.

Gewirtz tells me using MidJourney along with Adobe Photoshop's new AI-powered tools to create images for his wife's e-commerce company has "proven hugely helpful in providing those images for social media posts and newsletters."

Also: The best AI art generators

The second most common use of generative AI was creating avatar profile pictures, which 46% of content creators reported doing.

Remini AI has recently garnered attention in social media platforms like TikTok for generating headshots. Another example is Photo AI, an AI tool singlehandedly created by Pieter Levels to create AI models based on photos of a person to generate new images.

A user uploads at least 30 photos of themselves into the site to train the model. The system learns the facial patterns from the images and can create a model, which you can name and generate new images following your prompts. This could enable you to create professional headshots without ever having to hire a professional photographer or capture the perfect Instagram influencer aesthetic without even looking at a camera lens.

"Photo AI can help content creators save time and money as they'll no longer need to travel to different locations or hire expensive photographers to do photoshoots," according to Levels. "After creating your AI model, you can take photos of yourself anywhere from your laptop or phone, 24 hours a day, seven days a week."

The AI-generated images from Photo AI (the three at the top) compared to three of the photos of myself I used to train the model.

Testing it out myself, I can see the feature is still in its growing phases, as it's not as accurate as a real camera, but it's still impressive. The most remarkable part of Photo AI to me is that, while the images don't always precisely capture every single feature of a person, the delicate subtleties that make you stand out seep through the photos. It could be a crook under an eye or slight imperfection — but the promise of what could be accomplished is incredibly stunning.

2. Social media content AI tools

A good creator can combine the excellent generative AI tools available and use them as instruments to more easily create social media content, like text for their Instagram posts or even some graphics for their photos.

Other companies have opted to become a one-stop-shop solution for content creators. Microsoft recently launched the Designer app, which uses AI to generate graphics you can edit. To use it, you need to enter a prompt, a description of the design you want, like an Instagram post about a hair product launch, including a photo you uploaded, like a ChatGPT-powered Canva.

Typeface is another multimodal tool that uses generative AI to create content using personalized product shots, social media posts, e-commerce websites, product descriptions, creative briefs, and more. ZDNET spoke to Vishal Sood, founding member and head of product of Typeface, and he explained the app brings the customer's brand and the foundational models together to create content in seconds.

"We let you bring your own assets into the system," Sood said. "You can bring examples of your voice — brand voice, to create part of your output."

Typeface lets users upload their product images and create personalized photos and marketing assets with the help of generative AI powered by OpenAI's GPT-4 and DALL-E, Microsoft Azure AI, Stable Diffusion, and Google Vertex AI.

This goes beyond an AI tool for marketing professionals, extending to the creator economy. According to Lightricks, 56% of content creators report they've been asked to use generative AI by brands they work with.

"With Typeface, we have unlocked the power to generate thousands of personalized, on-brand images, spanning over multiple diverse markets and significantly reducing our production time to amplify our content factory initiative. Typeface provides us the capability to create a unified 'one brand' approach, amplifying our cross-selling opportunities, consistent brand representation across various business units, and the ability to deliver tailored content to each target market," said David Kang, SVP of digital commerce and marketing solutions at LG Electronics.

3. Using AI for video production

From scriptwriting to video editing, AI can accompany a content creator throughout video production, as evidenced by the survey showing most creators use it to generate video and photo backgrounds.

Video creators can streamline their scriptwriting process using ChatGPT or another generative AI text tool to enter a prompt describing the video details they want to make.

That said, there aren't as many widely-available AI video generators yet — at least not ones capable of putting out realistic results to pass as human-created.

For post-production, different editing software programs have found their way of incorporating AI, like Adobe Premiere Pro using Content-Aware Fill and many of the AI tools in CapCut's editing library.

For example, Gewirtz explained, "I recently used Adobe's Generative Fill to repair a filming background. Adobe's podcast audio repair tool fixed a very damaged audio recording, which meant I didn't have to set back up and re-record it."

Also: These 3 AI tools made my two-minute how-to video way more fun and engaging

Among content creators, 71% found that their followers responded positively to their AI-generated content, while only 10% found they reacted negatively. "It's both a force multiplier and terrifying competition," Gewirtz says, adding his video audience seems to like the slightly higher production value tacked on by the AI tools.

"On one hand, it will improve productivity in certain areas (like where it fixed my audio). But it also will open a giant can of worms in terms of ownership, rights, and even whether something can be attributed to human work," according to Gewirtz. "In the near term, I think the biggest downside is that there has been a huge increase in spammy YouTube and social media content produced by AI-powered content farms. That wastes the viewers' time and creates a tougher competitive environment for human creators."

4. Blog automation and other AI writing assistance

Using generative AI to write content is a hot topic as we debate whether it will replace writers' jobs, among many other professions worldwide. In my completely biased opinion, I believe generative AI to be an outstanding instrument for writing, but no more than that. It's a tool, not a crutch. It's a key on your keyboard, not your entire keyboard.

If all the sites use AI to write content, eventually, all the content begins to sound the same, no matter how hard different teams tweak it. Ultimately, we'll end up craving the human voice behind the onscreen text, much like we desire simple answers over Google searches in ChatGPT.

But generative AI is still an excellent tool to keep in your arsenal — I know I keep it in mine to quickly get summaries of long bodies of texts and translate news from other languages.

"Specifically, in writing, I have found that using ChatGPT (more than Bard and Bing) is useful for brainstorming. I will often ask it to discuss a topic or provide me with a list of ideas to play with," says Gewirtz. "Sometimes, I'll dive into those brainstormed ideas with it to further spark my thoughts. But I don't ever use the literal results in my work."

These AI tools can work exceptionally well to summarize text and write blog posts for you. They can also write emails, briefs, job posts with detailed requirements, resumes, cover letters, etc. However, it's still recommended you do a final edit with a human set of eyes.

For content creators specifically, there are many available that can expedite the creation of social media posts and go as far as learning your brand's tone from your past posts and even automatically posting them for you, with minimal interaction on your end.

Tools like Narrato and Lately use AI to generate web copy and social media posts but also follow tone guidelines to ensure your posts follow a consistent voice that sounds true to your brand. They can use AI to generate new blog posts for you and publish them automatically after you give them a prompt and a scheduled date.

It takes less time in your week to schedule these SEO-powered AI posts, but as with all generative AI, I'd say don't set it and forget it. Follow up on it to make sure the system is doing a good job.

And generative AI isn't limited to blog posts. Wix, a user-friendly website creation platform, recently released a generative AI tool to help users generate websites. Giving the WIX AI assistant some prompts in a conversation on a chat window to describe what you want the site to look like, including photos you want on it, the type of website, and how you want the layout, will easily generate a design for you to edit.

Concerns about generative AI

It's easy to imagine how generative AI can become a double-edged sword for content creators.

Content creators are very concerned about the adverse effects of generative AI, as 74% report they're worried about creating deep fakes, which use deep learning algorithms to superimpose someone's face and voice on another person's body over video or a photo.

When you consider using apps and websites to create images of yourself in places you've never been, you must also wonder about the ethical implications. What's stopping someone else from taking photos from your social media accounts and training an AI model to create whatever pictures they want?

Levels, from Photo AI, addressed these concerns for me, explaining that his company's terms and conditions clearly state that users cannot produce content based on another person without their permission and that each user agrees to these terms and conditions when they use the site.

"You can only upload imagery and train a model of yourself or people you know if you have permission from them," Levels added.

Also: The AI boom will amplify social problems if we don't act now, says AI ethicist

Sood, from Typeface, explains that the platform has a built-in plagiarism checker to ensure the content is unique to each customer and customized with each brand's voice. Their models can quickly learn styles to adapt and create outstanding output for each brand.

Among content creators, however, 58% are concerned about copyright issues with generative AI, and 57% are worried about decreased content authenticity due to using it.

How generative AI and copyrighted content will look in the future and the regulations behind it will remain to be seen. Still, different authors, including Sarah Silverman, have sued OpenAI and Meta for copyright infringement.

The future landscape of the creator economy

The proliferation of generative AI has created a big fear of the loss of jobs due to automation. While this may be true in some form, it won't necessarily be in the way most people believe. If you think back to the industrial revolution when many jobs were automated, the change forced many people to adapt and find new trades or learn new machinery — we're at a similar crossroads, albeit a more minor one.

"As an individual creator, AI can save me time in several helpful ways," Gewirtz explains. "But it has the potential to be a much cheaper alternate solution to human-generated work and may well be used as a substitute by clients otherwise hiring creatives for a cheaper if somewhat lower quality solution."

While generative AI can be a time-saving tool to optimize a creator's workflow, it can yield lower-quality results. Many of the available AI services are free or cost a fraction of what an expert sound engineer, video editor, or writer with years of experience and skill would charge for their services. But the output from the generative AI tool could result in generic and low-quality stuff, especially when more people use it, and it all starts looking and sounding similar.

Levels believes, "AI won't replace creators like many scaremongers say. I think it will, however, become an essential part of any creator's toolkit. For example, in the future, we might see people mix elements of real photography and AI photography to create new content."

It's hard to say what the future of the creator economy will look like with things changing as fast as they do in social media. The YouTube and Instagram of 2023 are certainly not the same as just ten years ago. One thing is sure: AI is here to stay. But whether for work or use in our devices and software, we must use AI carefully.

Disclaimer: Using AI-generated images could lead to copyright violations, so people should be cautious if they're using the images for commercial purposes.

See also

Extending ChatGPT: Can AI chatbot plugins really change the game?

AI chat apps are seen on a mobile device.

Plugins have long been a fixture of complex software systems. As far back as the 1980s, I founded a company called Hyperpress that provided plugins for Apple's HyperCard (think of it as a web before there was a Web… or connectivity). My plugins added capabilities to HyperCard that weren't part of the original build.

Today, plugins enhance popular products like Photoshop and WordPress. WordPress offers roughly 60,000 plugins that modify its capabilities.

On the two main websites I operate, I use 49 separate plugins (which add roughly 49 new features) on one site, and 25 plugins (which add roughly 25 new features) on the other site. Neither of these sites would be the sophisticated professional e-commerce sites they are without the wide array of plugins that add features and value.

What are plugins?

Fundamentally, plugins are separate chunks of code that interact with the parent software platform. They do this via an API (application programming interface). All platforms that support plugins provide APIs that allow outside programmers to hook into the functionality of the platforms.

Often the breadth and reliability of the API determine the resilience and flexibility of the overall platform, especially when users rely on a great many plugins to expand the capabilities of the plugin-supporting platform they're using.

Done right, plugins benefit three constituencies: the platform provider (i..e., Adobe for Photoshop, the open source WordPress community, and OpenAI for ChatGPT), the developer of the plugin, and the users of the platform who get new capabilities.

Also: I asked ChatGPT to write a WordPress plugin I needed. It did it in less than 5 minutes

Platform providers often decide to compete with developers. If they see a plugin is particularly popular, they sometimes choose to include that functionality in the core product. They change APIs. Sometimes, if they offer a marketplace (like an app store, but for plugins), they are selective about who to allow into the marketplace and who to promote.

But when the dance between platform provider and plugin developer works, it's magical to behold. The original platform can be taken places nobody predicted, providing capabilities not otherwise possible.

Also: The best ChatGPT plugins (and how to make the most of them)

Plugin capabilities have been announced for Google's Bard, Microsoft's Bing AI, and ChatGPT. However, so far, only ChatGPT offers an extensive array of plugins you can actually put to use.

How plugins currently work in ChatGPT

Plugins for ChatGPT are only available to paying customers of ChatGPT Plus. That's a $20/month service that provides access to the GPT-4 dataset, plugins, and a special plugin called Code Interpreter (more on that later).

For that $20/month, you get to use a very beta, very unfinished product. It's still amazing, but it's also very annoying. You're limited to 25 queries in three hours, so if you're trying to get a job done, you may well run out of queries right in the middle of your work time. Yes, I speak from very frustrated experience. You'll also need to turn them on in Settings.

Now that you've got the plugins enabled, prepare to be impressed.

Plugins that move the needle

I experimented with a lot of plugins. Because you can only use three plugins at one time, you really need to pick and choose a core library of plugins that you use regularly. Here's the list I came up that powered the examples I'm about to show you:

  • Stories: This generates a story book from a prompt. I only used it once (because I don't have kids), but it was so impressive, it's something you need to see.
  • MixerBox WebSearchG: This makes the entire current web available to ChatGPT, and it does it quite well. This truly expands the value of ChatGPT.
  • World News: This scans news sources and provides up-to-date news summaries.
  • AskYourPDF: You can feed ChatGPT a PDF and interact with the data in the PDF document.
  • Code Interpreter: This is a special add-on/plugin provided by OpenAI. If you run this, you can't run any of the other plugins. It allows you to use code to talk to ChatGPT, but it also interprets complex requests and substantially extends the queries you can ask ChatGPT.

Note that I'm not providing URLs to each of these individual plugins. The ChatGPT Plugin Store doesn't offer unique URLs for each plugin. But they're pretty easy to find. Just go to the ChatGPT Plugin Store and do a search for their titles. If you want to know how to enable plugins and access the Plugin Store, ZDNET's Steven Vaughan-Nichols has a great tutorial.

The plugin that's the reason why writers are on strike

Let's start with Stories. In a previous article, I showed you how I got ChatGPT to write a short Star Trek story (and how it mostly succeeded). Stories takes that idea and makes it real.

Within ChatGPT, you can give Stories a prompt that describes the story you want to be told. Here's what I fed it:

Using stories, tell the story of a group of friends who live on a starship (three are human, one is a robot). Tell of their adventure traveling to a planet populated by sentient dinosaurs and where creativity performed by the sentient dinosaurs is illegal, and all stories, entertainment, music, etc is composed by generative AI systems.

Stories then goes on to generate a full story book based on this premise. While the generated graphics were a bit weak (they could benefit from Midjourney-quality tech), the overall production is quite good. I gave the plugin a premise that involved some moral choices, and the AI came up with not only a good little story, but one that held together.

But stories takes it further. You can read the story online in digital form. Or you can go to the Stories site and order a hardcover copy. You can even publish the story on Amazon.

You can see how this sort of thing puts fear into the hearts of professional creatives, particularly those currently on strike. In less than five minutes, I had a fully usable story book. Written and illustrated traditionally, this 50-page story book (with a paragraph on each page) could have taken weeks or months.

I just took a sip of coffee, made up a premise loosely based on a typical Star Trek theme, and fed it to the AI.

When ChatGPT can read the web

As we all know, ChatGPT's knowledge base ends in 2021. But with the MixerBox WebSearchG plugin, we get a taste of what will happen when an AI can read the web. This also produced that "woah" feeling I sometimes get when I know I'm touching the future.

On July 10, 2023, I published a ZDNET article about challenges I was having with my Google Cloud storage enterprise plan. In that article, I coined the term "infraquake" and mentioned it in the two concluding paragraphs of the article.

Then, on July 11, I asked ChatGPT this question:

How does Gewirtz define an "infraquake"

I didn't tell ChatGPT which Gewirtz, nor did I tell it that an article had been published on ZDNET. And yet, it came up with a very clear (and, I might add, accurate) description of the intent behind the newly coined word "infraquake."

Clearly, ChatGPT can now access current data and process it for meaning using the plugin.

Also: How does ChatGPT actually work?

You can also see how ChatGPT's ability to retain context is mixed with the plugin's ability to access the web in this next example. I asked it:

Tell me about David Gewirtz's dog Pixel

Not only did it pull the information, it showed a picture of the little guy, and linked to an article where I wrote about choosing his name.

Understanding current events

While working on this special feature, I found three "killer apps" that I'm going to use regularly as part of my job. Let's discuss the first one first: creating briefings explaining current news with full background information.

In my job, I track a lot of news. I'm often asked by clients to provide perspective on tech news, technology trends, and some geopolitical issues. That means I spend a few hours every day keeping up with my reading, making sure I get a good understanding of what's going on.

But there's no way I can keep up with everything, and I can't really take a lot of time for those topics that are not on my main "beat." Even so, I'd like to have a solid understanding of the issues.

Also: What is generative AI? Here's everything you need to know

One example of this is the Ukraine/Russia war. While I've written about how the attacks impact Ukrainian developers and even covered previous Ukrainian security issues going back almost a decade, I have not been "fully briefed" on the issue of Ukraine's bid to be a NATO member.

I could have dug through a bunch of news articles and done background research, but I don't really have time to allocate to it. Instead, I asked ChatGPT, MixerBox WebSearchG, and World News to prepare me with a comprehensive briefing. I asked two questions:

You are a US policy advisor briefing a CEO on the NATO situation. You've been asked to explain why President Biden says that Ukraine is not ready to join NATO. Use World News and MixerBox WebSearchG to provide a clear briefing on both sides of the argument so that your client has a deep and current understanding of the issues, complexities, and political ramifications. Relate your answer to the US political climate as well.

and

Using the same plugins, is there an opposing view by the GOP for Ukraine NATO membership?

From these two questions, the AI gave me a comprehensive briefing on the NATO membership question, the foundational and political issues behind membership, and the position of both parties with regard to this issue.

My wife subscribes to a service called Blinkist. The company describes its service as "Blinkist offers the key insights from top nonfiction in a made-for-mobile format." It's essentially audible or readable Cliff's Notes for currently popular nonfiction books, and it lets her come up to speed on topics she cares about in about half an hour.

By combining ChatGPT with MixerBox WebSearchG and World News, I can essentially get a Blinkest "key insights" briefing on any currently unfolding world news issue. This is powerful stuff, but as with all press coverage, it's important to be aware that there may be bias, omissions, and inaccuracies in what the AI presents to you.

Using PDFs as source material for analysis

I recently had an analysis project where I had to dig through very long and very dry white papers to try to understand the relationships between some key technologies. Using the AskYourPDF plugin, I fed those PDFs to ChatGPT, and then asked questions related to the content of the PDFs.

It was extraordinary. I was able to ask ChatGPT to analyze various concepts contained in the PDFs. I could even get it to draw a table comparing items discussed in the PDFs, but which had not been compared to each other directly in the source documents.

I would never use ChatGPT as a substitute for reading all the background information on a project I'm tasked with investigating. But much of the analysis I do for my own learning process requires a lot of very tedious clerical work to construct tables and charts to increase my understanding of relationships within documents.

I also used this to examine some contracts. I fed it a contract document we had with a service provider and asked it to show me how the limitations differed between the parties, something normally quite time consuming to extract and determine. Here's the prompt I used:

Create a table comparing the limitations, itemizing each limitation listed. Show only where the limitations differ between the parties, summarize each different limitation in 8 words or less

And here's the table I got back:

Of course, without question, the results from ChatGPT can't be trusted as completely accurate. But a quick scan can definitely save time.

Also: 5 ways to explore the use of generative AI at work

Using ChatGPT and AskYourPDF, clerical research tasks that would normally take me half a day or more were reduced to mere minutes. That's a killer app.

Doing large-scale data analysis

Finally, I want to explore an add-on for ChatGPT from OpenAI that runs on its own. It's called Code Interpreter, and it does so much more than interpret code. Code Interpreter allows you to upload data to it, which ChatGPT can then analyze.

WARNING: Do not use this feature for the first time if you have anything else to do that day. You will get sucked in. It's more disruptive to productivity than kitten and puppy videos.

Ask me how I know. I mean, it's hard to believe anything this fun can really be legal.

This tool allows you to import data files (Excel, CSV, etc) into ChatGPT. It can then perform analysis on it and even generate basic graphics. It is dangerously addictive. Hours after I started, I found myself downloading dataset after dataset from data.gov and laughing maniacally about the power at my fingertips. It was not a pretty sight.

Also: AI gold rush makes basic data security hygiene critical

I think this is, ultimately, why ChatGPT Plus limits you to 25 queries every three hours. It's not to reduce the load on their infrastructure. It's for our own darned good. I sure needed it. I probably wouldn't have eaten all day if I hadn't been forced to step away from the computer by the query limit.

I'll save you any more disturbing analytics visions inside The Mind of David and, instead, show you a simple application: of my email contacts, what big PR firms do I regularly correspond with, and what big tech companies have the most representation. To do this, I exported my contacts from Google Contacts.

Using the email-related fields, list the top 20 domain names represented

Of the top 20 domain names, which are related to PR companies

Then, I had the AI list and draw a pie chart of which tech companies I had the most contact with. Here's what I asked:

Looking at the dataset, create a pie chart showing the relative representation of large multibillion dollar technology companies

And here's what I got back.

The format of the pie chart isn't ideal, but the information is there. And, again, we're talking minutes instead of hours.

But…we're still in the early days

Plugins are available, but they're very new. Some, like the ones I demonstrated above, have big advantages. But because they're so new, they also have a bunch of disadvantages and annoyances:

  • There are almost 700 plugins in the ChatGPT plugin store. Most are uncurated (just about anything goes).
  • While you can search on a keyword, they're otherwise uncategorized. Plugins like Pluginpedia and PlugFinder claim to help with that, but they're just not that reliable.
  • Many plugins are…what's the word? Meh. Some do nothing more than access the provider's website. For example, there's a plugin for getting discount coupons. How is this better than RetailMeNot?
  • Many plugins seem to be branding or PR exercises to get a foothold in a new marketplace early on. For example, there's an "AI clone" plugin of a specific Silicon Valley small startup CEO. Supposedly, you can ask everything you ever wanted to know about him. So, yeah. This is not exactly something most of us are likely to use.
  • Many plugins don't work or don't do much. I tried to get the local food delivery plugin to tell me where I could get steak dinners in my town, and it recommended Subway. Yes, they have a steak sandwich, but I could have gotten better results from Yelp. I also found a bunch of plugins that just simply hung with no results.
  • You can only run three plugins at once, and if you want to swap out the plugin set, you have to launch a new chat session in ChatGPT, losing all of your current discussion history. This is a big limitation. But even with only three plugins, you saw how I did get the plugin interface to do some magical stuff.

According to Pam Baker, author of ChatGPT for Dummies, "It's harder to see the magic now, given the 3 plugin/25 queries caps and the questionable value of some of the current plugins. But the caps are necessary so OpenAI can better manage model stability and strengthen the guardrails as it assimilates more capabilities."

Also: Microsoft embraces OpenAI's ChatGPT plugin standard

To be fair, we're still very early. That's why I'm not sharing the names of plugins that fell short. There's a good chance they'll get a lot better over time.

So, do plugins really change the game?

Yeah. They do. They really do. We're just in the early stages here, where I had to take extra time to pick and choose four that I think I'll use all the time (plus Stories, which shows an additional level of generative AI potential).

I find that if I make my primary three plugin set a combination of MixerBox WebSearch G, World News, and AskYouTPDF, I can do a whole lot. I can switch to Code Interpreter if I want to do a more in-depth data analysis project.

As ChatGPT grows in its ability to handle plugins, as plugin discovery and curation get better, as we can use more and more plugins at once, it's pretty clear that the kind of chatbot we've just come to know in 2023 is in for a string of future upgrades, providing us with more and more assistance with our tasks and projects.

Also: GPT-3.5 vs GPT-4: Is ChatGPT Plus worth its subscription fee?

ChatGPT for Dummies author Baker shares her view of the future of plugins. She says, "Plugins add capabilities that will eventually enable us to mod ChatGPT on the fly." Her premise is that ChatGPT (and, by extension, other LLMs) will be able to grow their own capabilities.

She told ZDNET, "In an instant, ChatGPT will be able to morph into the perfect tool for any task. Eventually, ChatGPT will be able to automatically determine and select the plugins it needs to respond to every prompt. Where a plugin it needs doesn't yet exist, it will create it on the fly and seamlessly assimilate the new capability."

At this point, I can't tell whether we're creating Skynet or the Borg. Either way, enjoy the added capabilities plugins provide…while…you…still…can.

Resistance is futile.

You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter on Substack, and follow me on Twitter at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, and on YouTube at YouTube.com/DavidGewirtzTV.

More on AI tools

6 AI tools to supercharge your work and everyday life

Colorful image of a person's finger touching a smartphone

Since last year, artificial Intelligence has developed from a futuristic concept to a realistic tool capable of creating AI-generated art, producing human-like conversations via chatbots, and even identifying backyard birds solely based on their chirps.

That said, AI's recent boom has brought apps and tools to the surface with the potential to make our workflow — and even our lives — easier. And according to Gartner Analyst and AI expert Whit Andrews, AI has a leveling effect that is only intensifying, he told ZDNET.

"Think of all the people who find it daunting to express themselves in unfamiliar idioms: that could be drawing a picture, or drawing a map, or explaining a concept," Andrews said. "Generative AI and other AI applications now make that easy, and it has really advanced equality."

That intensification already has a solid foundation as more than 150 AI chatbot apps have been launched in 2023 so far. While most are familiar big-name apps and software like ChatGPT, I found the following six apps to be the most useful for time and budget management, a fitness routine tailored to your skillset, and even mindfulness. After integrating some of these apps into my own life, I think they're worth talking about.

Reclaim

A Google Chrome extension for time management

It never seems that there are enough hours in the day, but Reclaim.ai can find those time slots in your ever-changing schedule. The app uses AI to find time between your workday and your regular life to weave in your to-dos and healthy habits you're trying to commit to.

It can automatically schedule your meetings and block off time to focus on a specific project too, but I like this AI tool for its habit scheduling. Its AI builds flexibility into your schedule so that instead of saying you will go for a walk every day at 2 p.m., it will automatically schedule it to fit around the other events in your calendar or even reschedule it if a last-minute meeting pops up.

By using AI to automatically build the perfect schedule for your priorities each week, it helps you to stay on track with both your work tasks and the habits you wish to incorporate more into your life that you didn't think you had time for.

Currently, the app only works in conjunction with Google Calendar, but there may be Microsoft Office 365 integration in the future.

Cleo

An iOS and Android app for budget managing

This AI app was made for Millennials looking for an easy way to budget. Cleo communicates with users via chatting and uses emojis, memes, and GIFs to get the point across when you spend too much on ordering takeout.

Its AI integrates tough love humor, which it calls "Roast Mode" that playfully shames you on how much you've spent or how little you've saved due to certain repetitive habits even a robot knows you should break. Conversely, Cleo can also hype you up and praise you for your good habits.

Also: Unexpected bill? This AI bot can 'Haggle It' for you

Cleo also has a "Haggle It" feature that helps customers draft letters to help negotiate rent, credit card fees or interest rates, or car insurance rates. A survey conducted by Cleo even found that out of the customers who negotiated their credit card fees and interest, about 20% received reduced rates and fees, so it's definitely worth a try.

But overall, the app builds a budget around your real-life needs and spending habits to set you up for financial success.

Breathhh

An app and a Google Chrome extension for practicing screen-time mindfulness

This Chrome extension has become one of my favorite AI tools because it forces me to slow down and take a break during my workday. Breathhh gets to know your browsing history over time and keeps track of how long you've been in a Google spreadsheet (or how long you've been scrolling through social media).

The tool then suggests a practice or exercises like breathing or documenting your mood by offering it at the right time based on how long you've been on a website, what kind of website it is (i.e., for work or for entertainment), and by learning your browsing habits so it can intervene.

I've found this AI extension to remind me when to take a break and reset my mind so I can avoid work burnout.

ActiveFence Video Content Moderation

An AI tool that works behind the scenes of live streaming apps for managing cyber-bullying

95% of teens today have been exposed to violent subject matter online. This new AI tool from a partnership with Agora and ActiveFence isn't downloadable, but acts as an extension for social media sites or live streaming apps.

The content moderation technology works so that when an app developer activates the extension, it takes screenshot snippets in second-long intervals, passing the images to the ActiveFence content moderation system. Then, the AI flags illicit content in real-time and can even kick a user out.

"What we've done with this integration is provide a low code offering that you click a button, you integrate one or two lines of code, and you are good to go. And you have protected content being streamed through your platform," said Sid Sharma, vice president at Agora. "We really wanted to make the internet and, in specific live interactions on the internet, a very safe place of inclusivity where anybody can go and feel protected and not have to worry about the keyboard warriors."

The AI can identify specific abuse areas like terrorism, hate speech, child safety, self-harm, etc., and flag it to the server in a matter of milliseconds.

"The internet should be a place that people get excited about and feel safe about. So ideally, I would look at this [AI tool] as a necessity for most of the platforms out there," Sharma added. "I would be very hopeful to see a world where live streaming or live video calls are protected, and ensuring utmost safety for the mental health and not providing any distress to every single consumer out there."

Gymbuddy

An iOS and Android app for creating and maintaining workout routines

When it comes to a fitness routine, starting is often the hardest part. It's hard to know what cardio-weights mix is right for you, what exact plan to follow to achieve any specific goals, or what's even safe given your current skill set without enlisting the help of an expensive trainer or going down the TikTok/Pinterest rabbit hole. Gymbuddy uses AI to analyze your current self-assessed fitness level while taking into account body composition factors like height and weight to curate a specific workout plan and schedule in just 24 seconds.

You can also tell the app which body parts you want to focus on strengthening, and it'll keep track of your advancements and increase your difficulty level as you improve.

The app's handy workout scheduler also builds time into your schedule so you can actually complete the personalized workouts it creates for you during lunch breaks, after work, right when you wake up, etc.

Wordtune

An app and Google Chrome extension for content summarization

Sometimes, we just don't have five minutes to spare to watch an entire article on how to fix a pressing issue, or read an article (unless it's a ZDNET article, of course) on the latest tech or social media trend before jumping on your morning meeting. Wordtune, however, is a handy Chrome extension that uses AI to give you the critical points (or Sparknotes, if you will) of that article or video.

For example, a 3,500-word article turns into 24 simple focus points, so you can save about 10 minutes of reading but still come away with the article's most important information.

Also: How to use Wordtune AI to rewrite texts on your iPhone

Wordtune also has an app version for iOS and Android, and this mobile version can generate content like text messages and emails, photo captions, LinkedIn or Twitter posts, cover letters, blog posts, and more by a simple ask. You can ask Wordtune to write a cover letter applying for your dream job, and it'll generate multiple responses to choose from. Or, more simply, you can ask the AI to write a response to a text message when you just can't figure out how to reply.

Aside from these six tools, there's still a slew of AI applications available to help with productivity and workflow, learn how to develop a new skill, or even create a professional headshot free of charge. And given Andrews' insight, we're only on the cusp of seeing AI's full capabilities.

"A generation from now, people will not remember life before this moment when AI made so many things more equitable," he said. "There are all kinds of things that we'll be able to do, and I love that about AI," said Andrews.

More on AI tools

MLPerf Releases Latest Inference Results and New Storage Benchmark

MLPerf Releases Latest Inference Results and New Storage Benchmark September 14, 2023 by John Russell

MLCommons this week issued the results of its latest MLPerf Inference (v3.1) benchmark exercise. Nvidia was again the top performing accelerator, but Intel (Xeon CPU) and Habana (Gaudi1 and 2) performed well. Google provided a peak at its new TPU (v5e) performance. MLCommons also debuted a new MLPerf Storage (v0.5) benchmark intended to measure storage performance under ML training workloads. Submitters in the first Storage run included: Argonne National Laboratory (ANL), DDN, Micron, Nutanix, and Weka.

Digging through the latest Inference results – more than 12,000 performance and 5,000 power inferencing results from 120 systems – is a challenge. There were a more modest 28 results in the storage category. From a usefulness perspective, MLCommons provides direct access to results spreadsheets that permit potential system users/buyers to drill down onto specific system configurations and benchmark tests for comparison. (Links to Inference Datacenter and Edge v3.1 results and Storage v0.5 results)

In the past, HPCwire has tended to try to cover the full exercise in a single article. The rising number of results and introduction of a new category make this less tenable. Instead, we’ll present a broad overview in this article and drill deeper into some vendor-specific results in separate articles (Nvidia and Intel/Habana). By now, you may be familiar with the MLPerf release cadence which is twice yearly for training and inference, with each released on alternate quarters. — so, inference results are released in spring and (early) fall; training results are released in winter and summer. The HPC Training benchmark is released just once yearly, close to the annual SC conference.

Broadly, inferencing and training are the foundational pieces of ML applications, with training deemed the more computational-intense of the two (i.e. think of training LLMs with trillions of parameters). Inferencing, though, is the volume workhorse, sitting behind every chatbot and similar applications.

MLPerf Inference v3.1 introduced two new benchmarks to the suite. The first is a large language model (LLM) using the GPT-J reference model to summarize CNN news articles; it garnered results from 15 different submitters, reflecting the rapid adoption of generative AI. The second change is an updated recommender, meant be more representative of industry practices, using the DLRM-DCNv2 reference model and larger datasets; it had 9 submissions. These new tests, say MLCommons, help advance AI by ensuring that industry-standard benchmarks represent the latest trends in AI adoption to help guide customers, vendors, and researchers, says MLCommons.

In a pre-briefing, David Kanter, MLCommons executive director, said, “We added our first generation recommender a couple of years ago and are now updating it. The LLM (inference) benchmark is brand new and reflects the explosion of interest in what people are calling generative AI, large language models.” An LLM had been added to the MLPerf Training benchmark in the spring (see HPCwire coverage, MLPerf Training 3.0 Showcases LLM; Nvidia Dominates, Intel/Habana Also Impress)

No ML benchmarking effort today would be complete without LLM coverage and MLCommon (parent organization for MLPerf) now has that.

“It’s important to understand large language models operate on tokens. A token is typically a piece of a word. An LLM simply takes a set of tokens as input and predicts the next token. Now, you can chain this together to actually build a predicted sentence. In practice, LLM s are used in a wide variety of applications. You can use them in search and in generating content, like essays or summaries. Summarization is what we do here,” said Kanter.

The MLPerf LLM inference benchmark is quite different from the training benchmark, he emphasized.

“One of the critical differences is the inference LLM is fundamentally performing a generative task. It's writing fairly lengthy sentences, multiple sentences, [but] it’s also actually a different and smaller model,” he said. “Many folks simply don't have the compute or the data to really support a really large model. The actual task we're performing with our inference benchmark is text summarization. So we feed in an article and then tell the language model to summarize the article.”

As is MLCommons’ practice, submitting organizations are invited to submit brief statements on their submissions. These range in quality from pure marketing to providing more granular technical descriptions of a submission's distinguishing features. Given the high number of results, a fast review of the vendor statements can be informative in conjunction with consulting the spreadsheet.

Both Inference and storage submitter statements are appended to the end of this article. As examples, here are a few snippets from a few vendor statements in MLPerf Inference v3.1 exercise:

Azure promoted its online versus on premise showing access to H100 instances. “Azure was the only submitter to publish results for virtual machines in the cloud, while matching the performance of on premises and bare metal offerings. This has been possible thanks to innovative technologies including: AI supercomputing GPUs: Equipped with eight NVIDIA H100 Tensor Core GPUs, these VMs promise significantly faster AI model performance than previous generations, empowering businesses with unmatched computational power; Next-generation computer processing unit (CPU): Understanding the criticality of CPU performance for AI training and inference, we have chosen the 4th Gen Intel Xeon Scalable processors as the foundation of these VMs, ensuring optimal processing speed."

CTuning Foundation, the non-profit ML tool developer, noted that it “[delivered the new version of the open-source MLCommons CM automation language, CK playground and modular inference library (MIL) that became the 1st and only workflow automation enabling mass submission of more than 12000 performance results in a single MLPerf inference submission round with more than 1900 power results across more than 120 different system configurations.”

Google touted its new TPU v5e. “TPU v5e systems use multiple accelerators linked together by a high-speed interconnect and can be configured with a topology ranging from 1×1 to 16×16 (256 chips), giving the user the flexibility to choose the system that best meets their needs. This wide range of topology options offered by TPU systems allows users to run and scale AI inference workloads cost-effectively, without compromising on performance."

“In this submission, Google Cloud used a TPU v5e system with a 2×2 topology (4 TPU chips) to run the 6-billion-parameter GPTJ benchmark. This benchmark demonstrates both the ease of scaling and the cost-efficiency offered by the TPU v5e systems for inference of large language models. Users can easily add more TPU v5e instances to achieve higher total queries per second (QPS), while maintaining the same performance per dollar advantage.”

HPE reported, “In the datacenter category, HPE Cray systems with eight (8) NVIDIA GPUs led our portfolio in performance, delivering more than 340,000 samples per second throughput for ResNet-50 Computer Vision, and more than 28,000 samples per second throughput for Bert 99.0 NLP. HPE also submitted for the first time the newly available HPE ProLiant DL380a Gen11 and HPE ProLiant DL320 Gen11 servers with NVIDIA H100 and L4 GPUs. The HPE ProLiant DL380a Gen11 with four (4) NVIDIA H100 GPUs is ideal for NLP and LLM inference. The HPE ProLiant DL320 Gen11 with four (4) NVIDIA L4 GPUs is a 1U server positioned for computer vision inference.”

Intel discussed Gaudi2 accelerators, 4th Gen Intel Xeon Scalable processors and Intel Xeon CPU Max Series. “Gaudi2 performance on both GPT-J-99 and GPT-J-99.9 for server queries and offline samples are 78.58/second and 84.08/second, respectively. These outstanding inference performance results complement our June training results and show continued validation of Gaudi2 performance on large language models. Performance and model coverage will continue to advance in the coming benchmarks as Gaudi2 software is updated continually with releases every six to eight weeks.

“Intel remains the only server CPU vendor to submit MLPerf results. Our submission for 4th Gen Intel Xeon Scalable processors with Intel AMX validates that CPUs have great performance for general purpose AI workloads, as demonstrated with MLPerf models, and the new and larger DLRM v2 recommendation and GPT-J models.”

You get the general flavor. It’s necessary to dig into the spreadsheet for meaningful comparisons.

MLPerf Debuts Storage Benchmark

The new storage benchmark (v0.5) has been in the works for two years. MLCommons says, “It’s the first open-source AI/ML benchmark suite that measures the performance of storage for ML training workloads. The benchmark was created through a collaboration spanning more than a dozen leading industry and academic organizations and includes a variety of storage setups including: parallel file systems, local storage, and software defined storage. The MLPerf Storage Benchmark will be an effective tool for purchasing, configuring, and optimizing storage for machine learning applications, as well as for designing next-generation systems and technologies.”

Although it’s being introduced along with the latest inference results, storage performance in ML is typically a more sensitive system element in training. MLCommons notes, “Training neural networks is both a compute and data-intensive workload that demands high-performance storage to sustain good overall system performance and availability. For many customers developing the next generation of ML models, it is a challenge to find the right balance between storage and compute resources while making sure that both are efficiently utilized.”

MLPerf Storage is intended to help overcome this problem by accurately modeling the I/O patterns posed by ML workloads, providing the flexibility to mix and match different storage systems with different accelerator types. The new benchmark reports results in sample/s and MB/s. Of course, the choice of storage hardware, protocol/filesystem, and network all influence performance.

The MLPerf Storage benchmark suite is built on the codebase of DLIO, a benchmark designed for I/O measurement in high performance computing, adapted to meet current storage needs.

Talking about the motivation and goals for the new benchmark, Kanter said “I’d heard about pretty large hyperscalers, who deployed really large training clusters, that could not hit their peak utilization because they didn't have enough storage. That [suggested] there's fundamentally a hard problem in storage and one that's under appreciated. Most hyperscalers that are buying 1000s, or tens of 1000s of accelerators also have engineers on staff to design proper storage subsystems."

“The key accomplishment is we created a tool that represents ML training IO patterns, that doesn't require having any compute or accelerators,” said Kanter. “That's important, because if you want to size a storage subsystem for 1000 accelerators, you don't want to have to buy 1000 accelerators. Another interesting thing is it’s a dynamic tool that is coupled to compute. The metric for MLPerf storage is how many samples per second can be streamed out, for a given compute utilization; so we model a compute subsystem. If your storage falls behind too much, the compute subsystem will be idle, and we only allow 10% idle due to storage.”

If the storage system us too slow, you can’t run the benchmark, said Kanter. Obviously, these are early days for MLPerf Storage and it will take some time for the community take its full measure. There are already plans for additions. Given its newness, it’s best look through MLCommon’s documentation. (Link to MLPerf Storage Benchmark Rules)

Link to MLCommons, https://mlcommons.org/en/

VENDOR SUPPLEMENT STATEMENTS ON INFERENCING RESULTS (Unedited)

ASUSTeK

ASUStek recently benchmarked its new AI servers using the MLPerf Inference v3.1 suite, aiming to highlight its performance across varied deep learning tasks. Our results exhibit our system's competency in inferencing some of the most demanding models with remarkable efficiency.

In the modern era of AI, speed and efficiency in deploying machine learning models to production are paramount. Enter ASUS GPU Server portfolios — designed to redefine the standards of inference, as validated by our recent MLPerf Inference benchmarks. Harness the power of AI frameworks like TensorFlow, PyTorch, and more. ASUS servers are not just about raw power; they're about smart power. Optimized software-hardware integrations ensure that you get the most out of every tensor operation. Power doesn’t have to come at the cost of the planet. ASUS GPU servers not only boast top-tier performance metrics but do so with impressive energy efficiency ratings, as highlighted in the MLPerf power efficiency results. Seamlessly scale your AI workloads. With our multi-GPU configurations and optimized in hardware and software, ASUS GPU servers are built to handle increasing data demands, ensuring you’re always ahead of the curve.

System Configuration:

Hardware: ASUS flagship AI Server ESC8000A-E12 with Dual AMD Genoa CPU up to 8 NVIDIA H100 GPUs, and ESC4000A-E12 with Dual AMD Genoa CPU up to 8 L4 GPUs

The results signify the DL system's enhanced performance and capability to address contemporary deep learning challenges, making it an apt choice for researchers and industries requiring accelerated inferencing workloads.

Azure

Microsoft Azure announced the general availability of the ND H100 v5-series for Generative AI at scale. These series of virtual machines vary in sizes ranging from eight to thousands of NVIDIA H100 GPUs interconnected by NVIDIA Quantum-2 InfiniBand networking. Azure was the only submitter to publish results for virtual machines in the cloud, while matching the performance of on premises and bare metal offerings. This has been possible thanks to innovative technologies including:

  • AI supercomputing GPUs: Equipped with eight NVIDIA H100 Tensor Core GPUs, these VMs promise significantly faster AI model performance than previous generations, empowering businesses with unmatched computational power.
  • Next-generation computer processing unit (CPU): Understanding the criticality of CPU performance for AI training and inference, we have chosen the 4th Gen Intel Xeon Scalable processors as the foundation of these VMs, ensuring optimal processing speed.
  • Low-latency networking: The inclusion of NVIDIA Quantum-2 ConnectX-7 InfiniBand with 400Gb/s per GPU with 3.2 Tb/s per VM of cross-node bandwidth ensures seamless performance across the GPUs, matching the capabilities of top-performing supercomputers globally.
  • Optimized host to GPU performance: With PCIe Gen5 providing 64GB/s bandwidth per GPU, Azure achieves significant performance advantages between CPU and GPU.
  • Large scale memory and memory bandwidth: DDR5 memory is at the core of these VMs, delivering greater data transfer speeds and efficiency, making them ideal for workloads with larger datasets.
  • These VMs have proven their performance prowess, with up to six times more speedup in matrix multiplication operations when using the new 8-bit FP8 floating point data type compared to FP16 in previous generations. The ND H100 v5 VMs achieve up to two times more speedup in large language models like BLOOM 175B end-to-end model inference, demonstrating their potential to optimize AI applications further.

The ND H100 v5 is now available in the East United States and South Central United States Azure regions. Enterprises can register their interest in access to the new VMs or review technical details on the ND H100 v5 VM series at Microsoft Learn.

CTuning

As a founding member of MLCommons, cTuning.org is committed to democratizing MLPerf benchmarks and making them accessible to everyone to deliver the most efficient AI solutions while reducing all development, benchmarking and optimization costs.

We are proud to deliver the new version of the open-source MLCommons CM automation language, CK playground and modular inference library (MIL) that became the 1st and only workflow automation enabling mass submission of more than 12000 performance results in a single MLPerf inference submission round with more than 1900 power results across more than 120 different system configurations from different vendors (different implementations, all reference models and support for DeepSparse Zoo, Hugging Face Hub and BERT pruners from the NeurIPS paper, main frameworks and diverse software/hardware stacks) in both open and closed divisions!

This remarkable achievement became possible thanks to open and transparent development of this technology as an official MLCommons project with public Discord discussions, important feedback from Neural Magic, TTA, One Stop Systems, Nutanix, Collabora, Deelvin, AMD and NVIDIA, and contributions from students, researchers and even school children from all over the world via our public MLPerf challenges. Special thanks to cKnowledge for sponsoring our developments and submissions, to One Stop Systems for showcasing the 1st MLPerf results on Rigel Edge Supercomputer, and to TTA for sharing their platforms with us to add CM automation for DLRMv2 available to everyone.

Since it’s impossible to describe all the compelling performance and power-efficient results achieved by our collaborators in a short press-release, we will make them available with various derived metrics (power efficiency, cost, etc) and reproducibility reports at the MLCommons CK playground (x.cKnowledge.org), github.com/mlcommons/ck_mlperf_results and github.com/mlcommons/ck/blob/master/docs/news-mlperf-v3.1.md shortly after official release.

We continue enhancing the MLCommons CM/CK technology to help everyone

automatically co-design the most efficient end-to-end AI solutions

based on their requirements and constraints. We welcome all submitters to join our public MLCommons Task Force on Automation and Reproducibility if you want to automate your future MLPerf submissions at scale.

Connect Tech Inc

As a new member of MLCommons, Connect Tech ran performance and accuracy benchmarks in the Inference: Edge category in its recent MLPerf submission. Using Connect Tech’s feature-rich Hadron carrier board with the NVIDIA Jetson Orin NX, a high-performance, energy-efficient platform, showcased remarkable levels of performance across various AI workloads.

Connect Tech additionally supports NVIDIA Jetson Orin NX with Photon and Boson carrier boards, and system devices like Polaris and Rudi-NX. By deploying on Connect Tech’s production-ready hardware, customers can take immediate advantage of Jetson Orin NX for performance improvements and enhanced user experience with robotics and other edge AI applications.

Connect Tech's involvement in MLCommons signifies more than just technical achievement. It reflects the company's commitment to pushing the envelope of what's possible in the world of AI at the edge. The seamless integration of Connect Tech's hardware with NVIDIA's cutting-edge technology presents engineers and scientists with the tools to drive AI and machine learning innovations across diverse industries, including robotics, industrial automation, and healthcare.

Connect Tech is a hardware design and manufacturing company, specializing in rugged, small form factor solutions. As an Elite NVIDIA Jetson ecosystem partner, Connect Tech designs carrier boards, enclosures, and embedded systems for each Jetson generation. With a rich history of innovation, Connect Tech integrates edge AI solutions within various industries, empowering engineers and scientists to harness the potential of machine learning.

Connect Tech remains at the forefront as the world delves deeper into AI and machine learning. Navigating the complex landscape of embedded AI computing is made easier by using NVIDIA and Connect Tech’s innovative products.

Dell

Enterprise IT is bracing for the most transformative technology trend in decades: generative AI. Dell Technologies is ready to meet this demand with the world’s broadest Generative AI solutions portfolio from desktop to edge to data center to cloud, all in one place.

For the MLPerf inferencing v3.1 benchmark testing, Dell submitted 230 results, including the new GPT-J and DLRMv2 benchmark results, across 20 system configurations. Dell Technologies works with customers and collaborators, including NVIDIA, Intel, and Qualcomm, to optimize performance and efficiency, boosting inferencing workloads, including generative AI.

The Dell PowerEdge XE accelerated server family continues to deliver tremendous performance gains across several benchmarks. Here are some of the latest highlights:

  • The PowerEdge XE9680 with 8 NVIDIA H100 SXM GPUs continues to deliver Dell’s best performance results, up to ~16% better performance than the previous MLPerf 3.0 benchmark results in Image classification, Speech-to-text, Language processing, and Recommendation.
  • Stellar results for the complete PowerEdge XE server family, including the Direct Liquid Cooled PowerEdge XE9640 that packs either 4 NVIDIA H100 SXM GPUs or 4Intel Data Center GPU Max OAM GPUs in an ultra-dense 2RU profile.
  • New NVIDIA L40 GPU results in the PowerEdge R760xa, Dell’s most versatile accelerated server with the best price-to-GPU performance ratio.
  • New Intel Xeon CPU-based machine learning and GPT-J results for the PowerEdge R760.
  • The rugged, edge-optimized PowerEdge XR5610 yielded impressive power efficiency relative to performance results with the NVIDIA L4 GPU.
  • Updated Qualcomm results featuring the compact and energy-efficient PowerEdge XR4520c compute sled with Qualcomm Cloud AI 100 Standard accelerator cards.

Generate higher quality, faster time-to-value predictions and outputs while accelerating decision-making with powerful solutions from Dell Technologies. Take a test drive in one of our worldwide Customer Solution Centers. Collaborate with our Innovation Lab and tap into one of our Centers of Excellence.

Fujitsu

Fujitsu offers a fantastic blend of systems, solutions, and expertise to guarantee maximum productivity, efficiency, and flexibility delivering confidence and reliability. Since 2020, we have been actively participating in and submitting to inference and training rounds for both data center and edge divisions.

In this round, Fujitsu demonstrated the performance of PRIMERGY CDI with four A100-PCIe-80GB GPUs installed in an external PCIe BOX and measured the benchmark program only for the data center closed division. Fujitsu Server PRIMERGY CDI is expertly engineered to deploy the necessary resources according to each customer's unique workload, releasing them when no longer needed. CDI stands for Composable Disaggregated Infrastructure, a next-generation technology that supports the diversification of data processing. This results in an efficient operation that maximizes resource utilization, while providing user-friendly services that eliminate the drawbacks of traditional physical servers.

As demonstrated by the impressive results of this round, the PRIMERGY CDI confirms that even with GPUs mounted in an external PCIe BOX, it delivers outstanding performance and remarkable scalability for PCIe components.

Our purpose is to make the world more sustainable by building trust in society through innovation. With a rich heritage of driving innovation and expertise, we are dedicated to contributing to the growth of society and our valued customers. Therefore, we will continue to meet the demands of our customers and strive to provide attractive server systems through the activities of MLCommons.

Giga Computing

Giga Computing Technology, a subsidiary wholly owned by GIGABYTE, is the enterprise unit that split off from GIGABYTE that designs, manufactures, and sells servers, server motherboards, immersion solutions, and workstations. As the GIGABYTE brand is widely recognized, Giga Computing will continue to use and promote it, and that includes at expos where we will join as GIGABYTE. Although the company name has changed, our customers can still expect the same quality and services as before. Giga Computing strives to do better and that includes greater push for efficiency and cooling with immersion and DLC technology. As well as providing public AI benchmarks.

As one of the founding members of MLCommons, GIGABYTE has continued to support the community’s efforts in benchmarking server solutions for various AI training & inference workloads. In the latest round of MLPerf Inference v3.1, Giga Computing submitted a powerful GIGABYTE system for platforms: Intel Xeon & NVIDIA H100 SXM5, and the results speak for themselves while showing great efficiency as measured in performance/watt. We did find that our system achieved excellent performance in some tests such as rnnt-Server and bert99-offline. We would have liked to have more benchmarks, but due to resource limitations we are not able; however, we found that our partners NVIDIA, Qualcomm, and Krai chose our GIGABYTE servers to do their own testing.

Google

Google Cloud recently launched an expansion to its AI infrastructure portfolio — Cloud TPU v5e — and is proud to announce its performance results in this round of MLPerf Inference (data center category). TPU v5e systems use multiple accelerators linked together by a high-speed interconnect and can be configured with a topology ranging from 1×1 to 16×16 (256 chips), giving the user the flexibility to choose the system that best meets their needs. This wide range of topology options offered by TPU systems allows users to run and scale AI inference workloads cost-effectively, without compromising on performance.

In this submission, Google Cloud used a TPU v5e system with a 2×2 topology (4 TPU chips) to run the 6-billion-parameter GPTJ benchmark. This benchmark demonstrates both the ease of scaling and the cost-efficiency offered by the TPU v5e systems for inference of large language models. Users can easily add more TPU v5e instances to achieve higher total queries per second (QPS), while maintaining the same performance per dollar advantage.

We are looking forward to seeing what Google Cloud customers achieve with the new TPU v5e systems.

HPE

HPE successfully submitted results in partnership with Intel, NVIDIA, Qualcomm, and Krai. HPE demonstrated a range of high-performing inference systems for both the datacenter and edge in Computer Vision, natural language processing (NLP), and large language models (LLM).

In the datacenter category, HPE Cray systems with eight (8) NVIDIA GPUs led our portfolio in performance, delivering more than 340,000 samples per second throughput for ResNet-50 Computer Vision, and more than 28,000 samples per second throughput for Bert 99.0 NLP.

HPE also submitted for the first time the newly available HPE ProLiant DL380a Gen11 and HPE ProLiant DL320 Gen11 servers with NVIDIA H100 and L4 GPUs. The HPE ProLiant DL380a Gen11 with four (4) NVIDIA H100 GPUs is ideal for NLP and LLM inference. The HPE ProLiant DL320 Gen11 with four (4) NVIDIA L4 GPUs is a 1U server positioned for computer vision inference. The HPE ProLiant DL380a Gen11 showed strong inference performance using 4th Gen. Intel Xeon Scalable Processors in CPU-only inference scenarios. The HPE ProLiant DL385 Gen10 Plus v2 with eight (8) Qualcomm Cloud AI 100 Standard accelerators remained well balanced for over-network inference compared to offline datacenter performance. Qualcomm Cloud AI 100 Standard is ideal for both computer vision and NLP inference.

In the Edge category, HPE Edgeline e920d powered by four (4) Qualcomm Cloud AI 100 Standard accelerators remains one of the lowest latency systems in the Edge category for SingleStream and MultiStream inference scenarios. The HPE Edgeline e920d also achieved strong performance improvements in throughput and energy efficiency.

Many thanks to Krai’s collaboration in achieving high-performance and energy efficiency for Qualcomm Cloud AI 100 accelerators.

IEI

IEI Industry Co., LTD is a leading provider of data center infrastructure, cloud computing, and AI solutions, ranking among the world’s top 3 server manufacturers. Through engineering and innovation, IEI delivers cutting-edge computing hardware design and extensive product offerings to address important technology arenas like open computing, cloud data center, AI, and deep learning.

In MLCommons Inference v3.1, IEI submitted the NF5468M6 system.

NF5468M6 is a highly versatile 4U AI server supporting between 4 and 16 NVIDIA single and double-width GPUs, making it ideal for a wide range of AI applications including AI cloud, IVA, video processing and much more. NF5468M6 offers ultra-high storage capacity and the unique function of switching topologies between Balance, Common and Cascade in one click, which helps to flexibly adapt to various needs for AI application performance optimization.

Intel

Intel is pleased to report MLPerf Inference v3.1 performance results for our Gaudi2 accelerators, 4th Gen Intel Xeon Scalable processors and Intel Xeon CPU Max Series. These results reinforce Intel’s commitment to delivering the full spectrum of products to address wide-ranging customer AI requirements.

Gaudi2 performance on both GPT-J-99 and GPT-J-99.9 for server queries and offline samples are 78.58/second and 84.08/second, respectively. These outstanding inference performance results complement our June training results and show continued validation of Gaudi2 performance on large language models. Performance and model coverage will continue to advance in the coming benchmarks as Gaudi2 software is updated continually with releases every six to eight weeks.

Intel remains the only server CPU vendor to submit MLPerf results. Our submission for 4th Gen Intel Xeon Scalable processors with Intel AMX validates that CPUs have great performance for general purpose AI workloads, as demonstrated with MLPerf models, and the new and larger DLRM v2 recommendation and GPT-J models.

The results confirm that 4th Gen Intel Xeon Scalable processor with optimized data pre-processing, modeling and deployment tools and optimizations, is an ideal solution to build and deploy general purpose AI workloads with the most popular open source AI frameworks and libraries.

For the GPT-J 100-word summarization task of a news article of approximately 1,000 to 1,500 words, 4th Gen Intel Xeon processors summarized two paragraphs per second in offline mode and one paragraph per second in real-time server mode.

This is the first time we’ve submitted MLPerf results for our Intel Xeon CPU Max Series, which provides up to 64GB of high-bandwidth memory. For GPT-J, it was the only CPU able to achieve 99.9% accuracy, which is critical for usages for which the highest accuracy is of paramount importance.

With our ongoing software updates, we expect continued advances in performance and productivity, and reporting new training metrics with the November training cycle.

For more details, please see MLCommons.org.

Notices & Disclaimers

Performance varies by use, configuration and other factors. Learn more at www.Intel.com/PerformanceIndex .

Performance results are based on testing as of dates shown in configurations and may not reflect all publicly available updates. See

backup for configuration details. No product or component can be absolutely secure. Your costs and results may vary.

Intel technologies may require enabled hardware, software or service activation.

© Intel Corporation. Intel, the Intel logo, and other Intel marks are trademarks of Intel Corporation or its subsidiaries. Other names and brands may be claimed as the property of others.

Krai

Established in 2020 in Cambridge, UK, KRAI comprises a dynamic group of exceptional engineers committed to transforming the practice of AI development and deployment on cutting-edge hardware. Acknowledging the growing significance of AI, and the limitations of computer systems, KRAI was founded with the mission to help propel the adoption of AI in a responsible way. We take pride in our collaborations with industry leaders such as Qualcomm, HPE, Dell, Lenovo, and more.

We are thrilled to power many v3.1 submissions with our KRAI X automation technology. In v3.0, we previewed KRAI X through a single-digit number of submissions. In v3.1, we have fully transitioned from using our acclaimed automation technology based on Collective Knowledge v1 to using KRAI X, having enabled a triple-digit number of submissions with compelling outcomes.

Notably, we dedicated considerable effort to optimizing energy efficiency. For example, on a server equipped with 16 Qualcomm Cloud AI accelerators, we achieved 237.0 QPS/W for ResNet50 and 9.2 QPS/W for BERT-99, compared with 210.7 QPS/W and 7.7 QPS/W, respectively, in the previous round (up to 20%).

Finally, we unveiled support for several new targets to our open-source KRAI Inference Library Technology (KILT): TensorRT for GPUs; ONNX Runtime for CPUs and GPUs; and Snapdragon Neural Processing Engine (SNPE) for Qualcomm's CPUs, GPUs, DSPs and NPUs.

Our team remains steadfast in enhancing KRAI technologies to deliver fully optimized end-to-end AI solutions tailored to specific constraints and performance objectives. Our technologies empower system designers to accelerate the design and deployment of AI by eliminating laborious manual processes. Our mission is to assist companies in the development, benchmarking, optimization and deployment at scale of their AI solutions.

Moffett

Moffett AI is a leader in sparse AI computing, dedicated to providing AI computing platforms and services, with a mission to continue evolving the frontiers of AI performance using sparse computing.

Following the outstanding performance in the MLPerf Inference v2.1 and v 3.0 benchmarks, Moffett AI has once again submitted impressive results for the S30 Accelerator. These results were achieved on GPT J-99 model in offline mode in the Data Center Open Division.

In MLPerf Inference v3.1, MLPerf's first introduction of large model inference, Moffett AI was the only one in the Data Center Open Division to submit results of large model GPT J inference.

The S30 Accelerator is powered by Moffett's Antoum processor — the world's First Al computing Accelerators with 32x sparsity. Moffett's patented deep dual sparsity algorithm and the co-designed Antoum chip and software platform architecture drastically increase the computing performance of the accelerators, increasing throughput, reducing latency and power, maintaining accuracy, while also significantly reducing the Total Cost of Ownership (TCO).

Moffett Al's submissions showcase the remarkable benefits of Moffett's Antoum processor, especially the advantages of sparse computing for large model inference with software and hardware co-design:

  • Enabling extraordinarily low TCO: Moffett AI accelerators perform computing workloads with less infrastructure, simplified deployment, and reduced Total Cost of Ownership (TCO) to achieve overall cost reduction and efficiency.
  • Excellent performance on large model inference: The S30 accelerators in 8-card mode achieve impressively high performance (170.59 Sample/s) on GPT J-99.
  • Scalable highly performant across data center ecosystem: The S30 accelerator performed exceptionally well in single-card, 4-card, 8-card mode across different servers.

Moffett's deep sparse industry leading performance is optimal for generative AI workloads such as GPTJ. These AI models are massively increasing in size and the demand for computing performance is skyrocketing, as well as the need to reduce power, latency and TCO.

Neural Magic

While pursuing research at MIT, Nir Shavit and Alexander Matveev encountered the limiting barriers of GPUs and other prevailing hardware options in the realm of deep learning. This frustration drove them to develop software that unfetters AI innovation from GPUs, which ultimately led to the founding of Neural Magic in 2018. Now, enterprises and communities alike, can use Neural Magic software and algorithms to get performant and accurate AI deployments on commodity CPUs.

Neural Magic’s DeepSparse is a sparsity-aware inference runtime that delivers power-efficient AI performance on commodity CPUs, from the cloud to the edge. Our open-source compression framework, SparseML, unifies state-of-the-art sparsification algorithms for easy use and application across ML use cases like computer vision, natural language processing, and generative AI.

In partnership with the cTuning Foundation, Neural Magic leveraged the open-source CK technology to automate and reproduce results across all 74 benchmarks in MLPerf Inference v3.1. We're excited to share our DeepSparse CPU benchmarks, which showcase the performance of an array of sparsified BERT question answering models. These models are readily available for deployment from Neural Magic’s SparseZoo. They have been tested across a diverse range of platforms — like Intel and AMD, across GCP and AWS. Our benchmarks span both x86 and ARM architectures, to ensure comprehensive coverage between edge devices and robust cloud infrastructures.

As we continue to refine our algorithms to finesse performance and accuracy, particularly for large model compression, we’ve appreciated the early enthusiasm shown towards our nascent research, including GPTQ and SparseGPT. We’re excited to see our research included in initiatives that drive a new era of generative AI pursuits by partners. As we continue to push boundaries with model optimization, we are motivated by our work with customers, to redefine what's possible in the space of deep learning, all without the burden of specialized hardware, operational complexities, or daunting costs.

NVIDIA

In MLPerf Inference 3.1, we are thrilled to have made our first submission using the NVIDIA GH200 Grace Hopper Superchip, which combines NVIDIA Grace, our first data center CPU, with the NVIDIA Hopper GPU to create a processor for the era of generative AI and accelerated computing. It ran every workload in the datacenter category, including the new GPT-J and DLRMv2 tests, and extended the leading performance of the H100 Tensor Core GPU across the board.

We were also pleased to make our first available submission using the L4 Tensor Core GPU powered by the NVIDIA Ada Lovelace architecture. With a single-slot, low-profile design and low thermal design power, it can bring the performance and versatility of the NVIDIA platform in AI, video, and graphics to any server.

And on our Jetson platforms for edge AI and robotics, powered by the NVIDIA Orin system-on-chip (SoC), we were thrilled to deliver up to 85% more performance compared to our prior submissions thanks to software advances and use of the second-generation Programmable Vision Accelerator (PVA) built into Orin.

The NVIDIA AI platform delivers innovation across the full stack, accelerates the entire AI workflow end-to-end – from data preparation to model training to deployed inference from cloud to edge – and achieves great performance across a broad range of AI models. It’s also available from every major cloud and server maker, and offers the quickest path to production AI and enterprise-grade support with NVIDIA AI Enterprise.

We are thrilled to see 13 NVIDIA partners submit great inference results, with both on-prem and cloud solutions spanning the breadth of our data center GPU portfolio.

We also wish to commend the ongoing work MLCommons is doing to bring benchmarking best practices to computing, enabling peer-reviewed apples-to-apples comparisons of AI and HPC platforms to better understand and compare product performance across diverse workloads.

Nutanix

The Nutanix Cloud Platform solution is a hybrid multicloud platform that provides a software stack to enable the full lifecycle of AI/ML applications, regardless of where the hardware is deployed. The consistent operating model helps enable ease of management whether data is being collected, transformed, or fed into models for inferencing at the edge, or models are being fine-tuned at the core data center or in the public cloud.

Nutanix is pleased to announce the first set of published results with the MLPerf benchmark. The benchmark was executed in a lab environment on a Nutanix NX-3155-G8 node with two NVIDIA A100 Tensor Core GPU 80GB. It was fully virtualized on the AHV hypervisor.

Oracle

Oracle Cloud Infrastructure (OCI) offers AI Infrastructure, AI Services, ML Services, and AI in our Fusion Applications. Our AI infrastructure portfolio includes bare metal instances and virtual machines powered by NVIDIA H100 (coming soon), NVIDIA A100, and NVIDIA A10 GPUs.

The inference benchmark results for the high-end BM.GPU.H100.8 and BM.GPU.A100-v2.8 instances demonstrate that OCI provides high performance that is comparable to other deployments on both on-premises and cloud infrastructure. These instances provide eight NVIDIA GPUs per node. In addition to inferencing, for training workloads each node can be clustered using a high performance RDMA network to tens of thousands of GPUs.

The results for GPU.A10.4 indicate superior performance for smaller AI model inferencing workloads. The evaluated bare metal instance featured four NVIDIA A10 GPUs. OCI also offers virtual machines with one or two NVIDIA A10 GPUs. All three OCI instance types based on the NVIDIA A10 GPU are priced lower than their OCI counterparts with NVIDIA A100 or H100 GPUs.

Qualcomm Technologies, Inc.

The Qualcomm Cloud AI inference accelerators leverage the Company’s heritage in advanced signal processing and power efficiency to deliver high throughput, low power AI inference processing in the cloud and at the Edge. The Cloud AI products support all AI Inferencing workloads, spanning ML frameworks and network models, including large networks with multi-accelerator aggregation.

Qualcomm MLPerf v3.1 inference results demonstrate optimization of all the ML Benchmarks across the board. All our submissions show incremental upside in performance, power efficiency and lower latencies for NLP and computer-vision networks. Qualcomm partner Lenovo has added a new datacenter server platform ThinkSystem SR665v1 with five Qualcomm Cloud AI 100 PCIe Pro AI accelerators. Qualcomm and HPE have added RetinaNet network benchmarks to its Network division submissions in addition to BERT. Our partner Dell has made submission with Cloud AI 100 Standard Accelerator with PowerEdge XR4520c Server. Qualcomm AI software partner Krai.ai has made several Cloud AI accelerator-based platform submissions in Edge Category and Open Division.

Qualcomm MLPerf v3.1 inference benchmark results have surpassed its own previous records of peak offline performance, power efficiency and lower latencies in several categories. The 2U datacenter server platform with 16x Qualcomm Cloud AI 100 PCIe Pro (75W TDP) accelerators has further improved its power efficiency by another 15-20% across NLP and CV networks. The Gloria Highend, Edge appliance platform has peaked its performance and power efficiency for all three neural network categories we have submitted to. The RetinaNet performance across all platforms has been optimized by additional ~12%

Qualcomm continues to innovate and optimize AI solutions across all submissions. The Network division applicable for datacenter has been extended to RetinaNet Network in addition to BERT. All Network division submission results achieved nearly the same results as Closed division.

All submissions are powered by the KRAI X and KILT technologies.

Qualcomm is a trademark or registered trademark of Qualcomm Incorporated.

Qualcomm Cloud AI is product of Qualcomm Technologies, Inc. and/or its subsidiaries

Quanta Cloud Technology

Quanta Cloud Technology (QCT) is a global datacenter solution provider that enables diverse HPC and AI workloads, was named amongst MLPerf inference list in the latest MLPerf results released by MLCommons.

QCT participated in the latest round of MLPerf Inference v3.1 and submitted results to the data center closed division for four different system configurations.

One of the configurations showcased QCT's cutting-edge platform, QCT submitted its state-of-the-art platform QuantaGrid D54U-3U in preview with four NVIDIA H100 PCIe GPU. QuantaGrid D54U-3U is an acceleration server designed for AI/HPC. Supporting two 4th Gen Intel Xeon Scalable processors with up to 350W and 32x DIMM slots, this 3U system features support for four dual width accelerator cards or up to eight single width accelerator cards to provide a comprehensive and flexible architecture that can be optimized for various AI/HPC applications.

Additionally, QCT presented results for the QuantaGrid-D54Q-2U System, equipped with four NVIDIA L4 Tensor Core GPUs. Thanks to innovative hardware design, meticulous system tuning, and software optimization, QCT achieved outstanding performance in MLPerf Inference v3.1.

Going forward, QCT remains committed to delivering comprehensive hardware systems, solutions, and services to both academic and industrial users. The company will continue to share its MLPerf results with the MLCommons community, contributing to the advancement of MLPerf inference and training benchmarks.

SiMa

SiMa.ai prides itself on being at the forefront of edge AI technology, consistently pushing the boundaries of performance and energy efficiency. We are extremely excited to share our results in this latest MLPerf benchmark, where we achieved a 20% improvement in our scores since April 2023.

In the edge AI sector, where both performance and energy efficiency are paramount, the standout metric is frames per second per watt. This metric is a critical indicator of how many frames our system can process per watt of electricity consumed and vital to edge AI workloads. SiMa.ai’s custom-made ML Accelerator is the linchpin of our success in achieving unparalleled power efficiency without compromising performance.

Our 20% improvement since the April 2023 submission is one of the most thrilling aspects of SiMa’s latest MLPerf results. Substantial enhancements in our compiler technology and memory management systems drove this increase. These foundational improvements optimize code execution and resource allocation, boosting the overall performance of our hardware. What makes this even more compelling is the translation of these improvements beyond the benchmarks to real-world use cases. We’re able to enhance the performance of all models running on our hardware and offer our customers more value and versatility across a wide range of applications.

SiMa.ai’s consistent participation and performance in MLPerf is part of a broader growth strategy where we are transitioning from the 16nm process to more advanced technologies in future generations. This move is not merely a technical upgrade; it's a strategic evolution to ensure we continue to lead in performance, efficiency, and innovation. As we look to the future, our focus remains clear: to continue pushing the boundaries of what's possible in edge AI, particularly in performance and ease of use.

Supermicro

Supermicro has a long history of designing a wide range of products for various AI use cases. In MLPerf Inference v3.1, Supermicro has submitted four systems in the closed division, datacenter category.

Supermicro’s mission is to provide application-optimized systems for a broad range of workloads. For example, Supermicro designs and manufactures four types of systems for the NVIDIA HGX H100 8GPU and 4GPU platforms, with 16 lanes of connectivity so that nothing stands in the way of the flow of data to the accelerators, which is ideal for AI workloads that are very I/O intensive and that need a balance of CPU and GPU performance.

Supermicro offers a range of CPUs and quantities of GPU in various form factors for customers who may have differing compute and environmental requirements. Furthermore, Supermicro provides upgraded power supplies to give customers choices on using cost-effective power supplies or genuine N+N redundancy to maximize the TCO. Supermicro also offers liquid cooling options for NVIDIA HGX systems to help deployments use higher TDP CPUs and GPUs without thermal throttling.

For those in pursuit of PCIe Gen5 platforms, Supermicro presents an array of compelling alternatives. For the MLPerf v3.1 Inference, Supermicro submitted H100 results for the GPU SuperServer SYS-521GE-TNRT, a compact high-performance server with a 5U rackmount form factor. The system is currently shipping worldwide.

Supermicro’s GPU A+ Server, the AS-8125GS-TNHR (AMD CPU), and SuperServer, the SYS-821GE-TNHR (Intel CPU), both have 8 H100 SXM5 GPUs and GPU-GPU interconnect with NVLink and NVSwitch. Furthermore, the dual root configuration features directly attached 8 GPUs to achieve the lowest latency possible and improve performance, which is hugely beneficial for demanding scenarios our customers face with machine learning (ML) and HPC workloads. The models achieved outstanding results and are an excellent choice for AI development.

Supermicro's commitment extends beyond these specific models, spanning a vast assortment of GPU-based servers tailored for any conceivable environment. This steadfast commitment is evidenced by the impressive performance exhibited across a gamut of MLPerf tests. Supermicro pledges to continue raising the bar for exceptional performance, serving users' distinct needs across various workstations and servers.

TTA

Telecommunications Technology Association (TTA) is a non-profit organization that was established in 1988. Its primary mission revolves around the standardization of information and communication technology (ICT), alongside the testing and certification of various ICT products, services, and data.

TTA also supports global marketing and market share expansion for Korean server/storage hardware vendors. TTA recognizes the value of publishing performance results through platforms like MLCommons, and on this occasion, we test the server KR580S1 manufactured by KTNF Co, Ltd.

Since its founding in 2001 under the motto of Korea Technology aNd Future, KTNF has emerged as a server industry leader with the best technology and expertise

KTNF is a trusted company that not only provides server circuit and system technology for services optimized for cloud and edge computing environments, but one that also supports your business in response to the rapidly changing fourth industrial revolution environment based on expertise and excellent service.

The KR580S1 server used in this MLCommons inference v3.1 is a good system for AI, Cloud service workload in Datacenter.

This server is an Intel Xeon-SP based server that can install up to two double slot GPUs. It support DDR4 DIMM memory and High performance NVMe SSDs. As a hot tolerance server, it is monitoring the GPU card to achieve maximum performance.

TTA carried out the tests using two different GPUs: the NVIDIA A100 (40G) for Edge and the NVIDIA Tesla T4 for Edge and Datacenter. With Tesla T4, the results showed a relatively decent performance score for algorithms like resnet50, and bert-99, affirming that the server’s suitability for deployment in both the Edge Server Market and the Datacenter Server Market.

Thanks to CTuning for helping to automate our submissions using the MLCommons CM automation language and CK playground.

xFusion

xFusion Digital Technology Co., Ltd. is committed to becoming the world's leading provider of computing power infrastructure and services. We adhere to the core values of "customer-centric, striver-oriented, long-term hard work, and win-win cooperation", as we continue to create value for customers and partners, and accelerate the digital transformation of the industry.

In this performance competition of MLPerf Inference v3.1, we used a new generation of GPU server product, FusionServer G5500 V7 to conduct performance tests on all benchmarks under various GPU configurations and achieved excellent results.

FusionServer G5500 V7 (G5500 V7) is a new-generation 4U 2-socket GPU server. It supports a maximum of 10 x double-width GPU cards. We use Intel Xeon Platinum 6458Q CPU x2 and 8 to 10 A30 or L40 GPU configurations to test all evaluation items. It has made excellent achievements for bert, dlrm-v2 and gptj models. 62 test results can achieve the best performance under the same GPU hardware configuration.

FusionServer G5500 V7 features high performance, flexible architecture, high reliability, easy deployment, and simplified management. It accelerates applications such as AI training, AI inference, high-performance computing (HPC), image and video analysis, and database, and supports enterprise and public cloud deployment.

VENDOR SUPPLEMENT STATEMENTS ON STORAGE RESULTS (Unedited)

ANL

We evaluated the MLPerf Storage Benchmark on the Polaris supercomputer at the Argonne Leadership Computing Facility (ALCF), a US Department of Energy Office of Science User Facility. We assessed the performance of the storage system on Polaris for two distinct MLPerf Storage AI workloads: UNet3D and Bert.

Polaris is a HPE Nvidia, 44 petaflops system, which consists of 560 NVIDIA DGX A100 nodes interconnected with HPE Slingshot. Each node is equipped with 2 NVMe drives each of 1.60 TB. Eagle is a Lustre parallel file system residing on an HPE ClusterStor E1000 platform equipped with 100 Petabytes of usable capacity across 8480 disk drives. It has 160 Object Storage Targets and 40 Meta Data Targets with an aggregate transfer rate of 650 GB/s.

We ran the storage benchmark with datasets hosted on both the Eagle Lustre file system and node-local NVMe SSDs to emulate the behavior of how production users tend to use the system for AI workloads. Notably, our findings revealed a remarkable linear increase in I/O throughput for both UNet3D and Bert, as we scale to 2048 accelerators. The efficient I/O handling on Polaris allows data transfer to overlap with the computation, resulting in an impressively high accelerator utilization rate, close to 100%. For UNet3D, which is an I/O intensive workload, we observed a peak throughput of 200 GB/s when utilizing the Eagle parallel file system. When leveraging the node-local NVMe SSDs, the I/O throughput achieves 800 GB/s. In the case of Bert, which is not as I/O intensive as UNet3D, we observed the same ideal scaling pattern in I/O throughput. These outcomes clearly demonstrated that the storage system deployed at the ALCF supports efficient I/O operations for the state-of-the-art AI applications.

DDN

Organizations need a reliable and scalable storage platform to achieve their AI goals. DDN enables end-to-end Accelerated Computing with a simple but powerful appliance proven in the largest and most demanding AI deployments worldwide.

For the inaugural MLPerf storage benchmark, DDN is pleased to submit two configurations using the AI400X2 appliance in the CLOSED division to evaluate storage performance in typical small-scale machine learning deployments. The benchmark results measure the number of GPUs saturated by a single AI400X2 appliance during prescribed machine-learning tasks, the first using a single GPU compute system and the second with a cluster of GPU compute systems.

  • In the single compute node benchmark, one DDN AI400X2 NVMe appliance running DDN’s EXAScaler 6.2 parallel filesystem served 40 accelerators at a throughput of 16.2 GB/s.
  • In the multi-node benchmark, one DDN AI400X2 NVMe appliance served 160 accelerators across ten GPU compute nodes at a throughput of 61.6 GB/s.
  • We note that this second benchmark submission was limited by the performance of the compute clients, rather than the single AI400X2 system, demonstrating the superior efficiency of the 2U appliance.

The AI400X2 appliances scale out linearly, adding performance or capacity to meet the requirements of the most ambitious AI projects.

We are excited to support the ongoing work of MLCommons to establish benchmarking best practices for AI/ML systems and allow the AI/ML community to make informed decisions based on standardized comparative measures. We look forward to future versions of the MLPerf Storage benchmark with additional workload models.

To learn more about DDN’s AI400X2, configurations, capabilities, and use cases, visit https://ddn.com/a3i

Micron

The Micron 9400 is designed to manage the most demanding data center workloads, particularly in artificial intelligence (AI) training, machine learning (ML) and high-performance computing (HPC) applications. The drive delivers an industry-leading 30.72 terabytes (TB) of storage capacity and 77% improved input/output operations per second (IOPS). The Micron 9400 is one of the world’s fastest PCIe Gen4 data center U.3 drive shipping and delivers consistently low latency at all capacity points.

Micron is proud to announce the first MLPerf storage benchmark results on the Micron 9400 NVMe SSD.

A single 7.68TB 9400 Pro is capable of supporting 17 accelerators with a throughput of 6.1 GB/s.

The Micron 9400’s capacity and performance enable larger datasets and accelerated epoch time leading to more efficient utilization of graphics processing units (GPUs).

While many SSDs are designed for pure read or write use cases, the Micron 9400 was designed with real-world applications in mind.

The Micron 9400 SSD is available in a U.3 form factor that is backwards-compatible with U.2 sockets and comes in capacities ranging from 6.4TB to 30.72TB. These options provide data center operators the flexibility to deploy the most energy efficient storage while matching their workloads with the right blend of performance, capacity and endurance.7 This versatile SSD is built to manage critical workloads whether in on-premises server farms or in a multi-tenant shared cloud infrastructure, and can be flexibly deployed in hyperscale, cloud, data center, OEM and system integrator designs.

Nutanix

The Nutanix Cloud Platform solution is a hybrid multicloud platform that provides a software stack to enable the full lifecycle of AI/ML applications. The consistent operating model helps enable ease of management whether data is being collected or fed into models for inferencing at the edge, or models are being fine-tuned at the core data center or in the public cloud.

Nutanix Files Storage provides distributed, scale-out file storage on the Nutanix Cloud Platform, with NFS and SMB support. With integrated cybersecurity and ransomware protection, Nutanix Files Storage provides high performance and low latency, along with native snapshots and disaster recovery.

Nutanix Objects Storage provides distributed, scale-out S3-compatible object storage on the Nutanix Cloud Platform. With integrated cyber resilience, encryption, and replication, Nutanix provides high-performance Objects Storage for cloud-native, analytics, AI/ML, and archive applications.

Nutanix is pleased to announce the first set of published results with the MLPerf storage benchmark.

Here are some highlights:

  • We achieved 65 accelerators and a throughput of 25 GB/s with Unet3d ML training workload using five Nutanix NX-8170-G8 nodes with Nutanix Files Storage, using the standard NFS protocol
  • We achieved 32 accelerators and a throughput of 13 GB/s with Unet3d ML training workload using four Nutanix NX-8150-G8 nodes with Nutanix Objects Storage, using the standard S3 protocol
  • Delivered in our standard software-defined solution running Files Storage Version 4.3 on the Nutanix Cloud Platform running AOS Version 6.7 and AHV 9 hypervisor, and Objects Storage Version 4.0 on the Nutanix Cloud Platform running AOS Version 6.6.2.6 and AHV 9 hypervisor.

The benchmark was executed in a lab environment using conventional dual port 100Gb/s data center networking infrastructure. The setup could be expanded to up to 16 nodes to service more accelerators.

Weka

WEKA is a data platform software provider for AI and other performance-intensive workloads.

The WEKA Data Platform is purpose-built for modern data stacks in the cloud and AI era. It transforms stagnant data silos into dynamic data pipelines that power GPUs efficiently and fuel AI, ML, and HPC workloads seamlessly and sustainably. Its advanced, cloud-native architecture is optimized to solve complex data challenges at scale, delivering 10-100x performance improvements, whether running on-premises, in the cloud, at the edge and in hybrid and multicloud environments. The WEKA platform’s capability to support all types of IOs, reads/writes, small/large in a low latency fashion with massive metadata performance is the foundation for these performance accelerations, as well as its multiprotocol support, which eliminates multiple copying of the data throughout the AI pipelines. Additionally, the WEKA platform requires zero tuning and is auto-tuned for any IO pattern, allowing a mix of multiple AI IO patterns on the same datasets in the same filesystems, and scales linearly with the number of compute and storage instances for additional capacity and performance.

WEKA’s initial effort with this benchmark on a single host generated an industry-leading 7.3 GB/s serving 20 accelerators with the UNET3D model while being able to serve 24 accelerators for the IO-intensive BERT model with a throughput of 2.8MB/s). This single-client performance test was limited by the core and networking capabilities of the available client gear — a client with more cores could have supported more accelerators.

WEKA is committed to supporting the MLCommons Storage benchmark as it develops and looks forward to providing more extensive scale-out submissions in the future.

Related

Stability AI just unveiled a text-to-music generator, and you can try it. Here’s how

waves-gettyimages-184286397

AI chatbots and image generators are all the craze, and soon, AI audio generators may join that list.

On Wednesday, Stability AI released Stable Audio, which uses generative AI techniques to create music, sound effects, and audio up to 90 seconds long based on user prompts. Stable Audio's release follows similar efforts from other industry giants like Google and Meta.

Also: Adobe Firefly, now out of beta, boasts fix for DALL-E's drawbacks

Users can test out the technology themselves by using the basic free version, which enables the generation of up to 20 45-second non-commercial tracks a month.

If you plan on using this tool as part of your workflow, consider investing in the Pro subscription, which delivers up to 500 90-second long tracks per month with a commercial license.

Stability AI said that the AI tool will be especially useful for musicians who can use it to create samples to incorporate into their music.

"Our hope is that Stable Audio will empower music enthusiasts and creative professionals to generate new content with the help of AI, and we look forward to the endless innovations it will inspire," said Emad Mostaque, CEO of Stability AI.

I tried the tool for myself, and here's how it went.

How to access Stable Audio

To try it out for yourself, all you need to do is visit the Stable Audio website and create or sign into an account. Because the tool is free, a credit card is not required.

Currently, because of the high interest in — and traffic volume to — the application, you may be greeted by an error message or blank page when going to the site.

Also: How Apple weaved AI into the iPhone 15 and Apple Watch 9

Stability AI acknowledged the issue on X (formerly Twitter), noting that the high demand pushed their servers to total capacity. If you can't get in, they recommend you try again in 24 hours.

When I was met with the error message, I just kept trying to access the site on different incognito browsers, and eventually, it went through. I don't know if that's a foolproof approach but it's worth a try since it worked for me.

Generating Music

Once you are in, all you need to do is type in a prompt for the sound you'd like to be generated, and after about 15 seconds, you can expect your result.

For writing the prompt, you can include as much detail as you'd like to get the best results, including details such as genre, mood, length, BPM, and more.

Also: The 10 best ChatGPT plugins of 2023 (and how to make the most of them)

Stability AI's example prompts included "Post-Rock, Guitars, Drum Kit, Bass, Strings, Euphoric, Up-Lifting, Moody, Flowing, Raw, Epic, Sentimental, 125 BPM" or something as simple as "car passing by."

I used the prompt "Happy, beachy, mellow country track," as seen by the photo above, and was impressed with the output, which had the exact vibe I was searching for. Once the track is generated, you have the option to download the track as an mp3 file.

Artificial Intelligence

Arm after the IPO

Arm after the IPO Frederic Lardinois @fredericl / 10 hours

“The growth of AI, I believe, is the growth of Arm,” Arm EVP and Chief Commercial Officer Will Abbey told me this morning, minutes before the chip designer’s stock started trading on Nasdaq. While AI may not always be the first thing you think about when you hear about Arm, when I asked Abbey about what’s next for the company, he immediately jumped to AI. “When I think about what’s next, firstly, AI runs on Arm today — and AI is everywhere. So when I think about what’s next, what that means for us is as AI continues to be more pervasive, I think that demands for more compute, more power efficiency and a software ecosystem that’s relevant for AI will drive us to continue to deliver in those areas of compute, power efficiency and ecosystem.”

He cited Nvidia’s Grace Hopper superchip as an example for this. It features 72 of Arm’s Neoverse core, combined with Nvidia’s H100 Tensor Core GPU. It was, of course, this kind of synergy that led Nvidia to try to acquire Arm, even though the deal later fell through due to regulatory concerns. “We believe that whether it’s training which is taking place today, which will lead to inferencing downstream, the ARM architecture is ideal to enable AI at scale,” said Abbey.

Arm Holdings CEO Rene Haas poses for a photo with members of leadership outside of the Nasdaq MarketSite on September 14, 2023 in New York City. Arm, the chip design firm that supplies core technology to companies that include Apple and Nvidia, priced its initial public offering at $51 a share. (Photo by Michael M. Santiago/Getty Images)

With well over 250 billion Arm-based chips shipped so far, according to the company’s own data, the company obviously has a massive install base on the hardware side, but one topic Abbey came back to repeatedly during our conversation was the importance of a software ecosystem around these chips as a differentiator between Arm and its competitors.

That’s something the company has invested in heavily in recent years. Indeed, in its latest factsheet, Arm highlights that it invested 10 million engineering hours to create the base software and tools for chips with its Armv8 processors, it invested 30 million hours for the software tooling around its Armv9 chips.

“In the markets that we are focusing on, it’s not just about a hardware solution,” he said. “It’s the ability to make sure that the developers can access your architecture in an easy and accessible way. It’s the developer community which is critical. When you develop an application, you want to make sure that that application runs on as many devices as possible. For us, our ability to make sure that we’re developing best-in-class products, but also ensuring that developers can access us in an easy and accessible way — the moat between what we do and others is pretty, pretty huge.”

Talking about hiring, Abbey noted that this was never really a problem for Arm, but that the IPO would “elevate ARM to its rightful position as a tech employer” since it creates “a bigger shopping window for engineers to want to join Arm.”

“We are focusing more on software, we have 15 million software developers worldwide [in the ecosystem] and that’s an area that we’re going to continue to invest in. I think AI is going to create future opportunities for us and that will place future demands on Arm to build the right products and staff up appropriately.”

And while he does think that the IPO will create an opportunity for Arm to attract more talent, as for the IPO itself, Abbey echoed the familiar line that most executives use on an IPO day: “It’s just a moment in time — and it’s an important moment. It is a recognition of the pervasiveness of what we bring to market. We’re going to continue to invest in the three areas of power efficiency, ultimate performance and an ecosystem. The IPO doesn’t change our trajectory from an investment perspective.”

Could Arm be worth more than $51B?