Shanghai AI Lab Develops First Version of Karpathy’s AI Operating System

Shanghai AI Lab has introduced the first version of an AI Operating System inspired by Karpathy’s innovative model. The system, named FRIDAY (Fully Responsive Intelligence, Devoted to Assisting You) is an OS-Copilot which serves as a versatile agent, developed through a combination of Python code and GPT-4 language model prompts.

Shanghai 🇨🇳 AI Lab Achieved 1st Version of Karpathy's AI Operating System
> Nov 2023: @karpathy proposes LLM OS
> Feb 2024: Chinese (+ Princeton) team proposes self learning operating system
What did they do?
> Built an agent using a mix of Python code and GPT-4 language model… pic.twitter.com/1aJaCjwPgN

— Ate-a-Pi (@8teAPi) February 14, 2024

FRIDAY takes charge of Linux or Mac OS computers, navigating through applications like browsers, Excel, and PowerPoint to perform tasks. Notably, the system possesses the ability to self-improve, utilising GPT-4 as its underlying large language model.

Comprising several key components, FRIDAY exhibits a comprehensive approach to task execution. The Planner segment efficiently decomposes user requests into manageable tasks, while the Configurator acts as middleware, incorporating data from memory or tool repositories before passing tasks to the Executor.

The Declarative Memory holds user profiles and action histories, while the Tool Repository provides a library of available tools. The Working Memory keeps track of task progress and previous history, and the Executor generates executable commands. Lastly, the Critic assesses task completion, determining success or the need for iteration.

FRIDAY outperformed GPT-4 on benchmarks for web retrieval, Excel and Powerpoint usage. Agents like FRIDAY functioning as universal user interfaces, could potentially reshape the landscape of human-computer interactions.

Meanwhile Andrej Karpathy recently left OpenAI. Karpathy confirmed his departure from OpenAI in a post on X saying that it is not because of any other reason apart from his plan to work on personal projects.

The post Shanghai AI Lab Develops First Version of Karpathy’s AI Operating System appeared first on Analytics India Magazine.

FTC seeks to modify rule to combat deepfakes

FTC seeks to modify rule to combat deepfakes Kyle Wiggers 9 hours

Spurred by the growing threat of deepfakes, the FTC is seeking to modify an existing rule that bans the impersonation of businesses or government agencies to cover all consumers.

The revised rule — depending on the final language, and the public comments that the FTC receives — might also make it unlawful for a GenAI platform to provide goods or services that they know or have reason to know are being used to harm consumers through impersonation.

“Fraudsters are using AI tools to impersonate individuals with eerie precision and at a much wider scale,” FTC chair Lina Khan said in a press release. “With voice cloning and other AI-driven scams on the rise, protecting Americans from impersonator fraud is more critical than ever. Our proposed expansions to the final impersonation rule would do just that, strengthening the FTC’s toolkit to address AI-enabled scams impersonating individuals.”

1. Fraudsters are using voice cloning & other AI tools to impersonate individuals with eerie precision and at scale. @FTC proposes to expand its impersonation rule to cover impersonation of individuals, so these fraudsters would pay hefty penalties.https://t.co/8ON0G63ZjL

— Lina Khan (@linakhanFTC) February 15, 2024

It’s not just folks like Taylor Swift who have to worry about deepfakes. Online romance scams involving deepfakes are on the rise. And scammers are impersonating employees to extract cash from corporations.

In a recent poll from YouGov, 85% of Americans said they were very concerned or somewhat concerned about the spread of misleading video and audio deepfakes. A separate survey from The Associated Press-NORC Center for Public Affairs Research found that nearly 60% of adults think AI tools will increase the spread of false and misleading information during the 2024 U.S. election cycle.

Last week, my colleague Devin Coldewey covered the FCC’s move to make AI-voiced robocalls illegal by reinterpreting an existing rule that prohibits artificial and pre-recorded message spam. Timely in light of a phone campaign that employed a deepfaked President Biden to deter New Hampshire citizens from voting, the rule change — and the FTC’s step today — are the current extent of the federal government’s fight against deepfakes and deepfaking technology.

No federal law squarely bans deepfakes. High-profile victims like celebrities can theoretically turn to more traditional existing legal remedies to fight back, including copyright law, likeness rights and torts (e.g. invasion of privacy, intentional infliction of emotional distress). But these patchwork laws can be time-consuming — and laborious — to litigate.

In the absence of congressional action, 10 states around the country have enacted statutes criminalizing deepfakes — albeit mostly non-consensual porn. No doubt, we’ll see those laws amended to encompass a wider array of deepfakes — and more state-level laws passed — as deepfake-generating tools grow increasingly sophisticated. (Case in point, Minnesota’s law already targets deepfakes used in political campaigning.)

Amazon Demos the Largest text-to-speech AI Model,  Big Adaptive Streamable TTS with Emergent Abilities

Amazon shared BASE TTS, a text-to-speech model. It was trained on 100,000 hours of public domain speech data, mainly in English but also including German, Dutch, and Spanish, making it a new standard for natural speech.

The model uses a 1-billion-parameter Transformer and a convolution-based decoder for efficient text-to-speech conversion. This model introduces a new approach for analysing speech so as to distinguish between different voices. It also employs a technique called byte-pair encoding to reduce the size of the speech data to enhances the model’s efficiency and speed in processing and generating speech.

BASE TTS shows new or ‘emergent’ capabilities as it’s trained with more data. With over 10,000 hours of training, it understands text better, allowing it to produce speech that sounds right for the context. The model can also handle complex language features like compound nouns and emotional expressions, showing its versatility.

An example provided by the paper, ‘In the classroom, filled with the chatter of students sharing their holiday stories and the rustling of new textbooks, Mrs. Thompson, excited to embark on a new academic year, prepared a lesson that would challenge and inspire her students.’

The development of BASE TTS was developed from the idea that larger text-to-speech systems would get better with scale. BASE TTS not only has high-quality speech but also shows new skills, like pronouncing difficult texts correctly and using the right emotional tone. It performs better than other large text-to-speech systems, making it a leading model.

Another example where the audio changes the tone and whispers for the sentence, ‘A profound sense of realisation washed over Matty as he whispered, “You’ve been there for me all along, haven’t you? I never truly appreciated you until now.”’

BASE TTS could improve user experiences and help languages with few resources. It can mimic speaker characteristics with little reference audio, offering new ways to create synthetic voices for people who cannot speak. Amazon decided not to share BASE TTS openly to avoid misuse, highlighting ethical considerations in using advanced AI.

These capabilities which eluded speech models until now seems possible as demonstrated by BASE TTS. The research team also highlights the importance of diverse speech data in representing different languages, ethnicities, dialects, and genders. They call for more research on how data affects the model and ways to make voice technology more inclusive.

Another similar model is MetaVoice, an open source 1.2B parameter foundational model for TTS.

The post Amazon Demos the Largest text-to-speech AI Model, Big Adaptive Streamable TTS with Emergent Abilities appeared first on Analytics India Magazine.

How to use Copilot Pro to write, edit, and analyze your Word documents

Word

Microsoft's Copilot Pro AI offers a few benefits for $20 per month. But the most helpful one is the AI-powered integration with the different Microsoft 365 apps. For those of you who use Microsoft Word, for instance, Copilot Pro can help you write and revise your text, provide summaries of your documents, and answer questions about any document.

First, you'll need a subscription to either Microsoft 365 Personal or Family. Priced at $70 per year, the Personal edition is geared for one individual signed into as many as five devices. At $100 per year, the Family edition is aimed at up to six people on as many as five devices. The core apps in the suite include Word, Excel, PowerPoint, Outlook, and OneNote.

Also: Microsoft Copilot vs. Copilot Pro: Is the subscription fee worth it?

Second, you'll need the subscription to Copilot Pro if you don't already have one. To sign up, head to the Copilot Pro website. Click the Get Copilot Pro button. Confirm the subscription and the payment. The next time you use Copilot on the website, in Windows, or with the mobile apps, the Pro version will be in effect.

How to use Copilot Pro in Word

Open Word.

Submit your request.

Review the response and your options.

Keep, regenerate, or remove the draft.

Alter the draft.

Revise existing text.

Review the different versions.

Replace or Insert.

Adjust the tone.

Turn text into a table.

Respond to the table.

Summarize a document.

Review the summary.

Revise the summary.

Ask questions about a document.

More how-tos

Jensen Huang & His Newfound Obsession

Jensen Huang, the CEO of the world’s third-most valued company, has been on a side quest to spread one meaningful message: Every nation, regardless of its size or resources, must develop its own ‘sovereign AI’.

The man in black has travelled over a dozen geographies, right from the Indian subcontinent to his latest stop, the UAE, for the recent World Government Summit in Dubai.

As Huang envisions it, sovereign AI isn’t just about having some fancy algorithms humming away in a server room. It’s about a fundamental shift in power, reclaiming national autonomy in generative AI. It’s about ensuring that the decisions made by AI, which will increasingly impact everything from healthcare to defence, reflect the priorities of each individual nation.

Huang is basically helping countries take control of how they can shape and build AI to fit their needs. To make sure that when AI makes decisions about things like healthcare and defence, those decisions match the values and priorities of each country. He believes AI could change not just how we live but also the character of different nations.

Money Matters

The trillion-dollar company offering graphic and digital media processors has clearly outlined plans to invest in and scale the operations of the countries Huang is visiting. Last year, NVIDIA signed several deals with those states.

In India, the company has joined hands with national conglomerates Reliance and Tata Group to build AI computing infrastructure and platforms powerful than the fastest supercomputer in India today. Interestingly, both the partnerships were announced on the same day.

Reliance will be provided with the latest tech to further develop the local infrastructure. On the other hand, TCS will upskill its 600,000-strong workforce apart from the technical aspect of the partnership. Huang’s efforts show that the AI visionary wants to help countries worldwide develop AI at the grassroot levels.

Even when he met Omar Al Olama, the Middle Eastern state’s minister of AI, Huang focused on explaining how AI can be trained on local data to protect cultural identities. “It’s not that costly; it is also not that hard,” Huang said. “The first thing that I would do, of course, is codify the language – the data of your culture into your own large language model.”

Over the past six months, the chief of the most powerful company in tech right now has also visited Canada, France, Japan, Malaysia, Singapore and Vietnam distributing similar wise words. By building locally (with NVIDIA), these nations can boost their economies while ensuring national security.

Empowering Competition

Huang’s message resonates with world leaders since large language models and other recent AI developments have kickstarted a competition not just among tech companies, but also nations. China has been ambitiously open-sourcing its contenders, India is going all in, and UAE is betting big on generative AI, too.

Other leaders in the field, too, are noticing the stark difference between generative AI and the technologies that prevailed earlier. “The thing that’s different this time”, continued Sri Elaprolu, global head of AWS Generative AI Innovation Center, “is that it has started to be hot across and not just one active geography. We’re supporting customers across all areas, including Latin America, Africa, and the Middle East — locations we normally don’t see a jump right into emerging tech.”

From nations large to small, Huang, the Midas of AI, is making sure no one is left untouched by his golden touch. His call is not merely a suggestion, but has a greater purpose of building technology from scratch without anyone being left behind.

We have seen technologies in the past get easily monopolised, causing economic disparity and even putting the security of nations at risk. But Huang wants to lay a better foundation for generative AI and is clearly (and literally) going miles to do so.

Where do you think Huang is headed next to spread the word about sovereign AI?

The post Jensen Huang & His Newfound Obsession appeared first on Analytics India Magazine.

OpenAI’s newest model Sora can generate videos — and they look decent

OpenAI’s newest model Sora can generate videos — and they look decent Kyle Wiggers 11 hours

OpenAI, following in the footsteps of startups like Runway and tech giants like Google and Meta, is getting into video generation.

OpenAI today unveiled Sora, a generative AI model that creates video from text. Given a brief — or detailed — description or a still image, Sora can generate 1080p movie-like scenes with multiple characters, different types of motion and background details, OpenAI claims.

Sora can also “extend” existing video clips — doing its best to fill in the missing details.

“Sora has a deep understanding of language, enabling it to accurately interpret prompts and generate compelling characters that express vibrant emotions,” OpenAI writes in a blog post. “The model understands not only what the user has asked for in the prompt, but also how those things exist in the physical world.”

Now, there’s a lot of bombast in OpenAI’s demo page for Sora — the above statement being an example. But the cherry-picked samples from the model do look rather impressive, at least compared to the other text-to-video technologies we’ve seen.

For starters, Sora can generate videos in a range of styles (e.g., photorealistic, animated, black and white) up to a minute long — far longer than most text-to-video models. And these videos maintain reasonable coherence in the sense that they don’t always succumb to what I like to call “AI weirdness,” like objects moving in physically impossible directions.

Check out this tour of an art gallery, all generated by Sora (ignore the graininess — compression from my video-GIF conversion tool):

OpenAI Sora

Image Credits: OpenAI

Or this animation of a flower blooming:

OpenAI Sora

Image Credits: OpenAI

I will say that some of Sora’s videos with a humanoid subject — a robot standing against a cityscape, for example, or a person walking down a snowy path — have a video game-y quality to them, perhaps because there’s not a lot going on in the background. AI weirdness manages to creep into many clips besides, like cars driving in one direction, then suddenly reversing or arms melting into a duvet cover.

OpenAI Sora

Image Credits: OpenAI

OpenAI — for all its superlatives — acknowledges the model isn’t perfect. It writes:

“[Sora] may struggle with accurately simulating the physics of a complex scene, and may not understand specific instances of cause and effect. For example, a person might take a bite out of a cookie, but afterward, the cookie may not have a bite mark. The model may also confuse spatial details of a prompt, for example, mixing up left and right, and may struggle with precise descriptions of events that take place over time, like following a specific camera trajectory.”

OpenAI’s very much positioning Sora as a research preview, revealing little about what data was used to train the model (short of ~10,000 hours of “high-quality” video) and refraining from making Sora generally available. Its rationale is the potential for abuse; OpenAI correctly points out that bad actors could misuse a model like Sora in myriad ways.

OpenAI says it’s working with experts to probe the model for exploits and building tools to detect whether a video was generated by Sora. The company also says that, should it choose to build the model into a public-facing product, it’ll ensure that provenance metadata is included in the generated outputs.

“We’ll be engaging policymakers, educators and artists around the world to understand their concerns and to identify positive use cases for this new technology,” OpenAI writes. “Despite extensive research and testing, we cannot predict all of the beneficial ways people will use our technology, nor all the ways people will abuse it. That’s why we believe that learning from real-world use is a critical component of creating and releasing increasingly safe AI systems over time.”

Nearly 75% Enterprises Pivoted to Text-Based Generative AI to Improve Operational Efficiency

According to the recent “Generative AI in CXM Survey Report” by Everest Group and WNS, 75% of enterprises are elevating their business strategies by piloting, deploying, or scaling up text-based generative AI solutions, followed by 62% for code generation and 52% for image generation.

The potential of this is widely recognised, with more than 90% of enterprises believing in the high potential of text generation, and around 70% for code and image generation. In terms of deployment, about 75% are piloting, deploying, or scaling up text generation solutions, and nearly 60% are doing the same for code generation.

Growth Drivers

The generative AI adoption in customer experience (CX) operations is driven by the need to enhance customer satisfaction and operational efficiency. Enterprises are increasingly integrating tech into CX management to personalise and customise customer interactions by understanding individual preferences and behaviors. This leads to more engaging and gratifying customer experiences.

“When generative AI tailors interactions to individual preferences, it improves operational efficiency by automating tasks and intelligent tools like agent assist and language translation. This helps improve productivity, reduces response time, and lets agents focus on value-added tasks, ultimately resulting in more satisfying customer experiences,” Sanjay Jain, chief business transformation officer, WNS, told AIM.

Furthermore, it boosts efficiency within CX operations. It equips customer support agents with intelligent tools such as agent assist, next-best-action recommendations, language translation, cross-sell, up-sell, and accent neutralisation.

Another driver is its ability to analyse vast amounts of customer data, enabling enterprises to extract meaningful insights for competitive advantage. This facilitates data-driven decision-making for strategy development, product improvements, and service enhancements.

However, 45% of enterprises report a shortage of internal technical expertise in generative AI which hinders their ability to come up with new models.

Talking about the AI talent gap, Jain said thatinvesting in strategic talent acquisition and development is crucial for maximising the potential of AI in CXM.

“Enterprises need to focus on recruiting AI and ML engineers, data scientists, and software developers, especially those with a mix of AI and software development skills to enhance teamwork and communication. Additionally, ongoing training programs are vital to ensure that current employees are upskilled,” said Jain.

He also added that AI, despite fears of job displacement, acts as an enhancer rather than a substitute for human skills. “While AI excels in accuracy and speed for execution, human critical thinking is pivotal for oversight,” added Jain, highlighting that this symbiotic relationship promises long-term benefits.

Key Findings of the Report

The survey highlights the expanding role of tools like OpenAI’s ChatGPT, Google’s Bard Gemini, and Microsoft’s Copilot in business transformation. These advancements are driving a significant shift in CXM, with generative AI’s content generation, data analysis, and insight extraction capabilities leading the way.

Key findings from the survey cover the awareness, perceived potential, and deployment areas for generative AI applications. It also covers the technology, process and enterprise readiness for foundational models.

Everest Group’s research included over 200 companies from diverse sectors such as telecom, media, BFSI, healthcare, retail, and technology, including FGT/Hi-Tech industries. These companies, predominantly from regions including Asia Pacific, have annual revenues exceeding $500 million.

Awareness and perceived potential of generative AI applications

  • Text generation capabilities: Over 75% of enterprises have high awareness.
  • Code generation capabilities: 62% awareness.
  • Image generation capabilities: 52% awareness.
  • Over 90% of enterprises believe in high potential for text generation.
  • About 70% see high potential for code and image generation capabilities.

Planned Deployment Areas

  • Text generation capabilities: About 75% of enterprises actively piloting, deploying, or scaling up solutions.
  • Code generation capabilities: Nearly 60% of enterprises have initiated pilot programs or implementation.

Readiness for Generative AI

  • Nearly 50% of enterprises are concerned about having sufficient computing power.
  • Over 70% say they have adequate cloud capabilities.
  • Approximately 40% express concerns about the availability of high-quality training data.
  • Over 60% express significant concerns regarding data security.
  • More than 45% of enterprises report a shortage of internal technical expertise.
  • Over 70% of BFSI and healthcare sectors noted regulatory compliance issues.
  • 40% identify cultural inertia as a major obstacle.
  • Only two-thirds have adequate capability for redundancy and failover measures.

Enterprise Readiness for generative AI by Industry

  • BFSI: Approximately 60% prepared across technology, people, process, and change management.
  • Healthcare: 60-70% readiness across technology, data, process, and change management.
  • Retail: Least prepared, with key challenges in computing power, training data, and talent.
  • Technology and FGT/Hi-Tech: Over 60% significantly ready across key parameters.
  • Telecom and media: Approximately 65% highly ready across favorable parameters for generative AI implementation.

Read more: Data Science Hiring Process at WNS

The post Nearly 75% Enterprises Pivoted to Text-Based Generative AI to Improve Operational Efficiency appeared first on Analytics India Magazine.

CIOs: Consider 3 Best Practices When Using Emerging Technology to Achieve Business Results

It’s a great time to be a CIO. Business has never been so focused on using technology to drive growth. And technology has never been so accessible. The variety of technology options and faster delivery methods (thank you, cloud) have streamlined the process of identifying, buying and deploying technology and services. Generative AI is only one — granted, a big one — example of this.

But while the increased interest in new and emerging technology from business partners is exciting, it does bring with it some risk and could change how IT works with its business partners. Understanding those risks and guiding business leaders on how best to approach them is critical for CIOs when shaping the IT-business partnership.

Understand the risks of emerging technology

Did you ever think you would see so much interest in AI? What used to be a complex and mysterious technology to the average employee is now showing up on the nightly news and mentioned in nearly every meeting you attend. Generative AI has made the technology accessible and useful to a wide variety of business roles — from the CEO to across the entire organization. Nearly every function and role can see some benefit from genAI.

According to Forrester, 89% of AI decision-makers said their organization is expanding, experimenting with or exploring the use of genAI. But beyond those approved use cases, many workers are using ChatGPT or other non-sanctioned genAI tools in the office — in fact, Forrester recently coined the term BYOAI for bring-your-own-AI and predicts that 60% of workers will use their own AI to do their jobs in 2024.

That’s a lot of shadow AI, and it brings both privacy and brand concerns (i.e., What are employees putting into those genAI tools, and what proprietary IP might your firm be losing?). Unfortunately, the IT organization cannot manage technology it doesn’t know about. And the barrage of pitches we all receive from technology vendors and service providers can spark some of that shadow AI. Vendors will often say they have a technology solution for a specific process, but the reality is that you may have to fix the process before you can successfully apply a tech solution to it. Another issue that can result from this trend is overlapping functionality, as each of the various software platforms provides its own genAI capability.

DOWNLOAD: This customizable Shadow IT Policy from TechRepublic Premium

When evaluating emerging tech, there must be a balance. You want to be innovative and make the most of the right technologies, but avoid being distracted by “shiny object syndrome.” Given that applying the wrong technology to business processes or use cases can be a disastrous waste of time and money, this can have the opposite effect as intended — slowing productivity, frustrating workers and creating more technical debt.

And that’s where the IT organization can bring real value in today’s environment.

Use three best practices when relying on emerging tech to achieve business outcomes

As more technology spending is initiated outside the IT organization, it is changing the role of both the CIO and IT teams. As CIOs step into the new role of trusted advisor to their business partners, they are helping them understand new and emerging technologies and providing guidance on how they can best leverage those technologies to solve real business challenges. This culminates in three key best practices.

1. Tie technology to a business case

Whether it’s a technology project initiated outside of IT or within the IT organization, the decision should start with these questions:

  • What business problem are we trying to solve?
  • What is the business use case for this technology?
  • Does this technology align with and propel our business initiatives forward?

This is where IT-business alignment is critical. When IT is aligned with the business at all levels — understanding its goals and desired outcomes — the technology evaluation process becomes more streamlined, and there is clarity about the right choice for each use case. When IT is not aligned with the business objectives or less engaged with business partners, there is a higher risk of deploying technology for technology’s sake and building up technical debt.

2. Align IT members with specific business functions

How do you move “business and IT alignment” from a bullet on a slide to an action in your organization? To use Forrester as an example, we aligned IT team members to specific business functions so they could develop expertise in that function’s objectives and document vital business processes. When a technology evaluation discussion arises in that function, the business is more likely to consult with IT because there is a higher level of trust and familiarity between the two organizations.

3. Streamline access to strategies with a plan-on-a-page approach

Forrester’s plan-on-a-page approach is a best-in-class method for ensuring strategy alignment. Having direct visibility into the most important objectives for our various business partners allows us to align our work and budgets with those objectives. Without that level of visibility, there is a risk that non-prioritized business projects could absorb vital IT resources and budgets.

Conclusion

Ultimately, if you reach a point where IT becomes an enabler for the business to achieve its most critical goals, the risk of the business veering off course due to shiny object syndrome becomes much lower. Additionally, business leaders will have the confidence to consult IT first and get their input before pursuing emerging tech solutions.

Profile photo of Mike Kasparian.
Mike Kasparian. Image: Forrester

This article was written by Mike Kasparian, who serves as chief information officer at Forrester, leading a global IT organization that manages all internal business technology. He is responsible for leveraging technology to drive productivity and deliver business value, ensuring network and data security, and optimizing technology operations.

Prior to this role, Mike held various positions across Forrester’s organization over the past 18 years, including vice president of business technology strategy and vice president of product for research and analytics. Mike also worked in product marketing, operations, and was part of the team that launched the project consulting business. He began his career at Forrester in sales working with the company’s largest global accounts.

Mike is a summa cum laude graduate of the University of Massachusetts, Amherst.

Meet the Creator of Microsoft Phi-2

Meet the Creator of Microsoft Phi-2

A few days back, Microsoft released a blog on ‘Three Big AI Trends to Watch in 2024’, which highlighted the impact of small language models (SLMs) such as Orca and Phi, along with multimodal AI, and AI in science. AIM got in touch with Harkirat Behl, a senior researcher in the Physics of AGI team of Microsoft Research, who is also one of the creators of Phi-1, Phi-1.5, and Phi-2.

Behl said that his team is currently working on the next version of Phi-2 and making it more capable. “Phi-1.5 started showing great coding capabilities, Phi-2 was code with common sense abilities, and the next one would be even more capable,” he said.

Behl is also working in video generation as it is one of the newest trends catching up in the AI industry. He recently published a paper called PEEKABOO, which is focused on creating such digital content, highlighting Microsoft’s interest in multimodal AI.

How is Phi-2 better?

“One of the things that makes Phi-2 better than Meta’s Llama 2 7B and other models is that its 2.7 billion parameter size is very well suited for fitting on a phone,” said Behl. He appreciated that a lot of Indic models are currently being built on top of Llama 2 and while he acknowledged Meta’s great work on Llama 2, he also encouraged people to build on top of Phi-2.

“When GPT-3 came out, everyone, including Google, started making big models. But then began the discussions around the scaling laws and how efficient these bigger models are, giving rise to smaller language models for specific tasks,” said Behl.

He said that scaling laws are not necessarily true. “You don’t need a specific size or number of parameters for a model to get good at coding,” said Behl, saying that you do not need large models to instil intelligence. “All you need is a small amount of high quality data, aka textbook quality data.”

Citing Phi-2, Behl said that training models on synthetic data reduces the size of the model, and also brings in a lot of capabilities within them, which is different from how GPT-3 was trained. “Textbooks are written by experts in the field, unlike the internet where anybody can write and post, which is how GPT-3 is trained,” said Behl.

Calls for Indic open source moment

With smaller and open language models such as Meta’s Llama 2 and Microsoft’s Phi-2 performing on par with their larger counterparts on specific tasks and ranking on top of Open LLM Leaderboard, the conversation has completely shifted to building smaller models for specialised use cases and domains for maximum outcomes and efficiency.

“I believe that India should release its own foundational models. If there are models from China on top of the leaderboard, why aren’t there any from India?” said Behl, adding that Indian tech companies should focus on this and also partner with academic institutions to build models.

Behl emphasised that centralisation of compute would solve a lot of the issues with building models in India. “I think all IITs should come together and fund the resources. This way we would have enough resources for building a foundational model,” said Behl, giving an example of how the UK government did a similar thing for building AI capabilities.

“It is necessary to build Indic and local language models as everyone would be able to use them. At the same time, it would also be great to have an Indian model to compete on benchmarks against other countries. India should definitely do that,” emphasised Behl.

A science lover

Behl started his journey in AI from IIT Kanpur. Even before that he had worked on robotics projects and built an autonomous underwater vehicle that would recognise objects underwater. “My interest was always in computer vision, and that is when I applied for an internship at Oxford,” said Behl, where he collaborated with Google DeepMind to learn about safety in AI. He later went on to do an internship with Microsoft Research.

For almost two years, Behl focused on working on automation and how it can be democratised to everyone. “The scope is that everybody should be able to do [automation], not only machine learning experts,” said Behl. He then worked on training models that could find synthetic data automatically, which was also his focus during PhD.

Apart from working on generative AI and Microsoft, Behl is very keen to learn more about science and how AI is solving a lot of scientific problems. There are a lot of scientific problems that need solving. Behl said that Google’s AlphaFold paper for predicting protein structures was a great example of what AI can do in the scientific world.

The post Meet the Creator of Microsoft Phi-2 appeared first on Analytics India Magazine.

This German nonprofit is building an open voice assistant that anyone can use

This German nonprofit is building an open voice assistant that anyone can use Kyle Wiggers 8 hours

There’s been many attempts at open source AI-powered voice assistants (see Rhasspy, Mycroft and Jasper, to name a few) — all established with the goal of creating privacy-preserving, offline experiences that don’t compromise on functionality. But development’s proven to be extraordinarily slow. That’s because, in addition to all the usual challenges attendant with open source projects, programming an assistant is hard. Tech like Google Assistant, Siri and Alexa have years, if not decades, of R&D behind them — and enormous infrastructure to boot.

But that’s not deterring the folks at Large-scale Artificial Intelligence Open Network (LAION), the German nonprofit responsible for maintaining some of the world’s most popular AI training data sets. This month, LAION announced a new initiative, BUD-E, that seeks to build a “fully open” voice assistant capable of running on consumer hardware.

Why launch a whole new voice assistant project when there’s countless out there in various states of abandonment? Wieland Brendel, a fellow at the Ellis Institute and a contributor to BUD-E, believes there isn’t an open assistant with an architecture extensible enough to take full advantage of emerging GenAI technologies, particularly large language models (LLMs) along the lines of OpenAI’s ChatGPT.

“Most interactions with [assistants] rely on chat interfaces that are rather cumbersome to interact with, [and] the dialogues with those systems feel stilted and unnatural,” Brendel told TechCrunch in an email interview. “Those systems are OK to convey commands to control your music or turn on the light, but they’re not a basis for long and engaging conversations. The goal of BUD-E is to provide the basis for a voice assistant that feels much more natural to humans and that mimics the natural speech patterns of human dialogues and remembers past conversations.”

Brendel added that LAION also wants to ensure that every component of BUD-E can eventually be integrated with apps and services license-free, even commercially — which isn’t necessarily the case for other open assistant efforts.

A collaboration with Ellis Institute in Tübingen, tech consultancy Collabora and the Tübingen AI Center, BUD-E — recursive shorthand for “Buddy for Understanding and Digital Empathy” — has an ambitious roadmap. In a blog post, the LAION team lays out what they hope to accomplish in the next few months, chiefly building “emotional intelligence” into BUD-E and ensuring it can handle conversations involving multiple speakers at once.

“There’s a big need for a well-working natural voice assistant,” Brendel said. “LAION has shown in the past that it’s great at building communities, and the ELLIS Institute Tübingen and the Tübingen AI Center are committed to provide the resources to develop the assistant.”

BUD-E is up and running — you can download and install it today from GitHub on a Ubuntu or Windows PC (macOS is coming) — but it’s very clearly in the early stages.

LAION patched together several open models to assemble an MVP, including Microsoft’s Phi-2 LLM, Columbia’s text-to-speech StyleTTS2 and Nvidia’s FastConformer for speech-to-text. As such, the experience is a bit unoptimized. Getting BUD-E to respond to commands within about 500 milliseconds — in the range of commercial voice assistants such as Google Assistant and Alexa — requires a beefy GPU like Nvidia’s RTX 4090.

Collabora is working pro bono to adapt its open source speech recognition and text-to-speech models, WhisperLive and WhisperSpeech, for BUD-E.

“Building the text-to-speech and speech recognition solutions ourselves means we can customize them to a degree that isn’t possible with closed models exposed through APIs,” Jakub Piotr Cłapa, an AI researcher at Collabora and BUD-E team member, said in an email. “Collabora initially started working on [open assistants] partially because we struggled to find a good text-to-speech solution for an LLM-based voice agent for one of our customers. We decided to join forces with the wider open source community to make our models more widely accessible and useful.”

In the near term, LAION says it’ll work to make BUD-E’s hardware requirements less onerous and reduce the assistant’s latency. A longer-horizon undertaking is building a data set of dialogs to fine-tune BUD-E — as well as a memory mechanism to allow BUD-E to store information from previous conversations and a speech processing pipeline that can keep track of several people talking at once.

I asked the team whether accessibility was a priority, considering speech recognition systems historically haven’t performed well with languages that aren’t English and accents that aren’t Transatlantic. One Stanford study found that speech recognition systems from Amazon, IBM, Google, Microsoft and Apple were almost twice as likely to mishear Black speakers versus white speakers of the same age and gender.

Brendel said that LAION’s not ignoring accessibility — but that it’s not an “immediate focus” for BUD-E.

“The first focus is on really redefining the experience of how we interact with voice assistants before generalizing that experience to more diverse accents and languages,” Brendel said.

To that end, LAION has some pretty out-there ideas for BUD-E, ranging from an animated avatar to personify the assistant to support for analyzing users’ faces through webcams to account for their emotional state.

The ethics of that last bit — facial analysis — are a bit dicey needless to say the least. But Robert Kaczmarczyk, a LAION co-founder, stressed that LAION will remain committed to safety.

“[We] adhere strictly to the safety and ethical guidelines formulated by the EU AI Act,” he told TechCrunch via email — referring to the legal framework governing the sale and use of AI in the EU. The EU AI Act allows European Union member countries to adopt more restrictive rules and safeguards for “high-risk” AI including emotion classifiers.

“This commitment to transparency not only facilitates the early identification and correction of potential biases, but also aids the cause of scientific integrity,” Kaczmarczyk added. “By making our data sets accessible, we enable the broader scientific community to engage in research that upholds the highest standards of reproducibility.”

LAION’s previous work hasn’t been pristine in the ethical sense, and it’s pursuing a somewhat controversial separate project at the moment on emotion detection. But perhaps BUD-E will be different; we’ll have to wait and see.