Top 5 AI Coding Assistants You Must Try

Top 5 AI Coding Assistants You Must Try
Image by Author

AI coding assistants have become an essential part of the development process, as they assist in code generation, comprehension, item search, and performing various tasks using prompts or code. Even cloud IDE platforms like Google Colab and Deepnote offer AI-assisted coding that can help you generate code and resolve issues.

In this post, I will share the top 5 AI coding assistants that are worth checking out. They all come with VSCode extensions and are easy to set up. All you need to do is install them and start experiencing the newer and easier way of developing programs.

1. GitHub Copilot

GitHub Copilot is a tool that uses artificial intelligence to help programmers write code more efficiently. By installing the Copilot extension in VS Code, developers can generate code, learn from the code, autocomplete, and configure their editor.

Top 5 AI Coding Assistants You Must Try
Image from GitHub Copilot

Copilot is a mature product that provides the latest and most accurate suggestions compared to other tools. With the new chat feature, developers can generate, modify, and improve code on the go by using the natural language. Furthermore, the inline chat allows you to generate code directly in the text editor.

The only downside of GitHub Copilot is that it is a paid tool. However, if you are a full-time developer or software engineer, paying $10 per month is a bargain.

2. Codeium

Codeium is a widely known and free tool that has gained popularity recently. It offers most of the features that GitHub Copilot provides, and the best part is, it's free of cost for individuals.

Top 5 AI Coding Assistants You Must Try
Screenshot by Author

You can use Codeium to ask questions related to the file, and it will read it and provide you with context-aware answers. Additionally, you can ask it to refactor, explain, improve, and resolve errors in the code.

It also includes autocomplete, but I'd suggest that you stick to your old Python autocomplete as Codeium's autocomplete isn't always accurate. However, the only downside of Codeium is that it may not always generate the latest version of the code.

3. Cody

Cody is an AI-powered coding assistant that aims to help you write and understand code faster. It offers similar features to Codeium, such as chat, commands, code explanations, and autocomplete. It's available in both free and pro versions.

Top 5 AI Coding Assistants You Must Try
Screenshot by Author

I have been using Cody for nearly two months now and it's been a smooth journey, until I started using it for my data and machine learning projects. Unfortunately, I've noticed that it's not context-aware, and as a result, produces inaccurate code.

However, in my personal experience using Codeium and Cody, I've found that Cody sometimes fails to understand the code and produces inaccurate autocomplete suggestions. That's why I prefer Codeium over Cody.

4. Code GPT

I recently discovered Code GPT: Chat & AI Agents and I was impressed by how it integrates any state-of-the-art large language model and offers a wide range of features. This extension provides AI chat assistance, code explanation, error-checking, auto-completion, and much more. If you have access to OpenAI API or any other platform, you can use this extension for free.

Top 5 AI Coding Assistants You Must Try
Screenshot by Author

I tested it using Google AI, Anthiopic AI, and OpenAI API. Although Claude 2.1 API is fast, I was not impressed with its accuracy. To make it more usable, developers need to work on reducing the lag in autocomplete and fixing the issue of generating wrong answers. One possible solution is to use Codeium for autocomplete and CodeGPT for code generation and issue resolution.

5. Tabnine

Tabnine is an AI coding assistant that can help you speed up and simplify your software development process. It also ensures that your code remains private, secure, and compliant. Tabnine is currently being used by over one million developers across various industries and it has seven million downloads on VSCode.

Top 5 AI Coding Assistants You Must Try
Gif from Tabnine

Although Tabnine's free version is available, it may not be as effective as the Pro version. If you want to experience better coding assistance, it may be worth investing in the Pro version. However, the free version's autocomplete function is still quite fast and accurate.

If you're interested in trying Tabnine, you can take advantage of their 90-day trial period. Keep in mind that you'll need to add your payment details to access the trial.

Conclusion

AI-powered coding assistants are transforming software development by increasing the efficiency and productivity of programmers. In this post, we have covered the top 5 AI coding assistants that I think should be used by anyone who struggles with code logic, formatting, and testing.

Integrating one or more of these assistants into your workflow can boost your productivity, generate and understand the code, resolve issues quicker, and focus more on coding. Ultimately, these AI helpers allow developers to spend less time fighting with code so they can instead create amazing software. Give them a try during your next project.

Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master's degree in Technology Management and a bachelor's degree in Telecommunication Engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.

More On This Topic

  • The 5 Best Vector Databases You Must Try in 2024
  • 5 Must Try Awesome Python Data Visualization Libraries
  • 15 Python Coding Interview Questions You Must Know For Data Science
  • KDnuggets News, May 4: 9 Free Harvard Courses to Learn Data…
  • 7 Must-Know Python Tips for Coding Interviews
  • 5 Amazing & Free LLMs Playgrounds You Need to Try in 2023

The most promising AI smart glasses are from a brand you’ve never heard of

Frame by Brilliant Labs

Brilliant Labs today is announcing the launch of Frame, what the company calls the world's first glasses with an integrated multimodal AI assistant. If your definition of smart glasses includes placing visual overlays in your environment and seeing text floating as you wander about, then the Frame — not the dozens of screen-mirroring glasses — is the one you've been waiting for.

Also: 5 top mobile trends in 2024: On-device AI, the 'new' smartphone, and more

Frame, much like the Humane AI Pin and Rabbit R1, lets you navigate and interact with the physical world through multimodal, generative AI agents. You can ask the always-on AI assistant, "Noa," questions about what's in front of you, how many calories you're about to consume, or what's written on that foreign signage, and it'll find the most appropriate AI model to answer, from GPT-4V for visual-based queries, to Stable Diffusion model for image generation, to Perplexity AI for search.

The Perplexity integration is arguably the most notable AI partnership of the Frame, with its search capabilities rivaling Google and being able to generate fast and reliable results. "On a more technical level, we've tried a lot of stuff, and there's just no one as fast. Speed matters when you're in that moment and only have a few seconds to know about something before moving to the next thing," Bobak Tavangar, CEO of Brilliant Labs, tells me.

The form factor of the Frame is based on retro, "off-beat style" glasses, as Bobak calls it, that the late Steve Jobs, John Lennon, and Gandhi wore. "This (technology) is so new and unfamiliar, so we tried to reference something that looks familiar but at the same time has a lot of lineage to our shared pop culture." Hidden within the lenses is a projector that can beam out text and images to roughly 20 degrees diagonal of field of view. It's more than sufficient to display around the AI output use case and feels like an iPad Pro at arm's length, I'm told. See below for a first look at the Frame glasses.

It's difficult not to draw comparisons between the Frame and the Meta Ray-Ban smart glasses; both wearables are purposed to put an AI assistant on your face, available at all times. But there's one thing the Frame isn't doing, and you may or may not like it: It won't capture and store images and videos. There's only one front-facing sensor on the Frame, and as soon as it captures what's in front of you when answering visual-based questions, like "What would you recommend on this menu for vegans?", the data gets discarded shortly after.

Also: This AI startup made a $199 gadget that replaces apps with 'rabbits' — and it might just work

However, like the Meta Ray-Ban, the Frame is powered by your smartphone and operates on the cellular or Wi-Fi network that it's connected to, meaning the AI wearable is not intended to replace your phone. That also means that the glasses will be just that, a fashion accessory and nothing more, when you're on the subway or in an area with very, very poor signal.

If there's one thing that Bobak wants to leave readers with, it's his belief in creating a device that's open-source and hackable (in the better sense). "It's essential for developers, hackers, artists, mad scientists to sink their teeth into this technology and really understand their implications, for all of us, collectively," he tells me.

The Frame is available for preorder today and retails for $349, with shipments beginning in April. Brilliant Labs has also partnered with AddOptics to provide "precision bonded" prescription lenses for those who need them.

Featured

Crux is building GenAI-powered business intelligence tools

Crux is building GenAI-powered business intelligence tools Kyle Wiggers 9 hours

GenAI might be one of the most exciting technologies today. But it’s also one of the fastest-moving — posing a challenge for enterprises keen to deploy it. With each new GenAI innovation, companies have to worry not only about staying on top of trends but validating what’s working, all while maintaining a semblance of accuracy, compliance and security.

There’s an entire cohort of startups tailoring GenAI tools to fit enterprise needs. Arcee, for instance, is creating solutions to help businesses securely train, benchmark and manage GenAI models, and Articul8 AI, an Intel spin-out, is building algorithm-powered enterprise software.

Now another upstart’s joining the crew.

Himank Jain, Atharva Padhye and Prabhat Singh are the co-founders of Crux, which creates AI models that answer questions about business data in plain language along the lines of OpenAI’s ChatGPT. Padhye explains it thusly:

“With a simple natural language command, executives can get any report, insight, root cause analysis or prediction at their fingertips,” he said via email. “For example, marketers using Hubspot can ask a question like ‘Why are my email campaigns to new users getting low conversions?'”

Crux — which Jain, Padhye and Singh pivoted nearly 15 times before arriving at the platform’s current form — converts schema, the structure of databases, into a “semantic layer” that AI models can understand. Beyond this, Crux allows customers to customize question-answering models to their business intelligence needs, terms and policies, thus improving the quality of the outputs, according to Padhye.

Crux

Crux’s platform for building, customizing and deploying enterprise-focused GenAI models and tools.

Crux leverages a multi-model framework that breaks questions posed by users down into individual parts, distributing these parts among specialized, purpose-built models. For example, one model, which Padhye calls the “clarification agent,” asks follow-up questions to get at a user’s intent.

Crux is deployed on-premise and doesn’t use customer data to train the models, Padhye says.

“Crux aims to launch [an] AI-powered decision-making engine for enterprises,” Padhye said. “Crux is challenging the incumbent business intelligence tools … by iterating faster and rethinking the analytics stack as a decision-making stack.”

Crux makes money by selling subscriptions plus charging setup and maintenance fees that depend on the type and size of a company’s deployment. It’s proven to be a profitable model; Crux’s annual recurring revenue reached $240,000 in four months with a four-company customer base, Padhye claims.

Crux aims to remain small for the foreseeable future, maxing out at a headcount of around 20 people by the end of the year. But with $2.6 million in capital it recently raised from Emergent Ventures, Y Combinator and a number of angel backers, Crux plans to expand “up-market,” Padhye said — focusing on acquiring new enterprise clients.

Added Emergent Ventures’ Anupam Rastogi via email:

“BI and data analytics … is on the brink of a major transformation driven by large language models and advanced AI. This sector is poised for exponential growth over the next decade, transitioning from static dashboards and manual reporting to real-time, accurate and actionable intelligence on demand. Crux has developed a groundbreaking product in this space which helps its customers drive incremental revenues quickly, and is already attracting significant customer interest.”

Google Releases Gemini Ultra, Rebrands Chatbot Bard to Gemini

Google has finally released the highly anticipated Google Gemini Ultra. The version featuring Ultra will be named Gemini Advanced, offering a new, enhanced experience that excels in reasoning, following instructions, coding, and creative collaboration.

Simultaneously, Google announced that Bard will now simply be called Gemini. “It is available in 40 languages on the web and will be coming to a new Gemini app on Android, as well as the Google app on iOS,” wrote Google Chief Sundar Pichai in a blog post.

The capabilities of Gemini Advanced extend to serving as a personal tutor, creating customised step-by-step instructions, sample quizzes, and engaging in back-and-forth discussions tailored to individual learning styles.

For coding enthusiasts, the chatbot proves invaluable in handling advanced coding scenarios, providing a sounding board for ideas, and aiding in the evaluation of different coding approaches.

To access Gemini Advanced, users can subscribe to the Google One AI Premium Plan, priced at $19.99/month, which includes a two-month trial at no cost. This plan provides users with Google’s latest AI advancements along with the benefits of the existing Google One Premium plan, featuring 2TB of storage.

AI Premium subscribers will also soon gain access to Gemini in Gmail, Docs, Slides, Sheets, and more, formerly known as Duet AI. Gemini Advanced is available today in more than 150 countries and territories in English.

Google has also introduced a mobile app for Gemini on Android and iOS.The Gemini app for Android is available on Google Play with support for English, but support for Korean and Japanese will be coming next week.

By downloading the Gemini app or opting in through Google Assistant, users gain access to the new overlay experience, available from the app or wherever Google Assistant is activated, such as through the power button, corner swiping on select phones, or using the “Hey Google” command.

On iOS, access to Gemini will be introduced within the Google app in the coming weeks. Users can simply toggle the Gemini feature and engage in conversations to boost creativity, create custom images, seek assistance in writing social posts, and even plan activities directly from the Google app.

The post Google Releases Gemini Ultra, Rebrands Chatbot Bard to Gemini appeared first on Analytics India Magazine.

India Needs Quality Data for Indic LLMs

The assertion that “AI models are only as good as the data they’re trained on” has been reiterated by experts countless times, and the statement affirms the undeniable truth.

ChatGPT, the popular chatbot by OpenAI, is so impressive because it has been trained on a dataset called WebText2, which is a library of over 45 terabytes of text data which includes articles, blogs, forums, and other publicly available written content.

However, for India to flourish in India, it needs to create AI models that understand India’s linguistic nuances and cultural complexities.

So far, we have already seen a host of Indic large language models (LLMs) come up within a short period of time from Tech Mahindra and Sarvam AI on an enterprise level.

Besides, various Indic adaptations of Meta’s widely used open-source LLM, LLama, have been created for languages like Tamil, Kannada, Telugu, and more. However, the critical challenge lies in the necessity for robust data on the diverse set of languages spoken across the country, a resource currently lacking in building effective Indic LLMs.

The necessity to develop good Indic datasets

India boasts linguistic diversity, encompassing hundreds of languages and thousands of dialects. Generating datasets for such a vast array of languages is a challenging endeavour. Presently, even though datasets are available for the 22 official languages, hundreds of other languages are actively spoken in the country.

“Large language models require very large amounts of high-quality data. For many Indian languages, we do not have this right now,” Pratyush Kumar, co-founder at Sarvam AI and AI4Bharat, an open-source AI research initiative focused on Indian language AI, told AIM.

However, several initiatives are in progress to collect datasets for Indic languages. One such effort is Project Vaani, a collaboration between the Indian Institute of Sciences (IISc) and Google, which focuses on gathering speech datasets to facilitate the training of Indic LLMs.

Additionally, AI4Bharat is also building Indic language datasets like ‘IndicCorp v2, which is the most extensive collection of texts for Indic languages, comprising 20.9 billion tokens, out of which 14.4 billion tokens cover 23 Indic languages. And then there is Bhashini, which is an Indian government’s initiative to create an open-source Indic language dataset.

However, even though these are good datasets, there is a need for more data. “The current volume of this type of data is relatively small; we need to collect even more,” Vivek Raghavan, co-founder of Sarvam AI, told AIM.

Collecting data is not easy

Building datasets remains the most challenging task for those wanting to build Indic LLMs. It will involve digitisation of books, collaborating with linguists, engaging communities, organising content creation workshops, and partnering with local institutions. Surveys, interviews, and transcription services will also play a crucial role.

Tech Mahindra, which recently completed developing its Hindi LLM consisting of 539 million parameters and 10 billion Hindi+ dialect tokens, sent its crew to Northern India to collect data.

“We went to Madhya Pradesh, Rajasthan, and parts of Bihar. The team’s task was to collect Hindi and dialect data by interacting with professors and leveraging the Bhasha-dan portal available on ProjectIndus.in,” Nikhil Malhotra, global head at Makers Lab, Tech Mahindra and the brain behind Project Indus, told AIM.

He also highlighted the initiative to involve Tech Mahindra employees in contributing everyday interactions through sentences like ‘Main ghar se bahar jata hoon’ to the portal to gather diverse linguistic prompts.

Similarly, Swecha Telangana, an open-source advocacy group, recently collaborated with 25-30 colleges and over 10,000 students, who were involved in translation, correction and digitalisation of 40,000-45,000 pages of Telugu folk tales. The dataset was then used to train a 7 billion parameter Telugu small language model (SLM) called ‘AI Chandamama Kathalu’.

Bhashini, too, is making similar efforts to build Indic language datasets. “We lack digital data on low-resource languages, such as Bodo or Sindhi. For these languages, we approach individuals proficient in both Bodo and English,” Amitabh Nag, CEO at Bhashini, told AIM. He explained that they help create a training dataset by providing parallel text in English corresponding to the content written in Bodo.

These efforts stand as a testament to the arduous task ahead for those involved in building good Indic LLMs. Nonetheless, Krutrim, which is an AI startup by Ola founder Bhavesh Aggarwal, claims to have built a model from scratch which understands the 22 official languages of India.

Similarly, earlier this month, Qx Lab AI launched Ask Qx, which also understands 22 official languages. However, neither company has revealed the details of the datasets their models have been trained on.

It’s worth highlighting that ChatGPT, leveraging the capabilities of both GPT-3.5 and GPT-4, exhibits multilingual proficiency, including an understanding of Indic languages like Hindi, Bengali, and Assamese.

However, the chatbot occasionally encounters challenges in distinguishing between Bengali and Assamese. This issue arises due to the quality of datasets on which the models have been trained, particularly for these specific languages.

An ecosystem needs to develop

Kumar believes building good datasets for the 22 official languages will take a reasonable while and, “We should as an ecosystem continue to focus on that because we want to support these languages,” he said.

What India truly requires is a unified and collaborative endeavour, where individuals, organisations, and enterprises join forces, extending support whenever necessary. The encouraging aspect is that a significant portion of the ongoing initiatives is adopting an open-source approach.

This not only promotes transparency and accessibility but also facilitates a collective and inclusive effort towards the common goal of advancing linguistic technologies in the country.

“I believe you will observe a notable shift, particularly in languages spoken more frequently, gaining momentum sooner. However, with time, I am confident that we have the capability to encompass all languages, including their various dialects and nuances,” Raghavan said.

The post India Needs Quality Data for Indic LLMs appeared first on Analytics India Magazine.

Free Data Science Interview Book to Land Your Dream Job

Free Data Science Interview Book to Land Your Dream Job
Image by Author

If you're preparing for data science interviews, you know how overwhelming it can be to go through all the available resources online. One can easily get lost in the details. That's why I'm excited to introduce you to a hidden gem of a resource: "The Data Science Interview Book" by Dip Ranjan Chatterjee.

This freely available web-based book covers all the essential topics you need to know for data science interviews, from statistics and model building to algorithms, neural networks, and business intelligence. But what makes it different from other resources is its focus on providing only the relevant information to get you ready for the interview. This makes it the perfect resource for busy data scientists who need to brush up on a wide range of concepts quickly. Here are a few things that I believe make this book unique:

  1. Real-world interview questions: This book includes real-world interview questions from companies like Google, DoorDash, and Airbnb, along with detailed solutions and case studies.
  2. Updated content: The book is continually updated with new sections, questions, and richer content.
  3. Cheatsheets and references: The book includes cheatsheets for quick reference guides for various topics, as well as additional references for those who want to study topics more deeply.

Overview of Content

Don’t panic if you encounter a section followed by a ?? symbol. This simply indicates that those sections are still being worked on and are subject to change. Here are the major sections covered in this book:

1. Statistics

This section covers the fundamentals of statistics, which are essential for data analysis and model building. Topics include probability basics, probability distributions, central limit theorem, Bayesian vs. frequentist reasoning, hypothesis testing, and A/B testing.

2. Model Building

This section of the book will guide you through the process of creating a successful model, from data gathering to model selection. It also teaches you the data preprocessing techniques essential for any data scientist, including feature scaling, handling outliers, dealing with missing values, and encoding categorical variables. It also has a subsection on hyperparameter optimization and some famous open-source tools used for it.

3. Algorithms

Algorithms are fundamental to data science, and understanding them is crucial for acing a data science interview. This section covers various machine-learning algorithms and also provides you a practical advice on how to choose the right algorithm for your use case. This section starts with the basics of bias-variance tradeoff, and generative vs discriminative models. Then, it proceeds to advanced concepts of regression, classification, clustering, decision trees, random forests, ensemble learning, and boosting. Additionally, the section also discusses time series analysis and anomaly detection. Finally, it concludes with a comprehensive table on Big O analysis, which covers the time and space complexities of different machine learning algorithms.

4. Python

Python is a versatile language used in data science for various tasks. This section has the following sub-sections:

  1. Theoretical: It covers some fundamental concepts in Python such as mesh grid, statistical methods, range vs xrange, switch case, and lambda functions.
  2. Basics: There are some common programming techniques that you must be familiar with to solve Python questions during an interview like lists, tuples, and dictionaries, and understanding control flow using loops and conditionals.
  3. Coding Algorithms from Scratch: Often, companies ask candidates to code algorithms from scratch during a coding demo round. The general steps for coding an algorithm from scratch are discussed here.
  4. Questions: It covers some sample questions related to statistics, data manipulation, and NLP.

5. SQL

In data science interviews, SQL queries are often used to evaluate a candidate's ability to work with data and solve complex problems. This section covers the basics of SQL, including joins, temp tables vs table variables vs CTE, window functions, time functions, stored procedures, indexing, and performance tuning. The Temp Table vs Table Variable vs CTE section explains the differences between these three temporary data structures and when to use each one. You will also learn how to create and use stored procedures. The Performance Tuning section covers various tips to optimize your SQL queries. Overall, it will provide you with a solid foundation in SQL.

6. Analytical Thinking

While the book includes several ongoing sections like Excel, Neural Networks, NLP, Machine Learning Frameworks, Business Intelligence, etc., I'd like to highlight this one specifically. I think it is unique because it covers business scenarios and behavioral management-related questions, which are becoming increasingly important in data science interviews. Companies are not just looking for technical expertise, but also for candidates who can think strategically and communicate effectively.

For example, here is a question that Salesforce asked in one of their interviews:

"As a data scientist at Salesforce, you are speaking with a Product Manager who wants to understand the user base of Salesforce. What would be your approach?"

By going over these scenario-based questions, you will be well-prepared for your interviews.

7. Cheatsheets

Instead of spending hours searching for cheatsheets online, you can find quick and comprehensive guides for topics such as Numpy, Pandas, SQL, statistics, RegEx, Git, PowerBI, Python basics, Keras, and R basics all in one place. These guides are perfect for a quick refresh before an interview or for referencing during a coding challenge.

Conclusion

I completely understand the importance of having a reliable and comprehensive resource to prepare for interviews, and I believe that this book fits the bill. I am sure it will help you succeed. I wish you all the best for your data science preparation journey! In case of any questions, please feel free to reach out to me.

Kanwal Mehreen is an aspiring software developer with a keen interest in data science and applications of AI in medicine. Kanwal was selected as the Google Generation Scholar 2022 for the APAC region. Kanwal loves to share technical knowledge by writing articles on trending topics, and is passionate about improving the representation of women in tech industry.

More On This Topic

  • A Data Science Portfolio That Will Land You The Job
  • Unable to Land a Data Science Job? Here’s Why
  • KDnuggets™ News 22:n05, Feb 2: 7 Steps to Mastering Machine…
  • Data Science Projects That Will Land You The Job in 2022
  • A Data Science Portfolio That Will Land You The Job in 2022
  • 3 Data Science Projects Guaranteed to Land You That Job

Kyndryl Expands Partnership to bring Google Gemini for Generative AI Solutions

Kyndryl, the leading provider of IT infrastructure services globally, today announced an expanded collaboration with Google Cloud aimed at developing responsible generative AI solutions and accelerating adoption by customers.

Since 2021, Kyndryl and Google Cloud have collaborated to facilitate the transformation of global enterprises through the utilization of Google Cloud’s sophisticated AI capabilities and reliable infrastructure. The upcoming stage of this partnership will concentrate on integrating Google Cloud’s internal AI technologies, such as Gemini, its most powerful large language model (LLM), with Kyndryl’s proficiency and managed services to create and implement generative AI solutions for clients.

Kyndryl’s advisory and implementation services will help clients identify optimal generative AI use cases and data foundations, leveraging their expertise in Google Cloud technology for business transformation. Kyndryl also plans to offer its new LLMOps Framework to Google Cloud customers, facilitating responsible and cost-effective solutions for common challenges in generative AI adoption.

Kyndryl intends to utilise the Google Cloud Cortex Framework to enhance the business value derived from customers’ Enterprise Resource Planning (ERP) data on Google Cloud, aiming to enhance productivity, foster innovation, and offer improved business insights, ultimately driving new outcomes for customers.

“Given Kyndryl’s data services expertise, along with our more than 30 years of experience managing large enterprise environments, Kyndryl understands the complexities in moving a generative AI solution from an idea into production. By combining this unique perspective and the Kyndryl Responsible AI Principles with Google’s AI history and in-house generative AI capabilities, we can quickly and responsibly bring this new generation of AI to customers and drive their business value,” said Nicolas Sekkaki, Kyndryl’s Global Applications, Data and AI Practice Leader.

In an earlier exclusive interview with AIM, Naveen Kamat, Vice President & CTO of Data and AI Services at Kyndryl, spoke about how the company is charting its course in IT and AI. “Apart from our skills, experience and resources, we also have been investing into our partner ecosystem, which includes Microsoft, and AWS, apart from data platforms like CloudEra and Databricks. We bring in additional differentiation when we work with clients with our IP, assets and accelerators.”

The post Kyndryl Expands Partnership to bring Google Gemini for Generative AI Solutions appeared first on Analytics India Magazine.

AI is the Future, and the Future of AI is India

The sentiment in our headline was aptly captured by one of India’s seasoned journalists in her recent episode of ‘Vantage with Palki Sharma’, in the backdrop of the latest high-profile visits by Google and Microsoft executives to the country.

Microsoft chief Satya Nadella is currently on his trip to India, and plans to engage with promising AI startups. Nadella has high hopes for the country in terms of its potential contribution to the development of AI. “If the Indian economy is going to be, let’s say, 5 trillion, the AI-driven part of it could be something like, maybe, 10% of it—perhaps 500 billion,” he said.

A week ago, Google hosted its first-ever Research@ Bangalore event, bringing together leading academic researchers, developers, and startups. Dr Jeff Dean, chief scientist of Google DeepMind and Google Research, attended the event in person, along with several other Google leaders, to discuss the future of AI and its implications in India.

“India is well-positioned because you have an incredibly strong culture of respecting engineering, science, and computer science. Many amazing computer scientists have come from India and are studying computer science,” he said.

“The young students that I meet in India are amazing, and they’re all excited and eager to understand the shift from traditional computer science to learning-based approaches for solving all kinds of problems,” added Dean.

While AI startups in the west are looking for young talent, India has no dearth of it. Recently, many local foundational models have popped up in India, and interestingly, they are created by young individuals. For instance, Tamil Llama by Abhinand Balachandran and Kannada Llama by Adarsh Shirawalmath, who is a 2nd-year B.Tech student at Vellore Institute of Technology. Similarly, Telugu Llama was crafted by Ravi Theja, who is under 30.

Moreover, many new startups are emerging in India, with a focus on creating bilingual LLMs catering to the Indian market. One such startup is Sarvam AI, which recently introduced OpenHathi and Microsoft has partnered with them.

“I am excited to see innovative startups in India. I had the chance to meet founders Pratyush and Vivek from Sarvam AI. Pratyush previously worked at Microsoft Research. We are excited to support them and build out their LLM, which is trained in all the Indic languages. The demo they showed me is tremendous,” said Nadella during his visit to Bengaluru.

Microsoft is also backing another AI startup BrainSightAI which uses AI and machine learning to map the human brain which aids thousands of clinicians, including neurosurgeons, psychiatrists, and neurologists, across the country.

Lately, Indian startups are experiencing a funding surge. India witnessed its first AI Unicorn, Krutrim AI, which raised $50 million in funding from Matrix Partners. Meanwhile, Google is looking to invest approximately $4 million in the Indian conversational AI startup CoRover.

Interestingly, nearly 84% of Indian CEOs are raising new capital or reallocating budgets to invest in generative AI, compared to 70% globally.

The Rise of AI Engineers

Indian IT giants have been making every possible effort to train their employees in generative AI. TCS recently announced that it has trained over 1 lakh-strong pool of GenAI ready consultants and prompt engineers who are engaged in hundreds of GenAI projects for clients across segments.

Similarly, last year, Infosys announced plans to train over 1 lakh employees in partnership with NVIDIA & Google Cloud. Wipro, on the other hand, has globally trained 210,000 employees in AI and is integrating generative AI across its platforms.

HCLTech has internally trained close to 42,000 people across various areas of generative AI, while Accenture claims to have trained over 600,000 employees in AI. .

Indian IT firms are seeing a combined ~450 GenAI projects currently in the pipeline.

“Software developers who are already working will want to pick up these (generative AI) kinds of skills if they don’t already have them. This is something that will transform basically every software developer’s workflow and the way they approach problems. This holds true not just in India but all across the world,” said Dean. Similarly, Nadella spoke how ‘copilot’ has increased the productivity of Indian IT employees across Infosys, HCL Tech and MindTree.

‘Unique’ Data

One of the primary attractions for big tech companies in India is its wealth of data. During his visit to India, Nadella praised the work being done by Bhashini and Karya. Karya creates datasets in several Indian languages to train AI models and for research, simultaneously generating job opportunities for Indians, particularly in rural areas.

On the other hand, Bhashini is the Indian government’s initiative to create an open-source Indic language dataset and has also partnered with Karya. AI4Bharat is another initiative from IIT Madras dedicated to developing open-source resources for Indian languages, including datasets, models, and applications.

Google and Microsoft, along with other companies, are actively seeking to acquire both datasets and the companies producing them. Their aim is to ensure that the LLMs they develop surpass the ones currently trained on publicly available internet data.

“The next wave of merger and acquisitions might be led by companies like Google and Microsoft and Facebook (now Meta) looking at these (data) companies saying can they be viable inputs to my large language models or to my other machine learning and AI models” said Chamath Palihapitiya ,Social Capital Founder & CEO.

Moreover, India is generating terabytes of enterprise data monthly due to increased 5G penetration, OTT services, and social media usage. With an abundance of data available in India, there is a high likelihood that the next major hyperscaler could emerge from the country. Recently, laptop maker Lenovo announced its plans to start manufacturing servers locally to support its data center business.

The post AI is the Future, and the Future of AI is India appeared first on Analytics India Magazine.

Microsoft to Extend Shiksha CoPilot to 100 Schools

Microsoft

Microsoft aims to expand the reach of Microsoft Research India’s initiative, the AI copilot, Shiksha CoPilot, to 100 schools by the end of the academic year.

The tech giant announced Shiksha CoPilot in November last year and was being tested in 10 schools in Bengaluru, India.

Shiksha copilot was built on Microsoft Azure OpenAI Service and harnessed Azure Cognitive Services to ingest the content in textbooks, including how the content is organised.

The project, implemented in collaboration with the Sikshana Foundation, an NGO dedicated to enhancing the quality of public education, has been initially deployed at several public schools in Karnataka.

Sikshana found the copilot cut down lesson plan preparation time from an hour and more to just 5 minutes to 15 minutes. A majority of the teachers said they only needed to make minor modifications, if any, to the lesson plans generated.

This initiative is part of Project VeLLM (Universal Empowerment with Large Language Models) at Microsoft Research India, which primarily aims to address the digital divide by overcoming language, income, digital literacy, and information access barriers with Large Langauge Models (LLMs).

The post Microsoft to Extend Shiksha CoPilot to 100 Schools appeared first on Analytics India Magazine.

Meet Lag-Llama, First Open-Source Foundation Model for Time Series Forecasting

A collaborative team from Université de Montréal, the CERC-AAI lab, ServiceNow, and Morgan Stanley has introduced Lag-Llama, an open-source foundation model for time series forecasting.

While foundation models have revolutionised natural language processing and computer vision in recent years, their application to time series forecasting has been notably absent. Lag-Llama aims to bridge this gap by presenting a unique, decoder-only transformer architecture that employs lags as covariates in univariate probabilistic time series forecasting.

Lag-Llama’s pretraining on a diverse corpus of time series data from various domains showcases its remarkable zero-shot generalisation capabilities. The pretraining corpus consists of 27 datasets across six different domains, including energy, transportation, economics, nature, air quality, and cloud operations, with close to 8K univariate time series and 352M tokens.

Its tokenization method uses lag features (specific points in the history) and date-time features (indicating the timestamp). The team has developed stratified batch sampling and normalisation strategies which prove effective during pre training.

Fine-tuning on previously unseen datasets propels Lag-Llama to a state-of-the-art position, surpassing prior deep learning approaches. This solidifies its status as the leading general-purpose model on average, heralding a significant breakthrough in time series forecasting.

Beyond its impact on machine learning, Lag-Llama holds promise for diverse applications in computational medicine, natural sciences, finance, climate, retail, ecology, energy, and more. Its adaptability to scenarios with limited data availability positions it as a crucial asset for applications requiring a transfer from a general-purpose pretrained model.

Forecasts on the unseen datasets closely match the ground truth. pic.twitter.com/vV4rnEN5yd

— Arjun Ashok (@arjunashok37) February 7, 2024

The post Meet Lag-Llama, First Open-Source Foundation Model for Time Series Forecasting appeared first on Analytics India Magazine.