Meet தமிழ் Llama

Meta’s Llama 2 has reached the land of temples. Thanks to Kaggle Master, and a young tech thalaiva Abhinand Balachandran, Llama can now converse in Tamil, marking the introduction of Tamil-Llama (தமிழ் Llama).

In an exclusive interview with AIM Balachandran said that he got the inspiration to build Tamil Llama from another Chinese model called Chinese Llama Alpaca. “Chinese is a bit of a complex language, but if they can make it work for Chinese, then surely we will also be able to make it work for Indian languages, right? So that was the motivation,”said Balachandran.

Balachandran shared that when he began working on Tamil Llama, there weren’t any language models for Indian languages. He started the project for research and later published a paper. “Since the model is still pretty young, it’s not really very, very great or something, but it is good. It can be used as a very good starting point.” added humble Balachandran.

Tamil Llama is trained with an additional 16,000 Tamil tokens, aiming to achieve superior text generation and comprehension in the Tamil language. This model serves as an extension of the LLaMA model and has been enhanced by incorporating extra Tamil tokens, utilising the LoRA methodology for efficient training.

“That step was actually crucial because the original Llama model did not have enough words in its vocabulary to accurately represent or even understand any aspect of common language,” said Balachandran.

Advantage over ChatGPT

Balachandran believes that Tamil Llama could be used in a RAG-based system to incorporate a variety of Tamil books or literature from common sources. “This way, it can be utilised for conversational purposes relevant to the specified period. Moreover, being bilingual, it can serve as a tool for learning English, particularly beneficial for individuals in remote areas lacking access to quality content.” he said.

“It can even generate a bit of code and explain concepts in Tamil. Use Cases that actually illustrate the kind of fine-tuning people are going to do with such models after they come out to the public will be really interesting.” he added.

Speaking of the advantage it has over GPT-4, Balachandran said that fine-tuning for OpenAI’s models is very expensive. Moreover, he mentioned that Tamil Llama is going to be a lot leaner than that. “You could even host it on your own systems, or maybe you could have a contextual version running on your laptop to interact with it for your day-to-day use case,” he said.

“These are 7-billion and 13-billion models, making them accessible to entry-level laptops. Even if you have something like 16 GB of RAM or 8 GB RAM with a dedicated GPU, it can easily run.” he added.

Moreover, Balachandran mentioned that other models available in the market, such as GPT-3.5 and 4, are mostly English-centric even though they can generate text in multiple languages. “However, as robust as these models, like LLaMA and Mistral, might be, their proficiency in generating coherent text in Tamil and several other Indian languages remains noticeably deficient.” he said.

A fundamental limitation lies in their minimal vocabulary of Tamil characters, which is essential for effective text encoding and generation” he added.

Jarvis Powers Tamil Llama

Balachandran said that since the beginning of the project, he tried to keep the cost as low as possible.

To train Tamil Llama, Balachandran mentioned that he collaborated with a startup called Jarvis AI Labs. “There is a startup called Jarvis AI Labs, and they provide GPUs at very low costs,” he said.

He initiated the experimental stage there, which lasted approximately 30 minutes, during which he addressed debugging issues and monitored the proceedings in the Jarvis Labs instance, periodically checking for any anomalies.

After the experimental stage, he transitioned to Microsoft Azure. “I shifted to Azure, utilising an 800-instance setup with 80 GB of RAM, all configured as spot instances. Fortunately, during this period, spot prices were exceptionally low—approximately $0.95 per hour for training,” he said.

“So, for both the 7 billion and 13 billion models, the total cost, since it’s a spot instance, becomes a challenge due to occasional deallocation. I had mechanisms in place to resume training even after it gets reallocated, allowing me to keep the cost significantly lower than opting for reserved instances.” he explained.

Main Challenges

Balachandran explained that the model is trained in two stages. According to him, the primary challenges were related to data availability, especially for Tamil text that could be found online.

“The first stage is pre-training, during which we expose the model to a substantial amount of data from the internet,” he said. For pre-training, he utilised a dataset called CulturaX, which offered a substantial Tamil corpus. “However, in terms of quality compared to English, it remains a challenge not only for Tamil but also for many Indian languages.” he added.

The second stage posed challenges due to insufficient data. “I had to either translate existing content from English to Tamil or explore the possibility of using GPT-4 to generate more Tamil content. Dealing with the data aspect was particularly challenging in this phase.” explained Balachandran.

Limitations

Tamil Llama can answer questions related to Tamil culture or even basic questions correctly. However, it still lacks the ability to fully comprehend Tamil culture. “Since it is just trained on internet documents and some translated instruction data, one of the main weaknesses is that it doesn’t know the cultural aspects of the region. So, that is also another thing I’m actually trying to improve in the next version,”said Balachandran.

Another limitation is that it can respond to malicious instructions as well. “I didn’t do the alignment step because it would have cost a little bit more, and that’s the reason I skipped that part. However, in the future, maybe I can implement that as well.” said Balachandran.

What’s Next?

Balachandran mentioned that he plans to build a multilingual LLM next. “I am currently experimenting with a multilingual model which is also one of my goals going into 2024 — to develop a multilingual LLM for Indian languages, a model that can work for Hindi, Telugu, and Tamil.” he said.

“There is a ‘Tamil Computing Conference’ happening in February 2024. I am also invited, and I’m excited for that,” he concluded. Balachandran will also be attending India’s biggest generative AI conference, MLDS 2023, on February 1-2, 2023, in Bengaluru. Register now.

The post Meet தமிழ் Llama appeared first on Analytics India Magazine.

How This Startup is Making You Eat Healthy with Generative AI

Recently, at the Ignite conference, Khosla Ventures-backed Indian health and fitness startup Healthify (formerly known as Healthifyme) introduced Ria 2.0, a generative AI-powered personal health coach with multimodal, multilingual conversational capabilities.

The chatbot health advice adjusts to individual lifestyles and goals by considering factors like schedule, meals (captured via photos), blood sugar levels, activity, and dietary preferences.

Healthify’s ‘Snap’ feature, introduced this year, revolutionises nutrition tracking by allowing users to photograph meals for instant identification of various food types and calculation of nutritional content, including calorie count. The company claims that it is 40% more accurate than the original version, recognising over one million global foods.

Two other announcements were also made at the event. Firstly, the fintech company has entered a commercial partnership with food commerce platform Swiggy to allow users to order meals according to Ria’s smart dietary recommendations through the Healthify app, facilitating adherence to a healthy diet.

Secondly, Healthify’s “Coach Co-Pilot”, launched in 2019, which combines AI with human coaching, was validated by a Stanford study for a 70% increase in weight-loss effectiveness.

Inside Healthify’s Generative AI Game

“Our approach involves using multiple foundational models to contextualise information based on specific tasks and not relying on one for all,” Tushar Vashisht, cofounder and chief executive officer told AIM during the event.

Healthify’s team uses solution models based on diverse statistical models, including AWS, Meta, OpenAI, and open-source options, with a focus on domain knowledge for new insights.

“Leveraging an extensive dataset of over 350 million messages and diet plans, their approach involves assembling different models tailored to specific outputs, such as enhancing conversational engagement in voice interactions with user data analysis,” Abhijit Khasnis, vice president – technology, told AIM. When standard foundation models fall short, the team strategically leverages specific use cases to interact more effectively with real-world data.

Another important point is that the company places a strong emphasis on user control and confidentiality in handling health data. Customers can delete their data and decide whether to share it with their coach. “We never participate in data partnerships or selling user data, aligning with regulatory and ethical considerations and customers have full control over it,” noted Vashisht.

Growing with AI Through Years

According to the company, one significant aspect of its journey has been the shift from older AI systems like the original Ria (launched in 2019) to the new generative AI systems.

“Our initial AI could only handle a limited scope of queries and data points. However, with advancements, we’ve expanded our data vectors to include a wide array of inputs like calendar, geolocation, heart rate, nutrition, and fitness, greatly enhancing the comprehensiveness of our service,” he added.

Agreeing with Vashisht, Anjan Bhojarajan, chief business officer recounted that another turning point came when the team first considered monetisation for their AI services back in 2022 when ChatGPT was born.

“Initially, we thought it would only improve conversations, but it quickly became apparent that this new mode of communication would improve app interfaces, leading to the development of a new app with a chat messaging system, marking a shift from traditional input methods to a more unified and interactive interface,” Bhojarajan added.

Currently, Healthify is preparing for international expansion, and this financial year is set to be the most profitable yet, despite restructuring efforts and layoffs last fiscal year, according to Vashisht.

“For FY24, we are expecting double-digit growth, crossing $30 million in revenues and approaching a $40 million run rate while maintaining lower costs than ever before,” he concluded.

The post How This Startup is Making You Eat Healthy with Generative AI appeared first on Analytics India Magazine.

Top Tech Movers and Shakers of 2023

This year had its fair share of drama in the tech sector. New roles were created for AI-specific positions, and there were layoffs as researchers and executives moved on to new ventures. Some of these moves were dramatic, to say the least. Here is a list of the top moves made by some of the brightest minds in technology this year:

Sam Altman and Greg Brockman

Most notably, last month, Sam Altman and Greg Brockman, both co-founders of OpenAI, joined Microsoft after Altman was abruptly fired, and Brockman quit in protest. This move followed a dramatic series of events at OpenAI, which included leadership changes and internal turmoil. These events occurred amid significant changes and internal challenges at OpenAI. This upheaval at OpenAI, which also involved replacing a newly appointed CEO within days, led to both joining Microsoft to lead a new advanced AI research team and then returning to OpenAI within days.

Mistral founders

The three founders of Mistral, the AI startup—Arthur Mensch (former researcher at Google DeepMind), Timothee Lacroix (former researcher at Facebook AI), and Guillaume Lample (former research scientist at Facebook AI)—moved on to build Mistral AI in May this year.

The founders, Arthur, Guillaume Lample, and Timothée Lacroix, met as students at École Polytechnique and École Normale Supérieure. They bring a blend of deep technical expertise and experience from working in leading AI labs. Arthur contributed to significant projects at DeepMind, while Guillaume and Timothée were instrumental in developing the LLaMa large language models.

Mistral AI focuses on foundational models with an open technology approach. The company released its first model, Mistral 7B, under an open-source Apache 2.0 license, available for free download.

Alexis Black Bjorlin

Alexis Black Bjorlin, previously Meta Platforms’ Vice President of Infrastructure overseeing AI chip development, joined Nvidia. She is leading Nvidia’s DGX Cloud business, renting servers with Nvidia GPUs to customers.

Bjorlin, a notable figure in the semiconductor and cloud sector, will report directly to Nvidia CEO Jensen Huang. Her move follows departure from Meta in September this year. Bjorlin also serves on the boards of Digital Realty and the Global Semiconductor Association.

Ruth Porat

Ruth Porat, Alphabet and Google’s long-standing CFO, is transitioning to a new role overseeing the company’s “Other Bets” portfolio, as revealed in the recent second-quarter earnings report. Starting September 1, 2023, Porat will become the President and Chief Investment Officer of Alphabet and Google.

She will work closely with CEO Sundar Pichai, focusing on global investments and technological economic growth. She will continue as CFO in the interim, overseeing the 2024 and long-term capital planning, while a successor is selected.

Peter Deng

Peter Deng, who has a history of working at notable companies as Product Head, left Airtable to join OpenAI this year. He joined OpenAI as VP of Consumer Product, leading the product, design, and engineering teams responsible for ChatGPT. His goal is to make AI useful, accessible, and beneficial to everyone.

From developing new interaction paradigms to creating assistive capabilities that enhance productivity and creativity, he saw the opportunity to explore untapped areas.

Neal Mohan

Neal Mohan replaced Susan Wojcicki as the CEO of YouTube earlier this year. He served as the Chief Product Officer at YouTube for seven years. Neal discussed the emerging potential of Artificial Intelligence (AI), stating that AI was just beginning to surface and had the capability to ‘reinvent video.’ He added that while the platform was eager to develop these features, it planned to take its time to ’embrace this technology responsibly.’

Jerome Pesenti

Jerome Pesenti, Meta’s VP of AI for four years in 2022. He left the company as it integrated AI teams across various product groups, moving away from a centralized AI organization. Meta CTO Andrew Bosworth announced the reorganization, aiming to “leverage the newest AI technology at scale” within the company.

This year he founded Sizzle AI, a New York-based company focused on AI-driven direct-to-learner products. Sizzle AI recently secured $7.5 million in seed funding. Pesenti’s background includes BenevolentAI, and IBM Watson, co-founding Vivisimo apart from his role at Meta. He also co-chaired a UK government-commissioned review on AI growth.

Jim Keller

Tenstorrent, a Toronto-based AI startup, reshuffled its leadership, appointing Jim Keller as CEO. The company, valued at $1 billion and backed by venture funding, views this change as aligning with Keller’s role.

Keller, with experience in CPU core designs including AMD’s Zen and Tesla’s FSD, joined Tenstorrent in 2016. His leadership focuses on developing AI accelerators, such as the Grayskull chip, and entering the RISC-V space.

Reed Hastings

Reed Hastings, co-CEO and founder of Netflix, announced in January this year that he will step down from his role. The announcement, made via a company blog post, comes as part of a planned leadership transition. Netflix had appointed Ted Sarandos as co-CEO with Hastings in 2020, a move that formalized the company’s existing operational structure.

Following Hastings’ departure, Netflix will maintain its co-CEO structure with COO Greg Peters joining Sarandos. Hastings has voiced his trust in Sarandos and Peters to drive the company’s growth. He will become Executive Chairman of the Board, a role similar to those undertaken by other tech company founders like Jeff Bezos and Bill Gates.

The post Top Tech Movers and Shakers of 2023 appeared first on Analytics India Magazine.

OpenAI Finally Takes on Apple

OpenAI Finally Takes on Apple

A few months ago, there were reports that Sam Altman is joining Jony Ive, the legendary iPhone designer to start their own AI hardware project for AI devices. Now, the move is reported to take further steps as both of the veterans have approached outgoing Apple executive Tang Tan, vice president of iPhone and watch design, to join the project.

According to reports, Tan will join LoveFrom, the company created by Ive in 2019, which is a design studio to create AI devices. Tan would lead the hardware engineering project as he has stepped down from the current position at Apple. People with knowledge of the matter said that his role is already divided with others at Apple, and he would leave by February.

Apple has witnessed a significant departure of design expertise with Tan’s exit, highlighting a broader trend of talent leaving the company. Since 2019, approximately 14 individuals from Ive’s original team have departed, leaving only around half a dozen of the designers who were once under Ive’s supervision still employed at Apple.

Meanwhile, SoftBank CEO and investor Masayoshi Son is also involved in this new development and has held talks with both Altman and Ive about the idea, but it is unclear if Son will remain involved in the longer run.

It all makes sense for OpenAI

LoveFrom has successfully gathered a notable clientele, including prominent names like Airbnb Inc., Ferrari NV, and Moncler SpA. Additionally, the company had a three-year consulting deal with Apple, which concluded in 2022. Remarkably, over 20 former Apple employees are now part of LoveFrom’s team.

Though the plans for the exact form factor that LoveFrom is working on is unclear, speculations around a home device like an Alexa Mini, or something similar to Humane Ai are floating around. OpenAI’s technology and Ive’s hardware design prowess are bound to make something beautiful happen.

For Altman, this decision of investing into other startups was one of the reasons that OpenAI’s board members were unhappy, which eventually led to him getting fired. Though now that he is back as the CEO, he wants people to be able to better use the company’s technology in different form factors.

In 2020, Altman also invested $30 million in series A funding of Humane, which has announced its GPT-powered Ai Pin. One fascinating fact about Humane is that it was founded by former Apple employees Imran Chaudhari and Bethany Bongiorno. Imran spent over 20 years at Apple working on products like iPod, iPad and Apple Watch and iPhone.

Until the Humane Ai Pin was announced, Meta’s Ray-Ban seemed like a gadget of the future, that everyone would be wearing, and recording the world around them. But now, investing in a pair of glasses seems to be outdated altogether, and OpenAI knows that.

Humane is not the only thing. OpenAI also recently partnered with WHOOP to introduce WHOOP Coach suggests that SoftBank might soon become part of OpenAI’s endeavours. Leveraging OpenAI’s latest technology, WHOOP Coach instantly generates personalised and conversational answers to your inquiries about health, fitness, and well-being.

OpenAI’s upgrades to GPT-4 with vision and audio capabilities all lead to this, and it makes sense for the company to embed this technology in a hardware device.

What about Apple?

In other news, OpenAI has also announced that it is planning to raise more funds at a valuation of $100 billion. This includes discussion with Abu Dhabi-based G42 chip venture fund. This along with the investment in RainAI for developing neuromorphic processing units, OpenAI is going all in on hardware.

Altman’s plan is working out, meanwhile Apple is headed the OpenAI way and is increasingly focusing on building generative AI capabilities for its hardware. Most recently, Apple also approached news publications such as Conde Nast, NBC News, and IAC, to help them build Apple’s generative AI capabilities in a multiyear deal worth at least $50 million.

The company also has been releasing open source projects for working AI on its silicon, which includes multimodal LLMs such as Ferret. It is clear that OpenAI is ready for the AI hardware challenge and Apple is ready with its generative AI capabilities on software.

Right now, it seems that OpenAI and Altman, with Humane, LoveFrom, and all the new Apple employees on their team, have the upper hand over their competitors.

The post OpenAI Finally Takes on Apple appeared first on Analytics India Magazine.

Top Most Powerful 5 AI Chips Released in 2023

In August this year, Gartner forecasted that the worldwide AI chips market would reach a whopping $53 billion in revenue. The major key drivers fueling this growth include the widespread adoption of AI across diverse industries and the escalating demand for efficient, specialized hardware to support AI workloads, particularly in healthcare, finance, and retail applications.

With the arrival of generative AI, the chip market is experiencing unprecedented growth, underscored by robust revenue projections and dynamic trends. Major players like NVIDIA, Intel, Qualcomm, AMD, and Alphabet are driving innovation, investing substantially in cutting-edge AI chip architectures.

Noteworthy trends include a shift towards diverse chip architectures, with GPUs maintaining dominance but facing competition from specialised AI chips and evolving CPUs. Custom-designed AI chips tailored to specific needs are on the rise, and the focus on edge computing is driving the development of low-power, efficient AI chips.

Technological trends, including neuromorphic computing and the nascent field of quantum computing, promise to elevate AI chip capabilities further. Moreover, the emphasis is on developing energy-efficient AI chips and advancing memory technologies

underscores the industry’s commitment to sustainability and performance optimisation.

Here is a list of the top 5 AI chips that made it to the market this year:

Gaudi 3

Intel’s Gaudi3 AI Accelerator, unveiled in December 2023, is reshaping the landscape of AI acceleration with a focus on generative AI. Crafted on an advanced 5-nanometer process, Gaudi3 boasts heightened performance and efficiency compared to its predecessor, Gaudi2.

Tailored for text-to-image and image-to-image processing tasks, this chip is a powerhouse for creating realistic art, manipulating photos, and designing products. Intel claims substantial performance improvements over Gaudi 2, potentially challenging competitors like NVIDIA’s H100 accelerator. Gaudi3’s scalability allows integration into systems with multiple chips, offering the ability to scale up performance for complex generative AI tasks.

AMD MI300

Launched in November 2023, the AMD MI300 family is set to disrupt the AI accelerator landscape by directly challenging Nvidia’s H100. Featuring a chipset architecture, the MI300 takes a modular approach, offering flexibility and scalability by mixing and matching different chipsets for computing, memory, and I/O.

Advanced packaging with 3D chip stacking enhances communication efficiency. Dedicated AI accelerators, including Matrix Cores optimised for AI workloads, multi-precision support, and ample memory capacity, position the MI300 competitively. While it may not match the H100’s raw performance in certain benchmarks, the MI300 shines in key AI tasks, particularly generative AI.

Google TPU v5e

The Google TPU v5e, released in August 2023, emerges as a powerhouse in AI hardware, specifically tailored for large language models and generative AI. Boasting a remarkable 2x improvement in training performance per dollar and a 2.5x uplift in inference performance per dollar compared to its predecessor, the TPU v5e offers substantial cost savings. Its groundbreaking multislice architecture enables the seamless connection of tens of thousands of chips, breaking previous limitations and opening avenues for tackling massive AI tasks.

With eight different virtual machine configurations, the v5e caters to diverse AI needs, from small-scale research to large-scale enterprise deployments. Technical specifications, including custom AI accelerators, HBM3 memory, and a robust inter-chip communication network, highlight its cutting-edge capabilities.

Amazon’s Trainium2

Amazon’s Trainium2, unveiled at re:Invent 2023, emerges as a cutting-edge AI chip tailored for the training and execution of large language models (LLMs) and tasks related to natural language processing (NLP) and generative AI. The architecture boasts dedicated AI-optimized cores, high memory bandwidth, and seamless inter-chip communication, resulting in a remarkable 4x performance improvement compared to its predecessor.

With a potential for 65 exaflops of performance in large clusters, Trainium2 is poised to handle unprecedentedly complex AI tasks. Its applications span LLM training, NLP tasks, and generative AI, paving the way for advancements in natural language understanding and creative text generation. Currently in limited preview through Amazon EC2, Trainium2’s wider availability is anticipated, with ongoing collaborations for optimized tools and libraries.

Azure Maia AI Accelerator

Azure Maia, Microsoft’s AI Accelerator, emerges as a groundbreaking advancement in AI hardware, purpose-built to meet the demands of complex AI workloads. Manufactured using a cutting-edge 5-nanometer TSMC process, the Maia AI chip boasts 105 billion transistors.

Its architecture positions it as a formidable tool for large language model training and inferencing, featuring specialised matrix multiplication units, high-bandwidth memory, and scalable inter-chip communication. With the promise of up to 5x faster training times and real-time inference capabilities, Azure Maia stands to accelerate AI development and improve energy efficiency. Its applications span large language models, generative AI, and various natural language processing tasks.

The post Top Most Powerful 5 AI Chips Released in 2023 appeared first on Analytics India Magazine.

DiffSeg : Unsupervised Zero-Shot Segmentation using Stable Diffusion

DiffSeg : Unsupervised Zero-Shot Segmentation using Stable Diffusion

One of the core challenges in computer vision-based models is the generation of high-quality segmentation masks. Recent advancements in large-scale supervised training have enabled zero-shot segmentation across various image styles. Additionally, unsupervised training has simplified segmentation without the need for extensive annotations. Despite these developments, constructing a computer vision framework capable of segmenting anything in a zero-shot setting without annotations remains a complex task. Semantic segmentation, a fundamental concept in computer vision models, involves dividing an image into smaller regions with uniform semantics. This technique lays the groundwork for numerous downstream tasks, such as medical imaging, image editing, autonomous driving, and more.

To advance the development of computer vision models, it's crucial that image segmentation isn't confined to a fixed dataset with limited categories. Instead, it should act as a versatile foundational task for various other applications. However, the high cost of collecting labels on a per-pixel basis presents a significant challenge, limiting the progress of zero-shot and supervised segmentation methods that require no annotations and lack prior access to the target. This article will discuss how self-attention layers in stable diffusion models can facilitate the creation of a model capable of segmenting any input in a zero-shot setting, even without proper annotations. These self-attention layers inherently understand object concepts learned by a pre-trained stable diffusion model.

DiffSeg : An Enhanced Zero-Shot Segmentation Algorithm

Semantic segmentation is a process that divides an image into various sections, with each section sharing similar semantics. This technique forms the foundation for numerous downstream tasks. Traditionally, zero-shot computer vision tasks have depended on supervised semantic segmentation, utilizing large datasets with annotated and labeled categories. However, implementing unsupervised semantic segmentation in a zero-shot setting remains a challenge. While traditional supervised methods are effective, their per-pixel labeling cost is often prohibitive, highlighting the need for developing unsupervised segmentation methods in a less restrictive zero-shot setting, where the model neither requires annotated data nor prior knowledge of the data.

To address this limitation, DiffSeg introduces a novel post-processing strategy, leveraging the capabilities of the Stable Diffusion framework to build a generic segmentation model capable of zero-shot transfer on any image. Stable Diffusion frameworks have proven their efficacy in generating high-resolution images based on prompt conditions. For generated images, these frameworks can produce segmentation masks using corresponding text prompts, typically including only dominant foreground objects.

Contrastingly, DiffSeg is an innovative post-processing method that creates segmentation masks by utilizing attention tensors from the self-attention layers in a diffusion model. The DiffSeg algorithm is composed of three key components: iterative attention merging, attention aggregation, and non-maximum suppression, as illustrated in the following image.

The DiffSeg algorithm preserves visual information across multiple resolutions by aggregating the 4D attention tensors with spatial consistency, and utilizing an iterative merging process by sampling anchor points. These anchors serve as the launchpad for the merging attention masks with same object anchors absorbed eventually. The DiffSeg framework controls the merging process with the help of KL divergence method to measure the similarity between two attention maps.

When compared with clustering-based unsupervised segmentation methods, developers do not have to specify the number of clusters beforehand in the DiffSeg algorithm, and even without any prior knowledge, the DiffSeg algorithm can produce segmentation without utilizing additional resources. Overall, the DiffSeg algorithm is “A novel unsupervised and zero-shot segmentation method that makes use of a pre-trained Stable Diffusion model, and can segment images without any additional resources, or prior knowledge.”

DiffSeg : Foundational Concepts

DiffSeg is a novel algorithm that builds on the learnings of Diffusion Models, Unsupervised Segmentation, and Zero-Shot Segmentation.

Diffusion Models

The DiffSeg algorithm builds on the learnings from pre-trained diffusion models. Diffusion models is one of the most popular generative frameworks for computer vision models, and it learns the forward and reverse diffusion process from a sampled isotropic Gaussian noise image to generate an image. Stable Diffusion is the most popular variant of diffusion models, and it is used to perform a wide array of tasks including supervised segmentation, zero-shot classification, semantic-correspondence matching, label-efficient segmentation, and open-vocabulary segmentation. However, the only issue with diffusion models is that they rely on high-dimensional visual features to perform these tasks, and they often require additional training to take complete advantage of these features.

Unsupervised Segmentation

The DiffSeg algorithm is closely related to unsupervised segmentation, a modern AI practice that aims to generate dense segmentation masks without employing any annotations. However, to deliver good performance, unsupervised segmentation models do need some prior unsupervised training on the target dataset. Unsupervised segmentation based AI frameworks can be characterized into two categories: clustering using pre-trained models, and clustering based on invariance. In the first category, the frameworks make use of the discriminative features learned by pre-trained models to generate segmentation masks whereas frameworks finding themselves in the second category use a generic clustering algorithm that optimizes the mutual information between two images to segment images into semantic clusters and avoid degenerate segmentation.

Zero-Shot Segmentation

The DiffSeg algorithm is closely related to zero-shot segmentation frameworks, a method with the capability to segment anything without any prior training or knowledge of the data. Zero-shot segmentation models have demonstrated exceptional zero-shot transfer capabilities in recent times although they require some text input and prompts. In contrast, the DiffSeg algorithm employs a diffusion model to generate segmentation without querying and synthesizing multiple images and without knowing the contents of the object.

DiffSeg : Method and Architecture

The DiffSeg algorithm makes use of the self-attention layers in a pre-trained stable diffusion model to generate high-quality segmentation tasks.

Stable Diffusion Model

Stable Diffusion is one of the fundamental concepts in the DiffSeg framework. Stable Diffusion is a generative AI framework, and one of the most popular diffusion models. One of the main characteristics of a diffusion model is a forward and a reverse pass. In the forward pass, a small amount of Gaussian noise is added to an image iteratively at every time step until the image becomes an isotropic Gaussian noise image. On the other hand, in the reverse pass, the diffusion model iteratively removes the noise in the isotropic Gaussian noise image to recover the original image without any Gaussian noise.

The Stable Diffusion framework employs an encoder-decoder, and a U-Net design with attention layer where it uses an encoder to first compress an image into a latent space with smaller spatial dimensions, and utilizes the decoder to decompress the image. The U-Net architecture consists of a stack of modular blocks, where each block is composed of either of the following two components: a Transformer Layer, and a ResNet layer.

Components and Architecture

Self-attention layers in diffusion models grouping information of inherent objects in the form of spatial attention maps, and DiffSeg is a novel post-processing method to merge attention tensors into a valid segmentation mask with the pipeline consisting of three main components: attention aggregation, non-maximum suppression, and iterative attention.

Attention Aggregation

For an input image that passes through the U-Net layers, and the Encoder, the Stable Diffusion model generates a total of 16 attention tensors, with 5 tensors for each of the dimensions. The primary goal of generating 16 tensors is to aggregate these attention tensors with different resolutions into a tensor with the highest possible resolution. To achieve this, the DiffSeg algorithm treats the 4 dimensions differently from one another.

Out of the four dimensions, the last 2 dimensions in the attention sensors have different resolutions yet they are spatially consistent since the 2D spatial map of the DiffSeg framework corresponds to the correlation between the locations and the spatial locations. Resultantly, the DiffSeg framework samples these two dimensions of all attention maps to the highest resolution of them all, 64 x 64. On the other hand, the first 2 dimensions indicate the location reference of the attention maps as demonstrated in the following image.

As these dimensions refer to the location of the attention maps, the attention maps need to be aggregated accordingly. Additionally, to ensure that the aggregated attention map has a valid distribution, the framework normalizes the distribution after aggregation with every attention map being assigned a weight proportional to its resolution.

Iterative Attention Merging

While the primary goal of attention aggregation was to compute an attention tensor, the primary aim is to merge the attention maps in the tensor to a stack of object proposals where each individual proposal contains either the stuff category or the activation of a single object. The proposed solution to achieve this is by implementing a K-Means algorithm on the valid distribution of the tensors to find the clusters of the objects. However, using K-Means is not the optimal solution because K-Means clustering requires users to specify the number of clusters beforehand. Furthermore, implementing a K-Means algorithm might result in different results for the same image since its stochastically dependent on the initialization. To overcome the hurdle, the DiffSeg framework proposes to generate a sampling grid to create the proposals by merging attention maps iteratively.

Non-Maximum Suppression

The previous step of iterative attention merging yields a list of object proposals in the form of probability ot attention maps where each object proposal contains the activation of the object. The framework makes use of non-maximum suppression to convert the list of object proposals into a valid segmentation mask, and the process is an effective approach since each element in the list is already a map of the probability distribution. For every spatial location across all maps, the algorithm takes the index of the largest probability, and assigns a membership on the basis of the index of the corresponding map.

DiffSeg : Experiments and Results

Frameworks working on unsupervised segmentation make use of two segmentation benchmarks namely Cityscapes, and COCO-stuff-27. The Cityscapes benchmark is a self-driving dataset with 27 mid-level categories whereas the COCO-stuff-27 benchmark is a curated version of the original COCO-stuff dataset that merges 80 things and 91 categories into 27 categories. Furthermore, to analyze the segmentation performance, the DiffSeg framework uses mean intersection over union or mIoU and pixel accuracy or ACC, and since the DiffSeg algorithm is unable to provide a semantic label, it uses the Hungarian matching algorithm to assign a ground truth mask with each predicted mask. In case the number of predicted masks exceeds the number of ground truth masks, the framework will take into account the unmatched predicted tasks as false negatives.

Additionally, the DiffSeg framework also emphasizes on the following three works to run interference: Language Dependency or LD, Unsupervised Adaptation or UA, and Auxiliary Image or AX. Language Dependency means that the method needs descriptive text inputs to facilitate segmentation for the image, Unsupervised Adaptation refers to the requirement for the method to to use unsupervised training on the target dataset whereas Auxiliary Image refers that the method needs additional input either as synthetic images, or as a pool of reference images.

Results

On the COCO benchmark, the DiffSeg framework includes two k-means baselines, K-Means-S and K-Means-C. The K-Means-C benchmark includes 6 clusters that it calculated by averaging the number of objects in the images it evaluates whereas the K-Means-S benchmark uses a specific number of clusters for each image on the basis of the number of objects present in the ground truth of the image, and the results on both these benchmarks are demonstrated in the following image.

As it can be seen, the K-Means baseline outperforms existing methods, thus demonstrating the benefit of using self-attention tensors. What’s interesting is that the K-Means-S benchmark outperforms the K-Means-C benchmark that indicates that the number of clusters is a fundamental hyper-parameter, and tuning it is important for every image. Furthermore, even when relying on the same attention tensors, the DiffSeg framework outperforms the K-Means baselines that proves the ability of the DiffSeg framework to not only provide better segmentation, but also avoid the disadvantages posed by using K-Means baselines.

On the Cityscapes dataset, the DiffSeg framework delivers results similar to the frameworks utilizing input with lower 320-resolution while outperforming frameworks that take higher 512-resolution inputs across accuracy and mIoU.

As mentioned before, the DiffSeg framework employs several hyper-parameters as demonstrated in the following image.

Attention aggregation is one of the fundamental concepts employed in the DiffSeg framework, and the effects of using different aggregation weights is demonstrated in the following image with the resolution of the image being constant.

As it can be observed, high-resolution maps in Fig (b) with 64 x 64 maps yield most detailed segmentations although the segmentations do have some visible fractures whereas lower resolution 32 x 32 maps tends to over-segment details although it does result in enhanced coherent segmentations. In Fig (d), low resolution maps fail to generate any segmentation as the entire image is merged into a singular object with the existing hyper-parameter settings. Finally, Fig (a) that makes use of proportional aggregation strategy results in enhanced details and balanced consistency.

Final Thoughts

Zero-shot unsupervised segmentation is still one of the greatest hurdles for computer vision frameworks, and existing models either rely on non zero-shot unsupervised adaptation or on external resources. To overcome this hurdle, we have talked about how self-attention layers in stable diffusion models can enable the construction of a model capable of segmenting any input in a zero-shot setting without proper annotations as these self-attention layers hold the inherent concepts of the object that a pre-trained stable diffusion model learns. We have also talked about DiffSeg, a novel post-pressing strategy, aims to harness the potential of the Stable Diffusion framework to construct a generic segmentation model that can implement zero-shot transfer on any image. The algorithm relies on Inter-Attention Similarity and Intra-Attention Similarity to merge attention maps iteratively into valid segmentation masks to achieve state of the art performance on popular benchmarks.

The Most Read AIM Stories of 2023

In this comprehensive article, we dive into the most read stories of 2023 from Analytics India Magazine. Covering a range of topics from the challenges of AI in job markets to OpenAI’s strategic moves, each story provides a unique insight into the evolving world of AI and technology. From Hugging Face’s bold steps to challenge OpenAI, to the surprising inefficacies of ChatGPT in software engineering queries, these stories highlight key developments and debates in the AI industry.

Read: The Best AIM Stories of 2022

“An Entire Generation is Studying for Jobs that Won’t Exist” by Mohit Pandey, published on January 21, 2023. This article discusses the impact of AI on job markets and the potential obsolescence of certain professions due to technological advancements. Read more.

“When ChatGPT Attempted UPSC Exam” by Pritam Bordoloi, published on February 28, 2023. This piece details ChatGPT’s attempt at India’s UPSC exam, highlighting its limitations in such competitive examinations. Read more.

“Infosys Announces Free AI Training Program for Upskilling” by Shritama Saha, published on June 23, 2023. The article reports on Infosys’ initiative to provide free AI training, focusing on data science and related fields. Read more.

“OpenAI Releases Paid ChatGPT Professional” by Shritama Saha, published on January 21, 2023. This article details the launch of a paid version of ChatGPT, discussing its features and potential impacts. Read more.

“No Need to Study Maths Anymore” by Mohit Pandey, published on June 8, 2023. This piece examines the evolving role of mathematics in AI and machine learning. Read more.

“Google Fools Everyone with Gemini” by Siddharth Jindal, published on December 8, 2023. This article discusses Google’s hurried release of the Gemini AI model, suggesting pressure from competitors, and examines its technical aspects and performance comparisons. Read more.

“Now Everyone’s a Developer, Thanks to Microsoft” by Mohit Pandey, published on May 24, 2023. The article talks about Microsoft’s role in democratizing development through AI, hinting at the evolution of programming and the impact on developers. Read more.

“OpenAI Launches Residency Program with $210,000 Annual Salary” by Mohit Pandey, published on October 4, 2023. This article highlights OpenAI’s new residency program aimed at nurturing AI talent, detailing its objectives, structure, and salary offerings. Read more.

“6 Brilliant New Free Courses by Andrew Ng on Generative AI” by Shritama Saha, published on September 7, 2023. The piece outlines Andrew Ng’s six new AI courses, their content, and partnerships, aimed at enhancing AI literacy. Read more.

“Hugging Face Makes OpenAI’s Worst Nightmare Come True” by Poulomi Chatterjee, published on March 3, 2023. The article discusses Hugging Face’s partnership with AWS and its implications for OpenAI, highlighting contrasting approaches to democratizing AI and the competition in generative AI technologies. Read more.

“Stack Overflow Snatches the Spot from ChatGPT” by Tasmia Ansari, published on August 17, 2023. This article explores the reliability of OpenAI’s ChatGPT in answering software engineering questions, with research showing over 50% of its responses to be inaccurate compared to Stack Overflow. Read more.

“Obsolete Code: 10 Programming Languages That Vanished Over Time” by K L Krithika, published on July 13, 2023. The article provides an overview of 10 once-popular programming languages that have become obsolete, exploring reasons behind their decline and impact on the programming landscape. Read more.

“OpenAI Likely To Pull the Plug on ChatGPT” by Mohit Pandey, published on August 17, 2023. The article speculates on the future of ChatGPT, considering operational costs and strategic moves by OpenAI, suggesting potential discontinuation or significant changes to the platform. Read more.

“Upskill with These Free Generative AI Courses Offered by Big Techs” by Siddharth Jindal, published on June 29, 2023. The article lists and describes free generative AI courses offered by major tech companies like Google, AWS, Microsoft, and Infosys, aimed at enhancing AI literacy and skills. Read more.

“Chromecast Joins Google Graveyard” by Tasmia Ansari, published on June 1, 2023. This article announces the end of support for the original Chromecast by Google, detailing its impact on users and its place in Google’s history of discontinued products. Read more.

“It’s Time OpenAI Launched GPT-5” by Mohit Pandey, published on July 16, 2023. The article discusses the competitive landscape in AI and urges OpenAI to accelerate the development of GPT-5 to stay ahead of emerging competitors in the field. Read more.

The post The Most Read AIM Stories of 2023 appeared first on Analytics India Magazine.

How Moreh is Making AI Software Better with AMD

How Moreh is Making AI Software Better with AMD

AMD is rising rapidly up the AI market ever since it announced its MI300X and the software updates for ROCm. It has been building partnerships with various AI companies for testing out and delivering its products. One such company led us to Korean-based Moreh.

One of the problems with AMD GPUs earlier was that the software stack was not well established, but now the company has ROCm, its alternative to CUDA. But still, it was not well suited for large GPU clusters. “Our software enables people to use AMD GPUs without any more codes, they just can run their or larger language models without any further engineering,” Junghwan Lim, head of AI group at Moreh, told AIM in an exclusive interview.

Lim has also worked as a data scientist at PUBG Corporation and has also worked with Samsung in South Korea. Currently, Lim is focused on developing better software for AI workloads at Moreh and making language models smaller for better efficiency.

AMD is being embraced way too much, and that’s good

Moreh’s flagship AI software, known as MoAI, is positioned similarly to NVIDIA’s CUDA but boasts compatibility with existing machine learning frameworks like Meta’s PyTorch and Google’s TensorFlow, and even OpenAI’s Triton. Now, the company is helping AMD boost up its ROCm performance.

In August, Moreh announced that it has been using AMD MI250 for the longest time and it is outperforming NVIDIA. According to Moreh, AMD’s MI250 Instinct accelerator, when powered by the MoAI platform, achieved 116% higher GPU throughput than NVIDIA’s A100.

“We have been using more than 400 MI250 GPUs along with a few MI300X for training AI models,” said Lim. The company also plans to buy more MI300X from AMD in the future.

“If anyone wants to use AMD GPUs, they can come to us without any code, and just use our software to run on them,” said Lim. This is somewhat similar to what Lamini has been doing in partnership with AMD, but Lim says that there is still a difference in the requirement of codes, as companies can also use their existing GPUs to run their models.

A key aspect of Moreh’s success lies in its software being built on AMD GPU infrastructure, showcasing performance that surpasses NVIDIA GPUs in AI model development. The MoAI platform, a comprehensive software offering, is not tied to a specific hardware vendor and supports various device backends, including AMD GPUs.

“Every configuration or any technique that the customer wants to apply should be easily applied without the need of any programmer. That is what we are aiming to do, and that is what differentiates us,” Lim said.

“If you are building an AI model on a thousand GPUs, there can be issues such as a GPU malfunction because of hardware or software issues,” explained Lim. “The training suddenly stops and it takes hours, or sometimes days to start it again.” He explains that Moreh is also building software and devising techniques to reduce this by parallelising computers to different fragments.

Large language models for the win

Moreh has also finished training its own LLM for the Korean language which consists of 221 billion parameters and wishes to make upcoming models open source.

“We are developing something similar to GPT or Gemini, but it is going to be open source.” Lim says that the current model is too big to open source, so Moreh is also planning to release smaller models soon. “Our models would include the code, weights, inference code, and everything else,” Lim highlighted about the recent trend in open source models which come with some or the other restrictions.

In October, AMD and Korean telecommunications (KT) invested a $22 million series B round in Moreh, bringing the valuation of the startup to $30 million. The company also projected its revenue to reach $30 million by the end of 2023.

KT has also bought one of the largest AMD GPU cluster in the world and are also building AI models using them. “We are supporting all those clusters and cloud systems for KT,” said Lim. KT is starting to focus on GPU cloud provider business, also providing APIs for language models, but which would just be focused on Korean language, not English.

KT, which has been working with Moreh since 2021, claims that Moreh’s technology has demonstrated superior performance compared to NVIDIA’s DGX, specifically in terms of speed and GPU memory capacity.

“People believe that AMD GPUs are not that suited for machine learning, but the company has been increasingly proving them wrong,” concluded Lim.

The post How Moreh is Making AI Software Better with AMD appeared first on Analytics India Magazine.

Top 10 Videos of AIM in 2023: A Year of Data and Insights

Analytics India Magazine (AIM) has been at the forefront of providing valuable content and insights. As we bid farewell to 2023, let’s take a moment to reflect on the top video highlights from AIM that captured the essence of this data-centric year.

  1. Google Data Leaders Exchange | Panel Discussion: Unified, flexible, & accessible: The future of data
    • Premiered: January 6, 2023
    • Views: 28,329
    • Description: Data Leaders Exchange, in association with Google Cloud, brought together leading technology experts to discuss the future of data and best practices for data-driven businesses. Watch the discussion here.
  2. Creating Sustainable Water From Air | Uravu Labs
    • Premiered: June 5, 2023
    • Views: 25,018
    • Description: Uravu Labs, a Bengaluru-based deep-tech startup, uses data and analytics to create sustainable water from the air. Learn how they are making the world more sustainable here.
  3. MachineCon 2023 | Official Aftermovie
    • Premiered: June 28, 2023
    • Views: 24,104
    • Description: MachineCon, India’s largest gathering of analytics and AI leaders, was a grand success. Watch the aftermovie to relive the experience here.
  4. The Women of Genpact
    • Premiered: March 20, 2023
    • Views: 18,151
    • Description: Meet the inspiring women of Genpact and hear their stories of overcoming challenges and achieving success in an organization that values diversity. Watch their stories here.
  5. Meet the Data Scientists at Genpact
    • Premiered: January 16, 2023
    • Views: 11,986
    • Description: Explore the world of data science with Genpact’s top data scientists. This series provides insights into how data science can drive business impact. Watch the series here.
  6. Sureshkumar Rajasekar – We Need To Promote Data Sharing
    • Premiered: January 19, 2023
    • Views: 11,495
    • Description: Sureshkumar Rajasekar, VP Technology at Optum Global Solutions (India), discusses the importance of data sharing for harnessing the power of Machine Learning. Watch the conversation here.
  7. Data-driven Transformation | Google Cloud’s Data Leaders Exchange
    • Premiered: January 25, 2023
    • Views: 11,316
    • Description: AIM Fireside Chat with Google Cloud on “Data-driven Transformation” featuring Raja Sekhar Kommu and Kunal Mathuria. Explore how organizations can harness data for insights and value creation here.
  8. TheMathCompany is certified as a Best Firm For Data Scientists
    • Premiered: January 10, 2023
    • Views: 11,052
    • Description: TheMathCompany’s certification as a Best Firm For Data Scientists showcases their commitment to creating an outstanding employee experience. Learn more here.
  9. Women in Data Science Conference 2023 @intuit
    • Premiered: July 18, 2023
    • Views: 9,969
    • Description: Relive the impactful Women in Data Science conference organized by Intuit, where data scientists shared their expertise and empowered aspiring professionals. Watch the highlights here.
  10. How to use DragGAN AI | Demo
    • Premiered: June 12, 2023
    • Views: 8,558
    • Description: Traditional photo editing tools go as far as to manipulate existing pixels and make changes to what is already present. But with this latest AI-powered tool, it might be safe to say photo editing might go through a never seen before upgrade! Introducing DragGAN AI, the futuristic photo editor! How does it work and how is it different from conventional photo editing apps? Take a look at this video and find out more here.

These videos encapsulate the diverse and dynamic world of data, analytics, and technology. They have educated, inspired, and empowered professionals in their data journeys throughout 2023. Here’s to another year of data-driven excellence with AIM!

Stay tuned for more insights in 2024!

Watch AIM’s YouTube Channel

The post Top 10 Videos of AIM in 2023: A Year of Data and Insights appeared first on Analytics India Magazine.

How to use Leonardo AI to generate stunning artwork and images

Leonard AI conjures up two aliens eating a sandwich

With all the buzz around generative AI, a variety of AI-driven image creators have cropped up this year. One tool you may want to try is Leonardo. Named after the celebrated Italian artist, Leonardo offers a host of options to help you generate the images you want.

Also: Thanks to my 5 favorite AI tools, I'm working smarter now

You can control the image size and dimensions, choose an image that's photorealistic, and specify the number of images you want. Select a specific image, and you're able to tweak it, refine it, copy it, and download it.

How to use Leonardo AI to generate images

What you need: The basic version of Leonardo is free but does impose daily tokens that get used up with each new prompt. Paid plans with more features and fewer restrictions start at $10 a month, but most people should be fine with the free flavor. Here's how to create images with Leonardo.

More on AI tools