6 Trending Computer Vision Models on GitHub

Humans can spot things super quick – all they need is just a glance. Computer scientists are teaching computers to do the same through object detection, classification and image recognition in AI. They’re getting machines to look at pictures or videos, figure out what’s in them, and slap labels on the details.

New paradigms of image recognition in AI are being explored since real-world use cases are on the rise. So here are six tools to help you build better computer vision AI.

YOLO

YOLO, short for ‘You Only Look Once’, is a widely adopted real-time object detection algorithm in computer vision, embraced by major tech players in commercial products. Introduced in 2016, the original model revolutionised object detection by outpacing its counterparts in speed.

Since then, various iterations, including YOLOv4, have emerged, each enhancing performance and efficiency. YOLOv7, unveiled in July 2022 by Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao, stands out as one of the fastest and most accurate real-time object detection models.

Notably, crafted by Ultralytics, YOLOv8 prioritises speed, accuracy, and user-friendliness, making it a top choice for tasks like object detection, tracking, instance segmentation, image classification, and pose estimation.

With innovations like Mosaic data enhancement, self-adversarial training, and cross-mini-batch normalisation, these YOLO iterations continue to advance the capabilities of computer vision systems.

Here’s the GitHub repository.

ImageAI

ImageAI is an open-source Python library built to empower developers to build applications and systems with self-contained capabilities using simple and few lines of code.

Created by Moses Olafenwa, the library empowers programmers with all levels of expertise to easily integrate state-of-the-art computer vision features, train/deploy custom image/video AI models to detect and recognise custom objects.

The library has been installed over 400,000 times and has 7,000+ starts. Since 2018, Olafenwa has released more open source projects for AI inference and solving AI data problems with plans to build and release more to facilitate AI democratisation and access.

Some of the projects are IdenProf, FireNET, ActionNET, DeepStack_ExDark and TrafficNET.

Here’s the GitHub repository.

PaddleClas

PaddleClas, developed by PaddlePaddle, is a robust image classification and recognition toolset, catering to both industry and academia within image recognition.

Tailored for training top-tier computer vision models, it supports diverse image classification models like those from ImageNet1k and PULC datasets, offering Python wheel packages for predictions. PaddleClas accommodates various network structures such as ResNet, MobileNet, and ShuffleNet with a range of documentation, including tutorials and application examples.

Its versatility extends to evaluation environments for both CPU and GPU, making it an invaluable resource for developers and researchers engaged in image classification and recognition endeavours.

Here’s the GitHub repository.

Emgu CV

Emgu CV is a cross-platform .NET wrapper for the OpenCV image-processing library, facilitating the invocation of OpenCV functions from .NET compatible languages like C#, VB, VC++, and IronPython. Crafted entirely in C#, it seamlessly compiles in Mono, rendering compatibility across platforms supported by Mono—Windows, Linux, Mac OS X, iOS, and Android.

Boasting features like a generic image class, automatic garbage collection, XML serializable images, and Intellisense support, Emgu CV streamlines image-processing tasks. It supports generic pixel operations and arrives with illustrative code snippets. The current iteration is conveniently accessible as a NuGet package.

Here’s the GitHub repository.

SOD Embedded

SOD was created to establish a unified foundation for computer vision applications, fostering the widespread adoption of machine perception in both open-source and commercial products.

This advanced, embedded, cross-platform computer vision and machine learning software library provides APIs for deep learning, sophisticated media analysis, and real-time, multi-class object detection.

Specifically designed for embedded systems with constrained computational resources and IoT devices, SOD encompasses a diverse array of classic and cutting-edge deep neural networks, complete with their pre-trained models. It is a versatile solution for accelerating machine perception across various applications and platforms.

Here’s the GitHub repository.

MILVUS Bootcamp

This model is made to help with unstructured data like finding pictures, searching for audio or molecules, analysing videos, and working on questions and answers using natural language. It’s not a complete training program but has examples for developers and researchers to use with Milvus for different tasks.

The repository includes things that go along with Milvus Lite, a simpler version. You can find helpful examples and materials here if you’re trying to work on more straightforward Milvus-based solutions.

Here’s the GitHub repository.

The post 6 Trending Computer Vision Models on GitHub appeared first on Analytics India Magazine.

AIM launches the 3rd Edition of Data Engineering Summit. May 30-31, Bengaluru

The realm of data engineering stands at the precipice of a transformative era, marked by rapid advancements and evolving challenges. Recognizing this pivotal moment, Analytics India Magazine proudly announces the 3rd edition of the Data Engineering Summit (DES) 2024, slated for May 30-31, in Bengaluru, India. This two-day conference is poised to be a cornerstone event, focusing on the innovation in data engineering and offering a unique blend of knowledge, expertise, and networking opportunities.

A Gathering of Visionaries and Innovators

DES 2024 promises to be an assembly of over 1000 professionals, featuring more than 50 speakers who are at the forefront of data engineering and technology. This event is a community coming together to share insights, challenges, and successes. Among past illustrious speakers, we had Prakash Padukone, former badminton player and co-founder of Olympic Gold Quest, who shareed his insights on leveraging data in sports analytics. Hari Prasad Reddy from Zee Entertainment, Susmita Chaudhury of Deloitte, Baxish Mission from Infocepts, Dmitry Ustalov of Toloka, and Ramanuj Vidyanta from AWS also graced the stage in 2023, discussing a range of topics from cloud data engineering to AI business development.

Exploring the Cutting Edge of Data Engineering

The summit’s agenda is meticulously crafted to cover the spectrum of data engineering, from the architecture of machine learning systems to the latest data frameworks and solutions tailored for business use cases. Key topics include the future of Large Language Models operations, architecting data pipelines for generative AI models, and operationalizing foundational models. These discussions are aimed at equipping attendees with the knowledge and tools to navigate the complexities of modern data engineering.

Why DES 2024 is a Must-Attend Event

DES 2024 stands out as India’s first and only conference dedicated entirely to the field of data engineering, making it an unparalleled opportunity for professionals to enhance their skills, knowledge, and networks. The summit offers:

  • Direct Access to Industry Leaders: Engage with top engineers and innovators from leading tech companies.
  • Comprehensive Learning Tracks: From enlightening keynote speeches to interactive workshops and panel discussions, the summit offers a full spectrum of learning opportunities.
  • Networking Opportunities: Connect with peers, industry leaders, and potential collaborators.
  • Exhibition Showcases: Explore the latest technologies, tools, and services in data engineering.

Early Bird and Regular Passes

To ensure participation is within reach for everyone, DES 2024 offers tiered pricing with early bird, regular, and late passes. These passes provide all-access to the two-day event, with group discounts available to encourage team participation.

Book your passes for DES, here.

Data Engineering Awards 2024

A highlight of the summit is the Data Engineering Awards, celebrating teams and individuals who have demonstrated exceptional innovation and achievement in analytics and AI. These awards are a testament to the pioneering work being done in the field and serve as an inspiration for all attendees.

Nominations Open for Data Engineering Awards 2024.

A Testament to Excellence

Reflecting on the previous editions, attendees like Yashwanth Kumar from Fivetran, Sunil K. from Genpact, and Dipanjan Deb from Wells Fargo have shared their enriching experiences, emphasizing the value of knowledge sharing and the impact of new technologies in scaling AI and data engineering efforts.

Join Us at DES 2024

We invite data engineers, AI specialists, business leaders, and anyone passionate about the future of technology to join us at DES 2024. This summit is not just an event but a milestone in the data engineering community’s journey towards innovation, excellence, and collaboration. Book your passes now and be part of shaping the future of data engineering.

For partnership inquiries or to express interest in speaking at the Data Engineering Summit 2024, please reach out to us at info@analyticsindiamag.com. Let’s embark on this exciting journey together, exploring the infinite possibilities that data engineering holds for the future.

The post AIM launches the 3rd Edition of Data Engineering Summit. May 30-31, Bengaluru appeared first on Analytics India Magazine.

A Roadmap For Your Data Career

Careers in data are not for everyone — you need patience to work with evolving business, security, and infrastructure requirements and a good amount of mental endurance to work with endless data issues and changes. But they can also be the most interesting jobs in the world. Every month has a new puzzle to put together in a new way, and new technologies and innovations keep the field continually new. Data careers are also going to be even more vital as AI demands larger and more diverse training datasets.

There is no one way to enter or continue a career in data, but let’s start with a reference roadmap that will serve as the framework for this article, and then we can talk about the decision points and phases that you will find along the way. Whether you are a data scientist, data engineer, software developer, business analyst, or infrastructure/security specialist, these paths could apply to you.

A Roadmap For Your Data Career
Roadmap to a Data Career

Next, before we explore the roadmap, let’s look at some of the reasons that you might be reading this article. Do any of these common complaints match your current experience?

  • “I’m not sure how to get started, and if a university education is worth the investment.”
  • “I can’t get in the door to get my first job.”
  • “My job is getting tedious and a bit boring. I’m ready for something different.”
  • “At my job, I’m stuck on old technologies or programing languages. How do I upgrade my skills?”
  • “I feel like a little gear in a big machine. I wish I had a role with more influence and visibility.”
  • “I want to get into management, but I need a path.”

We can group those concerns into the following career phases:

A Roadmap For Your Data Career
Data professional concerns by career phase.

In the roadmap, we can find key questions and considerations for each of the phases.

Career Launch

Your perspective on how to launch a career will be affected by a variety of factors.

  • Your family culture related to formal education and careers — Did your parents go to college?
  • Resources and preparation — Do you have financial help? How well did you do in high school?
  • Economic conditions — When was the last recession? How high is unemployment?
  • Career orientation — How do you define the right work/life balance? How driven are you to climb the promotion ladder?
  • Interest in theory versus application — Do you gravitate towards practical technical skills that you can use immediately? Or would you rather dive deep into computer science algorithms or statistics theory?

Based on those factors, you may decide that spending 3 months and a few hundred dollars on a few certifications in a practical, technical area like Microsoft Entra ID is right for you. On the other end of the spectrum, you may decide that 8 years to get a Bachelor’s degree and PhD in Statistics fits your aspirations. There are dozens of combinations of paths, but the main question is whether you want to make the effort to build a broad, solid foundation or if you feel like you have limited time and money and need a job ASAP.

In the launch phase, it is important to research the future roles that appeal to you and then study the job descriptions for those positions. Based on the job descriptions, you can reverse engineer the steps that you will need to take to get there. What are the the skills, programming languages, and experience needed for those roles? With those items in mind, you can create goals to upskill in those specific areas.

Regardless of your path, one of the biggest challenges is to get in the first door at a company in your target field. Having taken some random online courses won’t help in this situation. You need to find a way to add to your resume some clear credentials, an official degree, or documented accomplishments in a recognized project-experience forum like Kaggle. As you select between your options, take a close look at completion rates at the various programs and make sure that you feel like you can fully commit to finishing before you make any payments or go into debt

Finally, be creative in how you expand your network. Join a school technology club; attend meet-ups with a local tech group; look for an upcoming industry conference with networking events. And if you have attended a school with a career center, become best friends with everyone in the office and ask them for help to find the open positions that other people might overlook. You might find a niche position, inside track, or unadvertised opening, that won’t have 500 other online applicants.

Career Shift

As time passes, you will need to be prepared to shift and reskill as your company changes directions, your job function is reassigned, or you move to a new job at a new company. In the roadmap, these minor shifts are referred to as needing to “take the next step.”

If those steps require you to work with new technologies or learn a new skill that is adjacent to your current skillset, then online courses are an excellent way to help navigate the shifts.

Online courses, such as those offered by DataCamp or Pluralsight, are practical, targeted, and cost effective, and they can be assembled into a various certifications. The big challenge with these courses, though, is that their completion rate is abysmal (often under 10%). If your company offers training funds, or if you can negotiate for funds, make sure to use those funds wisely and complete your courses.

Career Progression

Career progression is often a bigger challenge for technical data professionals than managing career shifts. If you decide to seek promotion beyond the senior analyst/developer/engineer level, the challenges come in three forms:

— The skills that make a great developer, analyst or engineer are different than those needed to be a great team lead or manager (i.e. what got you here won’t get you there)

— The business awareness and context for management work is difficult to acquire while doing technical data work

— The network of relationships you have developed as an individual contributor is different than the network you will need as a manager.

A formal university master’s degree is a great way to solve all three of those issues. New masters programs are designed to be shorter, more targeted, and more flexible for the time constraints of people with a full-time job. Online degrees are not as effective at building a network of relationships, but you might decide that the flexibility is worth the tradeoff.

The NULL Path — Doing Nothing

At the center of the roadmap is the comfort zone labeled with “null.” Unfortunately, this is where a lot of us end up. We let our manager decide which skills or projects we are assigned to develop. We let company layoffs determine the timing of our next career shift. Or we look at options to upgrade our career, but decide they are too intimidating or expensive.

Your career is going to last 40–45 years; isn’t it worth being proactive about the path we want to take?

Stan Pugsley is a freelance data engineering and analytics consultant based in Salt Lake City, UT. He is also a lecturer at the University of Utah Eccles School of Business. You can reach the author via email.

More On This Topic

  • KDnuggets News, August 31: The Complete Data Science Study Roadmap…
  • Learning Data Science and Machine Learning: First Steps After The Roadmap
  • The Complete Data Science Study Roadmap
  • The Complete Data Engineering Study Roadmap
  • MLOps And Machine Learning Roadmap
  • How to get Python PCAP Certification: Roadmap, Resources, Tips For…

Groq’s LPU Demonstrates Remarkable Speed, Running Mixtral at Nearly 500 tok/s

Groq

Groq recently introduced the Language Processing Unit (LPU), a new type of end-to-end processing unit system. It offers the fastest inference for computationally intensive applications with a sequential component, such as LLMs.

It has taken the internet by storm with its extremely low latency, serving at an unprecedented speed of almost 500 T/s.

The first public demo using Groq: a lightning-fast AI Answers Engine.
It writes factual, cited answers with hundreds of words in less than a second.
More than 3/4 of the time is spent searching, not generating!
The LLM runs in a fraction of a second.https://t.co/dVUPyh3XGV https://t.co/mNV78XkoVB pic.twitter.com/QaDXixgSzp

— Matt Shumer (@mattshumer_) February 19, 2024

This technology aims to address the limitations of traditional CPUs and GPUs for handling the intensive computational demands of LLMs. It promises faster inference and lower power consumption compared to existing solutions.

Wow, that's a lot of tweets tonight! FAQs responses.
• We're faster because we designed our chip & systems
• It's an LPU, Language Processing Unit (not a GPU)
• We use open-source models, but we don't train them
• We are increasing access capacity weekly, stay tuned pic.twitter.com/nFlFXETKUP

— Groq Inc (@GroqInc) February 19, 2024

Groq’s LPU marks a departure from the conventional SIMD (Single Instruction, Multiple Data) model employed by GPUs. Unlike GPUs, which are designed for parallel processing with hundreds of cores primarily for graphics rendering, LPUs are architected to deliver deterministic performance for AI computations.

Energy efficiency is another noteworthy advantage of LPUs over GPUs. By reducing the overhead associated with managing multiple threads and avoiding core underutilisation, LPUs can deliver more computations per watt, positioning them as a greener alternative.

Groq’s LPU has the potential to improve the performance and affordability of various LLM-based applications, including chatbot interactions, personalised content generation, and machine translation. They could act as an alternative to NVIDIA GPUs especially since A100s and H100s are in such high demand.

Groq was founded in 2016 by its chief Jonathan Ross. He initially began what became Google’s TPU (Tensor Processing Unit) project as a 20% project and later joined Google X’s Rapid Eval Team before founding Groq.

The post Groq’s LPU Demonstrates Remarkable Speed, Running Mixtral at Nearly 500 tok/s appeared first on Analytics India Magazine.

SoftBank Bets $100 Bn on AI Chips, Takes Aim at NVIDIA

Competing in the AI dominance race, SoftBank’s CEO, Masayoshi Son, has announced a plan to gather a whopping $100 billion for his AI project. This move is like a direct challenge to the world’s third largest company NVIDIA, known for its computer graphic chips. According to Bloomberg sources, SoftBank wants to put $30 billion of its own money into this project and is asking for another $70 billion from investors in the Middle East.

The main goal of this huge investment is to make SoftBank a big player in the AI business and maybe shake things up for Nvidia. The plan is to work closely with a UK company called Arm, in which SoftBank already owns a 90 per cent stake following its IPO.

Recently, Arm got a lot of attention when Nvidia said they’re putting $147.3 million into it, with the blessing of SoftBank. This move shows how the tech industry is changing, with big companies trying to be on top in new areas like AI and making computer chips.

SoftBank’s journey with Arm has been kind of like a rollercoaster. They bought Arm for $32 billion in 2016, but when they tried to sell it to Nvidia for $40 billion in 2022, they ran into problems with the rules that say whether such deals are allowed. It’s a reminder that these kinds of big money moves are not easy.

But Masayoshi Son isn’t the only one with big plans in the AI world. Sam Altman who runs a company called OpenAI is another ambitious industry insider. He’s talking to investors, including the government of the UAE, to get money for making chips due to scarcity. The amount he has demanded is a shocking $7 trillion.

At several instances, Altman has repeatedly emphasised the need for additional GPU prowess to support OpenAI’s objectives in AI research and development. With global chip sales reaching $527 billion last year and projected to surpass $1 trillion annually by 2030, the race is set to intensify.

The post SoftBank Bets $100 Bn on AI Chips, Takes Aim at NVIDIA appeared first on Analytics India Magazine.

OpenAI Sora Ignites Physics Debate  

Last week, the internet went berserk with OpenAI’s first video-generation model Sora. However, at the same time, a flurry of AI experts and researchers from competitor companies were quick to dissect and criticise Sora’s transformer model, igniting a physics debate of sorts.

AI scientist Gary Marcus was one among the many to criticise not just the accuracy of the videos generated by Sora, but also the generative AI model used for video synthesis.

Competitors Unite

In a move that seemed to undermine Sora’s diffusion model structure, Meta and Google dissed the model’s understanding of the physical world.

Meta chief Yann LeCun said, “The generation of mostly realistic-looking videos from prompts does not indicate that a system understands the physical world. Generation is very different from causal prediction from a world model. The space of plausible videos is very large, and a video generation system merely needs to produce one sample to succeed.”

LeCun explains further in order to differentiate Sora from Meta’s latest AI model offering, V-JEPA (Video Joint Embedding Predictive Architecture), a model that analyses interactions between objects in videos. He said, “That is the whole point behind the JEPA (Joint Embedding Predictive Architecture), which is not generative and makes predictions in representation space” – a push to make V-JEPA’s self-supervised model seem superior to Sora’s diffusion transformer model.

Researcher and entrepreneur Eric Xing chimed in to support LeCun’s views. “An agent model that can reason based on understanding must go beyond LLMs or DMs,” he said.

The timing of the Gemini Pro 1.5 announcement couldn’t have been better. The videos generated by Sora were made to run on Gemini 1.5 Pro, where the model critiqued the inconsistencies in the video suggesting that “it is not a real-life scene”.

Elon Musk was not far behind. He called Tesla’s video-generation capabilities superior to OpenAI’s with respect to predicting accurate physics.

Source: X

While the experts have been quick to dismiss the generative model’s capabilities, the understanding of the ‘physics’ behind the model has been overlooked.

The Physics of Things

Sora uses a transformer architecture similar to GPT models, and OpenAI believes that the foundation will ‘understand and simulate the real world’, which will help towards achieving AGI. Though not called a physics engine, it is possible that Unreal Engine 5’s generated data may have been used to train Sora’s underlying model.

Senior research scientist at NVIDIA, Jim Fan, clarified OpenAI’s Sora model by explaining a data-driven physics engine. “Sora learns a physics engine implicitly in the neural parameters by gradient descent through massive amounts of videos,” he said, referring to Sora as a learnable simulator or world model.

Fan also expressed his disapproval of Sora’s reductionist views. “I see some vocal objections: ‘Sora is not learning physics, it’s just manipulating pixels in 2D’. I respectfully disagree with this reductionist view. It’s similar to saying, ‘GPT-4 doesn’t learn coding, it’s just sampling strings’. Well, what transformers do is just manipulate a sequence of integers (token IDs). What neural networks do is just manipulating floating numbers. That’s not the right argument,” he said.

Sora is at the GPT-3 Moment

Perplexity founder Aravind Srinivas, who has been vocal on social media of late, also spoke in support of LeCun. “Reality is Sora, while being amazing, is still not ready yet to model physics accurately,” he said.

Interestingly, OpenAI themselves have called out the limitations of the model before anyone could point them out. The company blog states that Sora may struggle with accurately simulating the physics of a complex scene, where it may not understand specific instances of cause and effect. It can also get confused with spatial details of a prompt, such as following a specific camera trajectory, and more.

Fan has also likened Sora with the ‘GPT-3 moment’ in 2020, when the model required ‘heavy prompting and babysitting’. However, it was the ‘first compelling demonstration of in-context learning as an emergent property’.

The current limitations do not cloud the quality of output generated. When OpenAI acquired Global Illumination, a digital product company that created open-source game Biomes (that resembles Minecraft) in August last year, the scope of video generation and building simulation-model platforms via auto agents were some of the speculations.

Now, with the release of Sora, the possibilities to disrupt the video game industry has only escalated. If Sora is at the GPT-3 moment, GPT-4 of the model will be incomprehensible. Until then, sceptics will continue to debate and probably teach one another a thing or two.

Source: X

The post OpenAI Sora Ignites Physics Debate appeared first on Analytics India Magazine.

NVIDIA Researchers Make Indic AI Model to Talk to their Spouses’ Indian Parents

These NVIDIA Researchers Made an Indic AI Model to Talk to their Spouses’ Indian Parents

NVIDIA researchers, Akshit Arora and Rafael Valle, wanted to speak to their wives’ families in their native languages. Arora, a senior data scientist supporting one of NVIDIA’s major clients, speaks Punjabi, while his wife and her family are Tamil speakers, a divide he has long sought to bridge. Valle, originally from Brazil, faced a similar challenge as his wife and family speak Gujarati.

“We’ve tried many products to help us have clearer conversations,” said Valle. This motivation led them to build multilingual text-to-speech models that could convert their voice into different languages in real time, which led them to winning competitions.

Arora, in an exclusive interview with AIM, shed more light on this. “When this competition came to our radar, it occurred to us that one of the models that we had been working on called P-Flow, would be perfect for this kind of a competition,” said Arora.

Arora and Valle, along with Sungwon Kim and Rohan Badlani, triumphed in the LIMMITS ’24 challenge, which tasked participants with replicating a speaker’s voice in real-time in different languages. Their innovative AI model achieved this feat using only a brief three-second speech sample.

Fortunately, Kim, a deep learning researcher at NVIDIA’s Seoul office, had been working on an AI model well-suited for the challenge for some time. For Badlani, residing in seven different Indian states, each with its own dominant language, inspired his involvement in the field.

The Signal Processing, Interpretation, and Representation (SPIRE) Laboratory at IISc in Bangalore orchestrated the MMITS-VC challenge, which stood as one of the major challenges within the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2024.

In this challenge, a total of 80 hours of Text-to-Speech (TTS) data were made available for Bengali, Chhattisgarhi, English, and Kannada languages. This additional dataset complemented the Telugu, Hindi, and Marathi data previously released during LIMMITS 23.

Never seen before

The competition included three tracks where the models were tested. “One of the models got into the top leaderboard on all the counts for one of the tracks,” said Arora. “In these kinds of competitions, not even a single model performs best on both the tracks. Every model is good at certain things and not good at other things,” explained Arora.

NVIDIA’s strategy for Tracks 1 and 2 revolves around the utilisation of RAD-MMM for few-shot TTS. RAD-MMM works by disentangling attributes such as speaker, accent, and language. This disentanglement enables the model to generate speech for a specific speaker, language, and accent without the need for bilingual data.

In Track 3, NVIDIA opted for P-Flow, a rapid and data-efficient zero-shot TTS model. P-Flow utilises speech prompts for speaker adaptation, enabling it to produce speech for unseen speakers with only a brief audio sample. Part of Kim’s research, P-Flow models borrow the technique large language models employ of using short voice samples as prompts so they can respond to new inputs without retraining.

One of the unique things about P-Flow is its zero-shot capabilities. “Our zero shot TTS model happened to perform the best in the zero shot category on the speaker similarity and naturalness course,” said Arora. They would also be presenting this model at GTC 2024.

A long project

Last year, the researchers also used RAD-MMM, developed by NVIDIA Applied Deep Learning Research Team, and developed “VANI” or “वाणी”, which a very lightweight multi-lingual accent controllable speech synthesis system. This was also used in the competition.

The journey began nearly two years ago when Arora and Badlani formed the team to tackle a different version of the challenge slated for 2023. Although they had developed a functional code base for the so-called Indic languages, winning in January required an intense sprint, as the 2024 challenge came onto their radar just 15 days before the deadline.

P-Flow is set to become a part of NVIDIA Riva, a framework for developing multilingual speech and translation AI software, included in the NVIDIA AI Enterprise software platform. This new capability will enable users to deploy the technology within their data centres, personal systems, or through public or private cloud services.

Arora expressed hope that their customers would be inspired to explore this technology further. “I enjoy being able to showcase in challenges like this one the work we do every day,” said Arora.

The post NVIDIA Researchers Make Indic AI Model to Talk to their Spouses’ Indian Parents appeared first on Analytics India Magazine.

Why ai.com Redirects to MKBHD YouTube’s Page?

In the past year, the ownership of the domain name ai.com has changed hands three times, with its latest buyer being Marques Keith Brownlee, also known as MKBHD—a surprise entry to the field of AI.

MKBHD, a renowned American tech YouTuber and professional ultimate frisbee player, has gained popularity for his tech-focused videos and podcasts. With a following of over 20 million subscribers and 3.93 billion views on YouTube, MKBHD is a leading figure in tech reviews, earning appreciation from industry leaders like Apple CEO Tim Cook and former Google executive Vic Gundotra.

In addition to creating videos which are characterised by easy explanation and high-quality production, the thirty year old collaborates with major tech companies for his podcast Waveform. Launched in July 2019, the podcast explores consumer electronics, featuring the latest news, product reviews, and interviews with influential figures, such as Elon Musk and Sundar Pichai.

Apart from his online presence, MKBHD is an accomplished ultimate frisbee player, competing professionally for the New York Empire in the American Ultimate Disc League. He won AUDL championships in 2019, 2022, and 2023 and achieved international success by winning the WFDF World Ultimate Club Championship with New York PoNy in 2022.

Growth Story of ai.com

Prior to MKBHD, the ownership of the domain name was under Elon Musk’s X (formerly Twitter). In November, xAI joined the chatbot competition with its generative AI chatbot Grok, developed as a prototype in less than four months.

However, the former founding member of OpenAI acquired the domain name from the company’s CEO Sam Altman in August 2023. Altman secured the username in February 2023, aiming to make AI synonymous with ChatGPT. This happened after many users on X were trolling ChatGPT over its two syllables after Google released its LLM chatbot Gemini (previously known as Bard).

OpenAI is currently basking in the success of its latest video generation model Sora which has been gaining popularity for its hyperrealistic unique features.

Likely, the ai.com domain has potentially lapsed in its renewal process, pointing towards a situation where the requisite renewal fees were not settled. This non-renewal scenario often triggers a release of the domain, making it available for acquisition by interested parties. And MKBHD was the one who ended up buying the domain. However, the real owner of the domain is still not revealed since there is no proof of a transaction yet.

The post Why ai.com Redirects to MKBHD YouTube’s Page? appeared first on Analytics India Magazine.

Tech YouTuber MKBHD Buys AI.com from Elon Musk

In the past year, the ownership of the domain name ai.com has changed hands three times, with its latest buyer being Marques Keith Brownlee, also known as MKBHD—a surprise entry to the field of AI.

MKBHD, a renowned American tech YouTuber and professional ultimate frisbee player, has gained popularity for his tech-focused videos and podcasts. With a following of over 20 million subscribers and 3.93 billion views on YouTube, MKBHD is a leading figure in tech reviews, earning appreciation from industry leaders like Apple CEO Tim Cook and former Google executive Vic Gundotra.

In addition to creating videos which are characterised by easy explanation and high-quality production, the thirty year old collaborates with major tech companies for his podcast Waveform. Launched in July 2019, the podcast explores consumer electronics, featuring the latest news, product reviews, and interviews with influential figures, such as Elon Musk and Sundar Pichai.

Apart from his online presence, MKBHD is an accomplished ultimate frisbee player, competing professionally for the New York Empire in the American Ultimate Disc League. He won AUDL championships in 2019, 2022, and 2023 and achieved international success by winning the WFDF World Ultimate Club Championship with New York PoNy in 2022.

Growth Story of ai.com

Prior to MKBHD, the ownership of the domain name was under Elon Musk’s X (formerly Twitter). In November, xAI joined the chatbot competition with its generative AI chatbot Grok, developed as a prototype in less than four months.

However, the former founding member of OpenAI acquired the domain name from the company’s CEO Sam Altman in August 2023. Altman secured the username in February 2023, aiming to make AI synonymous with ChatGPT. This happened after many users on X were trolling ChatGPT over its two syllables after Google released its LLM chatbot Gemini (previously known as Bard).

OpenAI is currently basking in the success of its latest video generation model Sora which has been gaining popularity for its hyperrealistic unique features.

However, the real owner of the domain is still not revealed since there is no proof of a transaction yet.

The post Tech YouTuber MKBHD Buys AI.com from Elon Musk appeared first on Analytics India Magazine.

This Visually-Impaired Co-Founder Duo is Challenging Disability with AI 

Growing up in Lajpat Nagar, New Delhi, visually impaired Kartik Sawhney did not have the option to study science for Classes 11 and 12 in CBSE. He had to advocate for his right to study the subject he desired and even when the opportunity was finally available to him, he encountered difficulties in accessing the necessary resources.

“Over 96% of the content available today is incompatible with various assistive technologies that people with disabilities use. So, chances that a PDF you’re going to find online is going to be accessible is only 4%,” Sawhney told AIM.

Often his mother had to translate the curriculum into Braille, to make it accessible for Kartik. This was not easy, but he, at a very early stage learnt how technology could come to his aid. He is now an engineer and entrepreneur with a background in AI and has worked with companies like Microsoft, Uber and IBM.

Seeking to help individuals like him, Sawhney, along with Shakul Sonker, who is also visually impaired, turned to AI to overcome the systematic barriers the community faces.

In 2018, they co-founded Inclusive Stem (I-Stem) an advocacy group to help their peers who were studying Science, Technology, Engineering, and Mathematics (STEM), to provide them support and mentorship. But, it was during the COVID Pandemic, that the duo decided to tap into their background and leverage AI at scale.

Sawhney, along with Sonker, who is also a computer engineer and has experience working as a machine learning engineer, started developing computer vision models.

Sonker told AIM that I-Stem’s technology is being leveraged by various universities, educational institutions and corporations across India and globally.

“For the Telangana government, our initiative involves converting K-12 books into accessible formats in both English and Telugu,” he continued, “Other customers include IIT Delhi, Ashoka University, Washington State’s Department of Services for the Blind, Google, and Intel. Ourtechnology is also being used by the United Nations (UN) and UNICEF is their first investor.”

Microsoft, GSMA, Bosch, and National Geographic Society are among the list of investors, according to Sonker. “Moreover, our partners and supporters include the Nudge Institute, Stanford StartX, Morgan Stanley, Oracle, Amazon, and Goldman Sachs.”

Going beyond Optical Character Recognition

Optical Character Recognition (OCR) served as the initial step, but its limitations are known. Although effective for text extraction from images, OCR falls short for diverse content. The goal for the duo was to go beyond OCR and leverage AI to translate complex STEM content.

“OCR gives you gibberish if you, for instance, use it for a maths problem. So, we thought, why not leverage AI to deeply comprehend the structure and then make it available in a way that is compatible with the screen readers, which visually impaired individuals use,” Sawhney said.

Most of the documents are not compatible with the screen readers and that is the big problem, according to Sawhney. What I-Stem developed is a deep learning computer vision model trained on diverse data, which includes academic papers, pamphlets, posters, presentations, receipts, invoices, menus, etc., ensuring comprehensive coverage for conversion into accessible formats.

The model has an accuracy of 92% and can read complex documents. The system excels with STEM, multicolumn, and finance documents—representing the majority of the training data. As the name of the company suggests, the duo has also emphasised making STEM content accessible, addressing a field that was previously less available to certain individuals.

Here, Sonker adds that since the accuracy is not fully 100%, the startup has a manual team of remediators who ensure the documents converted are accurate. This helps ensure that when a screen reader reads it, the content is accurate and is in a format which is optimised for the best user experience.

Moreover, the startup is now working towards leveraging generative AI models like GPT-V, the vision model from OpenAI. Currently, conventional AI models can identify an image but struggle with detailed descriptions. But foundational vision models make comprehensive image descriptions possible.

“Additionally, we enable users to ask questions using visual question answering, enhancing understanding of specific details in diagrams,” Sawhney said.

Changing the discourse on disability

The duo also clearly understands that ensuring content accessibility is insufficient; it’s merely the beginning of the challenges the community encounters. Hence, a crucial part of their efforts includes community programmes.

To make their technology widely accessible, I-Stem has teamed up with non-profits that deliver the technology on the ground and ensure they provide support in classrooms.

In the present scenario, disability is often linked with low expectations. By collaborating with individuals with disabilities and showcasing their accomplishments, Sawhney and Sonker want to challenge the misconception that disability equates to limited expectations.

“Demonstrating that people with disabilities can hold positions of power helps transform the overall narrative around disability. With universities, we administer a fellowship programme targeting individuals with high potential and disabilities.

“We believe that by assisting these individuals in securing meaningful roles within high-growth industries, they can serve as influential role models. This not only paves the way for their success but also contributes to reshaping the narrative and discourse surrounding disability,” Sawhney said.

Moreover, over 70% of the disabled community in India is unemployed, which is more than the population of Sri Lanka, Sawhney said who also believes this is contributing to an annual loss of nearly USD 12 billion to the Indian economy.

To solve this problem, I-Stem has launched a hiring portal that helps disabled candidates connect with employers based on their past experience, work preferences, aspirations and disability.

The post This Visually-Impaired Co-Founder Duo is Challenging Disability with AI appeared first on Analytics India Magazine.