In 2023, AI revolutionized manufacturing operations by automating costly processes. In 2024, AI factories will revolutionize manufacturing by managing and analyzing data from sensors, boosting productivity, quality, and setup costs through technologies like NVIDIA Tesla in the cloud.
Here are some of the examples:
NVIDIA x Foxconn
Foxconn recently announced partnering with NVIDIA to develop AI factories, a GPU computing infrastructure for processing and transforming vast amounts of data into valuable AI models and tokens. The partnership will also include the development of smart solution platforms like Foxconn Smart EV, built on NVIDIA DRIVE Hyperion 9, NVIDIA DRIVE Thor, NVIDIA Isaac autonomous mobile robot platform, and Foxconn Smart City, which incorporates the NVIDIA Metropolis intelligent video analytics platform.
The aim is to help the industry move faster into the new AI era, as stated by Foxconn Chairman and CEO Young Liu. Key NVIDIA technologies used include NVIDIA HGX reference designs, NVIDIA GH200 Superchips, NVIDIA OVX reference designs, and NVIDIA networking. Foxconn is also considering its own AI factory, using the NVIDIA Omniverse platform and Isaac and Metropolis frameworks to meet strict production and quality standards in the electronics industry.
NVIDIA x BMW
NVIDIA CEO Jensen Huang presented a vision of the future of manufacturing factories, blending reality, virtual reality, robotics, and AI to manage BMW’s automotive factories. The AI factory demo showcased the NVIDIA Omniverse Enterprise, the first technology platform enabling global 3D design teams to work simultaneously across multiple software suites in a shared virtual space. The AI factory demo included the NVIDIA Isaac platform for robotics, the NVIDIA EGX edge computing platform, and the NVIDIA Aerial software development kit.
BMW’s factory lines can produce up to 10 different cars, with over 100 options for each car and more than 40 BMW models. The technology revolutionizes BMW’s planning processes, allowing workers to work in a perfect simulation without having to travel. The use of NVIDIA Omniverse and NVIDIA AI allows for the simulation of 31 factories in a production network, enabling AI-enabled use cases such as virtual factory planning, autonomous robots, predictive maintenance, and big data analytics.
NVIDIA x SIEMENS
Siemens Energy, a leading supplier of power plant technology, is using the NVIDIA Omniverse platform to create digital twins for predictive maintenance of power plants. This is part of a wave of companies using digital twins to enhance their operations, including BMW Group and Ericsson. The global market for digital twin platforms is forecasted to reach $86 billion by 2028. Siemens Energy builds and services combined cycle power plants, including gas and steam turbines.
A 10% reduction in the industry’s average planned downtime for heat recovery steam generators could save $1.7 billion a year. Siemens Energy is using NVIDIA technology to develop a workflow to reduce planned shutdowns while maintaining safety. Real-time data is preprocessed to compute pressure, temperature, and velocity of water and steam, which is fed into a physics-ML model created with the NVIDIA Modulus framework. The flow conditions are visualized using NVIDIA Omniverse, a virtual world simulation and collaboration platform for 3D workflows. This helps Siemens Energy fine-tune maintenance needs without running the risk of failure.
The post 2024 is all about AI Factories appeared first on Analytics India Magazine.
As 2023 draws to a close, it’s time to look back and appreciate the art that has illuminated our understanding and enhanced our engagement with key concepts and stories.
In this special feature, “Top Illustrations of 2023 by AIM: Editor’s Pick,” we celebrate the creativity and talent behind the most striking and impactful illustrations created by our team this year. These visuals range from thought-provoking representations of technological advancements to nuanced depictions of societal themes, each embodying a unique blend of artistic skill and editorial insight.
Check Past Top Illustrations:
2022 | 2021
The post Top Illustrations of 2023 by AIM: Editor’s Pick appeared first on Analytics India Magazine.
With the resounding success of Chandrayaan 3, the successful launch of Aditya-L1—India’s mission to the L1 range to study the sun, and the release of India’s Space Policy 2023, this year has accelerated India’s growth in the space sector by an astounding margin and has brought global attention to the country’s efforts in the vast unexplored expanse of space.
However, the Indian Space Research Organisation is not in the mood for a breather and has a number of missions lined up for the upcoming year. These missions ranging from India’s manned mission to space to the exploration of Venus are ready to push past existing limits.
1. X-ray Polarimeter Satellite
The year is going to start with a bang with ISRO all set to launch the first mission on the very first day of the year.
ISRO’s XPoSat mission, set to launch via a Polar Satellite Launch Vehicle (PSLV) at 9:10 am, aims to unravel complexities in understanding celestial light. Over five years, the mission will gauge light wave vibrations to decode radiation mechanisms and celestial geometry, marking a significant leap in comprehension. The spacecraft houses two key scientific payloads: POLIX (Polarimeter Instrument in X-rays) and XSPECT (X-ray Spectroscopy and Timing). POLIX will measure polarisation in the medium X-ray range (8-30 keV), while XSPECT offers spectroscopic insight (0.8-15 keV). These payloads collectively enable the mission to delve into temporal, spectral, and polarimetric aspects of bright X-ray sources, aiding in comprehending the emission mechanisms of cosmic entities like black holes, neutron stars, and pulsar wind nebulae.
Additionally, the insertion of Aditya L1 in the Langragian point1 is also scheduled for January 6th, which is sure to make the first month of the year quite memorable.
2. INSAT 3DS
Following the XPoSat mission India is all set to launch the INSAT 3DS around January 12. INSAT-3DS is a meteorological satellite which will be launched using ISRO’s Geosynchronous Launch Vehicle (GSLV-F14). The INSAT-3DS mission, a collaboration between ISRO and IMD, aims to enhance climate services through a network of climate observatory satellites. It’s part of a multi-mission meteorological data system, launching three dedicated Earth observation satellites. This final launch includes INSAT-3DR, a meteorological satellite featuring a 19-channel sounder, a 6-channel imager, and specific payloads. Successfully launched in 2016, INSAT-3DR is the improved successor to INSAT-3D, omitting additional communication payloads.
The upcoming INSAT-3DS mission marks the seventh flight in the INSAT-3D series, undergoing vibration tests as of November 10, 2023, signalling its readiness for launch.
3. NISAR
The Nasa-Isro Synthetic Aperture Radar mission (NISAR), a collaborative effort between India and the US, is set for a probable launch in the first quarter of the upcoming year, around late February or early March. It’s a Low Earth Orbit (LEO) observatory designed to map the Earth every 12 days, offering consistent data for understanding ecosystem changes, ice mass variations, sea level rise, and natural hazards like earthquakes, tsunamis, and volcanoes.
This mission employs dual-band Synthetic Aperture Radar (SAR) in L and S bands, utilizing the Sweep SAR technique for high-resolution data over a wide area. The observatory consists of SAR payloads on the Integrated Radar Instrument Structure (IRIS) and the spacecraft bus. Jet Propulsion Laboratories and ISRO collaborate on the observatory, catering to national needs and supporting scientific research on surface deformation through repeat-pass InSAR.
In this partnership, NASA provides the L-Band SAR payload, while ISRO contributes the S-Band SAR payload, both using a common large reflector antenna. Additionally, NASA offers engineering payloads like the Payload Data Subsystem, High-rate Science Downlink System, GPS receivers, and a Solid State Recorder.
NISAR’s significance lies in being the first dual-frequency radar imaging mission using advanced Sweep SAR techniques, providing high-resolution L and S band SAR data. This capability aids in studying various complex phenomena like ecosystem changes, ice sheet behavior, and natural hazards, fostering microwave remote sensing applications in geosciences.
4. Gaganyaan 1
The Gaganyaan 1 project, scheduled for launch later this year, aims to demonstrate India’s human spaceflight capability by sending a crew of three members to orbit 400 km above Earth for a three-day mission, returning safely by landing in Indian sea waters.
Precursor missions like the Integrated AirDrop Test (IADT), Pad Abort Test (PAT), and Test Vehicle (TV) flights have been planned to demonstrate technology readiness levels before the actual human spaceflight mission. However, from this set, the Crew Escape System has already been validated. Unmanned missions preceding the manned mission will ensure the safety and reliability of all systems.
The Gaganyaan mission will utilise the Human Rated LVM3 rocket, derived from the well-proven LVM3 launcher. This launch vehicle is adapted to meet human rating requirements and features the Crew Escape System (CES) to ensure crew safety during emergencies.
The Orbital Module (OM) comprises the Crew Module (CM) and Service Module (SM). The CM provides a habitable space for the crew in space, equipped with life support systems, avionics, and deceleration systems for crew safety during descent. The SM supports the CM with thermal, propulsion, power, and avionics systems.
Several new technologies, engineering systems, and human-centric systems are being developed to ensure human safety during the mission. The Astronaut Training Facility in Bengaluru conducts diverse training covering academic courses, flight systems, microgravity familiarisation, aero-medical training, recovery, and survival techniques to prepare the crew for the Gaganyaan mission.
5. Mangalyaan 2
Following India’s historic entry into Mars’ orbit 9 years ago in 2014 with Mangalyaan-1, ISRO is preparing for the Mangalyaan-2 mission, slated for launch in 2024. France will collaborate with India for this mission, intending to probe Mars for signs of life and water. Mangalyaan-2 aims to study various aspects of Mars, including interplanetary dust and the Martian atmosphere.
This mission will carry four payloads: MODEX to analyse the origin and distribution of dust on Mars, a Radio Occultation experiment for studying the Martian atmosphere, an Energetic Ion Spectrometer to characterize solar particles, and a Langmuir Probe and Electric Field Experiment to understand Mars’ plasma environment.
During the then French President Francois Hollande’s visit to India, a partnership was formed between CNES and ISRO in 2016, signing agreements to jointly develop a Mars lander for exploration. The mission is aimed at discovering potential signs of life and water on Mars.
Mangalyaan-1 was launched in 2013, with the orbiter successfully entering Mars’ orbit after a nine-month journey. India’s accomplishment with Mangalyaan-1 marked the first time a probe reached Mars on its initial attempt, distinguishing it from previous European, American, and Russian missions that required multiple tries.
6. Shukrayaan 1
ISRO has also turned its attention to Venus, progressing with the development of the Venus mission, often called Shukrayaan-1. The mission aims to comprehensively study Venus’ surface and atmosphere, exploring its geological composition and atmospheric characteristics. The craft, derived from Sanskrit words meaning “Venus” and “craft,” is being configured with various payloads for this purpose.
Shukrayaan-1’s development started in 2012 when ISRO called for payload proposals from research institutes. While progressing steadily, key details such as the launch date are yet to be revealed by ISRO. Previous missions to Venus by agencies like the European Space Agency and Japan, along with flybys from NASA’s Parker Solar Probe, have contributed valuable data, including the recent capture of visible light images of Venus’ surface by NASA in February 2022 during a flyby.
Despite Venus’ harsh conditions with a dense atmosphere filled with acids and extremely high atmospheric pressure, akin to a hundred times that of Earth, scientists remain intrigued by the planet often referred to as “Earth’s twin.” There’s an ongoing debate about the possibility of microbial life in Venus’ upper atmosphere, despite doubts expressed by NASA about life on the planet.
The post ISRO’s Missions to Watchout for in 2024 appeared first on Analytics India Magazine.
2023 has been a vibrant year for AI and data science enthusiasts, marked by a series of hackathons that challenged minds and fostered innovation. Here’s an overview of the top 10 hackathons that left a significant mark this year.
1. Google Cloud BigQuery LTV Prediction Challenge
Total Registrations: 889
Focus: Participants used BigQuery and Colab to predict customer LTV from e-commerce data.
Details
2. Predict The Price Of Books Challenge
Total Registrations: 5443
Challenge: Predict book prices using a dataset featuring diverse genres and authors.
Details
3. oneAPI Hackathon: The LLM Challenge
Total Registrations: 1067
Objective: Develop a question-answering system using Large Language Models (LLMs).
Details
4. Analytics Olympiad 3.0
Total Registrations: 1051
Theme: Predicting Customer Loan Default in the BFSI sector.
Details
5. AI Blogathon Contest – “The AI Odyssey”
Total Registrations: 166
Format: Blogging contest focused on Generative AI.
Details
6. Genpact – Google Sustainability Hackathon
Total Registrations: 4076
Aim: Utilize AI/ML to model the impact of climate change on commodity prices.
Details
7. Data Science Student Championship 2023
Total Registrations: 857
Challenge: Predicting ‘total_fare’ for taxi rides.
Details
8. Cyber Threat Detection Hackathon by Shell
Total Registrations: 1033
Objective: Build a model to detect code in text to enhance web application security.
Details
9. LLM – Detect AI Generated Text on Kaggle
Focus: Develop models to distinguish between student and AI-generated essays.
Prize Pool: $110,000
Details
10. NFL Big Data Bowl 2024
Theme: Evaluate tackling tactics and strategy in American football.
Prize Pool: $100,000
Details
Each of these hackathons not only tested the skills and creativity of participants but also contributed to advancing the fields of AI and data science. From tackling real-world issues to enhancing theoretical understanding, these events have set new benchmarks in technological innovation and collaborative problem-solving.
The post Top 10 AI Hackathons of 2023: Showcasing Ingenuity and Technological Mastery appeared first on Analytics India Magazine.
Intel has a well thought out plan for the upcoming years when it comes to AI. And it was pretty evident at the recent AI Everywhere event hosted by the company, where the company announced a line of products like AI PCs, five semiconductor nodes, and the upcoming Gaudi3 AI accelerator.
“Shortly after my return to Intel Corporation, we set out a tremendously audacious goal – to manufacture five new process technology nodes in just four years,” Pat Gelsinger, CEO of Intel posted on LinkedIn.
Two nodes, Intel & and Intel Core Ultra were announced at the AI Everywhere event. While Intel 3 is going into production next year, Intel 20A which he calls a work of art, and the last 18A will be available by the end of next year and are developing in the fab. Gelsinger posed with them like a happy family.
“I like to just have one other little thing to show off here and they just brought it out of the lab,” just before the end of the conference, Gelsinger walked in with the first ever Gaudi3 AI accelerator.
“When we first announced this goal, many stated this would be nearly impossible. But we’re doing what we set out to do – a decade of semiconductor work in just four years. Two out of these five nodes are complete. The remaining three are underway and on-track.” Intel’s roadmap for AI looks great and a lot of it is already underway.
Gelsinger believes that the supercomputer that Intel is building will be the largest in Europe, powered by Gaudi3. Intel also announced the 5th generation of Xeon processors that would power Microsoft, Google, and IBM’s data centres.
The expansionist agenda
To ensure that the roadmap works well, Intel has secured a $3.2 billion grant from the Israeli government for the construction of a new $25 billion chip plant in southern Israel. This marks the largest investment ever made by a company in Israel. Intel’s expansion plan at its Kiryat Gat site, located 42 km from Hamas-controlled Gaza, is described by Intel as a crucial step in fostering a more resilient global supply chain.
Gelsinger has projected that the GPU market size would be around $400 billion by 2027. This definitely gives room for a lot of competitions to thrive, and thus there is a lot expected from Intel. Thus, he has led Intel’s substantial investments in building chip factories across three continents to regain dominance in chip manufacturing.
In Germany, Intel is set to invest over 30 billion euros ($33 billion) in establishing two chip-making plants in Magdeburg, as part of a multi-billion-dollar investment initiative in Europe to boost chip production. Germany has offered substantial subsidies to secure its largest-ever foreign investment.
Additionally, Intel also announced a plan in 2022 to invest up to $100 billion in building what could be the world’s largest chip-making complex in Ohio.
Intel also has a bunch of partnerships for its AI PCs goals. At the event, Gelsinger and his team announced the launch of Intel Core Ultra and Intel Arc GPUs for pushing the goal of making every PC in the world an AI PC. This will be achieved by its partnership with Dell, Lenovo, HP, Supermicro, and Microsoft, for bringing the hardware onto their devices.
Intel’s unique position in the industry, marked by openness, scalability, and end-to-end solutions, allows for the seamless infusion of AI capabilities into every platform. This is well established given that Intel was the first to produce the first commercial microprocessor chip in 1971, and has been the leader in developing components for PCs.
A long roadmap
During Intel’s recent conference call, Gelsinger stated, “Our Gaudi roadmap remains on track with Gaudi3 out of the fab, now in packaging and expected to launch next year. Looking ahead to 2025, Falcon Shores will integrate our GPU and Gaudi capabilities into a unified product.”
Gaudi3 is expected to arrive with a 5nm chip. The accelerators are set to provide a significant boost with up to 4 times the BFloat16 capabilities, double the compute power, 1.5 times the network bandwidth, and a 1.5 times increase in HBM capacities (144 GB compared to 96 GB).
In 2025, the successor to Gaudi3, Falcon Shores, will merge the AI capabilities of Gaudi with the powerful GPUs from Intel, all within a single package. This is something that would give Intel the edge over others.
Intel is also planning to onboard another version of the AI accelerator superchip, Falcon Shores 2, by 2026, which would be based on the Gaudi3 architecture. “We have a simplified roadmap as we bring together our GPU and accelerators into a single offering,” Gelsinger said. Though this is a far-out vision, the reveal at the AI Everywhere conference also gives out some hope for the company.
Intel has a long road ahead to compete with NVIDIA and AMD, and Gelsinger is definitely helping the company pave the path faster.
The post What is Intel’s AI Plan for 2024 appeared first on Analytics India Magazine.
In the heart of the seemingly-other-worldly expanse of Antarctica lies an outpost belonging to the Indian Space Research Organisation (ISRO), marking the agency’s presence beyond the habitable lanes of Bengaluru and Hyderabad—in an environment where temperatures range between -89 degrees Celsius in winter to -25 degrees Celsius in summer.
The Antarctica Ground Station for Earth Observation Satellites (AGEOS), situated at Bharati Station in the Larsemann Hills was commissioned in August 2013 and can house 72 personnel during summers.
It operates as a crucial link in ISRO’s data acquisition network, to enhance satellite connectivity (especially from polar orbits), data reception(enabling real-time reception) and processing from various Indian Remote Sensing (IRS) satellites like CARTOSAT-2 Series, SCATSAT-1, RESOURCESAT-2/2A, and CARTOSAT-1. This data is then seamlessly transmitted to the National Remote Sensing Centre (NRSC) in Shadnagar, near Hyderabad, marking a crucial link in India’s Earth Observation (EO) missions.
Why Antarctica?
Its strategic positioning provides a unique advantage, allowing up to 10 passes per day for satellite observation, offering comprehensive coverage crucial for various sectors, including agriculture, water resources, urban planning, disaster management, and more.
The ground station’s location in Antarctica is strategically chosen because of its unique position close to the South Pole—which enables increased satellite visibility and data reception up to 10 satellite passes per day from satellites in polar orbits.
As satellites orbit the Earth, those in polar orbits pass over the poles multiple times a day. Stations in Antarctica, being closer to the South Pole, can intercept these passes more effectively compared to stations located in higher latitudes. This increased frequency of satellite passes amplifies the ground station’s capacity to receive crucial data which proves vital across various sectors like agriculture, water resources, urban planning, and disaster management.
What else does AGEOS do?
The extended S/X/Ka antenna system at AGEOS plays a pivotal role in supporting ISRO’s remote sensing missions, enhancing the country’s capabilities in remote sensing. This tri-band approach also separates the national space agency from agencies or entities from other countries, elevating it into a league of its own. Manned by a team of ISRO engineers stationed at Bharati, the facility ensures seamless operations, contributing significantly to the success of ISRO’s endeavours in remote sensing.
The facility is proving to be a critical resource with the increasing number of launches and satellites in orbit while only a few ground stations exist. It is also used to track any launch aboard the PSLV through antennas.
Moreover, AGEOS not only serves ISRO’s remote sensing missions but also houses a C-Band station at the National Centre for Antarctica & Ocean Research (NCAOR) in Goa, India. This station serves as a dedicated communication link, facilitating round-the-clock operations and supporting vital applications like video conferencing, streaming, and internet browsing.
India has also commissioned the tender for an advanced Ka-band satellite link, significantly enhancing connectivity and enabling high-data-transfer satellite internet services. The colossal significance of this venture reverberates in its vital role for the upcoming NASA-ISRO Synthetic Aperture Radar (NISAR) space mission, a collaborative effort between Indian and United States scientists—which will amass extensive data regarding minute changes in ice sheets and rising sea levels, crucial for understanding the ramifications of global warming. The anticipated data load from this mission could surpass a staggering 80 terabytes per day, which will be shouldered by the Ka-band satellite link.
Apart from these primary functions, this station plays a key role in aiding India’s scientific community in conducting research at Maitri, another Indian station in Antarctica. This comprehensive infrastructure and connectivity demonstrate India’s commitment to scientific exploration and advancement even in the most extreme environments on Earth.
The relentless operations and maintenance of AGEOS are upheld by a team of ISRO engineers deputed to Bharati Station—who ensure round-the-clock monitoring, maintenance, and troubleshooting of the technological infrastructure to ensure optimal functionality.
Additionally, the new Ka-band installation at the Bharati station further elevates India’s capabilities in remote sensing and brings the other Indian station on the southern pole— Maitri, closer through stable internet, marking a significant leap in connectivity and scientific collaboration amidst Antarctica’s icy vastness and mainland India.
The post What’s ISRO Doing in the Middle of Antarctica? appeared first on Analytics India Magazine.
Chris Lattner, the man who walks around fixing programming languages and compilers, doesn’t spend any time worrying about AGI personally.
“People worry about what happens when a computer is smarter than us, but at the same time, we are surrounded by other people smarter than us. At least I am,” mused the CEO of Modular AI casually.
He believes that we’re a long way away from AI replacing the human capability of the world. “I have major questions about how much compute will take when we power it. From a pure technology level, I don’t think that’s on the cusp,” he remarked about the billions funnelled into building AI models in 2023 alone.
Lattner suggested that folks worried about this should take a step back and realise we already have super intelligence. “It’s made out of federated groups of humans with shared goals. That’s what I think has been true for hundreds of years, and that’s where progress is made,” he said.
Speaking of progress, the startup challenging NVIDIA, the top supplier of AI chips, recently secured $100 million from Silicon Valley heavyweight General Catalyst. The lean 76-person startup is building its toolkit, Mojo intended to be easier to use across different hardware and a software suite, the Modular Engine, which is avant-garde and allows users to customise and run their AI software.
Lattner wants people to be able to express themselves and create things. “I love that somebody doesn’t have to worry about the exact syntax because it is mechanical. That is not the point of coding,” he said. “That’s just something we must do because we need to express ourselves. It gets more and more people involved, making things more accessible. It’s just giving people superpowers,” the software genius added.
Not A Typical AI Company
There’s a clear message on Modular AI’s website: It focuses on AI because, as Lattner puts it, “that’s where most of the suffering is.” Lattner knows firsthand the struggles of programmers. “It bothers me,” he shares, showing a real understanding of the field.
For him, Modular is about making the complex world of AI more straightforward. He explained, “It is about making it possible for more people to participate in this ecosystem.” He points out that only big players like Google and OpenAI seem to dominate AI right now. The company has large teams of experts. But Lattner believes this approach leaves out many talented people worldwide. His vision for Modular AI is to change that.
Before launching Mojo through Modular AI, Lattner has contributed to open-source projects like the LLVM Compiler Infrastructure, Clang, and Swift, a programming language used in many Apple products.
Tech Now and Then
Growing up, Lattner started learning about computers and programming as a kid. Now, at 45, he’s spent decades creating tools to help other programmers build things. Around 2016, when AI was in its early stages, he got interested in the subject. He tried to get folks at Apple to understand why it was necessary. “There were always more important things, and they were not excited about this,” he said.
“I decided to go on a hero’s journey of understanding how all the technology worked and spent several years at Tesla, Google and other places, learning the fundamentals, over the last 5 to 7 years,” said Lattner, who has led teams at Apple, Tesla, Google and most recently, SiFive.
Looking to 2024, he thinks it takes quite a while for a technology base to build out to the point where the world understands it. Much of the AI out there is demo quality, said Lattner.
“2023 was the year of the language model demo. There wasn’t a tonne of language models impacting products. There were a lot of ChatGPT wrappers, but the impact was pretty low. One of the reasons for that is that technology is problematic. But it is disproportionately important, and there’s something real there, “he asserted.
Lattner predicts that 2024 will be the year of Generative AI getting into products.
Lattner further highlighted that Generative AI is not classical AI running on neural nets. “It’s a neural net embedded into larger applications where new data comes in. Language models have to be tokenized, a lot of this is non-traditional, and the stacks people have been writing on were never designed for that.”
For his company’s future, Lattner hinted that in the next few quarters, there will be continuous announcements about capabilities, new product features, APIs, ecosystem components, and even more things built on top of Max and Mojo.
“Right now, we’re very focused on inference; we’ll soon go into training. Training is a major sore point as models get larger; that’s a big deal for the world. We’ll support new hardware to support more of the AI workflow. We try to do everything to the best quality, so one of the things that is true about modular that annoys some people is that we move slowly,” he admitted.
The post Chris Lattner on Ending AI Suffering appeared first on Analytics India Magazine.
In a significant move signaling strengthened collaboration between the Adani Group and its backers in West Asia, a joint venture has been established with a unit of UAE’s International Holding Co (IHC). The conglomerates, led by Gautam Adani and UAE National Security Adviser Sheikh Tahnoon bin Zayed Al Nahyan, respectively, are set to explore AI and blockchain.
According to an exchange filing on Thursday, Adani Global Ltd. and IHC’s Sirius International Holding Ltd. will hold 49% and 51% stakes, respectively, in the newly formed entity named Sirius Digitech International Ltd., headquartered in Abu Dhabi.
The collaboration extends beyond AI, encompassing IoT and blockchain technologies, with both partners enjoying equal representation on the board.
Adani Global, a wholly owned subsidiary of flagship Adani Enterprises Ltd., has a history of incubating new businesses within its diverse portfolio, which ranges from ports to power. The joint venture’s primary objective is to capitalize on the vast $175 billion digitisation opportunity in India, spanning sectors such as Fintech, Healthtech, and Greentech, as outlined in the filing.
This partnership comes on the heels of IHC’s increased stake in the Adani Group to nearly 5% in October, highlighting renewed support for the conglomerate. The Adani Group, having weathered a short seller attack in January, has been steadily regaining investor and lender confidence.
IHC, a formidable $239 billion conglomerate led by Sheikh Tahnoon bin Zayed Al Nahyan, also oversees a $1.5 trillion empire, which includes G42—an Abu Dhabi firm playing a pivotal role in the UAE’s advancement into AI. The collaboration underscores the Adani Group’s resilience and its strategic alignment with influential partners as it seeks to leverage emerging technologies and capitalize on India’s burgeoning digitization landscape.
In January, Adani Group had also set up its AI lab in Tel Aviv.
In other news, Mukesh Ambani’s Reliance Jio Infocomm has forged an alliance with IIT Bombay for the ‘Bharat GPT’ initiative, according to revelations made by the company’s chairman, Akash Ambani. This collaboration, rooted in a partnership established in 2014, underscores a concerted effort to propel technological advancements.
The post Adani Groups Set an AI Joint Venture with UAE’s IHC appeared first on Analytics India Magazine.
ChatGPT and other generative AI programs spit out "hallucinations," assertions of falsehoods as fact, because the programs not being built to "know" anything; they are simply built to produce a string of characters that is a plausible continuation of whatever you've just typed.
"If I ask a question about medicine or legal or some technical question, the LLM [large language model] will not have that information, especially if that information is proprietary," said Edo Liberty, CEO and founder of startup Pinecone, in an interview recently with ZDNET. "So, it will just make up something, what we call hallucinations."
Liberty's company, a four-year-old, venture-backed software maker based in New York City, specializes in what's called a vector database. The company has received $138 million in financing for the quest to ground the merely plausible output of GenAI in something more authoritative, something resembling actual knowledge.
"The right thing to do, is, when you have the query, the prompt, go and fetch the relevant information from the vector database, put that into the context window, and suddenly your query or your interaction with the language model is a lot more effective," explained Liberty.
Vector databases are one corner of a rapidly expanding effort called "retrieval-augmented generation," or, RAG, whereby the LLMs seek outside input in the midst of forming their outputs in order to amplify what the neural network can do on its own.
Of all the RAG approaches, the vector database is among those with the deepest background in both research and industry. It has been around in a crude form for over a decade.
In his prior roles at huge tech companies, Liberty helped pioneer vector databases as an under-the-hood, skunkworks affair. He has served as head of research for Yahoo!, and as senior manager of research for the Amazon AWS SageMaker platform, and, later, head of Amazon AI Labs.
"If you look at shopping recommendations at Amazon or feed ranking at Facebook, or ad recommendations, or search at Google, they're all working behind the scenes with something that is effectively a vector database," Liberty told ZDNET.
For many years, vector databases were "still a kind of a well-kept secret" even within the database community, said Liberty. Such early vector databases weren't off-the-shelf products. "Every company had to build something internally to do this," he said. "I myself participated in building quite a few different platforms that require some vector database capabilities."
Liberty's insight in those years at Amazon was that using vectors couldn't simply be stuffed inside of an existing database. "It is a separate architecture, it is a separate database, a service — it is a new kind of database," he said.
It was clear, he said, "where the puck was going" with AI even before ChatGPT. "With language models such as Google's BERT, that was the first language model that started picking up steam with the average developer," referring to Google's generative AI system, introduced in 2018, a precursor to ChatGPT.
"When that starts happening, that's a phase transition in the market." It was a transition that he had to jump on, he said.
Also: Bill Gates predicts a 'massive technology boom' from AI coming soon
"I knew how hard it is, and how long it takes, to build foundational database layers, and that we had to start ahead of time, because we only had a couple of years before this would become used by thousands of companies."
Any database is defined by the ways that data are organized, such as the rows and columns of relational databases, and the means of access, such as the structured query language of relational.
In the case of a vector database, each piece of data is represented by what's called a vector embedding, a group of numbers that place the data in an abstract space — an "embedding space" — based on similarity. For example, the cities London and Paris are closer together in a space of geographic proximity than either is to New York. Vector embeddings are just an efficient numeric way to represent the relative similarity.
In an embedding space, any kind of data can be represented as closer or farther based on similarity. Text, for example, can be thought of as words that are close, such as "occupies" and "located," which are both closer together than they are near a word such as "founded." Images, sounds, program codes — all kinds of things can be reduced to numeric vectors that are then embedded by their similarity.
To access the database, the vector database turns the query into a vector, and that vector is compared with the vectors in the database based on how close it is to them in the embedding space, what's known as a "similarity search." The closest match is then the output, the answer to a query.
You can see how this has obvious relevance for the recommender engines: two kinds of vacuum cleaners might be closer to each other than either is to a third type of vacuum. A query for a vacuum cleaner might be matched for how close it is to any of the descriptions of the three vacuums. Broadening or narrowing the query can lead to a broader or finer search for similarity throughout the embedding space.
Also: Have 10 hours? IBM will train you in AI fundamentals — for free
But similarity search across vector embeddings is not itself sufficient to make a database. At best, it is a simple index of vectors for very basic retrieval.
A vector database, Liberty contends, has to have a management system, just like a relational database, something to handle numerous challenges of which a user isn't even aware. That includes how to store the various vectors across the available storage media, and how to scale the storage across distributed systems, and how to update, add and delete vectors within the system.
"Those are very, very unique queries, and very hard to do, and when you do that at scale, you have to build the system to be highly specialized for that," said Liberty.
"And it has to be built from the ground up, in terms of algorithms and data structures and everything, and it has to be cloud-native, otherwise, honestly, you can't really get the cost, scale, performance trade-offs that make it feasible and reasonable in production."
Matching queries to vectors stored in a database obviously dovetails well with large language models such as GPT-4. Their main function is to match a query in vector form to their amassed training data, summarized as vectors, and to what you've previously typed, also represented as vectors.
"The way LLMs [large language models] access data, they actually access the data with the vector itself," explained Liberty. "It's not metadata, it's not an added field that is the primary way that the information is represented."
For example, "If you want to say, give me everything that looks like this, and I see an image — maybe I crop a face and say, okay, fetch everybody from the database that looks like that, out of all my images," explained Liberty.
"Or if it's audio, something that sounds like this, or if it's text, it's something that's relevant from this document." Those sorts of combined queries can all be a matter of different similarity searches across different vector embedding spaces. That could be particularly useful for the multi-modal future that is coming to GenAI, as ZDNET has related.
The whole point, again, is to reduce hallucinations.
Also: 8 ways to reduce ChatGPT hallucinations
"Say you are building an application for technical support: the LLM might have been trained on some random products, but not your product, and it definitely won't have the new release that you have coming up, the documentation that's not public yet." As a consequence, "It will just make up something." Instead, with a vector database, a prompt pertaining to the new product will be matched to that particular information.
There are other promising avenues being explored in the overall RAG effort. AI scientists, aware of the limitations of large language models, have been trying to approximate what a database can do. Numerous parties, including Microsoft, have experimented with directly attaching to the LLMs something like a primitive memory, as ZDNET has previously reported.
By expanding the "context window," the term for the amount of stuff that was previously typed into the prompt of a program such as ChatGPT, more can be recalled with each turn of a chat session.
That approach can only go so far, Liberty told ZDNET. "That context window might or might not contain the information needed to actually produce the right answer," he said, and in practice, he argues, "It almost certainly will not."
"If you're asking a question about medicine, you're not going to put in the context window all of the knowledge of medicine," he pointed out. In the worst-case scenario, such "context stuffing," as it's called, can actually exacerbate hallucinations, said Liberty, "because you're adding noise."
Of course, other database software and tools vendors have seen the virtues of searching for similarities between vectors, and are adding capabilities to their existing wares. That includes MongdoDB, one of the most popular non-relational database systems, which has added "vector search" to its Atlas cloud-managed database platform. It also includes small-footprint database vendor Couchbase.
"They don't work," said Liberty of the me-too efforts, "because they don't even have the right mechanisms in place."
The means of access of other database systems can't be bolted to vector similarity search, in his view. Liberty offered an example of recall. "If I ask you what is your most recent interview you've done, what happens in your brain is not an SQL query," he said, referring to the structured retrieval language of relational databases.
"You have connotations, you can fetch relevant information by context — that similarly or analogy is something vector databases can do because of the way they represent data" that other databases can't do because of their structure.
"We are highly specialized to do vector search extremely well, and we are built from the ground up, from algorithms, to data structures, to the data layout and query planning, to the architecture in the cloud, to do that extremely well."
What MongoDB, Couchbase, and the rest, he said "are trying to do, and, in some sense, successfully, is to muddy the waters on what a vector database even is," he said. "They know that, at scale, when it comes to building real-world applications with vector databases, there's going to be no competition."
The momentum is with Pinecone, argues Liberty, by virtue of having pursued his original insight with great focus.
"We have today thousands of companies using our product," said Liberty, "hundreds of thousands of developers have built stuff on Pinecone, our clients are being downloaded millions of times and used all over the place." Pinecone is "ranked as number one by God knows how many different surveys."
Going forward, said Liberty, the next several years for Pinecone will be about building a system that comes closer to what knowledge actually means.
Also: The promise and peril of AI at work in 2024
"I think the interesting question is how do we represent knowledge?" Liberty told ZDNET. "If you have an AI system that needs to be truly intelligent, it needs to know stuff."
The path to representing knowledge for AI, said Liberty, is definitely a vector database. "But that is not the end answer," he said. "That is the initial part of the answer." There's another "two, three, five, ten years worth of investment in the technology to make those systems integrate with one another better to represent data more accurately," he said.
"There is a huge roadmap ahead of us of making knowledge an integral part of every application."
Computer vision is one of the most discussed fields in the AI industry, thanks to its potential applications across a wide range of real-time tasks. In recent years, computer vision frameworks have advanced rapidly, with modern models now capable of analyzing facial features, objects, and much more in real-time scenarios. Despite these capabilities, human motion transfer remains a formidable challenge for computer vision models. This task involves retargeting facial and body motions from a source image or video to a target image or video. Human motion transfer is widely used in computer vision models for styling images or videos, editing multimedia content, digital human synthesis, and even generating data for perception-based frameworks.
In this article, we focus on MagicDance, a diffusion-based model designed to revolutionize human motion transfer. The MagicDance framework specifically aims to transfer 2D human facial expressions and motions onto challenging human dance videos. Its goal is to generate novel pose sequence-driven dance videos for specific target identities while maintaining the original identity. The MagicDance framework employs a two-stage training strategy, focusing on human motion disentanglement and appearance factors like skin tone, facial expressions, and clothing. We will delve into the MagicDance framework, exploring its architecture, functionality, and performance compared to other state-of-the-art human motion transfer frameworks. Let's dive in.
MagicDance : Realistic Human Motion Transfer
As mentioned earlier, human motion transfer is one of the most complex computer vision tasks because of the sheer complexity involved in transferring human motions and expressions from the source image or video to the target image or video. Traditionally, computer vision frameworks have achieved human motion transfer by training a task-specific generative model including GAN or Generative Adversarial Networks on target datasets for facial expressions and body poses. Although training and using generative models deliver satisfactory results in some cases, they usually suffer from two major limitations.
They rely heavily on an image warping component as a result of which they often struggle to interpolate body parts invisible in the source image either due to a change in perspective or self-occlusion.
They cannot generalize to other images sourced externally that limits their applications especially in real-time scenarios in the wild.
Modern diffusion models have demonstrated exceptional image generation capabilities across different conditions, and diffusion models are now capable of presenting powerful visuals on an array of downstream tasks such as video generation & image inpainting by learning from web-scale image datasets. Owing to their capabilities, diffusion models might be an ideal pick for human motion transfer tasks. Although diffusion models can be implemented for human motion transfer, it does have some limitations either in terms of the quality of the generated content, or in terms of identity preservation or suffering from temporal inconsistencies as a result of model design & training strategy limits. Furthermore, diffusion-based models demonstrate no significant advantage over GAN frameworks in terms of generalizability.
To overcome the hurdles faced by diffusion and GAN based frameworks on human motion transfer tasks, developers have introduced MagicDance, a novel framework that aims to exploit the potential of diffusion frameworks for human motion transfer demonstrating an unprecedented level of identity preservation, superior visual quality, and domain generalizability. At its core, the fundamental concept of the MagicDance framework is to split the problem into two stages : appearance control and motion control, two capabilities required by image diffusion frameworks to deliver accurate motion transfer outputs.
The above figure gives a brief overview of the MagicDance framework, and as it can be seen, the framework employs the Stable Diffusion model, and also deploys two additional components : Appearance Control Model and Pose ControlNet where the former provides appearance guidance to the SD model from a reference image via attention whereas the latter provides expression/pose guidance to the diffusion model from a conditioned image or video. The framework also employs a multi-stage training strategy to learn these sub-modules effectively to disentangle pose control and appearance.
In summary, the MagicDance framework is a
Novel and effective framework consisting of appearance-disentangled pose control, and appearance control pretraining.
The MagicDance framework is capable of generating realistic human facial expressions and human motion under the control of pose condition inputs and reference images or videos.
The MagicDance framework aims to generate appearance-consistent human content by introducing a Multi-Source Attention Module that offers accurate guidance for Stable Diffusion UNet framework.
The MagicDance framework can also be utilized as a convenient extension or plug-in for the Stable Diffusion framework, and also ensures compatibility with existing model weights as it does not require additional fine-tuning of the parameters.
Additionally, the MagicDance framework shows exceptional generalization capabilities for both appearance and motion generalization.
Appearance Generalization : The MagicDance framework demonstrates superior capabilities when it comes to generating diverse appearances.
Motion Generalization : The MagicDance framework also has the ability to generate a wide range of motions.
MagicDance : Objectives and Architecture
For a given reference image either of a real human or a stylized image, the primary objective of the MagicDance framework is to generate an output image or an output video conditioned on the input and the pose inputs {P, F} where P represents human pose skeleton and F represents the facial landmarks. The generated output image or video should be able to preserve the appearance and identity of the humans involved along with the background contents present in the reference image while retaining the pose and expressions defined by the pose inputs.
Architecture
During training, the MagicDance framework is trained as a frame reconstruction task to reconstruct the ground truth with the reference image and pose input sourced from the same reference video. During testing to achieve motion transfer, the pose input and the reference image is sourced from different sources.
The overall architecture of the MagicDance framework can be split into four categories: Preliminary stage, Appearance Control pretraining, Appearance-disentangled Pose Control, and Motion Module.
Preliminary Stage
Latent Diffusion Models or LDM represent uniquely designed diffusion models to operate within the latent space facilitated by the use of an autoencoder, and the Stable Diffusion framework is a notable instance of LDMs that employs a Vector Quantized-Variational AutoEncoder and temporal U-Net architecture. The Stable Diffusion model employs a CLIP-based transformer as a text encoder to process textual inputs by converting text inputs into embeddings. The training phase of the Stable Diffusion framework exposes the model to a text condition and an input image with the process involving the encoding of the image to a latent representation, and subjects it to a predefined sequence of diffusion steps directed by a Gaussian method. The resultant sequence yields a noisy latent representation that provides a standard normal distribution with the primary learning objective of the Stable Diffusion framework being denoising the noisy latent representations iteratively into latent representations.
Appearance Control Pretraining
A major issue with the original ControlNet framework is its inability to control appearance amongst spatially varying motions consistently although it tends to generate images with poses closely resembling those in the input image with the overall appearance being influenced predominantly by textual inputs. Although this method works, it is not suited for motion transfer involving tasks where its not the textual inputs but the reference image that serves as the primary source for appearance information.
The Appearance Control Pre-training module in the MagicDance framework is designed as an auxiliary branch to provide guidance for appearance control in a layer-by-layer approach. Rather than relying on text inputs, the overall module focuses on leveraging the appearance attributes from the reference image with the aim to enhance the framework’s ability to generate the appearance characteristics accurately particularly in scenarios involving complex motion dynamics. Furthermore, it is only the Appearance Control Model that is trainable during appearance control pre-training.
Appearance-disentangled Pose Control
A naive solution to control the pose in the output image is to integrate the pre-trained ControlNet model with the pre-trained Appearance Control Model directly without fine-tuning. However, the integration might result in the framework struggling with appearance-independent pose control that can lead to a discrepancy between the input poses and the generated poses. To tackle this discrepancy, the MagicDance framework fine-tunes the Pose ControlNet model jointly with the pre-trained Appearance Control Model.
Motion Module
When working together, the Appearance-disentangled Pose ControlNet and the Appearance Control Model can achieve accurate and effective image to motion transfer, although it might result in temporal inconsistency. To ensure temporal consistency, the framework integrates an additional motion module into the primary Stable Diffusion UNet architecture.
MagicDance : Pre-Training and Datasets
For pre-training, the MagicDance framework makes use of a TikTok dataset that consists of over 350 dance videos of varying lengths between 10 to 15 seconds capturing a single person dancing with a majority of these videos containing the face, and the upper-body of the human. The MagicDance framework extracts each individual video at 30 FPS, and runs OpenPose on each frame individually to infer the pose skeleton, hand poses, and facial landmarks.
For pre-training, the appearance control model is pre-trained with a batch size of 64 on 8 NVIDIA A100 GPUs for 10 thousand steps with an image size of 512 x 512 followed by jointly fine-tuning the pose control and appearance control models with a batch size of 16 for 20 thousand steps. During training, the MagicDance framework randomly samples two frames as the target and reference respectively with the images being cropped at the same position along the same height. During evaluation, the model crops the image centrally instead of cropping them randomly.
MagicDance : Results
The experimental results conducted on the MagicDance framework are demonstrated in the following image, and as it can be seen, the MagicDance framework outperforms existing frameworks like Disco and DreamPose for human motion transfer across all metrics. Frameworks consisting a “*” in front of their name uses the target image directly as the input, and includes more information compared to the other frameworks.
It is interesting to note that the MagicDance framework attains a Face-Cos score of 0.426, an improvement of 156.62% over the Disco framework, and nearly 400% increase when compared against the DreamPose framework. The results indicate the robust capacity of the MagicDance framework to preserve identity information, and the visible boost in performance indicates the superiority of the MagicDance framework over existing state-of-the-art methods.
The following figures compare the quality of human video generation between the MagicDance, Disco, and TPS frameworks. As it can be observed, the results generated by the GT, Disco, and TPS frameworks suffer from inconsistent human pose identity and facial expressions.
Furthermore, the following image demonstrates the visualization of facial expression and human pose transfer on the TikTok dataset with the MagicDance framework being able to generate realistic and vivid expressions and motions under diverse facial landmarks and pose skeleton inputs while accurately preserving identity information from the reference input image.
It is worth noting that the MagicDance framework boasts of exceptional generalization capabilities to out-of-domain reference images of unseen pose and styles with impressive appearance controllability even without any additional fine-tuning on the target domain with the results being demonstrated in the following image.
The following images demonstrate the visualization capabilities of MagicDance framework in terms of facial expression transfer and zero-shot human motion. As it can be seen, the MagicDance framework generalizes to in-the-wild human motions perfectly.
MagicDance : Limitations
OpenPose is an essential component of the MagicDance framework as it plays a crucial role for pose control, affecting the quality and temporal consistency of the generated images significantly. However, the MagicDance framework still finds it a bit challenging to detect facial landmarks and pose skeletons accurately, especially when the objects in the images are partially visible, or show rapid movement. These issues can result in artifacts in the generated image.
Conclusion
In this article, we have talked about MagicDance, a diffusion-based model that aims to revolutionize human motion transfer. The MagicDance framework tries to transfer 2D human facial expressions and motions on challenging human dance videos with the specific aim of generating novel pose sequence driven human dance videos for specific target identities while keeping the identity constant. The MagicDance framework is a two-stage training strategy for human motion disentanglement and appearance like skin tone, facial expressions, and clothes.
MagicDance is a novel approach to facilitate realistic human video generation by incorporating facial and motion expression transfer, and enabling consistent in the wild animation generation without needing any further fine-tuning that demonstrates significant advancement over existing methods. Furthermore, the MagicDance framework demonstrates exceptional generalization capabilities over complex motion sequences and diverse human identities, establishing the MagicDance framework as the lead runner in the field of AI assisted motion transfer and video generation.