Beware of Chinese Open-Source LLMs 

If you’ve seen or even heard of the most popular American comedy series Silicon Valley, you might have come across a shady Chinese app developer, Jian-Yang, who goes back to China to build a knock-off version of Pied Piper – a fictional cloud-based compression platform which allows users to compress and share their files between their devices.

A similar drama is unfolding at OpenAI, where the company has filed for a patent for “GPT-6” and “GPT-7” in China, not in the US, to avoid the Pied Piper situation, obviously.

This new development also highlights the advancements in open source AI research in China, which even OpenAI is concerned about.

When it comes to open source AI research, we have often heard many say that it is a risk to open source powerful AI models because Chinese competitors would have all the weights of the models, and would eventually be on top of all the others. It is increasingly becoming the case that large language models (LLMs) from China are topping the leaderboards.

Most recently, DeepSeek, a 67 billion parameter model outperformed Llama 2, Claude-2, and Grok-1 on various metrics. The best part is that the model from China is open source, and uses the same architecture as LLaMA. Companies like Abacus AI are ready to host the models on their platforms.

Topping the leaderboards

Given the geopolitical issues between the US and China, the regulations on chip exports to the country are increasing, making it difficult for the country to build AI models, and up their business. But the increasing number of open source models give a hint that China does not really rely on US technology to further its AI field.

For instance, the Open LLM Leaderboard on Hugging Face, which has been criticised several times for its benchmark and evaluations, currently hosts AI models from China; and they are topping the list.

Tigerbot-70b-chat-v2 is currently leading the pack. The model, available on GitHub and Hugging Face, is built on top of Llama 2 70b architecture, along with its weight. It seems like open source models such as Llama 2 are actually helping the AI community in China to build models better than what the US is doing at the moment.

Tiger Research is a research lab in China under Tigerobo, which is dedicated to building AI models to make the world and humankind a better place, saying that they “believe in open innovations.” Similarly, DeepSeek is also a research lab with the mission of “unravelling the mystery of AGI with curiosity.”

Another hint of China’s open source AI dominance is the Yi-34B model released by 01.AI startup, reaching a unicorn status after the release. The AI startup by Kai-Fu Lee is developing AI systems for the Chinese market. The interesting part is that the second and third models on the Open LLM Leaderboard are also models based on Yi-34B, combining them with Llama 2 and Mistral-7B.

Not just this, Alibaba, the chinese tech giant also released Qwen-72B with 3 trillion tokens, and a 32K context length. This along with a smaller Qwen-1.8B is also available on GitHub and Hugging Face, which requires just 3GB of GPU memory to run, making it amazing for the research community. Notably, Qwen is also an organisation building LLMs and large multimodal models (LMMs), and other AGI-related projects.

Is China open source a threat?

Clearly, the fear of China rising up against US AI models is becoming a reality. The models from the country are increasingly dominating the open source, and will continue to do so in the upcoming year. And regulations are clearly not making it any better for the US. But how big of a threat it really is?

The recent slew of releases of open source models from China highlight that the country does not need US assistance in its AI developments. Moreover, if the US continues to crush its open source ecosystem with regulations, China will rise up even more in this aspect. As long as China continues to open source its powerful AI models, there is no threat at the moment.

Furthermore, China leading in the AI realm is not a new phenomenon. When GPT-3.5 was announced by OpenAI, Baidu released its Ernie 3.0 model, which was almost double the size of the former. The case has been the same ever since Ernie 2.0 was released, to compete with the first version of OpenAI’s GPT in 2019.

Even though these models are on the top of the Open LLM Leaderboard, a lot of researchers have been pointing out that it is just because of the evaluation metrics used for benchmarking. A lot of researchers in China are also hired from the US.

Moreover, a lot of these models are extremely restrictive. Given the information control in the country, these models might be fast, but are extremely poor when it comes to implementation into real use cases. “Don’t use Chinese models. They are going to nerf them the Chinese way, which is to alter behaviour even worse than current US censored models,” said a user on X.

The post Beware of Chinese Open-Source LLMs appeared first on Analytics India Magazine.

TSMC: The Wizard Behind AI’s Curtain

At AWS re:Invent, the hyper scaler unveiled the next generation of two AWS-designed chip families—AWS Graviton4 and AWS Trainium2—bringing improvements in price performance and energy efficiency across various customer workloads.

Only a few weeks prior to the AWS announcement, Microsoft, which competes with AWS in the cloud space, also announced two homegrown chips—Microsoft Azure Maia 100 AI Accelerator and Azure Cobalt 100 CPU. Interestingly, both AWS and Microsoft’s homegrown chips will be developed by Taiwan Semiconductor Manufacturing Company Limited (TSMC).

The Taiwanese semiconductor giant also manufactures the chips for Google’s Tensor Processing Units (TPU), which the tech giant announced at Google I/O 2016. Moreover, Google is also working on its own custom chip to power its Pixel smartphones and will replace Samsung with TSMC for chip manufacturing.

Reportedly, Apple, the most popular smartphone maker, is also working towards introducing a host of generative AI features in iOS 18, again, built on TSMC’s N3E 3-nanometer node.

Revenue from AI to skyrocket

TSMC’s advanced manufacturing processes enable the production of chips with increased computational capabilities, meeting the requirements of generative AI workloads.

With a market cap of USD 511.12 billion as of December 2023, TSMC is the world’s 12th most valuable company. At present, approximately 6% of TSMC’s overall revenue ( USD 73.86 billion in 2022) is derived from AI. Nevertheless, the company envisions this figure doubling within the next four to five years. TSMC anticipates a substantial compound annual growth rate (CAGR) of nearly 50% in the AI sector from 2022 to 2027.

The recent developments by big tech underscores how important TSMC has emerged in the AI ecosystem. Not to forget, NVIDIA also relies on TSMC for the fabrication of its Graphics Processing Units (GPUs), which became the most sought-after product in the AI world this year. Funnily enough, Intel’s Gaudi2 processors, which compete with NVIDIA H100, are also based on TSMC’s 7 nm processors.

To keep up with the demand, TSMC is also planning to invest USD 2.87 billion to build a new plant that will handle the advanced packaging of high-performance semiconductors necessary for generative AI.

Moreover, despite having fabrication units in Taiwan, TSMC has also announced numerous expansion plans. In Arizona, US, TSMC is building a second semiconductor factory, with an increased investment from USD 12 billion to USD 40 billion. The new facility, known as Fab 21, is expected to start chip production on TSMC’s advanced N3 process technologies by 2026.

Recent reports indicate that TSMC is considering the establishment of a third fabrication facility in the US to develop 2 nm and 1 nm technology. Similar expansion plans are also being considered in Japan. Moreover, a second fab in Europe is also being evaluated.

Dependence on TSMC- A dangerous precedent?

TSMC’s 3 nm technology is highly important for AI chip companies. For AI applications, where processing large amounts of data with high precision is crucial, TSMC’s 3 nm process enables the development of more powerful and efficient AI chips as it allows for more transistors to be packed into a chip, resulting in improved performance, energy efficiency, and overall capabilities.

TSMC is one of the few companies in the world that can reliably build chips at the leading edge of semiconductor technology, including advanced AI chips. The company has already started high-volume production of its 3 nm technology in 2022, making it the industry’s most advanced semiconductor process.

However, the significant dependence of the AI industry on TSMC for chip manufacturing introduces potential risks reminiscent of previous challenges faced by industries relying heavily on specific suppliers.

For example, the industry heavily depends on NVIDIA for its GPUs. However, a supply shortage left numerous companies in a scramble to acquire these GPUs. Could the AI sector’s reliance on TSMC could set a similar precedent, given the semiconductor industry already heavily relies on Taiwan to meet global chip demands?

Notably, Samsung is the only other company with the 3 nm technology. However, TSMC holds a 60% market share of the third-party chip manufacturing business compared to Samsung Electronics, which holds 12%.

Interestingly, Qualcomm, which is looking to disrupt the smartphone industry with its AI processors, earlier hinted at a dual-foundry strategy with TSMC and Samsung manufacturing simultaneously. However, Qualcomm has officially declared its decision not to enlist Samsung for its upcoming processors, emphasising once more the significance of TSMC in the AI domain.

The post TSMC: The Wizard Behind AI’s Curtain appeared first on Analytics India Magazine.

Researcher Develops Domain-Specific Scientific Chatbot

In scientific research, collaboration and expert input are crucial, yet often challenging to obtain, especially in specialized fields. Addressing this, Kevin Yager, leader of the electronic nanomaterials group at the Center for Functional Nanomaterials (CFN), Brookhaven National Laboratory, has developed a game-changing solution: a specialized AI-powered chatbot.

This chatbot stands out from general-purpose chatbots due to its in-depth knowledge in nanomaterial science, made possible by advanced document retrieval techniques. It taps into a vast pool of scientific knowledge, making it an active participant in scientific brainstorming and ideation, unlike its more general counterparts.

Yager’s innovation harnesses the latest in AI and machine learning, tailored for the complexities of scientific domains. This AI tool transcends the traditional boundaries of collaboration, offering scientists a dynamic partner in their research endeavors.

The development of this specialized chatbot at CFN marks a significant milestone in digital transformation in science. It exemplifies the potential of AI in enhancing human intelligence and expanding the scope of scientific inquiry, heralding a new era of possibilities in research.

Kevin Yager (Jospeh Rubino/Brookhaven National Laboratory)

Embedding and Accuracy in AI

The unique strength of Kevin Yager's specialized chatbot lies in its technical foundation, particularly the use of embedding and document-retrieval methods. This approach ensures that the AI provides not only relevant but also factual responses, a critical aspect in the realm of scientific research.

Embedding in AI is a transformative process where words and phrases are converted into numerical values, creating an “embedding vector” that quantifies the text's meaning. This is pivotal for the chatbot's functioning. When a query is posed, the bot's machine learning (ML) embedding model computes its vector value. This vector then navigates a pre-computed database of text chunks from scientific publications, enabling the chatbot to pull semantically related snippets to better understand and respond to the question.

This method addresses a common challenge with AI language models: the tendency to generate plausible-sounding but inaccurate information, a phenomenon often referred to as ‘hallucinating' data. Yager's chatbot overcomes this by grounding its responses in scientifically verified texts. It operates like a digital librarian, adept at interpreting queries and retrieving the most relevant and factual information from a trusted corpus of documents.

The chatbot's ability to accurately interpret and contextually apply scientific information represents a significant advancement in AI technology. By integrating a curated set of scientific publications, Yager's AI model ensures that the chatbot's responses are not only relevant but also deeply rooted in the actual scientific discourse. This level of precision and reliability is what sets it apart from other general-purpose AI tools, making it a valuable asset in the scientific community for research and development.

Demo of chatbot (Brookhaven National Laboratory)

Practical Applications and Future Potential

The specialized AI chatbot developed by Kevin Yager at CFN offers a range of practical applications that could significantly enhance the efficiency and depth of scientific research. Its ability to classify and organize documents, summarize publications, highlight relevant information, and quickly familiarize users with new topical areas stands to revolutionize how scientists manage and interact with information.

Yager envisions numerous roles for this AI tool. It could act as a virtual assistant, helping researchers navigate through the ever-expanding sea of scientific literature. By efficiently summarizing large documents and pointing out key information, the chatbot reduces the time and effort traditionally required for literature review. This capability is especially valuable for keeping up with the latest developments in fast-evolving fields like nanomaterial science.

Another potential application is in brainstorming and ideation. The chatbot’s ability to provide informed, context-sensitive insights can spark new ideas and approaches, potentially leading to breakthroughs in research. Its capacity to quickly process and analyze scientific texts allows it to suggest novel connections and hypotheses that might not be immediately apparent to human researchers.

Looking to the future, Yager is optimistic about the possibilities: “We never could have imagined where we are now three years ago, and I'm looking forward to where we'll be three years from now.”

The development of this chatbot is just the beginning of a broader exploration into the integration of AI in scientific research. As these technologies continue to advance, they promise not only to augment the capabilities of human researchers but also to open up new avenues for discovery and innovation in the scientific world.

Balancing AI Innovation with Ethical Considerations

The integration of AI in scientific research necessitates a balance between technological advancement and ethical considerations. Ensuring the accuracy and reliability of AI-generated data is paramount, especially in fields where precision is crucial. Yager's approach of basing the chatbot's responses on verified scientific texts addresses concerns about data integrity and the potential for AI to produce inaccurate information.

Ethical discussions also revolve around AI as an augmentative tool rather than a replacement for human intelligence. AI initiatives at CFN, including this chatbot, aim to enhance the capabilities of researchers, allowing them to focus on more complex and innovative aspects of their work while AI handles routine tasks.

Data privacy and security remain critical, particularly with sensitive research data. Maintaining robust security measures and responsible data handling is essential for the integrity of scientific research involving AI.

As AI technology evolves, responsible and ethical development and deployment become crucial. Yager's vision emphasizes not just technological advancement but also a commitment to ethical AI practices in research, ensuring these innovations benefit the field while adhering to high ethical standards.

You can find the published research here.

This adorable motion-tracking camera has proven to be indispensable in my smart home

Eufy Security Indoor Cam S350

The Eufy Security Indoor Cam S350 is a pan/tilt dual security camera with 360 degrees of coverage. It tracks objects in motion using two cameras, captures footage with 4K resolution, and supports dual-band Wi-Fi 6.

Also: This no-fee video doorbell can guard your packages this holiday season

I've had my Eufy Security Indoor Cam S350 for over a month and it's proven itself to be indispensable to my home security.

ZDNET RECOMMENDS

Eufy Security Indoor Cam S350

The Eufy Security Indoor Cam S350 has two cameras that work together to track motion and capture footage in 4K resolution.

View at Amazon

What truly makes this indoor security device stand out is the fact that it houses two cameras in one body; one with a wide-angle lens and the other with a telephoto lens. The wide-angle lens camera has a 130-degree field of view when static and records in 4K.

The telephoto lens camera records in 2K, features 3x optical zoom, so you can zoom in on people or objects without sacrificing image quality, and an 8x digital zoom. Both cameras have a low aperture, at f/1.6, to better capture lighting and record in low-light conditions.

Also: Eufy's new Floodlight Cam E340 is the hardest working security camera I've tested

And then the device can pan, tilt, and swivel in any direction smoothly, giving you a full view of its surroundings. The Eufy Cam S350 smoothly tracks people in motion, keeping the target in view with no staggering, and the Local AI in the HomeBase 3 can distinguish humans and pets.

We paired this camera with our HomeBase 3, which we had at home and use for the rest of our Eufy security system. The HomeBase 3 retails for $150; it's a hub for Eufy security devices that supports local storage for its cameras, so we can enjoy the full benefit of a security camera without monthly fees. However, you don't need to buy a HomeBase 3 to use this camera.

The Indoor Cam S350 can store footage on the HomeBase 3, but also on an inserted microSD card, which is not included, of up to 128GB. This card would still give you the local storage you need to avoid paying monthly cloud fees and it would support the live notifications of detected motion.

Also: EufyCam 3 and HomeBase 3 review: Why I'm not getting rid of these cameras yet

These pan/tilt indoor cameras are commonly marketed as baby monitors or pet cams because of their ability to track motion in real time and to cover the lens when the privacy mode is turned on. While this is a perfect camera for any of these options, I use it as an indoor security camera in conjunction with our security system.

My home's Eufy security system includes some motion sensors that are spread around downstairs, several outdoor security cameras, a keypad, and the HomeBase 3, which sounds an alarm when the system is triggered.

I don't use indoor cameras in our living room, where I have a motion detector, because I don't like the idea of having a camera looking at my family's every move. If a sensor detects motion and the alarm goes off at three in the morning, I have no way to see what the motion is. So, I do have a smaller pan/tilt Ezviz camera that is always set to privacy mode, just so I can look at it whenever motion is detected. But this process takes time.

Also: Missed Cyber Monday? The Schlage Encode smart lever lock is still $70 off right now

The Eufy Indoor Cam S350 has become a better solution to this problem, as it works with my existing system and gives me a live view within the same app as the rest of my security devices.

ZDNET's buying advice

The Eufy Security Indoor Cam S350 works seamlessly each time motion is detected, it doesn't get distracted by false alerts often, notifications are fast to appear on my mobile device when triggered, and it can quickly hide its lenses when privacy mode is turned on.

However, I'm left wishing for two things. First, I'd love to schedule when the camera goes to privacy mode, so I can set a schedule to have it automatically guard each night at bedtime and stop watching during the day. I also wish it had a button to manually engage privacy mode, so the camera can stop capturing motion when I press it.

These additions to the camera's privacy mode would only be a bonus, as the camera performs and records so well that I don't end up missing them, but they would be nice to have. As it is right now, I'd recommend the Indoor Cam S350 to anyone looking for an indoor security camera, pet cam, or baby monitor camera, as it's proven itself to be capable of performing these tasks with clear, high-resolution images.

Featured reviews

A Different AI Scenario: AI and Justice in a Brave New World – Part 1

Slide1

The recent upheavals at OpenAI and OpenAI’s Chief Scientist’s apprehensions regarding the “safety” of AI have ignited a fresh wave of concerns and fears about the march towards Artificial General Intelligence (AGI) and “Super Intelligence.”

AI safety concerns the development of AI systems aligned with human values and do not cause harm to humans. Some of the AI safety concerns that Ilya Sutskever, OpenAI’s Chief Scientist and one of the leading AI scientists behind ChatGPT, has expressed are:

  • The potential for AI to replace most human jobs and create massive unemployment
  • The risk of AI systems becoming misaligned with human goals or developing their own goals that conflict with ours
  • The challenge of ensuring that AI systems are transparent, accountable, and controllable by humans
  • The ethical and social implications of creating and deploying powerful AI systems that can negatively and incomprehensibly affect hundreds of millions of lives

Sutskever reportedly led the board coup at OpenAI that ousted CEO Sam Altman (Altman has since been reinstated at OpenAI) over these AI safety concerns, arguing that Altman was pushing for commercialization and growth too fast and too far without adequate attention to the safety and social impact of OpenAI’s AI technologies.

The wild agitation about AI “becoming misaligned with human goals or developing their own goals that conflict with ours” is an overwhelmingly essential and critical topic. If you want a refresher on how AI can go wrong, watch the movie “Eagle Eye.” In the movie, an AI model (ARIIA) was designed to protect American lives, but that “desired outcome” wasn’t counter-balanced against other desired outcomes to prevent unintended consequences.

With these concerns about AI models going rogue, I’m going to present a multi-part series on what would be required to create a counter scenario – “AI and Justice in a Brave New World” – that not only addresses the AI concerns and fears but establishes a more impartial and transparent government, legal, and judicial system that promotes a fair playing field in which everyone can participate (have a voice), benefit, and prosper.

AI’s Role: Laws and Regulations Enforcement, not Creation

The Challenge. In this AI scenario, elected officials would still make society’s laws and regulations. Unfortunately, these laws and regulations are not enforced fairly, consistently, or unbiasedly in today’s legal and judicial systems. Some people and organizations use their money and influence to bend the laws and regulations in their favor, while others face discrimination and injustice because of their identity, background, or position in society. This creates a system of inequality and distrust that leads to social disenfranchisement and political extremism. We need to ensure that the laws and regulations are enforced with fairness, equity, consistency, and transparency for everyone.

The Solution. To address the unfair, inconsistent, and biased enforcement issue, we could build AI models that would be responsible for enforcing the laws and regulations legislated by elected officials. With AI in charge of enforcement, there will be no room for prejudice or biased decision-making when enforcing these laws and regulations. Instead, AI systems would ensure that the laws and regulations are applied equitably, fairly, and transparently to everyone. Finally, justice would truly be blind.

This requires that our legislators define the laws and regulations and collaborate across a diverse set of stakeholders to define the desired outcomes, the importance or value of these desired outcomes, and the variables and metrics against which adherence to those outcomes could be measured. This would necessitate substantial forethought in defining the laws and regulations and identifying the potential unintended consequences of these laws and regulations to ensure that the variables and metrics to flag, avoid, or mitigate those unintended consequences are clearly articulated. Yes, our legislators would need to think more carefully and holistically about including variables and metrics across a more robust set of measurement criteria that can help ensure that the AI models deliver more relevant, meaningful, responsible, and ethical outcomes (Figure 1).

Slide2-1

Figure 1: Economic Value Definition Wheel

Developing a Healthy AI Utility Function

Here is the good news: those desired outcomes, their weights, and the metrics and variables against which the AI model will use to make its decisions and actions could be embedded into the AI Utility Function, the beating heart of the AI model.

The AI Utility Function is a mathematical function that evaluates different actions or decisions that the AI system can take and chooses the one that maximizes the expected utility or value of the desired outcomes.

Slide3

Figure 2: The AI Utility Function

Analyzing the performance of the AI Utility Function provides a means to audit the AI model’s performance and provide the ultimate transparency into why the AI model made the decisions or actions that it did. However, several deficiencies exist today that limit the AI Utility Function’s ability to enforce society’s laws and regulations.

The table below outlines some AI Utility Function challenges and what AI engineers can do to address these challenges in designing and developing a healthy AI Utility Function.

AI Utility Function Challenge Potential AI Utility Function Design Options
An AI utility function is not always well-defined, measurable, or achievable. – Clarify the objective and desired outcomes that the AI model is trying to optimize and ensure alignment with the problem statement, stakeholder needs, and ethical principles (“Thinking Like a Data Scientist” methodology)
– Choose the relevant data, algorithms, and KPIs/metrics to quantify the objective or desired outcomes and validate their quality, relevance, and reliability.
– Evaluate the feasibility, practicality, and desirability of the objective or criterion and consider the trade-offs, constraints, and risks involved in achieving it.
An AI utility function is not always consistent, stable, or optimal. – Monitor the objective or criterion that the AI model is trying to optimize and update it as needed to reflect the changes in the operating environment.
– Use a multi-objective optimization approach to balance conflicting objectives and outcomes using Pareto-optimal solutions representing the best compromise.
– Incorporate uncertainty and robustness into the objective or criterion and use stochastic or adaptive algorithms that handle noise, variability, and unpredictability.
An AI utility function is not always transparent, understandable, or explainable. – Document the objective or desired outcomes the AI model aims to optimize and explain its rationale, assumptions, limitations, and implications.
– Demand interpretable, understandable, and explainable data, algorithms, and metrics that can reveal the logic, reasoning, and evidence behind the model’s actions.
– Mandate a feedback loop that evaluates outcomes’ effectiveness, continuously learns and adapts the AI Utility Function based upon actual versus predicted outcomes, and enables oversight and intervention.

Table 1: Addressing AI Utility Function Challenges

Summary: AI and Justice in a Brave New World Part 1

Maybe I’m too Pollyannish. But I refuse to turn over control of this conversation to folks who are only focused on what could go wrong. Instead, let’s consider those concerns and create something that addresses those concerns and creates something that benefits everyone.

Part 2 will examine a real-world case study about using AI to deliver consistent, unbiased, and fair outcomes…robot umpires. We will also discuss how we can create AI models that act more human. Part 3 will try to pull it all together by highlighting the potential of AI-powered governance to ensure AI’s responsible and ethical application to society’s laws and regulations.

AI Assists Production in Indian Film Industry

The silent movie ‘Le voyage dans la lune’ (A Trip to The Moon), released in 1902, was considered the first science fiction movie in the world. Nothing like today, the movie’s ‘special effects’ were all human-made. 120 years later, we are standing at the crux of creative exploration, where AI can pretty much make a movie.

With the surge in AI video generation tools, the actual usage of these AI tools is yet to be measured in big production houses. Tools such as Stable Diffusion, Runway and the recent Pika may look promising at the moment, but there are still limitations that can restrict a creator from fully utilising them. However, production houses have already started experimenting with it, not just from a creative aspect but from an angle of operational efficiency.

AI Upgrades Pre-Production

While AI has been used at various stages of movie production, post-production (editing) is where the latest advancements in AI tools are implemented. From analysing clips for finding the best shot, splicing, de-ageing and other visual effects, including deepfakes, AI has found a crucial place to save time and effort. It has now found its way, in the pre-production stage , (planning) as well, where the biggest use case is storyboarding.

One of the crucial steps at the planning stage of a movie production is storyboarding. A storyboard visually outlines the sequence of shots that will unfold in a video. If done incorrectly, production houses not only face wastage of manpower but also face huge financial losses.

Realising the potential for simplifying storytelling and breaking barriers for those who wish to enter filmmaking with financial constraints and technical dearth, d.cult studio , a startup in Trivandrum, is out to ‘democratise the art of filmmaking’ through AI. The company is already working with studios in the Malayalam movie industry including Festival Studios and major production house Prithviraj Productions.

“What used to be sketches and infographic-based, is now changing into a cinematic-style storyboarding,” said Gokul A, founder of d.cult and co-founder of Accubits Technologies.

​​https://twitter.com/dcultstudio/status/1706578659195146533

Essence of Time and Cost

With accurate storyboarding, time and cost are saved during actual shoots. There’s a 50-80% overall reduction in time for the entire storyboarding process, and around 40-70% overall reduction in costs, after factoring in reduced labour and material expenses.

“In a 50 day shoot, where the per day shoot cost is 5 lakhs, even if five days are reduced, owing to AI storyboarding, 25 lakhs are saved from the production cost,” said Gokul. For a storyboard creator, such a tool is believed to augment their work. “Initially, where an experienced storyboard creator could make only two images in a day, such a tool will help the creator make upto 100 images in a day.“

The company uses a mix of models including proprietary and open-source ones, along with NLP algorithms. Idea to story/screenplay generation uses BUD LLMs – GenZ model, with natural intelligence algorithms. For auto music suggestion and background music integration, music generation models such as OpenAI’s Jukebox is used. Similarly, deep learning and computer vision models for AI video analysis and storyboard visualisation, respectively.

Treading Around Copyrighting

Storyboarding is only one part of the AI solution that d.cult studio offers. They also provide options of converting scripts to storyboard, collaborative whiteboard for idea discussion, video analysis for story ideas, pre-visualisation, photo generation, among other features.

With copyright issues causing a churn in companies producing AI-generated content, especially AI images, d.cult studio ensures strict measures on the same. “We are diligent in using only those materials that are either originally created by our team, fall within the public domain, or are available under creative commons licences for AI inputs,” said Gokul.

Animation Studios Are Not All Yay

Hayao Miyazaki- Studio Ghibli. Source: LA Times

While companies are experimenting with AI to ease movie production steps, there are some who will not jump on the wagon just yet. Signature style production houses such as Studio Ghibli, a Japanese animation studio that still uses traditional hand-drawn animation methods, will not embrace AI all too quickly. Co-founder and director Hayao Miyazaki has previously said that AI art is ‘utterly disgusting’ and that it will never succeed in giving artwork the genuine feel of a human being. He even said that he will never incorporate such technology in his work as he feels that it will be ‘an insult to life itself.’

On one end when you have animation studios shying away from embracing AI, there are others who are refining animation styles with the same. The LoRA training model simplifies the process of training Stable Diffusion on various concepts, including characters or specific styles. Once trained, these models can be exported for use in generating content by others. Thereby, enabling control levels for AI animation.

Furthermore, many are experimenting with a mix of models to create their own anime. Entrepreneur Varun Mayya said that production costs for an episode of anime costs from $140,000 to $180,000. He believes that making anime on their own with a mix of tools can probably drop the price by a 100 or 1000. He created one with a string of 20 different AI tools.

While people are experimenting with generative AI tools to create their own anime, movies, and more, the scaling of these techniques is yet to be seen. While human creativity will continue to play a significant role in movie production houses, thereby limiting complete adoption, the implementation of AI for operational and cost efficiency is where the actual use-case will lie.

The post AI Assists Production in Indian Film Industry appeared first on Analytics India Magazine.

9 Must-Know Open Source Models From Meta in 2023

Meta has been synonymous with open source ecosystems. Recently, its research arm, FAIR, completed 10 years of its contribution in the field of artificial intelligence and human-level intelligence.

“If you look back at Meta’s history, we’ve been a huge proponent of open-source,” believes Ahmed Al-Dahle, Meta’s vice president for generative AI.

Currently, Meta has over 900 GitHub repositories.

This year, Meta’s commendable effort towards open source has resulted in the introduction of state of the art models. Here’s a list of top must-known open-source AI models developed by Meta AI.

Llama

Meta has unexpectedly become the Robinhood of the LLM community, breaking the barrier of entry for a lot of developers to experiment with large language models ever since its launch in February 2023. Llama 2, which was launched later in partnership with Microsoft, in July 2023, changed the course of open-source language models once and for all. It is said that the company will be launching Llama 3 early next year. The best part is that it would be open source for researchers and commercial use as well.

Learn more about Llama here.

Seamless Communication

Meta has introduced “Seamless,” a cross-lingual communication system, anchored by two core models—SeamlessExpressive for preserving expression in speech-to-speech translation, and SeamlessStreaming, a high-performing streaming translation model with an impressive two-second latency—these innovations are built on the foundation of SeamlessM4T v2, Meta’s latest foundational model. SeamlessM4T v2 showcases notable performance enhancements in automatic speech recognition and various speech and text translation capabilities. This release represents a substantial leap forward in real-time cross-lingual communication, underscoring Meta’s commitment to pushing the boundaries of language and expression.

Learn more about Seamless here.

AudioCraft

Meta has unveiled the AudioCraft family of models, a design of generative audio models, offering both a user-friendly interface and the flexibility for users to explore their creative boundaries.

Comprising three distinct models—MusicGen, AudioGen, and EnCodec—AudioCraft stands out for its ability to produce high-quality audio with enduring consistency.BMeta is excited to make the pre-trained AudioGen model, EnCodec decoder, and all AudioCraft model weights and code available for research purposes, providing researchers and practitioners the opportunity to train their own models with personalized datasets and contribute to advancing the state of the art in generative audio technology.

Learn more about Audiocraft here.

DINOv2

Meta AI has unveiled DINOv2, a method for training high-performance computer vision models. With delivering a robust performance without the need for fine-tuning, it is positioning as a versatile backbone for various computer vision tasks. The release of Meta’s DINOv2 signals is marking a new era in computer vision methodologies, promising to redefine the landscape of high-performance model training.

Learn more about DINOv2 here.

XLS-R

Meta has fostered global inclusivity in voice technology, by releasing XLS-R that represents a major advancement in speech tasks, significantly outpacing previous multilingual models. This model sets a new benchmark, achieving state-of-the-art performance on renowned benchmarks like BABEL, CommonVoice, VoxPopuli for speech recognition, CoVoST-2 for foreign-to-English translation, and VoxLingua107 for language identification. With this release Meta is not only focusing to address the language gap in speech technology but also sets the stage for future developments, promising an enhanced and more inclusive voice technology landscape in the metaverse and beyond.

Learn more about XLS-R here.

Detectron2

Facebook AI Research (FAIR) has unveiled Detectron2, a cutting-edge library representing the next generation in state-of-the-art detection and segmentation algorithms. Positioned as the successor to Detectron, Detectron2 marks a significant leap forward in supporting various computer vision research projects and production applications within Facebook. Detectron2 is poised to redefine the landscape of detection and segmentation algorithms, setting new standards for innovation and performance in the realm of artificial intelligence.

Learn more about Detectron 2 here.

DensePose

Meta has stepped into a significant stride towards advancing human understanding by intorucing DensePose, a groundbreaking real-time approach that maps all human pixels in 2D RGB images to a comprehensive 3D surface-based model of the body. DensePose accelerates connections with augmented and virtual reality, operating at multiple frames per second on a single GPU and capable of handling multiple humans simultaneously. Meta is aiming for the future with the introduction of DensePose-COCO, a large-scale ground-truth dataset, further enhances the model’s accuracy and applicability, marking a pivotal advancement in the realm of human-centric image interpretation.

Learn more about DensePose here.

Wav2vec 2.0

Facebook AI Research has introduced wav2vec 2.0, the highly anticipated successor to wav2vec, accompanied by pretrained models and code. This advanced model revolutionizes self-supervised learning by mastering basic speech units, predicting masked portions of audio to simultaneously learn the correct speech units. This release marks a significant leap forward in the democratization of speech recognition technology, making it more accessible and effective across a broader linguistic spectrum.

Learn more about Wav2vec 2.0 here.

VizSeq

Embarking on a new era of efficiency in text generation tasks, Meta has introduced VizSeq, a Python toolkit, has emerged as a game-changer in visual analysis. Offering a user-friendly interface and harnessing the latest advancements in Natural Language Processing (NLP), VizSeq elevates user productivity by providing visualization capabilities in Jupyter Notebook and a web app. This release signifies a pivotal step toward more effective and comprehensive visual analysis in the realm of text generation tasks by Meta.

Learn more about VizSeq here.

How to Create Charts with Google Bard

Google recently updated Bard, an experimental large language model chatbot, to add the ability to create charts. With a few prompts, you may use Bard not only to gather data but also to display it in a chart; this can be useful for research and learning, such as when you want to understand market share, visualize financial data or compare pricing.

To get started, sign in to Bard with your personal Google account, then experiment with variations on the sequences below. Make sure to follow any guidelines and policies that apply when you use any AI tool in an organizational setting.

Visit Google Bard

Jump to:

  • Option 1: Prompt Bard for a table, then a chart
  • Option 2: Prompt Bard with data for a chart
  • How to regenerate or modify a Bard chart
  • How to save and share a Bard chart
  • Bard charts not working? Try a different Google account

Option 1: Prompt Bard for a table, then a chart

Most often, you’ll want to go through a two-phase process to make a chart with Bard: Create a table and then create a chart.

In your initial prompt, detail the data you seek and make sure to specify you want the information in a table. For example, try a prompt such as:

What are the largest 7 laptop makers in the U.S. by market share?

Sometimes this might return a numbered list of the manufacturers (Figure A). If it doesn’t include a table, you might follow up with:

Put that information in a table.

Figure A

Google Bard prompt for a table.
Enter a Bard prompt that produces a table, such as the table that contains laptop manufacturer data here. Image: Andy Wolber/TechRepublic

This sequence returns a list of seven laptop maker company names in one column, with market share in a different column. Now, prompt Bard to:

Create a chart.

In this case, Bard generates a pie chart and a title for the chart (Figure B). You may also ask for a specific type of chart, such as a bar chart or line graph.

Figure B

Google Bard Create A Chart.
After you have elicited a table from Bard, prompt the system to “Create a chart” as shown here. Image: Andy Wolber/TechRepublic

Option 2: Prompt Bard with data for a chart

Bard can create a chart with data you provide. For example, suppose you have a set of cells in a Google Sheet that contain your data (Figure C). Select the cells within the Sheet that contain the labels and data for your chart, as you typically would for any type of copy operation.

Figure C

Copy google sheet data for Google Bard chart.
You also may copy data in Google Sheet cells to use in a Bard prompt. Image: Andy Wolber/TechRepublic

Next, start a prompt with something such as “Create a chart from this data:” and paste the copied Sheets cell contents into the prompt (Figure D).

Figure D

Paste entries into Bard.
Begin a prompt that asks for a chart, then paste your data. Add spaces and commas to separate each entry. Image: Andy Wolber/TechRepublic

Edit the entry to ensure the data labels and values are appropriately separated by spaces and commas, then submit the prompt. Bard should generate a chart and a chart title from the data you provided (Figure E).

Figure E

Clean Up Prompt for generated chart.
Bard can generate a chart from supplied data. Note how the sample prompt was edited to clean up the Google Sheets source data content. Image: Andy Wolber/TechRepublic

How to regenerate or modify a Bard chart

If you’re not satisfied with the chart Bard builds, you have two options: Regenerate Draft or Modify.

Select the Regenerate Draft option in the upper-right region of the response area, and Bard will take another attempt to generate a response to your prompt. Sometimes this produces the same chart as the initial response; other times, it displays the data in a different type of chart, such as a bar graph instead of a pie chart.

Figure F

Regenerate or modify chart for Google Bard.
To change the chart, select either Regenerate Draft (upper right) or access the Modify menu (lower right). Image: Andy Wolber/TechRepublic

Select the Modify button in the lower-right region of the response area to access alternatives, which may vary based on your table and the current chart. For example, in one pie chart Bard created, the Modify menu lets you choose from two other chart types, Column Chart or Scatter Plot, or a Customize option; sometimes, Bar Chart and Line Chart options also may appear.

How to save and share a Bard chart

Take a screenshot of your Bard-created chart, save it and then share that image as you might any other image. The Share options won’t yet let you share a chart, so a screenshot is the best way to capture and convey chart content.

Bard charts not working? Try a different Google account

When you ask for a chart but Bard only displays code or charts found on the internet, switch to a personal Google account (instead of an organizational account) with access to Bard and give it another try. For example, Bard chart features fail to function for me when signed in with my Google Workspace Enterprise Standard account (which has both Duet AI and the Advanced Protection Program enabled), but work as described when signed in with a free Google account.

Visit Google Bard

2023 was a big year for AI: The top countries using it and which AI tools they prefer

3d low-poly people scattered around the Globe.

On the heels of ChatGPT's first anniversary, a new study on the use of the top 50 artificial intelligence tools available crowned OpenAI's chatbot as the most popular AI tool worldwide. ChatGPT accounted for 60% of traffic visits to the most popular tools between September 2022 and August 2023, an explosive time for the use of generative AI tools.

Also: Meta will enforce ban on AI-powered political ads in every nation, no exceptions

During this time, the United States was considered the leader in the use of AI tools. Writerbuddy, the company behind the study, scraped data from various directories that list AI tools and studied the use of over 3,000 AI tools. It then narrowed down the top 50 most used tools to investigate how people use AI.

What were the most popular AI tools?

It should come as no surprise that ChatGPT led with the most traffic visits, at 14.6 billion total visits since its launch in November 2022 through August 2023. The launch of ChatGPT triggered a generative AI boom that planted interest in other AI tools available.

  • Top gainers in the AI industry

ChatGPT, Character AI, and Google Bard were the top gainers in traffic.

Also: Think you can spot a fake AI-generated news story? Take this quiz to find out

"After the initial boom, the hype kept on climbing all the way until May 2023 when monthly visits peaked at around 4.1 billion. That's the first time we saw a pullback of 1.2 billion in the traffic of the industry," according to Sujan Sarkar with WriterBuddy.

  • Growth rate and monthly growth

The top 50 AI tools garnered over 24 billion visits combined and experienced a 10.7 times growth rate, with an average monthly growth of 236.3 million visits. Overall, the AI tools had an average of 2 billion visits monthly, with a surge to 3.3 billion in the last six months.

Also: 8 ways AI and 5G are pushing the boundaries of innovation together

Although Quillbot and Midjourney were in the top 5 most-visited AI tools, both experienced a drop in traffic in the past year. This is a sharp contrast with the explosive growth in traffic other AI tools experienced and could be attributed to the wider variety of tools created since the launch of ChatGPT. Midjourney lost 8.66 million visits in the past year, while Quillbot lost 5 million.

  • The average duration of AI use

The average visit lasted 12 minutes and 34 seconds, enough to ask ChatGPT about black holes and then for a recipe for a fun cocktail or two. Interestingly, ChatGPT was among the tools that failed to maintain user interest for longer than an average of 10 minutes, while tools like Character AI, Dezgo, and Janitor AI averaged 25 minutes per session.

  • Most popular AI categories

AI chatbots and AI writing were the most popular categories in AI tool use this past year, including ChatGPT, Character AI, and Google Bard in the top three. Image generators came in second to text generators, including Midjourney, DeepAI, and Craiyon.

Also: The best AI image generators of 2023: DALL-E 2 and alternatives

Less popular categories include design, video generators, voice and music generators, and others. Hugging Face, Taskade, and Cutout are among the "others" category stand outs.

What countries are using AI?

Countries like the United States and India and regions like Europe that have strong tech markets were found to lead AI use. The US led with 5.5 billion total users spread across different tools, which is about 22.6% of worldwide traffic to the AI industry. European countries comes in second with 3.9 billion users.

Also: How Google and OpenAI prompted GPT-4 to deliver more timely answers

Notably, India holds the second spot for most AI tool visits, with 2.1 billion visits, accounting for about 8.5% of the worldwide AI use in the past year.

  • Demographics of AI users

Men appear to be more inclined to use AI tools, as 69.5% of users were found to be male, while only 30.5% were female. The study didn't delve deep into the reasoning behind this gender gap, but it's worth noting that different gender roles in varied societies worldwide could influence this result.

Also: Generative AI can easily be made malicious despite guardrails, say scholars

Of all the top 50 AI tools, 80.5% were reached by direct access, which points "to a well-established user base or frequent repeat visitors," according to Sarkar. Organic search accounted for 11.4% of traffic sources, referrals took 6.7% of traffic, and organic social and others accounted for the rest. Only 0.12% of users reached the AI tools via a paid search.

This faux AI chatbot will judge your music taste and make you laugh (hopefully)

woman listening to music

On Wednesday, Spotify unveiled its Spotify Wrapped for 2023, and as a result, most social media feeds were swamped with people showing off their top songs and artists of the year. Although those Spotify Wrapped results might have been enough to impress your social media followers, this faux AI model will take a lot more convincing.

The digital publication The Pudding has a project called "How Bad Is Your Streaming Music?", a site that pretends to be a judgemental AI character that judges how bad your music taste is based on your Apple Music and Spotify streaming history.

Also: Apple names the 14 best apps and games of 2023. Your next favorite app may be on the list

Since it mimics the actions of a judgemental AI model, while looking at your music history, the website offers ruthless commentary, roasting many of your music choices and asking follow-up questions to better understand your music taste and ultimately deliver your final score.

Even though I personally felt comfortable with my music choices, I scored an embarrassing 21/100. If you are brave enough to try it yourself, follow the steps below.

Visit the website

The first step to getting your score is visiting the "How Bad Is Your Streaming Music?" website on your phone or computer. Once you are there, you can click on the "Find Out" button found in the center of the screen to get started.

Select what music streaming platform you use

To get you your results, the site will need access to your Spotify or Apple Music Account to see what music you listen to. The site reassures users that it will not change or post anything once it has access to their account.

Sign into your account

Once you select Spotify or Apple Music, you will be redirected to a new sign-in window. If you are hesitant about entering your log-in information, the site says that the sign-on information only gives them one-time access to your platform and that it doesn't see your username or password, or save that data.

Answer the different follow-up questions

Once you give the site access, it will ask you a series of follow-up questions based on the music you listen to. These questions are lowkey roasts — just roll with it.

Also, as a heads-up, in the follow-up questions, the site did use some foul language, so if that isn't your speed, you may want to avoid the quiz.

Get your score

Once you are done answering the questions, the platform will give you a score. Personal tip: prepare for the worst because the scores aren't very generous.

Also: Google Play crowns its best apps of 2023. Did your favorite make the list?

After I prompted the site to give me the final score, it briefly showed me a score out of 100 before moving onto a blank grey screen. However, others who have tried have been able to make it the entire way through without an issue. As a result, it seems like the issue is just the site being at capacity, with interest peaking during Spotify Wrapped time.

Even though I wasn't able to see the final screen, getting my score was still a fun experience.

Artificial Intelligence