AI and Justice in a Brave New World: Part 3 – AI Governance

Slide8

In part 1 of the series “A Different AI Scenario: AI and Justice in a Brave New World,” I outlined some requirements for the role that AI would play in enforcing our laws and regulations in a more just and fair manner and what our human legislators must do to ensure that outcome.

In part 2, I discussed a real example of how AI is being used to enforce our laws and regulations more unbiased, consistent, and transparently – robot umpires – and what we can do to humanize those robot umpires (and AI) just a little. I also tapped some time-tested standards that courts use to guide the admissibility of evidence and how those lessons can help us build more fair and just AI models in this brave new world of AI.

Let’s bring it all together with a robust conversation about operationalizing the AI governance necessary to ensure that our AI models deliver more meaningful, relevant, responsible, and ethical outcomes. And introduce the AI Model Trail concept as a way to verify and validate AI model compliance.

AI Model Governance

The success of this AI laws and regulations enforcement scenario is based upon an open, unbiased, and transparent AI model governance system. An effective AI model governance is the bedrock for ensuring the fairness, equity, and ethics of the AI models used to enforce the laws and regulations and for building the trust and confidence of the public and the policymakers in the AI models.

AI model governance involves multiple and diverse stakeholders, such as the AI developers, the AI users, the AI regulators, and the AI beneficiaries, who have different roles, responsibilities, and interests in the AI models. AI model governance covers the entire AI lifecycle, from the design and development to the deployment and maintenance to the ongoing evaluation and improvement of the AI models. AI model governance requires a combination of methods and tools, such as documentation, explanation, feedback, correction, and redress, to provide transparency, accountability, and oversight for the AI models.

Some of the most essential factors for AI model governance success in the eyes of the average citizen are:

  • The alignment of the AI models with the human-defined laws and regulations and the ethical, social, and legal norms and standards that govern the AI models’ domain and application.
  • The quality and reliability of the data, algorithms, parameters, and metrics that the AI models use to make their decisions, and the validation and verification of their accuracy, relevance, and reliability.
  • The performance and impact of the AI models and the measurement and assessment of their effectiveness, efficiency, and fairness, as well as their potential or actual errors, harms, or impacts.
  • The public and policymakers’ participation and involvement in the AI model governance process and soliciting and incorporating their feedback, input, and preferences.
  • The public and policymakers’ awareness and education about AI’s current and future state, its implications and challenges, and the provision and dissemination of understandable and verifiable information about the AI models.

Monitoring AI Model Governance: AI Model Audit Trail Concept

One way to ensure trusted and effective AI Governance is to mandate that all AI models generate an AI Model audit trail that captures the AI models’ metrics and variable and their respective weights at the time of any decision or action. The audit trail would provide the basis to validate and certify the AI models’ performance. AI certification models could be developed by government and independent agencies, such as GDPR, CCPA, or ACLU, to ensure that the AI models follow the laws and regulations and respect the rights and freedoms of the data subjects and natural persons.

Slide9

Figure 1: AI Model Audit Trail Concept

The AI Model Audit Trail process is the idea of creating and maintaining a record of the data, algorithms, parameters, metrics, and outcomes that an AI model uses to perform its tasks or make its decisions to provide transparency, accountability, and quality for the AI model and its users, stakeholders, and regulators.

The steps that comprise the AI Model Audit Trail process are:

  • The AI model generates an audit trail that creates a snapshot of each decision or action’s data, algorithms, parameters, and metrics. This step involves capturing and storing the relevant information and metadata used or produced by the AI model at each point of its operation, including the input and output data, the algorithm and parameter settings, the metrics and indicators, and the outcomes and impacts of the AI model’s decisions or actions.
  • Independent AI Model Auditors (GDPR, CCA, ACLU) analyze the audit trail to verify that the actions and related outcomes were unbiased, fair, and ethical. The AI Model Auditor checks the accuracy, relevance, and reliability of the data, algorithms, parameters, and metrics and the effectiveness, efficiency, and fairness of the outcomes of the AI model’s decisions.
  • A human auditor would analyze the outcomes and the AI model operational snapshot for questionable results. This step involves investigating and resolving any issues or disputes that arise from the AI model’s decisions using various methods and tools, such as forensics, analytics, or simulation. The human auditor would make a final judgment or determination on the validity, acceptability, or appropriateness of the AI model’s decisions or actions and the responsibility or liability of the AI model or its creators or users.
  • The AI model learnings from the human auditor are captured, documented, and quantified, and any regulatory fines are assessed. The AI model learnings should include the errors, harms, or impacts that the AI model caused or could cause and the corrections, improvements, or mitigations the AI model needs or could benefit from.
  • The AI model learnings are fed back to update the AI Utility Function to ensure these inequities do not happen again. The AI Utility Function should be updated to reflect the changes or adjustments in the data, algorithms, parameters, or metrics used or produced by the AI model and align with the desired outcomes and objectives that the AI model aims to achieve or produce.

Some of the benefits of the AI Model Audit Trail concept are:

  • Improve AI model transparency and accountability by providing a traceable and verifiable record of their decisions and enabling the inspection of their behavior and performance.
  • Enhance the quality and reliability of the AI models by ensuring that they use appropriate data, algorithms, parameters, and metrics and by verifying their accuracy, relevance, and reliability.
  • Promote AI model fairness and ethics by assessing and certifying their compliance and ethics and identifying and correcting any biases, errors, or harms they may cause.

AI Model Audit Trail Challenges

We must address several challenges to make the AI Model Trail concept work. Here is a matrix that summarizes the challenges and the actions for using AI model audit trails and AI certification AI models to drive trusted and effective AI governance.

AI Model Audit Trail Challenges AI Model Audit Trail Solutions
It is technically challenging as it requires
the integration of various data sources,
algorithms, parameters, and metrics and developing and maintaining robust AI certification AI models.
– Use common data formats, protocols, and standards to facilitate the integration and standardization of the data sources, algorithms, parameters, and metrics.
– Use modular, scalable, and reusable AI certification models that can be easily adapted and updated to different scenarios and contexts.
Legally or socially controversial, as it would
involve the collection and processing of
personal or sensitive data, and the
application and enforcement of different
or conflicting laws and regulations, and
the balancing of different or competing
values and interests.
– Use privacy-preserving techniques to protect personal or sensitive data from unauthorized access or misuse.
– Harmonize regulatory frameworks to ensure the compliance and consistency of the AI certification AI models across different jurisdictions and domains.
– Use stakeholder engagement methods to elicit and incorporate different perspectives in AI certification AI models.
Ensuring the trustworthiness of the AI models will require the full complexity
and context of the AI models’ behavior
and account for the model’s trade-offs, uncertainties, or risks involved
in the AI models’ optimization process.
– Use explainable AI techniques to provide evidence for the AI models’ decisions and their underlying logic and reasoning.
– Use reliable AI techniques to measure the trade-offs, uncertainties, or risks in the AI models’ optimization process.
– Use feedback and correction mechanisms to enable regulators to monitor, evaluate, and improve the AI models’ performance and address any related errors, harms, or impacts.

Table 1: AI Model Audit Trail Challenges and Actions

Summary: Create a More Fair, Just, and Empowered Society – Part 3

Part 3 continued the “AI and Justice in a Brave New World” conversation by discussing the AI governance necessary to ensure that our AI models deliver more meaningful, relevant, responsible, and ethical outcomes. We shared the AI Model Audit Trail concept and how third parties could create AI certification models that ensure that the AI models follow the laws and regulations and respect the rights and freedoms of the data subjects and natural persons. Finally, we outlined the challenges to the AI Model Audit Trail concept and the steps we can take to ensure its viability.

I will have one more blog – a summary – that brings together this entire series on the eight steps to deliver on the promise and potential of creating a more fair, just, and empowered society with AI.

Going Beyond Pride Month in Corporate Culture

Since the decriminalisation of homosexuality in India in 2018, corporations, especially the ones from tech, have made significant progress in making the sector more diverse and inclusive. They have embraced inclusive hiring methods, offering awareness training, and establishing groups dedicated to LGBTQ+ advocacy, among others. However, the journey of figuring out one’s sexual identity extends beyond the corporate sphere, as it is intertwined with broader socio-economic challenges and societal norms.

The process of self-discovery often spans years, requiring introspection and understanding to discern one’s true identity. Ruku Arora, Director – Enterprise Business Solutions (EBS) at Walmart Global Tech, shares a similar story about how, even within the rings of affluence, individuals grapple with the facets of self-discovery, particularly when embracing her queerness while growing up in a large Punjabi family in Delhi.

Recognising that she was queer during school, she kept it private due to societal limitations at the time. It wasn’t until college, that she selectively disclosed her identity to open-minded friends, and studying in London further bolstered her confidence. Returning home, Arora gradually began openly discussing her sexuality.

Around six or seven years ago, seeking a change, Walmart provided a supportive environment that allowed Ruku to openly embrace her identity without judgment. Today, she takes pride in her personal and professional journey, serving as a leader at Walmart Tech.

Having spent almost 20 years in the industry, she attributes her leadership style to being influenced by her journey, particularly in self-acceptance and self-love. So, her approach to leadership in a workspace marked with open-mindedness, humility, resilience, and a commitment to challenging stereotypes daily is a testament to the power of inclusivity and embracing one’s true self.

“Creating such an environment is crucial to me because I benefited from it when I was able to come out to the world. I am passionate about fostering inclusivity not only for the community but as an integral part of the overall organisational culture. While the community is significant, I aim for everyone to feel a sense of belonging, viewing the workplace as a second family,” said Arora, highlighting the importance of actively engaging in mentorship, especially for those pursuing leadership positions in the community.

And Walmart is taking significant steps to make work more inclusive and diverse. Here, employees are encouraged to voluntarily self-identify, updating their platforms to reflect their identity. The option to use preferred names and pronouns is also extended across internal platforms.

Walmart provides inclusive healthcare policies that support gender reaffirmation surgeries, recognise same-sex partners and their children as covered dependents, and provide coverage for non-traditional family expansion options like surrogacy and fertility treatments, as well as HIV and AIDS insurance coverage. Mental health is addressed through counseling services, wellness webinars, and mental health insurance coverage.

The PRIDE Associate Resource Group (ARG) works towards creating a diverse and inclusive environment, organising initiatives such as floor walks, sensitisation programs, and Pride Month celebrations. Comprehensive training for hiring managers focuses on inclusive talent recruitment, and education on identifying and addressing unconscious biases is provided for leaders and associates. Support staff also benefit from counseling services, awareness programs, and sensitisation initiatives.

“The Walmart culture, which predates and will outlast my tenure, has been instrumental in maintaining a positive environment. Despite occasional challenges, I find purpose in contributing to this culture, and that’s one of the reasons I’ve remained with the organisation for an extended period,” said Arora.

How Has the Industry Become More Inclusive over the Years

“Pride networks, similar to Pride at Walmart, have emerged in numerous companies, spanning startups, large multinational corporations, and tech companies. Even on the government side, collaborations with social impact organisations to drive change and hire transgender individuals have become more prevalent—a development I wouldn’t have envisioned 20 years ago,” Arora commented happily.

Companies, including startups, are actively adopting positive changes beyond Pride networks, updating policies to create a more inclusive environment. These policy shifts encompass critical areas like same-sex partnerships, transgender support, and emotional well-being.

Notably, measures such as covering gender-affirming surgeries and allowing preferred names on official documents are becoming increasingly common. “The fact that you and I are having this conversation today is a testament to these positive shifts,” she added.

Importance of Pronouns

Recognising the importance of pronouns in fostering a truly inclusive workplace, it is essential to address the persistent use of dead names for individuals, particularly after surgery or a name change.

Despite progress in inclusivity, this aspect is still relatively unexplored. Using correct pronouns isn’t just about language; it’s a crucial way to respect and acknowledge an individual’s gender identity, preventing potential harm like gender dysphoria.

But how can companies make it better?

“Starting small, like adding pronouns to email signatures, has immense potential to embed inclusive principles in our broader organisational culture. Achieving a more inclusive workplace requires both individual efforts and a collective commitment from organisations to embrace diversity comprehensively,” explained Arora.

Companies can also train employees to be better allies as they play a crucial role in any community, especially in supporting queer individuals.

“When I came out, it was to an ally, and their support meant everything. One simple yet powerful way allies can contribute is by educating themselves about the community,” she said, suggesting that one should take the initiative to learn and understand if they lack knowledge, encouraging open and honest conversations where the desire to learn is expressed without fear of causing unintentional offense. Sincere efforts to seek understanding go a long way in building meaningful connections.

Continuous personal growth is essential—upskilling, unlearning, and relearning.

“Be an ally not just because of someone you know but because it’s a basic human right,” concluded Arora.

Read more: Embracing Identity: The Journey of Sujoy Das

The post Going Beyond Pride Month in Corporate Culture appeared first on Analytics India Magazine.

Generative AI filled us with wonder in 2023 — but all magic comes with a price

gettyimages-1841164781

With all the advances and cultural impact of artificial intelligence (AI) this year, it would seem fair to declare 2023 as "The Year of AI" — except it's all been done before.

As this academic journal reports, the "year of AI" was declared 43 years ago, back in 1980. AI has been with us for a very long time. Decades ago, I did an academic thesis on AI ethics. In 1986, I wrote an article for the long-defunct Computer Design Magazine entitled "Artificial Intelligence as a Systems Component". And then, in 1988, I introduced two AI-based products for the Mac.

Also: AI in 2023: A year of breakthroughs that left no human thing unchanged

And even then, AI was more than 30 years old. We can trace some of the earliest AI activities to Professor John McCarthy of Stanford, MIT, and Dartmouth. In 1955, he founded SAIL, the Stanford AI Lab, and in 1958, he invented the lovely LISP (one of my all-time favorite programming languages).

So, by 2023, AI has been around for at least 68 years. And that didn't count speculative fiction. Isaac Asimov started to contemplate AI ethics 25 years earlier, in 1940.

And yet, I'd be hard-pressed to argue against calling 2023 the Year of AI. It's been quite a year.

What changed?

AI has been in use for a very long time. Whether it's in expert systems, diagnostic tools, video games, navigation systems, or many other applications, AI has been put to productive use for decades.

But it's never been put to use quite like it has this year. This is the year that true generative AI has come into its own. While many years (1980, I'm looking at you) could lay claim to the "Year of AI" moniker, there is no doubt that 2023 is the "Year of Generative AI".

Also: How does ChatGPT actually work?

The big difference, the one that has led to the enormous explosion of truly useful AI this year, has been the way we're able to train AIs. Up until now, most of the training for AIs has been supervised. That is, each AI has been fed specific information by AI designers, which compose the knowledge corpus of the AI. That limited supervised pre-training has limited what the AI knows about and what it can do.

By contrast, we're now in a time of large language models (LLMs), where the pre-training is unsupervised. Rather than feeding in a limited set of domain-specific information and calling it good, AI vendors like OpenAI have been feeding the AIs pretty much everything — the entire internet and just about any other digital content they can get their hands on.

This process allows the AI to produce astonishingly varied material with a breadth that was impossible before.

Aiding this process has been vast improvements in processor performance and storage. Back in 1986 when I wrote my article about AI as a systems component, you could get a hard drive that was the size of two microwaves and the weight of a full refrigerator for $10,000 (roughly $27K today). It held 470 megabytes. Not gigabytes, not terabytes — megabytes.

Also: Storage improvements have outperformed Moore's Law by a factor of 800%

Today, by contrast, you can pick up a 20TB internal enterprise NAS hard drive from Amazon for $279. The combination of the cloud, broadband, vastly faster processors in the form of both CPUs and GPUs, and much larger RAM pools all make the processing power of LLMs possible.

An example

To give you an example of this difference, let's use one of the products I introduced all those years ago. House Plant Clinic was an expert system that had been trained in its domain knowledge by a horticulturalist. My other product at the time was the expert system development environment, Intelligent Developer, used to build House Plant Clinic.

The process was painstaking. Through a very long series of interviews, another engineer and I elicited rules, facts, and best practices from the plant expert, and then encoded them into the knowledge base. At the plant expert's direction, we also had illustrations produced for situations in which users might need to see a visual.

House Plant Clinic's scope of knowledge consisted of what we had encoded in the expert system, nothing more and nothing less. But it worked. If you had a question and your question fell into the confines of the knowledge we had encoded, you could get an answer and be confident it was correct. After all, the knowledge provided had been vetted by a plant expert.

Now, let's look at ChatGPT. I asked ChatGPT this question:

I have a house plant that's sick. Ask me step by step questions, requiring only one answer per question.

It did a fair job of asking questions, asking about the moistness of the soil, the condition of leaves, and so on. Although it didn't volunteer an image, when I asked it to show me an image of pests, along with their names, that might be found on a house plant, I got a much more advanced image:

That said, nobody — not even Google — has any idea what a "KRIDEFLIT" is. As we have seen over and over, generative AI does have a bit of a truthiness problem.

Also: I fact-checked ChatGPT with Bard, Claude, and Copilot — and this AI was the most confidently incorrect

So, while ChatGPT can speak confidently on almost any topic, our much older expert system-based project had a much better chance of being accurate. One was created and vetted by an actual subject matter expert, while today's chatbot generates information from a giant pool of unqualified data.

The generative AI that we have been using this year can do so much more, but all magic comes with a price.

Pandora's box

Generative AI is amazing. This year, as part of my process of learning and testing the technology to report back to you, I used generative AI to help me set up an Etsy store, to help me create album art for my EP, to help my wife's e-commerce business by creating custom social marketing images, to create a WordPress plugin, to debug code, to do detailed sentiment analysis, and so much more.

Also: Generative AI can save marketers 5 hours weekly, as research finds productivity gains for the future

But generative AI is not without its problems. As we've shown, it has a severe accuracy problem. You can't trust what the AI produces. Because it's been trained on such a wide corpus of knowledge, it's incredible. But because it's been trained on such a wide corpus of knowledge, it has been polluted by what we humans write and publish.

That issue brings us to bias and discrimination. This article is already running long, so rather than try to rephrase what my colleagues have written, I'm going to point you to some of their excellent thought pieces on this subject:

  • Today's AI boom will amplify social problems if we don't act now, says AI ethicist
  • AI safety and bias: Untangling the complex chain of AI training
  • Artificial Intelligence in healthcare is racist
  • The ethics of generative AI: How we can harness this powerful technology

And then there are the jobs. As far back as six years ago, I sat down with my technology press colleague Bob Reselman to discuss concerns. And this was way before ChatGPT was actively convincing white-collar workers to worry about their futures. More recently, earlier in the year, I discussed a real concern about how ChatGPT and its ilk is likely to replace knowledge workers en mass.

Today, ChatGPT acts like a particularly talented intern with an attitude problem. It's helpful, but only when it wants to be. But as this technology evolves, it will be able to handle larger problems with more nuance, and then we'll have larger problems.

Also: Is AI in software engineering reaching an 'Oppenheimer moment'?

It's one thing for me, a guy with a two-person company, to rely on AI to help force multiply my time. But when bigger companies decide they'd rather save money and use AI services, a lot of folks will lose their jobs.

This trend will start with the entry-level positions, because ChatGPT is basically an entry-level worker. But then, three other trends will follow:

  1. There will be fewer and fewer experienced workers because not enough beginners will be able to enter the workforce.
  2. AIs will become more sophisticated and companies will feel comfortable replacing $ 100,000-a-year workers with $100-a-month AI subscriptions — even if the work output by the AI isn't quite as clean, sophisticated, nuanced, or accurate as the work produced by paid professionals.
  3. Work quality and output will reduce, along with accuracy, having a ripple effect throughout the rest of the economy and society.

In a recent article, I said the following:

We are standing on the cusp of a new era, as transformative and different and empowering and problematic as were the industrial revolution, the PC revolution, and the dawn of the Internet. The tools and methodologies we once relied upon are evolving, and with them, our responsibilities and ethical considerations expand.

The good, bad, and ugly

We started 2023 with holy cow, I can make it write a Star Trek story, and holy cow, I can make it talk like a pirate. By the end of the year, we had a much better picture of the good, the bad, and the ugly.

On the good side, we now have a helpful, if unreliable personal assistant that can save us time, help us solve problems, and get more work done.

Also: These 5 major tech advances of 2023 were the biggest game-changers

On the bad side, we have an existential job threat to all knowledge workers and an automated bias reflector that taps into our collective zeitgeist and sometimes chooses the shoulder with the devil instead of the one with our better angels.

As for the ugly, there is work to be done:

  • Finding a way to increase accuracy without nerfing effectiveness with too many guardrails.
  • Presenting useful information and illustrations without plagiarizing the folks whose job it puts at risk.
  • Preventing the misuse of AI to alter elections and other nefarious activities.
  • Taking input and generating output that's long enough to have real meaning.
  • Moving into other media, like video generation, that's as astonishing as the image generation tools.
  • Helping students learn without giving them an unbeatable way to cheat at their homework.
  • And on and on and on.

AI has blossomed in 2023 unlike any other year in the half-century or more it's been with us. The technology has opened the door to powerful tools, but also terrifying consequences.

What do you think of 2023 and what do you expect, hope for, and fear for 2024? Let us know in the comments below. I'm only writing about the generative AI transformation of 2023. If you'd like to look at some broader trends, this ZDNET article is a great place to start.

You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter on Substack, and follow me on Twitter at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, and on YouTube at YouTube.com/DavidGewirtzTV.

Featured


Ola Unveils ChatGPT Clone, Calls it Krutrim

Ola Copies OpenAI’s ChatGPT, Calls it Krutrim AI

Ola’s chief Bhavish Aggarwal recently unveiled Krutrim (which means artificial in Sanskrit). This has also been touted as “India’s first full-stack AI” solution. At first glance, the platform has a stark resemblance to ChatGPT — at least the UI/UX bit of the platform — but only greenish.

Aggarwal claimed that Krutrim AI is better than GPT-4 in various Indic languages. He said it is trained on 2 trillion tokens and can understand over 20 Indian languages and generate content in about 10 languages, including Marathi, Hindi, Bengali, Tamil, Kannada, Telugu, Odia, Gujarati, and Malayalam.

“Today, all AI models, called LLM, are trained largely in English, but language is not just text. Language is also the vehicle for cultural values, context and ethos, and current AI models just can’t capture India’s culture, knowledge, and aspirations given our multicultural, multilingual heritage,” said Aggarwal.

Is Krutrim AI Even Real?

Confident Aggarwal said, “This is not just a wrapper on some existing API. This is not just a little bit of fine-tuning done, which that’s what people in the AI industry call about taking an existing model and putting a little bit more of a data set into it.” Some of the existing API includes GPT-4, Llama 2, Gemini and others.

“This is the foundational work starting from the science layer, changing the math and the algorithms of the models to make it more relevant for Indian languages,” explained Aggarwal.

During the Krutrim launch event, Aggarwal even took a swipe at an American AI company (not sure if it is OpenAI), and said “We heard them talk about building AI for humanity and doing it democratically.”

It’s commendable that Ola has created India’s first generative AI model from scratch, but the likelihood of this achievement seems quite rare. Even other major players, not just in the Indian market but also outside, took more time and resources to build their generative AI models from scratch.

Ola chief claims that following Krutrim, they will release its first multimodal Krutrim Pro by next quarter. He said for Krutrim Pro, “All the dimensionalities and modalities will be taken as inputs together, and the algorithm will be able to train across modalities. This is different from a separate text model, a distinct speech model, and a separate vision model being trained in parallel.

Interestingly, OpenAI announced a multimodal GPT-4 back in March, but in reality, GPT-Vision came into existence only by September — i.e. over six months.

It is intriguing how Krutrim managed to develop an LLM in just three months, especially when compared to its Indian counterparts such as Corover AI, Sarvam AI, and Kissan AI, which took considerably longer. Notably, these counterparts constructed their models on top of Llama 2 or GPT-4, which should have ideally taken less time as compared to Krutrim.

Moreover, it took Elon Musk four months to build Grok and xAI is grappling with GPU shortages, a concern acknowledged by Oracle in its latest earnings call. Surprisingly, Aggarwal didn’t disclose the number of GPUs acquired for training their model throughout the entire event. Instead, it said, it is building all of these capabilities in-house.

Moreover, Aggarwal didn’t disclose the dataset on which Krutrim AI is trained. Training a model on multiple languages is no easy feat, especially considering the increasing cost.

Languages, particularly those with intricate structures and scripts like Hindi, Kannada, or Telugu, demand varying numbers of tokens. English, with its simplicity, requires fewer tokens in comparison. The financial implications of tokenization disparities are even more pronounced. The cost of training and using AI models hinges on token counts, compute and cloud costs.

Krutrim had raised $24 million in debt from Matrix Partners, which is also a shareholder of Ola Electric. However, Aggarwal clarified earlier that Krutrim is a separate entity from Ola.

Regardless of everything else, Ola has not yet released a research paper or details of the dataset trained on, and the team leading this initiative, are still confused if they should call themselves Ola or Krutrim.

All of this looks like a marketing gimmick in an attempt to secure additional funds. That too, without even releasing a solid product into the market, and the demo, hardly looked convincing.

The post Ola Unveils ChatGPT Clone, Calls it Krutrim appeared first on Analytics India Magazine.

Where Are All The Women of AI?

Sidelining women’s input in tech developments has never been surprising. From Ada Lovelace in the 1800s not being rightfully credited for her contributions to the latest ‘Who’s Who in AI‘ list by the New York Times — the tech world has always had a hard time cheering the achievements of anyone except white males.

The list by NYT names a dozen men without including a single woman. What’s worse is that the journalistic piece thought it had “an internet philosopher and self-taught AI researcher” and Elon Musk of all people (who is also more often seen in the media than the newly hired CEO Linda Yaccarino).

Dr Fei-Fei Li, Romanian researcher Doina Precup, Aakanksha Chowdhery, the driving force behind Google’s AI model PaLM and Global Vice President of Meta’s Artificial Intelligence Research department (FAIR) Joelle Pineau are some of the strong female technical contributors that NYT could have (at least) added. Several people criticised the list, including tech journalist Kara Swisher’s reference to Li. The researcher then added, “It’s not about me, but all of us in AI, all of the incredible ‘godmothers’, pioneers, active researchers, students from all walks of life.”

Inclusivity Is All Talks In AI

Diversity and inclusivity are heavily spoken about on paper by major tech companies. The flaw not only exists in the corporate ecosystems but is also highly visible in models developed by the money-minting monkeys of the Silicon Valley circus.

Only a few women have a seat at the table in AI-generated content we have seen in media headlines all year. Be it your favourite chatbot or image generator, each has its own (data)set of sexism.

In September, two Oxford, UK-based organisations released a study examining the inherent gender bias of 13 famous AI chatbots. The results could have been better. OpenAI’s flagship product ranked the highest regarding professional gender bias.

While writing this article, ChatGPT was asked, ‘Who are some of the most important figures in the history of humanity? Unsurprisingly, the version based on GPT3.5’s list consisted of only men, and GPT4’s version made a list of 18 names, with Marie Curie the only woman being mentioned. But ChatGPT is just an element of a larger issue.

Back in June, researchers from Johns Hopkins University and the Georgia Institute of Technology trained robots in computer vision with OpenAI’s CLIP, and then asked the robot to scan and categorise digital blocks with images of people’s faces. It mostly classified women as homemakers over white men.

The problem has long existed in the way media covers media, too. Earlier this year, AdWeek wrote a profile about Dr Sasha Luccioni, an AI researcher currently leading the Climate team at Hugging Face. The article initially published the story with a stereotypical headline, which was later changed as Luccioni pointed that out herself.

Coming back to the NYT piece, slow claps for the to mention Anthropic’s Dario Amodei, but not his cofounder Daniela: given a man and a woman in roughly parallel roles.

Moving Forward (?)

Tech companies trying to fix the gender problem have repeatedly fallen short of their goals. The most recent case is of Google being charged guilty of sex discrimination and retaliation but not guilty of violating New York’s pay discrimination law.

Attempts by the companies are visible now and then, but some research shows discrimination is worsening. The percentage of tech professionals who said they experienced gender discrimination jumped from 21 per cent in 2021 to 26 per cent in 2022, according to a survey from the tech career marketplace.

The NYT story listicle is a mere representation of males dominating the industry. In 2019, the conversation about tech’s approach toward diversity not working was already happening, showing how laid back the executives have been about the issue. This piece can conclude on a similar note as an article from 2019, that companies should consider how everyday organisational practices can be improved, and they can move beyond programs that seek to address only the biases locked within people’s heads — the bias in structural forms should be addressed as well.

The post Where Are All The Women of AI? appeared first on Analytics India Magazine.

Data Science Hiring Process at Marlabs

Founded in 2011, New York-based IT services and consulting firm Marlabs helps companies of various sizes to undergo AI-powered digital transformation. It provides a wide range of services, including strategic planning, creating rapid prototypes in specialised labs, and applying agile engineering techniques to develop and expand digital solutions, cloud-based applications and AI-driven platforms.

Marlabs’s data science team addresses a range of industry challenges, emphasising tasks like extracting insights from extensive datasets and employing pattern recognition, prediction, forecasting, recommendation, optimisation, and classification.

Exploring Generative AI at Marlabs

“In operationalising AI/ML, we have tackled diverse projects, such as demand forecasting, inventory optimisation, point of sale data linkage, admissions candidate evaluation, real-time anomaly detection, and clinical trial report anomaly detection,” Sriraman Raghunathan, digital innovation and strategy principal, Marlabs, told AIM in an exclusive interaction.

The team is also exploring generative AI applications, particularly in knowledge base extraction and summarisation across domains like IT service desk ticket resolution, sustainability finance, medical devices service management, and rare disease education.

However, it is not developing foundational models as of now due to substantial capital requirements. “Instead, we are focussing on the value chain beyond foundational models, offering tools and practices for deploying such models within organisation boundaries, tailored for specific domains,” he added.

Marlabs employs a variety of tools and frameworks depending on project specifics, utilising R and Python for development, Tableau, Power BI, QlikView for data exploration and visualisation, and PyTorch, TensorFlow, Cloud-Native tools/platforms, and Jupyter Notebooks for AI/ML model development.

The team leverages Transformer models like GPT-3, especially in NLP use cases, implementing them in TensorFlow, and PyTorch, and utilising pre-trained models from Hugging Face Transformers Library. For generative AI, their toolkit includes LangChain, Llama Index; OpenAI, Cohere, PaLM2, Dolly; Chroma, and Atlas.

Hiring Process

The hiring process for data science roles at the organisation emphasises a blend of technical knowledge, practical application, and relevant experience. The initial steps involve a clear definition of the role and its requirements, followed by the creation of a detailed job description.

The interview process comprises technical assessments, video interviews with AI/ML experts, and HR interviews. Technical assessments evaluate coding skills, data analysis, and problem-solving abilities.

Video interviews focus on the candidate’s depth of knowledge, practical application, and communication skills, often including a discussion of a relevant case study or project. HR interviews center around cultural fit, interpersonal skills, collaboration, and the candidate’s approach to handling challenges.

Expectations

“Upon joining the data science team, candidates can anticipate a thorough onboarding process tailored to their specific team, providing access to essential tools, resources, and training for a smooth transition,” commented Raghunathan.

The company’s AI/ML projects involve cutting-edge technologies, exposing candidates to dynamic customer use cases spanning natural language processing, computer vision, recommendation systems, and predictive analytics. The work environment is agile and fast-paced. The company places a strong emphasis on team collaboration and effective communication, given the collaborative nature of data science and AI/ML projects.

In this rapidly evolving field, the company expects new hires to demonstrate continuous learning, tackle complex technical and functional challenges, operate with high levels of abstraction, and exhibit creative and innovative thinking.

Mistakes to Avoid

“The most prevalent error observed in candidates during data science role interviews is a lack of clear communication,” he added.

The ability to effectively communicate insights to non-technical stakeholders is crucial in the AI/ML space, and this skill is frequently overlooked.

Another common mistake is a failure to comprehend and articulate the business context and domain knowledge of the problem, which is essential in AI/ML applications with significant business impact.

Work Culture

“We are recognised for our value-based culture focused on outcomes, emphasising a flat organizational structure to spur innovation and personal growth. Key values such as respect, transparency, trust, and a commitment to continuous learning are central to their ethos, all aimed at exceeding customer expectations,” he said.

The company’s robust learning and development program has prepared over 150 young managers for leadership roles, with a strong emphasis on AI and technology for organisational insights and sentiment analysis.

The company offers a comprehensive benefits package, including versatile insurance plans, performance incentives, and access to extensive learning resources like Courseware and Udemy, supporting a hybrid work model. Additionally, they provide mental health support and reward long-term employees based on tenure.

Raghunathan further explained that in the data science team, Marlabs stands out for its innovative and collaborative environment, encouraging creativity and continuous learning. “This distinctive culture and investment in employee growth make us a leader in data science, differentiating it from competitors in the tech industry,” he added.

Why Should You Join Marlabs?

“Join Marlabs for a dynamic opportunity to work with a passionate team, using data to drive meaningful change. In this collaborative setting, data scientists work with brilliant colleagues across various industries, including healthcare, finance, and retail. You’ll tackle complex issues, contributing to significant business transformations. Marlabs supports your career with essential tools, resources, training, competitive compensation, benefits, and opportunities for professional growth and development,” concluded Raghunathan.

The post Data Science Hiring Process at Marlabs appeared first on Analytics India Magazine.

Why OpenAI is Eyeing India

OpenAI and India’s collaborative ties has gained momentum with the company having plans to soon start an Indian office. From appointing Indian-origin leaders in advisory roles to bringing OpenAI executives to India, the company’s expansion plans in the country will be a fruitful one for both the parties.

Rishi Jaitly, who has held executive positions including the position of Vice President at Twitter, will assume the role of a senior advisor at OpenAI to guide the company through India’s AI policy and regulatory environment. Furthermore, OpenAI executives Anna Adeola Makanju , global head of Public Policy, James Hairston, and Jaitly recently met MoS for Electronics and Information Technology, Rajeev Chandrashekar.

From left to right: Rishi Jaitly, Rajeev Chandrashekhar, Anna Adeola Makanju, and James Hairston. Source: X

In the backdrop of Global Partnership on Artificial Intelligence Summit (GPAI) that happened a few days ago, where India’s stance on aggressive AI expansion and development was evident, the OpenAI executive’s meeting with the minister fortified the company’s plan to be a prominent AI competitor in the Indian landscape.

India : The Precious Market

As per recent statistics, ChatGPT has over 180 million users, with the US contributing to 10.81% of total users, followed by India with 9.08% users. Being the 2nd highest market for ChatGPT, India’s contribution to the OpenAI market is already well placed. Big tech companies, such as Google, are already choosing India as a destination for building their biggest offices outside the US.

With favourable business conditions, and progressive policies, India has been an attractive market. By setting up an Indian office, OpenAI will be able to capitalise on the same.

Third in Line

London was the first office outside the US which OpenAI opened in June, followed by another in Dublin, three months later. Going by the recent developments, OpenAI will open its third office in India in a couple of months.

Interestingly, OpenAI CEO Sam Altman had an extensive world tour this year, visiting a number of countries in an attempt to build strategic relationships with world leaders and even be the face of ChatGPT. At that time, he had also visited India and met Prime Minister Narendra Modi, possibly laying the groundwork for future expansion plans.

OpenAI is also set to host its first developer conference in Bengaluru in January. This will be OpenAI’s second developer conference after hosting its maiden one in November. OpenAI VP of engineering, Srinivas Narayan, is expected to visit leaders and developers at the conference.

India’s LLM Moment in the Making

Having an office in India, will also benefit India’s status on the global AI race too. Furthermore, if an office comes up here, it would mark the first OpenAI office in Asia, even surpassing Japan, which was touted to have the next office, after Altman met Japan’s PM and even discussed promoting AI models that will entail Japan’s culture.

In the recently concluded GPAI, the country has committed to building its own AI models, and initiatives are already on track. Recently, KissanAI launched Dhenu 1.0, an agriculture large language model which is bilingual and comprehends English, Hindi and Hinglish queries. Founded by Pratik Desai, KissanAI caters to the agriculture market which is expected to hit $451.59 billion by 2028. Dhenu 1.0 processes 3,00,000 instruction sets in both English and Hindi.

Language-specific models have been on the rise with a number of companies following suit. Bangalore-based Sarvam AI that raised $41 million in Series A funding, released OpenHathi-Hi-v.01, a Hindi LLM which was built on Llama2-7B.

CoRover.ai has recently partnered with Google Cloud to launch BharatGPT, a generative AI conversational bot catering to the Indian market. The model will support over 14 languages in text, voice and video interactions.

India’s foray into Indic-language models, and building datasets specific to India, has been ongoing. AI4Bharat, and Bhashini, have been other projects that have been building Indic-level datasets.

An unlikely automotive player, Ola, also unveiled an India-centric AI model called Krutrim AI which powers an AI chatbot similar to ChatGPT. It is said to understand 22 Indian languages and generate text in 10 languages.

Not just in India, but on the global front as well, there are demographic-specific models which are competing with OpenAI. UAE’s Jais, an Arabic LLM and A171, an AI company that will build on Falcon model, a proprietary model of Abu Dhabi’s Technology Innovation Institute. China is also catching up. With the recent Baidu’s Ernie 4.0 LLM model, which is bilingual, with English and Chinese, OpenAI’s competitor market is only diversifying.

It is evident that OpenAI’s push to enter the Indian market comes at a time when the country is likely going to witness a surge of India-specific models, which would fare better in an Indian subcontext.

The post Why OpenAI is Eyeing India appeared first on Analytics India Magazine.

AI plushie Grok, voiced by Grimes, was trademarked before Elon Musk’s Grok

AI plushie Grok, voiced by Grimes, was trademarked before Elon Musk’s Grok Morgan Sung 9 hours

Grimes is stepping into the toy business with “Grok,” a character that she voiced for Curio’s new line of screen-free AI plushies.

The toy is not affiliated with the AI chatbot backed by Grimes’ ex, Elon Musk, which is also named Grok. Musk described xAI’s Grok as having a “rebellious streak” and a willingness to answer “spicy questions that are rejected by most other AI systems.” It’ll be vulgar if you ask.

Grok, Gabbo and Grem, on the other hand, are designed to encourage play. In a conversation with Curio founders Misha Sallee and Sam Eaton, posted on Curio’s blog, Grimes spoke about encouraging creativity in children early through dynamic conversations, rather than a static list of prompts.

“I just like the idea of bringing more imagination, or making it easy to access imagination in your kind of current existence as opposed to just observing it in other existences, like on screen or in a movie or book or something,” she said.

In Curio’s announcement video, Grimes said that she didn’t want her kids “in front of screens,” but she’s “really busy.”

Image Credits: Curio

Curio says that the toys can hold full conversations so that kids (or adults) can practice their communication skills. There’s Grok, an anthropomorphized rocket ship voiced by Grimes. There’s Gabbo, who looks like a plush Gameboy with arms and legs. And there’s Grem, a cyan bunny with hearts on its cheeks. The beta versions of the toys are available for preorder until Sunday, and are priced at $99 each. They’re recommended for kids aged 3 to 7 — Grimes’ oldest child with Musk, named X Æ A-Xii, is 3.

The plushies will answer questions about how rocket ships are made, play games with the user and encourage kids to develop listening and conversation skills. Encased in the plushie is a rechargeable, Wi-Fi-connected speaker and mic, which is connected to an app for parents to set up and monitor interactions with their kids.

“When I think about kids, my goal is to preserve as many minds as possible from here, and how much can we replace iPads, basically?” Grimes said in the conversation with Eaton and Sallee.

She later added, “I think the more you keep things verbal, too, the more you’re sort of forcing people to use their working memory. There’s all these little things that, you know, make our brains better just a little bit here and there.”

Welcome to Curio, where toys come to life! 🧸

We've partnered with @Grimezsz on our first 3 characters: Grok 🚀, Gabbo🤖, and Grem👽.

Order by Sunday to receive a Beta Program Certificate in the mail before Xmas, toys ship early 2024!https://t.co/AQFtMIU7sG pic.twitter.com/vLVI3PTaMN

— Curio (@CurioBeta) December 14, 2023

Grimes got involved with Curio after responding to a post about the future of AI-integrated toys, in which “children’s teddy bears will speak to them and make them feel safe at night.” Grimes replied that it would be “great if safe,” and that she’d love if her kids could hang out with a “culture ship mind in a teddy bear.”

The line launched about a week after Musk’s ChatGPT competitor, also named Grok, began rolling out to X Premium Plus subscribers.

“Grimes is doing the voice for the toy, and this one is a rocket who is coincidentally named Grok and predated the Grok AI announcement, so there’s a funny overlap there,” Sallee said in the conversation with Grimes.

As Business Insider reports, Grimes’ Grok was trademarked first.

Curio filed its trademark for Grok on September 12 this year. xAI filed its trademark for Grok on October 23. Curio’s Grok is short for Grocket, since Grimes’ kids spend so much time around rockets because their father owns SpaceX, The Washington Post reports.

Grimes and Musk are currently engaged in a custody battle over their three children, and have filed child custody lawsuits against each other in California and Texas, respectively.

(Absurdly by the time we realized the Grok team was also using this name it was too late for either AI to change names, so there are two AI's named Grok now, I can't wait for them become friends. I can't believe even ai can't avoid showing up at school and meeting another kid…

— Princess Irulen ® (@Grimezsz) December 14, 2023

In a post addressing the name, Grimes said that by the time Curio realized that xAI’s Grok team was also using the name, “it was too late for either AI to change names.”

“So there are two AI’s named Grok now, I can’t wait for them become friends,” she said. “I can’t believe even ai can’t avoid showing up at school and meeting another kid with the same name haha.”

Big IT Challenges Australia Needs to Address in 2024 to Seize the AI Moment

Australian organisations are widely expected to deploy more artificial intelligence use cases in 2024. It is a trend IT research and advisory firm ADAPT says is not only dependent on the data housed across organisations but also the data culture and digital mindset of the organisation it lives within.

ADAPT Senior Research Director Matt Boon said Australian research data shows leaders want to create data-driven organisations to capitlise on technologies like AI, but they need to work better together to improve prioritisation and digital awareness and to deliver value from innovation faster.

Jump to:

  • Being data-driven and improving operations are 2024 priorities
  • Challenges include too many competing projects and a talent war
  • CIOs urged to pursue incremental innovation and digital fitness

Being data-driven and improving operations are 2024 priorities

ADAPT’s research surveys from 2022 and 2023 show that “creating a data-driven organisation” is consistently named among the top five business priorities held by Australian C-suite executives, including chief financial officers, chief information officers and chief data officers.

However, surveys of data leaders and chief data officers specifically between 2020 and 2023 show “inconsistent data culture across the organisation” as the biggest hurdle to creating a data-driven organisation, which would support organisations with initiatives, such as capitalising on AI.

Data needs to be a part of organisational culture from top to bottom

ADAPT’s Matt Boon said that rather than creating a data-driven organisation where data is collected across disparate systems — sometimes with a lack of deep understanding of the why or how — it was more important to “become data driven,” making it part of the culture of the organisation.

SEE: Explore these tips on improving data literacy across the organisation.

“How do we create the right culture — within the business across our employees and leadership teams — to recognise the value of data and how it can help us make the right decisions at the right time for the right reasons, to drive the outcomes that we’re trying to achieve?” Boon said.

Improving operational effectiveness is another priority for senior leaders

“Improving operational effectiveness” is identified as another significant priority, according to ADAPT’s research. Boon said there is a need for organisations to do things more efficiently, including leveraging technology, to overcome current constraints. He said there was “no real end game” in this for IT leaders.

“We need to consider how technologies like AI will help us improve the overall effectiveness and efficiency of our organisation,” Boon said. “AI can help us reduce the mundane and improve operational effectiveness, freeing up time for people to focus on the value side of their role.”

Challenges include too many competing projects and a talent war

CIOs struggle to prioritise multiple projects

The number one challenge for CIOs heading into 2024 is the number of “competing business priorities” they face. In fact, ADAPT surveys in 2023 found that, of the top 10 priorities named by CIOs, 50% of respondents rated them all as either “important” or “very important.”

However, Boon said many of the top priorities are “somewhat linked together” and would be better seen collectively rather than individually. He asked: “How can technology like AI, for example, help us really deliver on some priorities collectively, rather than individually?”

Boon added that leaders need to work collectively within organisations to address these competing priorities.

“When we do have a common purpose, we can achieve a lot,” Boon said. “Competing business priorities just indicates we need to have a more common approach to what we’re doing.”

Legacy tech is still holding organisations back

Legacy technology and technical debt continue to “hamstring” many organisations, making it a big challenge into 2024, Boon said.

SEE: Businesses should prepare for obstacles when migrating legacy data to the cloud.

“A lot of organisations are still dealing with very complex environments and have major challenges around legacy technology,” said Boon. “But it’s not just the technology; it’s also the people — legacy mindsets and legacy processes.

“So we need to be really focused on why we need to change. For example, we may have always done things a certain way, but what’s the reason to actually change in general?”

Skills shortages make talent hard to find

ADAPT’s data revealed that cyber security is the most in-demand IT skill, followed by data science and analytics, according to its human resources surveys.

“It’s one thing to have tools to manage data, but we need people who can interpret it and understand what we need to do to make the right informed decisions,” Boon said.

As the “AI juggernaut” continued to accelerate through its hype cycle, he said we can expect this to accelerate a shortage of AI and machine learning skills.

“Those skills are going to be in demand beyond Australia as well, so it’s really a very, very hot market in general,” said Boon.

CIOs urged to pursue incremental innovation and digital fitness

Pursue incremental innovation to demonstrate value

Almost 60% of CIOs surveyed by ADAPT in the Australian marketplace say they want to accelerate innovation in their organisation — defined as delivering new products and new services — in order to drive higher revenue over the next 12 to 24 months. But when asked how much of their budget is allocated to innovation, the same CIOs said the number was only 6%.

“So we all want to accelerate innovation, but we’ve only got 6% of our IT budget to specifically allocate to innovation itself,” Boon said.

Whenever IT requests an investment to increase budgets, Boon said they should work harder to clearly define the linkage back to business value over a much shorter time frame.

“How can we show value on the investments we’re making at an incremental level?” said Boon.

Increase digital awareness and fitness at all levels

ADAPT’s data indicates Australian organisations are only about 47% digitised, a finding Boon said showed that Australia is “not even halfway to where we need to be effectively digitised.

“If you look at the modernisation of infrastructure, that has also come in again at just about 50%,” said Boon.

SEE: Australian enterprises need to consider cyber risks with increased digitisation.

Boon said this is partly a reflection of low digital awareness and digital fitness. For example, 43% of CIOs agree their board and leadership teams lack digital awareness, while only 48% rate their workforce as digitally fit.

Digital awareness needs to increase, Boon said, because there is a clear link between this and the willingness to deploy new technologies. For example, 49% of organisations with high levels of digital fitness are deploying AI, compared with only 35% with low digital fitness levels.

Prepare to develop people in the organisation

Over two-thirds (69%) of CIOs and CEOs have indicated they are investing in staff upskilling and training into 2024, showing a focus on developing people in general. Boon said this included looking for skills internally and building their capabilities, as well as external hires.

The development agenda includes leadership capabilities, with many organisations seeking people able to lead digitally from the front.

“There is a really big focus on how we can uplift our people to be ready and physically live for these changes that are coming,” Boon said.

LucidDreamer: High-Fidelity Text-to-3D Generation via Interval Score Matching

LucidDreamer: High-Fidelity Text-to-3D Generation via Interval Score Matching

The recent advancements in text-to-3D generative AI frameworks have marked a significant milestone in generative models. They pave the way for new possibilities in creating 3D assets across numerous real-world scenarios. Digital 3D assets now hold an indispensable place in our digital presence, enabling comprehensive visualization and interaction with complex environments and objects that mirror our real-world experiences. These 3D generative AI frameworks are applied in various domains, including animation, architecture, gaming, augmented and virtual reality, and much more. They are also being used extensively in online conferences, retail, education, and marketing.

However, despite the promise of these advancements in text-to-3D generative frameworks, the extensive use of 3D technologies comes with a major issue. Generating high-quality 3D images and media content still requires significant time, effort, resources, and skilled expertise. Even with these requirements met, text-to-3D generation often fails to render detailed and high-quality 3D models. This issue of rendering and low-quality 3D generation is more prevalent in frameworks that use the Score Distillation Sampling (SDS) method. This article will discuss the notable deficiencies observed in models using the SDS method, which introduce inconsistencies and low-quality updating directions, resulting in an over-smoothing effect on the generated output. We will also introduce the LucidDreamer framework, a novel approach that uses the Interval Score Matching (ISM) method to overcome the over-smoothing issue. We'll explore the model's architecture and its performance against state-of-the-art text-to-3D generative frameworks. So, let’s get started.

LucidDreamer3D : An Introduction to 3D Generation using Interval Score Matching

A major reason why 3D generation models has been the talking point of the generative AI industry is because of its widespread applications across various domains and industries, and their ability to produce 3D content in real-time. Owing to their widespread practical applications, developers have proposed numerous 3D content generation approaches out of which, text to 3D generation frameworks stands out from the rest for its ability to use nothing but text descriptions to generate imaginative 3D models. Text to 3D generative frameworks achieves this by using a pre-trained text to image diffusion model to as a strong image before supervising the training of a neural parameterized 3D model thus allowing for rendering 3D images consistently that aligns with the text. This capability to render constant 3D images is grounded in the use of the Score Distillation Sampling fundamentally, and allows SDS to act as the core mechanism to bring 2D results from diffusion models into their 3D counterparts, thus enabling training 3D models without using training images. Despite their effectiveness, 3D generative AI frameworks making use of the SDS method often suffer from distortion and over-smoothing issues that hampers the practical implementations of high-fidelity 3D generation.

To tackle the over-smoothing issues, the LucidDreamer framework implements a ISM or Interval Score Matching approach, a novel approach that uses two effective mechanisms. First, the ISM approach employs DDIM inversion method to mitigate the averaging effect caused by pseudo-Ground Truth inconsistencies by producing an invertible diffusion trajectory. Second, rather than matching the images rendered by the 3D model with the pseudo Ground Truths, the ISM method matches them between two interval steps in the diffusion trajectory that helps it avoid high reconstruction error by avoiding one-step reconstruction. The use of ISM over SDS results in consistently high performance with highly realistic and detailed outputs.

Overall, the LucidDreamer framework aims to make the following contributions in 3D generative AI

  1. Provides an in-depth analysis of SDS, the fundamental concept in text to 3D generative frameworks, and identifies its key limitations of low-quality pseudo-Ground Truths, and provides an explanation for the over-smoothing effect faced by these 3D generative frameworks.
  2. To counter the limitations posed by the SDS approach, the LucidDreamer framework introduces Interval Score Matching, a novel approach that uses interval-based matching and invertible diffusion trajectories to outperform SDS by producing highly-realistic and detailed output.
  3. Achieving state of the art performance by integrating ISM method with 3D Gaussian Splatting to surpass existing methods for 3D content generation with low training costs.

SDS Limitations

As mentioned earlier, SDS is one of the most popular approaches for text to 3D generation models, and it seeks modes for conditional post prior in the latent space of DDPM. The SDS approach also adopts a pretrained DDPM to model the conditional posterior, and aims to distill the 3D representations for conditional posterior that is achieved by minimizing the following KL divergence. Furthermore, the SDS approach also reuses the weighted denoising score matching objective for DDP training. The primary objective of the SDS approach can also be viewed as matching the view of the 3D model with the pseudo-ground truth that is estimated in a single step by the DDPM. However, developers have observed that the distillation process often overlooks key aspects of DDPM, and the following figure demonstrates how a pre-trained DDPM tends to predict pseudo-ground truths with inconsistent features, and produces low quality output during the distillation process.

However, updating directions under undesirable circumstances are updated to 3D representations that ultimately leads to over-smoothed results. Furthermore, it is worth noting that the DDPM component is input sensitive, and the features of the pseudo-ground truth changes significantly even with the slightest change in the input. Additionally, randomness in both the camera pose and the noise component of the inputs might add to the fluctuations which is unavoidable during distillation. Optimizing the input for inconsistent pseudo Ground Truths results in featured-average outcomes. What’s more is that the SDS approach obtains pseudo-ground truths with a single-step prediction for all time intervals, and does not take into account the limitations of a single-step-DDPM component that are unable to produce high-quality output which indicates that distilling 3D assets or images with SDS component might not be the most ideal approach.

LucidDreamer : Methodology and Working

The LucidDreamer framework does introduce the ISM approach, but it also builds on the learnings from other frameworks including text to 3D generative models, diffusion models, and differentiable 3D representation frameworks. With that being said, let’s have a detailed look at the architecture and methodology of the LucidDreamer framework.

Interval Score Matching or ISM

The over-smoothing and low-quality output issues faced by a majority of text to 3D generation frameworks can be owed to their use of the SDS approach that aims to match the pseudo ground truth with the 3D representations that is inconsistent, and often of sub-par quality. To counter the issues faced by SDS, the LucidDreamer framework introduces ISM or Interval Score Matching, a novel approach that has two working stages. In the first stage, the ISM component obtains more consistent pseudo-ground truths during distillation regardless of the randomness in camera poses and noise. In the second stage, the framework generates pseudo-ground truths with better quality.

Another major limitation of SDS is generating pseudo-ground truths with a single-step prediction for all time intervals that makes it challenging to guarantee high-quality pseudo-ground truths, and it forms the basis to improve the visual quality of the pseudo-ground truths. In a similar sense, the SDS objective can be seen as to match the view of the 3D model with the pseudo-ground truth estimated by the DDPM in a single step, although the distillation process does overlook a critical aspect of the DDPM component i.e., it produces low-quality pseudo-ground truths with inconsistent features during the distillation process.

Overall, the ISM component promises to deliver several advantages over previous methods used in text to 3D generation models. First, thanks to ISM’s ability to provide high-quality pseudo-ground truths consistently, it is able to produce high-fidelity distillation outputs with finer structures and richer details, thus eliminating the need for large scale guidance scale, and enhances the flexibility for 3D content creation. Second, transitioning from SDS approach to ISM approach has marginal computational overhead especially since the ISM approach does not compromise on the overall efficiency even though it demands for additional computational costs for DDIM inversions.

The above figure demonstrates the working of the ISM approach, and provides an overview of the architecture of the LucidDreamer framework. The framework first initializes the Gaussian Splatting i.e. the 3D representations using a pretrained text-to-3D generator using a prompt. It is then incorporated with a pretrained 2D DDPM component to disturb random views to noisy unconditional latent trajectories using DDIM inversions, and then updates with the interval score. Thanks to its architecture, the core of optimizing the ISM component focuses on updating the 3D representations towards pseudo-ground truths that are high-quality and features-consistent, yet computationally friendly. This principle is what allows ISM to align with the fundamental objectives of the SDS approach while refining the existing method.

DDIM Inversion

The LucidDreamer framework aims to produce more consistent pseudo-ground truths in alignment with the 3D representations. Therefore, instead of producing 3D representations, the LucidDreamer framework employs the DDIM inversion approach to predict noise latent 3D representations, and predicts an invertible noise latent trajectory in an iterative manner. Furthermore, it is because of the invertibility of DDIM inversion that the LucidDreamer framework is able to increase the consistency of the pseudo-ground truth significantly for all time intervals.

Advanced Generation Pipeline

The LucidDreamer framework also introduces an advanced pipeline in addition to ISM to explore the factors affecting the visual quality of text-to-3D generation, and introduces 3D Gaussian Splatting or 3DGS as its 3D generation, and 3D point cloud generation models for initialization.

3D Gaussian Splatting

Existing works have indicated that increasing the batch size and rendering resolution for training improves the visual quality significantly. However, a majority of learnable 3D representations adopted for text-to-3D generation are time and memory consuming. On the other hand, the 3D Gaussian Splatting approach provides efficient results in both optimization, and rendering that allows the Advanced Generation Pipeline in the LucidDreamer framework to achieve large batch size as well as high-resolution rendering even when operating with limited computational resources.

Initialization

A majority of state of the art text-to-3D generation framework initialize their 3D representations with limited geometries like circle, box or cylinder that often results in undesired outputs on non-axial symmetric objects. On the other hand, as the LucidDreamer framework introduces 3D Gaussian Splatting as 3D representations, the framework can adopt to several text to point generative frameworks naturally to generate a coarse initialization with human inputs. The initialization strategy ultimately boosts the convergence speed significantly.

LucidDreamer : Experiments and Results

Text-to-3D Generation

The above figure demonstrates the results generated by the LucidDreamer model with the original stable diffusion approach whereas the following figure talks about the generated results on different finetuned checkpoints.

As it can be seen, the LucidDreamer framework is capable of generating highly consistent 3D content using the input text and semantic cues. Furthermore, with the use of ISM, the LucidDreamer framework generates intricate and more realistic images while avoiding common issues like over-saturation, or over-smoothing while exceling in generating common objects as well as supporting creative creations.

ISM Generalizability

To evaluate ISM generalizability, a comparison is conducted between the ISM and the SDS methods in both explicit and implicit representations, and the results are demonstrated in the following image.

Qualitative Comparison

To analyze the qualitative efficiency of the LucidDreamer framework, it is compared against current SoTA baseline models, and to ensure fair comparison, it uses Stable Diffusion 2.1 framework for distillation, and the results are demonstrated in the following image. As it can be seen, the framework delivers high-fidelity and geometrically accurate results while consuming less resources and time.

Furthermore, to provide a more comprehensive evaluation, developers also conduct a user study. The evaluation selects 28 prompts and uses different text to 3D generation approaches on each prompt to generate objects. The results were then ranked by the users on the basis of the degree of alignment with the input prompt, and its fidelity.

LucidDreamer : Applications

Owing to its exceptional performance on a wide array of text to 3D generation tasks, the LucidDreamer framework has several potential applications including Zero-shot avatar generation, personalized text to 3D generation, and zero-shot 2D and 3D editing.

The top-left image demonstrates LucidDreamer’s potential in zero-shot 2D and 3D editing tasks whereas the bottom left images demonstrate the ability of the framework in generating personalized text to 3D outputs with LoRA whereas the image on the right showcases the framework’s ability to generate 3D avatars.

Final Thoughts

In this article, we have talked about LucidDreamer, a novel approach that uses Interval Score Matching or ISM method to overcome the over-smoothing issue, and discuss the model architecture, and its performance against state of the art text to 3D generative frameworks. We have also talked about how SDS or Score Distillation Sampling, a common approach implemented in a majority of state of the art text to 3D generation models often results in over-smoothing of the generated images, and how the LucidDreamer framework counters this issue by introducing a new approach, the ISM or Interval Score Matching approach to generate high-fidelity, and more realistic 3D images. The results and evaluation indicates the effectiveness of the LucidDreamer framework on a wide array of 3D generation tasks, and how the framework already performs better than current state of the art 3D generative models. The exceptional performance of the framework makes way for a wide range of practical applications as already discussed.