Adobe Research and University of Oxford have come up with a new paper which introduces Continuous 3D Words, a method allowing users of text-to-image models to have fine-grained control over various attributes in an image.
By engineering special sets of input tokens, these attributes can be transformed continuously, enabling users to manipulate sliders for control alongside text prompts. The approach is demonstrated to provide continuous user control over 3D-aware attributes like illumination, bird wing orientation, dollyzoom effect, and object poses.
Why is this Development Important?
The current controls for image generation in diffusion models cannot recognise abstract, continuous attributes such as illumination direction or non-rigid shape changes.
The paper emphasises that while photography offers detailed control over composition and aesthetics, text prompts in text-to-image diffusion models are limited to high-level descriptions. On the other hand, 3D rendering engines allow precise control but are labor-intensive and require expertise.
This work aims to combine the advantages of both by expanding the vocabulary of text-to-image models with samples generated from rendering engines, creating Continuous 3D Words to enable fine-grained control during image generation.
Training Method
The core of the approach involves learning a continuous vocabulary, facilitating easier association between different attribute values and allowing interpolation during inference. Two training strategies are proposed to prevent degenerate solutions and enable generalisation to new objects.
The first strategy involves a two-stage training process, preventing the model from encoding each attribute value as a new object. The second strategy employs ControlNet with conditioned images to prevent overfitting to artificial backgrounds. The entire training process is carried out in a lightweight manner for efficiency.
The post Adobe Unveils Continous 3D Words for Text-to-Image Control appeared first on Analytics India Magazine.
Women In AI: Eva Maydell, member of European Parliament and EU AI Act advisor Kyle Wiggers 9 hours
To give AI-focused women academics and others their well-deserved — and overdue — time in the spotlight, TechCrunch is launching a series of interviews focusing on remarkable women who’ve contributed to the AI revolution. We’ll publish several pieces throughout the year as the AI boom continues, highlighting key work that often goes unrecognized. Read more profiles here.
Eva Maydell is a Bulgarian politician and a member of European Parliament. First elected to Parliament in 2014 at age 28, she was the youngest member serving at the time. In 2019, Maydell was re-elected to Parliament, where she continues to serve on the Committee on Economic and Monetary Affairs and on the Committee on Industry, Research and Energy (ITRE).
Maydell was the ITRE rapporteur for the EU AI Act, the proposed legal framework to govern the sale and use of AI in the European Union, and as such was in charge of drafting a report on the proposal of the European Commission — reflecting the opinion of ITRE members. Maydell — in consultation with outside experts and stakeholders — was also responsible for drafting compromise amendments.
Eva Maydell, member of European Parliament
Briefly, how did you get your start in AI? What attracted you to the field?
When I first became a member of the European Parliament, I was one of the few young female members of European Parliament (MEPs) that worked on tech issues. I’ve always been passionate about how Europe can better leverage the huge opportunities of tech innovation. The great thing about working on tech is that you’re always looking to the future. Having worked on cybersecurity, semiconductors and the digital agenda throughout my time in the Parliament, I knew I would find working on the AI Act incredibly interesting and be able to utilise my experience in those areas on this world first piece of regulation.
What work are you most proud of (in the AI field)?
I’m proud of the work we’ve done on the AI Act. We have laid out a common European vision for the future of this technology — one in which AI is more democratic, safe and innovative. Regulators and Parliaments naturally think about how to protect and prepare for worst-case scenarios and the risks; but I also pushed hard for competitiveness to be at the heart of this conversation. This included championing a research and open source exemption, an ambitious approach to regulatory sandboxes and aligning our work with our international partners as much as possible to reduce market frictions.
How do you navigate the challenges of the male-dominated tech industry, and, by extension, the male-dominated AI industry?
We’re slowly but surely seeing more women in tech and AI. I have female colleagues and friends that work in tech who are incredibly talented and really driving the tech agenda. It’s great that we have that network to support each other. I have also found that I have been embraced by the AI community and it’s what makes working on this issue so interesting and enjoyable.
What advice would you give to women seeking to enter the AI field?
Just go for it! Be yourself, don’t think you have to stick to the mould or be like other people. Everyone has something unique to offer. The more women keep sharing their ideas, visions and voice, the more they will inspire other women to step into the world of tech. Whenever I speak with student groups, or young MEPs, it’s wonderful to see so many women interested in entering this field — you can feel the change taking place.
What are some of the most pressing issues facing AI as it evolves?
The greatest challenge for any politician working on tech and AI is trying to regulate and prepare for the future with accuracy. Despite all the facts, figures and research, there’s a certain element of looking into a “crystal ball.” The big issues politicians will need to address are:
Firstly, how can this technology make our economies more competitive while ensuring wider social benefit? Secondly, how do we stop AI fuelling disinformation? And thirdly, how do we set international rules to ensure AI is developed and utilized according to democratic standards?
What are some issues AI users should be aware of?
The very serious challenge posed by AI as a vehicle to accelerate the spread of disinformation and deepfakes. This is particularly important this year, given 50% of the world will go to the polls to vote. We all need to use a critical eye on the images, videos and news articles we see. As the technology improves, we need to become more vigilant to being manipulated. This is an issue I’m working on extensively right now.
What is the best way to responsibly build AI?
If we want a future in which AI improves our lives and helps solve our most pressing challenges, then there’s one key ingredient: trust. We need trust in these technologies.
We can’t afford to rest on our laurels. The AI Act doesn’t mean we’re “one and done.” We need to keep asking ourselves what’s next — and that doesn’t necessarily mean more regulation. But it does mean keeping a constant eye on the big picture — how AI and the regulation is affecting our economy, security and lives.
How can investors better push for responsible AI?
Investing in AI or any innovative technology is no different to investing in any other product. Business, banks and corporations are aware of the fact that there are significant financial merits on being a positive force in the world around us. Ultimately, scaling AI in a responsible way is more likely to sustain success, reduce financial risks and failures, and therefore, create consumer and market confidence.
During Season 5 of Saturday Night Live, Steve Martin and Bill Murray performed a skit in which they comically pointed at something in the distance and asked, “What the hell is that…?” While the humor in the skit may have been somewhat obscure, their confusion accurately mirrors today’s confusion with the concept of “transparency.” What the heck is transparency?
Transparency is a subject that comes up frequently when talking about ensuring the responsible and ethical use of AI. But it’s a topic that seems only to get cursory attention from people in power while the dangers of unfettered AI models continue to threaten society.
Cathy O’Neil’s book “Weapons of Math Destruction” is a brilliant and eye-opening book that exposes the dangers of relying on algorithms to make crucial decisions that shape our lives and society. The book reveals the many challenges and pitfalls of algorithms, such as:
Algorithms are often black boxes, concealed from the scrutiny and understanding of the people they affect and those who implement them.
Algorithms can use flawed, incomplete, and biased data to generate unfair and harmful outcomes.
Algorithms can have far-reaching and pervasive effects on millions of people across various domains and sectors.
Algorithms can undermine people’s opportunities, rights, and well-being.
Algorithms can diminish trust and accountability and marginalize people.
Algorithms can destabilize social and economic systems.
Examples of areas where analytic algorithms or AI models are already making decisions that can adversely and unknowingly affect people are:
Employment: Algorithms can discriminate against applicants based on race, gender, age, or other factors. For instance, Amazon scrapped a hiring tool that favored men over women.
Health Care: Algorithms can deny or limit access to quality health care for certain groups of people. For example, a study found that a widely used algorithm was biased against black patients.
Housing: Algorithms can exclude or exploit potential tenants or buyers based on income, credit score, or neighborhood. For instance, Facebook was sued for allowing landlords and realtors to discriminate against minorities.
College Admissions: Algorithms can favor or disadvantage applicants based on their academic performance, extracurricular activities, or personal background. For example, a lawsuit alleged that Harvard discriminated against Asian-American applicants.
Lending & Financing: Algorithms can charge or refuse loans or credit cards to people based on their financial history, spending habits, or social network. For example, Apple was accused of giving women lower credit limits than men.
Criminal Justice: Algorithms can predict or influence the likelihood of recidivism, bail, sentencing, or parole for defendants or offenders. For example, a report found that a risk assessment tool was biased against black and Hispanic people.
Education: Algorithms can assess or affect the learning outcomes, feedback, or recommendations for students or teachers. For example, an algorithm that graded English exams was found to be unreliable and inconsistent.
Social Media: Algorithms can shape or manipulate users’ or influencers’ content, opinions, or behaviors. For example, a documentary revealed how social media algorithms can polarize and radicalize people.
So, what inalienable rights are mandatory in the 21st century, where more and more of the decisions and actions that impact us and society are being driven by AI models?
Five Inalienable Rights of AI and Data Transparency
Transparency is essential if organizations want to create AI models people can trust and will comfortably and confidently embrace. Understanding how the AI models reason and make decisions can determine whether people embrace them and have moral implications. Unfortunately, explanations of transparency can be too complicated and focus too much on technical details about data and AI technologies.
Consequently, here is my attempt to simplify the concept of transparency for everyone to understand:
Transparency is the ability to see and understand the rationale behind a decision or action
As a Citizen of Data Science, that means:
The Right to AI Awareness: Citizens have the right to know when decisions or actions that impact them are made with the help of AI. This transparency is essential if citizens are to develop trust and confidence in the results of the AI models.
The Right to Understand AI Decision Drivers: Citizens have the right to know the variables and metrics that determine the AI model’s decisions and actions. This helps individuals understand the reasoning behind actions affecting them, evaluate the fairness and relevance of criteria used, and demystify the decision-making process.
The Right to Data Integrity: Citizens have the right to understand the credibility, trustworthiness, and reliability of the data used in the AI model’s decision-making determination. This point seeks to address the potential for misinformation or bias to influence decisions, emphasizing the need for integrity and vigilance in guarding against the use of unbiased, engineered, or altered data.
The Right to Access and Correct Your Personal Data: Citizens have the right to access their personal data held by organizations, to request a copy of their data being processed, and, if the data is inaccurate, to request corrections.
The Right to Be Forgotten: Citizens have the right to request the removal of their personal data. Individuals should be able to request the deletion of their personal data if the data is no longer relevant or necessary or falls into specific categories specified by law.
Transparency In Action: Social Media
Social media emerged as a powerful platform for sharing news and ideas and fostering collaboration. Its potential for education and growth seemed limitless. However, the landscape has shifted. Radical individuals, fringe organizations, and nefarious actors now weaponize social media. Their intent? To disseminate lies, misinformation, and disruptive narratives that undermine societal trust and destabilize governments.
So, what if we applied the five inalienable rights for AI and data transparency to social media platforms? What might we expect?
Transparency Rights
Impact on Social Media Experience
1. The Right to Awareness
Social media platforms should clearly disclose when AI algorithms influence content presentation. Users should be informed when AI-driven decisions impact the content that is displayed to them. This transparency would build trust and helps users understand why certain posts appear or are suppressed.
2. The Right to Understand Decision Drivers
Social media companies should reveal the factors that determine content ranking and recommendations. Users have the right to know why a specific post or ad appears in their feed. Understanding the decision factors allows users to judge fairness and relevance.
3. The Right to Data Integrity
Social media platforms should ensure the credibility of data sources that underpin their content recommendations. They must proactively guard against misinformation, bias, and altered data that could skew viewing recommendations. Users should be confident that the information they see is reliable, unbiased, and accurate.
4. The Right to Access and Correct Their Data
Social media users should have the ability to review their personal data stored by the social media platform. Corrections can be requested if inaccuracies exist. This right empowers users to verify and manage their own data.
5. The Right to Be Forgotten
Individuals should be able to request removal of specific content or personal data from social media. This protects users’ privacy and enables users to control over their digital footprint.
Table 1: Five Inalienable Rights of Transparency Applied to Social Media
Heck, we could even leverage Data Science and AI to create a “BS Meter” (BS does not mean Bill Schmarzo in this case). Social media platforms are inundated with content, making it challenging for users to discern fact from fiction. A “BS Meter” could be a valuable tool to encourage critical thinking and empower users to evaluate information more effectively (Figure 1).
Figure 1: BS Meter
The meter could analyze various factors, including:
Source Credibility: Assess the reliability of the content creator or publisher.
Fact-Checking: Compare claims against verified information.
Consistency: Detect contradictions or inconsistencies.
Emotional Language: Identify sensationalism or bias.
User Feedback: Consider community ratings and reports.
The meter would assign a score or label (e.g., “Highly Reliable,” “Questionable,” or “Debunked”) that would help users make informed decisions about the content they consume, while enabling Social media platforms to enhance their accountability and trust.
AI and Data Transparency Summary
Transparency is not only a technical requirement but also a cultural one. It promotes a spirit of openness and collaboration to create a culture where data-driven insights are easily accessible and understandable. The goal is to foster an environment where decisions are made based on the clear, understandable, and justifiable use of data. This approach builds trust and encourages informed and engaged participation in data-driven initiatives.
Unfortunately, we don’t apply these five inalienable transparency rights to all of society and the policy decisions made by our politicians. Should we expect less from our leaders than we do from our AI models?
You're likely still in some state of recovery after America's great escape.
I refer, of course, to a Super Bowl full of eating, drinking, and other related drama.
Also: AI will have a big impact on jobs this year. Here's why that could be good news
It's worth asking, though, whether you've recovered from discovering just how AI will change your life.
Quite a few companies chose the Super Bowl to air their artificially intelligent wares before the maximum possible, suitably vulnerable audience.
It was moving, however, how some wanted to say those wares were AI-powered and how some declined the opportunity.
AI is good. It's really good
Microsoft attempted to inspire you with the sheer possibilities engendered by having an AI assistant. Your Copilot can do so many things for you — no, not just things you can't be bothered doing, but things that you genuinely believe you can't do at all.
The result? Your dreams, entrepreneurial or otherwise, will come true and you'll become the happiest you imaginable. It was all quite persuasive and placed AI at the very heart of the company's future.
Google, too, showed an exceptionally touching example of AI overtly improving someone's life.
Here, the company revealed how, with the new AI-powered Pixel 8, those with blindness or low vision can now frame photographs beautifully, as the phone's AI assistant guides them to the optimum framing.
The feature is quite clearly labeled as "Guided Frame with Google AI."
In both these cases, the message was similar: "AI isn't to be feared. It's something that truly improves lives."
AI? Did you say AI?
Yet as my eyes became increasingly squinty during the game, I began to notice that not every company wanted to crow overtly about AI and its heady future.
Cybersecurity company Crowdstrike, for example.
Perhaps you still recall — though I wouldn't blame you if you didn't — the Crowdstrike ad that featured an old wild western saloon in a little old wild western town.
Some oddly cyberpunkish characters ride into town — but not on futuristic Trojan horses. These are the so-called adversaries, leaders of the Cyberattack community. A lone young woman is prepared to fight them off. Not with a gun, but with a little bit of Crowdstrike software.
But listen to the sell from the voiceover: "Protecting your business from cyberattacks can be unrelenting. Today's adversaries move fast. Crowdstrike moves faster." As the cyberpunkists are routed — the voiceover continues: "Crowdstrike. We stop breaches."
Notice the missing word? Well, you'll find it if or when you get to Crowdstrike's website: "The world's leading AI-native cybersecurity platform."
It's interesting, then, that Crowdstrike didn't feel moved to insert the AI-native element into its ad.
Also: The 3 biggest risks from generative AI — and how to deal with them
Could it be that some companies are still concerned that the overt mention of AI may be offputting for some? Why, a recent Pew study offered this headline: "Growing public concern about the role of artificial intelligence in daily life."
It's Etsy, not Etsai
And then there was Etsy.
This certainly wasn't the first brand I'd associate with Super Bowl advertising. Or, indeed, any advertising. I rather thought of it as the very nice place to go where real people make real things for other real people to buy.
Also: This is why AI-powered misinformation is the top global risk
Yet here it was in the big game with a curious concoction.
It, too, chose a historical scenario to sell a modern problem. In this case, the American people have just received the Statue of Liberty from France. What can we gift the French in return? Why, they have everything, don't they? Well, everything worth eating, drinking, and enjoying. (Oh, except perhaps baseball.)
A wise American warrior suggests the president could use Etsy's new Gift Mode.
This requires you to prompt the Mode as to what sort of gift you seek.
The president, therefore, wonders what the French actually like, and then the whole thing takes a peculiar turn when a young servant suggests "Cheese."
So it is that the president uses this fine new Etsy tool to purchase a cheeseboard for the French.
You'll likely whisper that a cheeseboard is a fairly mundane offering in return for America's quintessential statue. I, though, will likely whisper that this Gift Mode — apparently — uses AI and human input to make suggestions.
Also: Generative AI in commerce: 5 ways industries are changing how they do business
No mention of AI in the ad, however.
Wouldn't one be a touch more fascinated by what an AI-powered suggestion tool might come up with? But no. Etsy chooses merely to say "Gift Easy with Gift Mode."
Perhaps you'll tell me I'm splitting hairs. Perhaps I'll tell you I'm bald.
But it remains interesting — to me, at least — which brands will use the promise of AI to overtly sell their AI-powered items and which will choose to downplay that little detail, focusing simply on the consumer benefit.
Or will there come a moment when, if a product or service isn't in some way powered by AI, it simply isn't any good?
It’s only been a couple of days since OpenAI released Sora, its new text-to-video model, which generates realistic videos. So far, it is available exclusively to red teamers for identifying potential issues and risks, as well as to artists, designers, and filmmakers for collecting feedback on improvements.“We’re sharing our research progress early to start working with and getting feedback from people outside of OpenAI and to give the public a sense of what AI capabilities are on the horizon,” the blog said.
Yet already, there are videos on social media that’ll make you stop to find out more. The videos, up to a minute long, maintain incredible visual quality and adherence to the user’s prompt.
The model can simulate complex scenes, featuring multiple characters, specific motions, and intricate details of the subject and background. Sora uses the recaptioning technique from DALL·E 3, which involves generating highly descriptive captions for the visual training data.
Here are some of the best examples of what Sora can create:
A bicycle race on the ocean
As soon as the model was released, Sam Altman was taking requests from users on X to create anything they’d like to see. Kunal Shah, the founder of CRED replied with the prompt “A bicycle race on the ocean with different animals as athletes riding the bicycles with a drone camera view”
His request was fulfilled with a video that features a shark, orca, penguin, turtle, and dolphin all racing on bicycles on water, as if by magic!
White SUV on a dirt road
Jessica Lessin, the editor in chief of The Information, has a favourite video from the ones initially previewed by OpenAI. It’s been created with the prompt, “The camera follows behind a white vintage SUV with a black roof rack as it speeds up a steep dirt road surrounded by pine trees on a steep mountain slope, dust kicks up from its tires, the sunlight shines on the SUV as it speeds along the dirt road, casting a warm glow over the scene. The dirt road curves gently into the distance, with no other cars or vehicles in sight. The trees on either side of the road are redwoods, with patches of greenery scattered throughout. The car is seen from the rear following the curve with ease, making it seem as if it is on a rugged drive through the rugged terrain. The dirt road itself is surrounded by steep hills and mountains, with a clear blue sky above with wispy clouds.”
The video follows the prompts to the T.
Pirates clash inside a coffee cup
Jim Fan, the Sr. Research Scientist & Lead of AI Agents at NVIDIA broke down the data-driven physics engine of Sora on X. He used the video that came out of the prompt “Photorealistic closeup video of two pirate ships battling each other as they sail inside a cup of coffee.”
He explained that the model is likely trained on synthetic data, possibly using Unreal Engine 5. The video animates the prompt realistically, simulating fluid dynamics like coffee movement, achieving photorealism, and applying real-world physics to fantastical scenarios.
Golden Retrievers and Snow
Srinivas Mohan, Indian visual effects designer, coordinator and supervisor who notably worked on Tamil, Telugu and Malayalam films like RRR, Bahubali, Enthiran posted on X an adorable close up video of three golden retrievers who after plonking their heads in the snow continued to play in it. This accurately captured the prompt, “A litter of golden retriever puppies playing in the snow. Their heads pop out of the snow, covered in it.”
A garden cat
Ben Nash, a full stack web designer and developer, posted a video on X a video of an orange tabby cat exploring a garden. The detailed prompt read, “ A white and orange tabby cat is seen happily darting through a dense garden, as if chasing something. Its eyes are wide and happy as it jogs forward, scanning the branches, flowers, and leaves as it walks. The path is narrow as it makes its way between all the plants. The scene is captured from a ground-level angle, following the cat closely, giving a low and intimate perspective. The image is cinematic with warm tones and a grainy texture. The scattered daylight between the leaves and plants above creates a warm contrast, accentuating the cat’s orange fur. The shot is clear and sharp, with a shallow depth of field.”
He then went on to compare the same prompt on Pika, Runway, Leonardo, FinalFrame and shared the comparative results. And they didn’t even come close!
Surfing indoors
The prompt reads, “In an ornate, historical hall, a massive tidal wave peaks and begins to crash. Two surfers, seizing the moment, skillfully navigate the face of the wave.” This bizarre request was stunningly captured by the model which was posted by the user Angry Penguin on X. He posted the video along with a few other equally bizarre options.
This Inception-esq video gives Hollywood a run for their money.
Snow and Sakura
Instead of planning a trip to Japan or even looking at pictures taken by travellers, Sora will give you a cinematic view of the scenery with a panoramic shot.
OpenAI’s demo which prompted, “Beautiful, snowy Tokyo city is bustling. The camera moves through the bustling city street, following several people enjoying the weather and shopping at nearby stalls. Gorgeous sakura petals are flying through the wind along with snowflakes.”
A stroll in Mumbai rains
Anu Aakash an AI content creator posted a video on X of a woman with bright purple overalls and cowboy boots taking a casual stroll in the streets of Mumbai during the monsoon season. Unsurprisingly that exactly was the prompt given to Sora. The image depth with a car moving towards the viewer with old Bombay buildings really captures the prompt.
Minecraft with Sora
In 2022 OpenAI trained a neural network to play minecraft. This time Sora is used to create a video of the game. Jim Fan posted the video on X with the title “Minecraft has been achieved internally” The video however accurate has an almost realistic sky which he pointed out saying, “It can’t resist the urge to make the sky look less pixelated ”
Attenborough and Sora
Debarghya Das, the head of product at Glean Assistant combined multiple AI tools to make a stunning video of sea creatures floating in the sky above a town. To this video he added David Attenborough’s voice on Eleven Labs, and sampled some nature music from Youtube on iMovie. The video transitions from a regular town to a magical one ‘through a portal,’ to show a stunning transformation.
He pointed out that, “People aren’t taking the ‘everyone will be filmmakers’ seriously enough”.
The post 10 Mind-Blowing Videos Created by Sora appeared first on Analytics India Magazine.
Argmax released WhisperKit, a software package which enables OpenAI’s Whisper speech recognition model to operate on Apple Watches. This integration leverages Apple’s CoreML framework, making it possible to use Whisper on other Apple devices using the framework.
Whisper running on WatchOS! > Powered by WhisperKit by @argmaxinc > Supports up to Whisper base > Leverages Neural Engine > Three lines of code 😉 > Works real-time! > MIT license Quite amazed by the speed with which Argmax is shipping. Possibly the fastest & reliable… pic.twitter.com/L5utrzIOzv
— Vaibhav (VB) Srivastav (@reach_vb) February 14, 2024
The software facilitates the implementation of speech recognition in WatchOS applications by harnessing the Apple Neural Engine for real-time voice data processing. Released under the MIT license, WhisperKit is designed for systems running macOS 14.0 or later and Xcode 15.0 or later.
Developers can incorporate WhisperKit into their Xcode projects, offering flexibility with audio formats and model selection. Another user went ahead and integrated the Whisperkit with his Vision Pro headset to transcribe his voice.
The fusion of WhisperKit with Apple Vision Pro's capabilities is no small feat. Finally got it running on my headset! pic.twitter.com/Q0d393SquE
— Ray Fernando (@RayFernando1337) February 14, 2024
Founded in 2020, Argmax Inc. specializes in fields like Natural Language Processing, Recommendation Systems, and Computer Vision. WhisperKit stands as an open-source project, encouraging contributions from developers to enhance its functionality and adaptability.
By making WhisperKit available, Argmax aims to broaden the accessibility of speech recognition technology across Apple’s ecosystem, facilitating its integration into applications.
Recent work has advanced the use of large language models on devices such as Apple Watches, aiming for local processing of complex tasks. This approach aims to decrease latency, improve privacy, and enhance user interaction.
Devices like the Rabbit R1 and Humane Ai Pin are notable for integrating AI in distinct ways. The Rabbit R1 uses a model for interfacing directly with user interfaces, while the Humane Ai Pin employs an AI-based operating system for quick service access without needing activation words, focusing on privacy and user initiation. These devices move the use of AI in smaller devices for more tailored and efficient use of the technology.
The post Now You Can Run OpenAI’s Whisper on Apple Watches appeared first on Analytics India Magazine.
TechCrunch highlights notable women in the field of AI
Kyle Wiggers Dominic-Madori Davis 11 hours
To give AI-focused women academics and others their well-deserved — and overdue — time in the spotlight, TechCrunch is launching a series of interviews focusing on remarkable women who’ve contributed to the AI revolution. We’ll publish several pieces throughout the year as the AI boom continues, highlighting key work that often goes unrecognized. Read more profiles here.
As a reader, if you see a name we’ve missed and feel should be on the list, please email us and we’ll seek to add them. Here’s some key people you should know:
Women In AI: Irene Solaiman, head of global policy at Hugging Face
The gender gap in AI
In a New York Times piece late last year, the Gray Lady broke down how the current boom in AI came to be — highlighting many of the usual suspects like Sam Altman, Elon Musk and Larry Page. The journalism went viral – not for what was reported, but instead for what it failed to mention: women.
The Times’ list featured 12 men — most of them leaders of AI or tech companies. Many had no training or education, formal or otherwise, in AI.
Contrary to the Times’ suggestion, the AI craze didn’t start with Musk sitting adjacent to Page at a mansion in the Bay. It began long before that, with academics, regulators, ethicists and hobbyists working tirelessly in relative obscurity to build the foundations for the AI and GenAI systems we have today.
Elaine Rich, a retired computer scientist formerly at the University of Texas at Austin, published one of the first textbooks on AI in 1983, and later went on to become the director of a corporate AI lab in 1988. Harvard professor Cynthia Dwork made waves decades ago in the fields of AI fairness, differential privacy and distributed computing. And Cynthia Breazeal, a roboticist and professor at MIT and the co-founder of Jibo, the robotics startup, worked to develop one of the earliest “social robots,” Kismet, in the late ’90s and early 2000s.
Despite the many ways in which women have advanced AI tech, they make up a tiny sliver of the global AI workforce. According to a 2021 Stanford study, just 16% of tenure-track faculty focused on AI are women. In a separate study released the same year by the World Economic Forum, the co-authors find that women only hold 26% of analytics-related and AI positions.
In worse news, the gender gap in AI is widening — not narrowing.
Nesta, the U.K.’s innovation agency for social good, conducted a 2019 analysis that concluded that the proportion of AI academic papers co-authored by at least one woman hadn’t improved since the 1990s. As of 2019, just 13.8% of the AI research papers on Arxiv.org, a repository for preprint scientific papers, were authored or co-authored by women, with the numbers steadily decreasing over the preceding decade.
Reasons for disparity
The reasons for the disparity are many. But a Deloitte survey of women in AI highlights a few of the more prominent (and obvious) ones, including judgment from male peers and discrimination as a result of not fitting into established male-dominated molds in AI.
It starts in college: 78% of women responding to the Deloitte survey said they didn’t have a chance to intern in AI or machine learning while they were undergraduates. Over half (58%) said they ended up leaving at least one employer because of how men and women were treated differently, while 73% considered leaving the tech industry altogether due to unequal pay and an inability to advance in their careers.
The lack of women is hurting the AI field.
Nesta’s analysis found that women are more likely than men to consider societal, ethical and political implications in their work on AI — which isn’t surprising considering women live in a world where they’re belittled on the basis of their gender, products in the market have been designed for men, and women with children are often expected to balance work with their role as primary caregivers.
With any luck, TechCrunch’s humble contribution — a series on accomplished women in AI — will help move the needle in the right direction. But there’s clearly a lot of work to be done.
The women we profile share many suggestions for those who wish to grow and evolve the AI field for the better. But a common thread runs throughout: strong mentorship, commitment and leading by example. Organizations can affect change by enacting policies — hiring, education or otherwise — that elevate women already in, or looking to break into, the AI industry. And decision-makers in positions of power can wield that power to push for more diverse, supportive workplaces for women.
Change won’t happen overnight. But every revolution begins with a small step.
Google’s Gemini 1.5 has raised questions about the authenticity of a video generated by OpenAI’s Sora, tagging it as fake and pointing out significant inconsistencies.
The critique comes after both tech giants, Google and OpenAI, unveiled their latest advancements – Gemini 1.5 Pro and Sora, respectively. The strategic timing of OpenAI’s Sora release has sparked speculations about a deliberate move to divert attention from Google’s Gemini 1.5.
@Google's Gemini 1.5 Pro watches and critiques a video made by @OpenAI SORA. pic.twitter.com/ZPBHHYUh3K
— Gabor Cselle (@gabor) February 17, 2024
In a counter move, Google took to its platform X to share a detailed analysis critiquing a video created by Sora.
Gemini 1.5 Pro dissected a scene featuring a snowy street in Japan adorned with cherry blossoms. The analysis pointed out several inconsistencies, casting doubt on the video’s authenticity.
According to Gemini 1.5 Pro, the juxtaposition of heavy snowfall and blooming cherry blossoms raised eyebrows, as cherry blossoms typically bloom in the spring, free from snow.
Further scrutiny revealed a uniform and unnatural pattern of snowfall, diverging from the irregularity seen in real-life scenarios. Additionally, despite the intense snowfall, the video characters were not wearing any winter clothing.
Gemini 1.5 concluded the analysis by saying , “Overall, the video is visually appealing, but the inconsistencies suggest that it is not a real-life scene.”
Sora is OpenAI’s all new super-cool, text-to-video tool can create videos of up to 60 seconds featuring highly detailed scenes, complex camera motion, and multiple characters with vibrant emotions. Many are also calling this a ChatGPT moment in video generation.
Google’s Gemini 1.5 features a staggering context window of 1M tokens, surpassing not only GPT-4 Turbo’s 128K but also Anthropic Claude 2.1’s 200K and it can process vast amounts of information in one go — including 1 hour of video, 11 hours of audio, and codebases with over 30,000 lines of code or over 700,000 words.
The post Google Gemini 1.5 Pro Calls OpenAI Sora Generated Video Fake appeared first on Analytics India Magazine.
Women In AI: Irene Solaiman, head of global policy at Hugging Face Kyle Wiggers 9 hours
To give AI-focused women academics and others their well-deserved — and overdue — time in the spotlight, TechCrunch is launching a series of interviews focusing on remarkable women who’ve contributed to the AI revolution. We’ll publish several pieces throughout the year as the AI boom continues, highlighting key work that often goes unrecognized. Read more profiles here.
Irene Solaiman began her career in AI as a researcher and public policy manager at OpenAI, where she led a new approach to the release of GPT-2, a predecessor to ChatGPT. After serving as an AI policy manager at Zillow for nearly a year, she joined Hugging Face as the head of global policy. Her responsibilities there range from building and leading company AI policy globally to conducting socio-technical research.
Solaiman also advises the Institute of Electrical and Electronics Engineers (IEEE), the professional association for electronics engineering, on AI issues, and is a recognized AI expert at the intergovernmental Organization for Economic Co-operation and Development (OECD).
Irene Solaiman, head of global policy at Hugging Face
Briefly, how did you get your start in AI? What attracted you to the field?
A thoroughly nonlinear career path is commonplace in AI. My budding interest started in the same way many teenagers with awkward social skills find their passions: through sci-fi media. I originally studied human rights policy and then took computer science courses, as I viewed AI as a means of working on human rights and building a better future. Being able to do technical research and lead policy in a field with so many unanswered questions and untaken paths keeps my work exciting.
What work are you most proud of (in the AI field)?
I’m most proud of when my expertise resonates with people across the AI field, especially my writing on release considerations in the complex landscape of AI system releases and openness. Seeing my paper on an AI Release Gradient frame technical deployment prompt discussions among scientists and used in government reports is affirming — and a good sign I’m working in the right direction! Personally, some of the work I’m most motivated by is on cultural value alignment, which is dedicated to ensuring that systems work best for the cultures in which they’re deployed. With my incredible co-author and now dear friend, Christy Dennison, working on a Process for Adapting Language Models to Society was a whole of heart (and many debugging hours) project that has shaped safety and alignment work today.
How do you navigate the challenges of the male-dominated tech industry, and, by extension, the male-dominated AI industry?
I’ve found, and am still finding, my people — from working with incredible company leadership who care deeply about the same issues that I prioritize to great research co-authors with whom I can start every working session with a mini therapy session. Affinity groups are hugely helpful in building community and sharing tips. Intersectionality is important to highlight here; my communities of Muslim and BIPOC researchers are continually inspiring.
What advice would you give to women seeking to enter the AI field?
Have a support group whose success is your success. In youth terms, I believe this is a “girl’s girl.” The same women and allies I entered this field with are my favorite coffee dates and late-night panicked calls ahead of a deadline. One of the best pieces of career advice I’ve read was from Arvind Narayan on the platform formerly known as Twitter establishing the “Liam Neeson Principle”of not being the smartest of them all, but having a particular set of skills.
What are some of the most pressing issues facing AI as it evolves?
The most pressing issues themselves evolve, so the meta answer is: International coordination for safer systems for all peoples. Peoples who use and are affected by systems, even in the same country, have varying preferences and ideas of what is safest for themselves. And the issues that arise will depend not only on how AI evolves, but on the environment into which they’re deployed; safety priorities and our definitions of capability differ regionally, such as a higher threat of cyberattacks to critical infrastructure in more digitized economies.
What are some issues AI users should be aware of?
Technical solutions rarely, if ever, address risks and harms holistically. While there are steps users can take to increase their AI literacy, it’s important to invest in a multitude of safeguards for risks as they evolve. For example, I’m excited about more research into watermarking as a technical tool, and we also need coordinated policymaker guidance on generated content distribution, especially on social media platforms.
What is the best way to responsibly build AI?
With the peoples affected and constantly re-evaluating our methods for assessing and implementing safety techniques. Both beneficial applications and potential harms constantly evolve and require iterative feedback. The means by which we improve AI safety should be collectively examined as a field. The most popular evaluations for models in 2024 are much more robust than those I was running in 2019. Today, I’m much more bullish about technical evaluations than I am about red-teaming. I find human evaluations extremely high utility, but as more evidence arises of the mental burden and disparate costs of human feedback, I’m increasingly bullish about standardizing evaluations.
How can investors better push for responsible AI?
They already are! I’m glad to see many investors and venture capital companies actively engaging in safety and policy conversations, including via open letters and Congressional testimonies. I’m eager to hear more from investors’ expertise on what stimulates small businesses across sectors, especially as we’re seeing more AI use from fields outside the core tech industries.
This list covers well-known as well as specialized libraries that I use rather frequently. Applications include GenAI, data animations, LLM, synthetic data generation and evaluation, ML optimization, scientific computing, statistics, web crawling, APIs, SQL, and more. I also mention my owns, and issues that I faced with standard libraries. In several instances, for instance sound generation, I did not use any library. In addition, included some functions that I regularly call. Many times, I explain why I had to create my home-made versions.
Synthetic Data
SDV is the most popular library to generate tabular synthetic data. The Fake library is integrated into it. Another one is CTGan. I was disappointed by the results: poor evaluation, poor synthetization. Thus, I created my owns: Genai-Evaluation and NoGAN-Synthesizer. As the name implies, it does not rely on neural networks. Thus, it is much faster, yet delivers better results.
Natural Language Processing
NLTK is well-known. The Stopwords module proved useless in my case, as I need a separate stopword list for each of the specialized LLM components in my multi-LLM architecture (xLLM): some basic words such as “even” cannot be a stopword depending on the sub-LLM. Thus, I need to create lists of words that cannot be stopwords. I experienced similar problems with Autocorrect and its Speller module. Idem with Singularize from the Pattern library.
I generate all n-grams, compute cosine similarity, create embeddings, deal with accented characters with just a few lines of home-made code, although some libraries also deal with that. That said, the above libraries are useful to many. Even for me, with some workarounds, they are useful.
Web Crawling
I am happy with Requests. You can even access password-protected content. Be careful about getting blocked by the websites that you crawl! I haven’t tested BeautifulSoup yet. If it can recursively retrieve specific navigations features (related pages, similar pages, indexes), which vary from website to website, I will definitely try it. I doubt that it will answer all my questions, as I try to reconstruct the underlying taxonomy of each website or repository that I crawl. Still, it could be handy.
Computer Vision
Of course, openCV is the most popular library. But I have been happy with Pillow. I also produce a lot of data videos, featuring a process evolving over time, or a continuous set of training sets or parameters slightly changing from one video frame to the next. The goal is to show 500 charts in a 1-minute video, as it is more compelling than displaying hundreds of them in separate images. Also great for model comparison, where multiple sub-videos run in parallel within a single video: one per model. I accomplish this very easily with the Moviepy library: see data animation below. Soundtracks are just wave files, so no special library is really needed. As for standard visualizations, I like Matplotlib. But Plotly is more sophisticated and better for special needs.
Deep Neural Networks
Like many programmers, my experience is with TensorFlow and Keras. Training can be very slow unless you switch to GPU, hyperparameters are anything but easy to tune, and results vary wildly from one dataset to another. They lead to non-replicable results unless you use a seed for each sub-module relying on random numbers. In addition, DNNs require a lot of pre- and post-processing (transformers, decoders, and so on). One of the issues is that the loss function is not a good proxy to the output quality. Disclaimer: I mostly used GANs in the context of synthetic tabular data generation; this may be the type of data where DNNs do worst. In the end, that’s why I created NoGAN: it solves all these issues and runs much faster.
Statistics and Machine Learning
Statsmodels is a popular library. Scipy, Numpy, and SKlearn also offer many statistical functions, while Seaborn focuses on visualizations. I work a lot with time series and geospatial data, with my own algorithms. For the latter, on occasions I used Pykrige (kriging to compare with my techniques) and Osmnx to add maps to geospatial data.
I created a very generic regression called cloud regression with Lagrange multipliers for regularization. Yet, I played quite a bit with the curve_fit function in Scipy. Indeed, I was about to create my own version until I realized that you could set bounds on the parameters of the target function, in the optimize module. For random forests (classification), I use SKlearn.
Miscellaneous
To create Web APIs, I use Streamlit. I also run automated SQL queries created on the fly with Python code. So far, I did it with Pandas. Of course, I import Numpy in most of my programs. Sometimes, just to call the random module. Specific functions such as quantiles do not sample outside the observation range and are univariate, so I have my own. Likewise for gradient descent (done with neural networks these days): I have my own version, although based on the gradient operator available in Numpy. I take care of stochastic descent on my own, but libraries exist.
To design faster vector search, I needed among other things to perform some interpolated binary search, but eventually settled for my own probabilistic search. On occasions, I played with Re for regular expressions. Finally, I used Gmpy2 and MPmath for scientific computing. It allows you to work with integers with trillions of digits in any base, complex numbers, and special mathematical functions (Bessel, Riemann Zeta, and so on). Especially useful in cryptography and number theory.
To learn more about my home-made functions, contact me. All are open-source, free, and well documented.
Author
Vincent Granville is a pioneering GenAI scientist and machine learning expert, co-founder of Data Science Central (acquired by a publicly traded company in 2020), Chief AI Scientist at MLTechniques.com and GenAItechLab.com, former VC-funded executive, author and patent owner — one related to LLM. Vincent’s past corporate experience includes Visa, Wells Fargo, eBay, NBC, Microsoft, and CNET. Follow Vincent on LinkedIn.