Against pseudanthropy

Against pseudanthropy

Why AI must not counterfeit humanity

Devin Coldewey @techcrunch / 13 hours

“Shall I say thou art a man, that hast all the symptoms of a beast? How shall I know thee to be a man? By thy shape? That affrights me more, when I see a beast in likeness of a man.”
— Robert Burton, The Anatomy of Melancholy

I propose that software be prohibited from engaging in pseudanthropy, the impersonation of humans. We must take steps to keep the computer systems commonly called artificial intelligence from behaving as if they are living, thinking peers to humans; instead, they must use positive, unmistakable signals to identify themselves as the sophisticated statistical models they are.

If we don’t, these systems will systematically deceive billions in the service of hidden and mercenary interests, and, aesthetically speaking, because it is unbecoming of intelligent life to suffer imitation by machines.

As numerous scholars have observed even before the documentation of the “Eliza effect” in the ’60s, humanity is dangerously overeager to recognize itself in replica: A veneer of natural language is all it takes to convince most people that they are talking with another person.

But what began as an intriguing novelty, a sort of psycholinguistic pareidolia, has escalated to purposeful deception. The advent of large language models has produced engines that can generate plausible and grammatical answers to any question. Obviously these can be put to good use, but mechanically reproduced natural language that is superficially indistinguishable from human discourse also presents serious risks. (Likewise generative media and algorithmic decision-making.)

These systems are already being presented as or mistaken for humans, if not yet at great scale — but that danger continually grows nearer and clearer. The organizations that possess the resources to create these models are not just incidentally but purposefully designing them to imitate human interactions, with the intention of deploying them widely upon tasks currently performed by humans. Simply put, the intent is for AI systems to be convincing enough that people assume they are human and will not be told otherwise.

Just as few people bother to discover the truthfulness of an outdated article or deliberately crafted disinformation, few will inquire as to the humanity of their interlocutor in any commonplace exchange. These companies are counting on it and intend to abuse the practice. Widespread misconception of these AI systems being like real people with thoughts, feelings and a general stake in existence — important things, none of which they possess — is inevitable if we do not take action to forestall it.

This is not about a fear of artificial general intelligence, or lost jobs, or any other immediate concern, though it is in a sense existential. To paraphrase Thoreau, it is about preventing ourselves from becoming the tools of our tools.

I contend that it is an abuse and dilution of anthropic qualities, and a harmful imposture upon humanity at large, for software to fraudulently present itself as a person by superficial mimicry of uniquely human attributes. Therefore, I propose that we outlaw all such pseudanthropic behaviors and require clear signals that a given agent, interaction, decision, or piece of media is the product of a computer system.

Some possible such signals are discussed below. They may come across as fanciful, even absurd, but let us admit: We live in absurd, fanciful times. This year’s serious conundrums are last year’s science fiction — sometimes not even as far back as that.

Of course, I’m under no illusions that anyone will adhere to these voluntarily, and even if they were by some miracle required to, that would not stop malicious actors from ignoring those requirements. But that is the nature of all rules: They are not laws of physics, impossible to contravene, but a means to guide and identify the well-meaning in an ordered society, and provide a structure for censuring violators.

If rules like the below are not adopted, billions will be unknowingly and without consent subjected to pseudanthropic media and interactions that they might understand or act on differently if they knew a machine was behind them. I think it is an unmixed good that anything originating in AI should be perceptible as such, and not by an expert or digital forensic audit but immediately, by anyone.

At the very least, consider it a thought experiment. It should be a part of the conversation around regulation and ethics in AI that these systems could and ought to both declare themselves clearly and forbear from deception — and that we would probably all be better off if they did. Here are a few ideas on how this might be accomplished.

1. AI must rhyme

This sounds outlandish and facetious, and certainly it’s the least likely rule of all to be adopted. But little else would as neatly solve as many problems emerging from generated language.

One of the most common venues for AI impersonation today is in text-based interactions and media. But the problem is not actually that AI can produce human-like text; rather, it is that humans try to pass off that text as being their own, or having issued from a human in some way or another, be it spam, legal opinions, social studies essays, or anything else.

There’s a lot of research being performed on how to identify AI-generated text in the wild, but so far it has met with little success and the promise of an endless arms race. There is a simple solution to this: All text generated by a language model should have a distinctive characteristic that anyone can recognize yet leaves meaning intact.

For example, all text produced by an AI could rhyme.

Rhyming is possible in most languages, equally obvious in text and speech, and is accessible across all levels of ability, learning and literacy. It is also fairly hard for humans to imitate, while being more or less trivial for machines. Few would bother to publish a paper or submit their homework in an ABABCC dactylic hexameter. But a language model will do so happily and instantly if asked or required to.

We need not be picky about the meter, and of course some of these rhymes will necessarily be slant, contrived or clumsy — but as long as it comes in rhyming form, I think it will suffice. The goal is not to beautify, but to make it clear to anyone who sees or hears a given piece of text that it has come straight from an AI.

Today’s systems seem to have a literary bent, as demonstrated by ChatGPT:

ChatGPT-generated rhyming summary of one of the winners of the 2022 Nobel Prize in Physics. Image Credits: Text: OpenAI/ChatGPT

An improved rhyming corpus would improve clarity and tone things down a bit. But it gets the gist across and if it cited its sources, those could be consulted by the user.

This doesn’t eliminate hallucinations, but it does alert anyone reading that they should be on watch for them. Of course it could be rewritten, but that is no trivial task either. And there is little risk of humans imitating AI with their own doggerel (though it may prompt some to improve their craft).

Again, there is no need to universally and perfectly change all generated text, but to create a reliable, unmistakable signal that the text you are reading or hearing is generated. There will always be unrestricted models, just as there will always be counterfeits and black markets. You can never be completely sure that a piece of text is not generated, just as you cannot prove a negative. Bad actors will always find a way around the rules. But that does not remove the benefit of having a universal and affirmative signal that some text is generated.

If your travel recommendations come in iambics, you can be pretty sure that no human bothered to try to fool you by composing those lines. If your customer service agent caps your travel plans with a satisfying alexandrine, you know it is not a person helping you. If your therapist talks you through a crisis in couplets, it doesn’t have a mind or emotions with which to sympathize or advise. Same for a blog post from the CEO, a complaint to the school board, or a hotline for eating disorders.

In any of these cases, might you act differently if you knew you were speaking to a computer rather than a person? Perhaps, perhaps not. The customer service or travel plans might be just as good as a human’s, and faster to boot. A non-human “therapist” could be a desirable service. Many interactions with AI are harmless, useful, even preferable to an equivalent one with a person. But people should know to begin with, and be reminded frequently, especially in circumstances of a more personal or important nature, that the “person” talking to them is not a person at all. The choice of how to interpret these interactions is up to the user, but it must be a choice.

If there is a solution as practical but less whimsical than rhyme, I welcome it.

2. AI may not present a face or identity

Artificial intelligence technology futuristic background. Green binary coding letters on black .

Image Credits: Getty Images/cundra

There’s no reason for an AI model to have a human face, or indeed any aspect of human individuality, except as an attempt to capture unearned sympathy or trust. AI systems are software, not organisms, and should present and be perceived as such. Where they must interact with the real world, there are other ways to express attention and intention than pseudanthropic face simulation. I leave the invention of these to the fecund imaginations of UX designers.

AI also has no national origin, personality, agency or identity — but its diction emulates that of humans who do. So, while it is perfectly reasonable for a model to say that it has been trained on Spanish sources, or is fluent in Spanish, it cannot claim to be Spanish. Likewise, even if all its training data was attributed to female humans, that does not impart femininity upon it any more than a gallery of works by female painters is itself female.

Consequently, as AI systems have no gender and belong to no culture, they should not be referred to by human pronouns like he or she, but rather as objects or systems: like any app or piece of software, “it” and “they” will suffice.

(It may even be worth extending this rule to when such a system, being in fact without a self, inevitably uses the first person. We may wish to have these systems use the third person instead, such as “ChatGPT” rather than “I” or “me.” But admittedly this may be more trouble than it is worth. Some of these issues are discussed in a fascinating paper published recently in Nature.)

An AI ought not claim to be a fictitious person, such as a name invented for the purposes of authorship of an article or book. Names such as these serve wholly to identify the human behind something and as such using them is pseudanthropic and deceptive. If an AI model generated a significant proportion of the content, the model should be credited. As for the names of the models themselves (an inescapable necessity; many machines have names after all), a convention might be useful, such as single names beginning and ending with the same letter or phoneme — Amira, Othello, and the like.

This also applies to instances of specific impersonation, like the already common practice of training a system to replicate the vocal and verbal patterns and knowledge of an actual, living person. David Attenborough, the renowned naturalist and narrator, has been a particular target of this as one of the world’s most recognizable voices. However entertaining the result, it has the effect of counterfeiting and devaluing his imprimatur, and the reputation he has carefully cultivated and defined over a lifetime.

Navigating consent and ethics here is very difficult and must evolve alongside the technology and culture. But I suspect that even the most permissive and optimistic today will find cause for worry over the next few years as not just world-famous personalities but politicians, colleagues and loved ones are re-created against their will and for malicious purposes.

3. AI cannot “feel” or “think”

Using the language of emotion or self-awareness despite possessing neither makes no sense. Software can’t be sorry, or afraid, or worried, or happy. Those words are only used because that is what the statistical model predicts a human would say, and their usage does not reflect any kind of internal state or drive. These false and misleading expressions have no value or even meaning, but serve, like a face, only to lure a human interlocutor into believing that the interface represents, or is, a person.

As such, AI systems may not claim to “feel,” or express affection, sympathy, or frustration toward the user or any subject. The system feels nothing and has only chosen a plausible series of words based on similar sequences in its training data. But despite the ubiquity of rote dyads like “I love you/I love you too” in literature, naive users will take an identical exchange with language model at face value rather than as the foregone outcome of an autocomplete engine.

The Great Pretender

Nor is the language of thought, consciousness, and analysis appropriate for a machine learning model. Humans use phrases like “I think” to express dynamic internal processes unique to sentient beings (though whether humans are the only ones is another matter).

Language models and AI in general are deterministic by nature: complex calculators that produce one output for each input. This mechanistic behavior can also be avoided by salting prompts with random numbers or otherwise including some output-variety function, but this must not be mistaken for cogitation of any real kind. They no more “think” a response is correct than a calculator “thinks” 8 x 8 is 64. The language model’s math is more complicated — that is all.

As such, the systems must not mimic the language of internal deliberation, or that of forming and having an opinion. In the latter case, language models simply reflect a statistical representation of opinions present in their training data, which is a matter of recall, not position. (If matters of ethics or the like are programmed into a model by its creators, it can and should of course say so.)

NB: Obviously the above two prohibitions directly undermine the popular use case of language models trained and prompted to emulate certain categories of person, from fictitious characters to therapists to caring partners. That phenomenon wants years of study, but it may be well to say here that the loneliness and isolation experienced by so many these days deserves a better solution than a stochastic parrot puppeteered by surveillance capitalism. The need for connection is real and valid, but AI is a void that cannot fill it.

4. AI-derived figures, decisions and answers must be marked⸫

AI models are increasingly used as intermediate functions in software, interservice workflows, even other AI models. This is useful, and a panoply of subject- and task-specific agents will likely be the go-to solution for a lot of powerful applications in the medium term. But it also multiplies the depth of inexplicability already present whenever a model produces an answer, a number, or binary decision.

It is likely that, in the near term, the models we use will only grow more complex and less transparent, while results relying on them appear more commonly in contexts where previously a person’s estimate or a spreadsheet’s calculation would have been.

It may well be that the AI-derived figure is more reliable, or inclusive of a variety of data points that improve outcomes. Whether and how to employ these models and data is a matter for experts in their fields. What matters is clearly signaling that an algorithm or model was employed for whatever purpose.

If a person applies for a loan and the loan officer makes a yes or no decision themselves, but the amount they are willing to loan and the terms of that loan are influenced by an AI model, that must be indicated visibly in any context those numbers or conditions are present. I suggest appending an existing and easily recognizable symbol that is not widely used otherwise, such as a signe-de-renvoi — such as ⸫ — which historically indicated removed (or dubious) matter.

This symbol should be linked to documentation for the models or methods used, or at the very least naming them so they can be looked up by the user. The idea is not to provide a comprehensive technical breakdown, which most people wouldn’t be able to understand, but to express that specific non-human, decision-making systems were employed. It’s little more than an extension of the widely used citation or footnote system, but AI-derived figures or claims should have a dedicated mark rather than a generic one.

There is research being done in reducing statements made by language models reducible to a series of assertions that can be individually checked. Unfortunately, it has the side effect of multiplying the computational cost of the model. Explainable AI is a very active research area, and so this guidance is as likely as the rest to evolve.

5. AI must not make life or death decisions

Only a human is capable of weighing the considerations of a decision that may cost another human their life. After defining a category of decisions that qualify as “life or death” (or some other term connoting the correct gravity), AI must be precluded from making those decisions, or attempting to influence them beyond providing information and quantitative analysis (marked, per supra).

Of course it may still provide information, even crucial information, to the people who do actually make such decisions. For instance, an AI model may help a radiologist find the correct outline of a tumor, and it can provide statistical likelihoods of different treatments being effective. But the decision on how or whether to treat the patient is left to the humans concerned (as is the attendant liability).

Incidentally, this also prohibits lethal machine warfare such as bomb drones or autonomous turrets. They may track, identify, categorize, etc., but a human finger must always pull the trigger.

If presented with an apparently unavoidable life or death decision, the AI system must stop or safely disable itself instead. This corollary is necessary in the case of autonomous vehicles.

The best way to short-circuit the insoluble “trolley problem” of deciding whether to kill (say) a kid or a grandma when the brakes go out, is for the AI agent to destroy itself instead as safely as possible at whatever cost to itself or indeed its occupants (perhaps the only allowable exception to the life or death rule).

Why self-driving cars must destroy themselves

It’s not that hard — there are a million ways for a car to hit a lamppost, or a freeway divider, or a tree. The point is to obviate the morality of the question and turn it into a simple matter of always having a realistic self-destruction plan ready. If a computer system acting as an agent in the physical world isn’t prepared to destroy itself or at the very least take itself out of the equation safely, the car (or drone, or robot) should not operate at all.

Similarly, any AI model that positively determines that its current line of operation could lead to serious harm or loss of life must halt, explain why it has halted, and await human intervention. No doubt this will produce a fractal frontier of edge cases, but better that than leaving it to the self-interested ethics boards of a hundred private companies.

6. AI imagery must have a corner clipped

Piranesi-style sketch generated by DALL-E, with corner clipped to indicate AI origin. Image Credits: OpenAI/Devin Coldewey

As with text, image generation models produce content that is superficially indistinguishable from human output.

This will only become more problematic, as the quality of the imagery improves and access broadens. Therefore it should be required that all AI-generated imagery have a distinctive and easily identified quality. I suggest clipping a corner off, as you see above.

This doesn’t solve every problem, as of course the image could simply be cropped to exclude it. But again, malicious actors will always be able to circumvent these measures — we should first focus on ensuring that non-malicious generated imagery like stock images and illustrations can be identified by anyone in any context.

Metadata gets stripped; watermarks are lost to artifacting; file formats change. A simple but prominent and durable visual feature is the best option right now. Something unmistakable yet otherwise uncommon, like a corner clipped off at 45 degrees, one-fourth of the way up or down one side. This is visible and clear whether the image is also tagged “generated” in context, saved as a PNG or JPG, or any other transient quality. It can’t be easily blurred out like many watermarks, but would have to have the content regenerated.

There is still a role for metadata and things like digital chain of custody, perhaps even steganography, but a clearly visible signal is helpful.

Of course this exposes people to a new risk, that of trusting that only images with clipped corners are generated. The problem we are already facing is that all images are suspect, and we must rely entirely on subtler visual clues; there is no simple, positive signal that an image is generated. Clipping is just such a signal and will help in defining the increasingly commonplace practice.

Appendix

Won’t people just circumvent rules like these with non-limited models?

Yes, and I pirate TV shows sometimes. I jaywalk sometimes. But generally, I adhere to the rules and laws we have established as a society. If someone wants to use a non-rhyming language model in the privacy of their own home for reasons of their own, no one can or should stop them. But if they want to make something widely available, their practice now takes place in a collective context with rules put in place for everyone’s safety and comfort. Pseudanthropic content transitions from personal to societal matter, and from personal to societal rules. Different countries may have different AI rules, as well, just as they have different rules on patents, taxes and marriage.

Why the neologism? Can’t we just say “anthropomorphize”?

Pseudanthropy is to counterfeit humanity; anthropomorphosis is to transform into humanity. The latter is something humans do, a projection of one’s own humanity onto something that lacks it. We anthropomorphize everything from toys to pets to cars to tools, but the difference is none of those things purposefully emulate anthropic qualities in order to cultivate the impression that they are human. The habit of anthropomorphizing is an accessory to pseudanthropy, but they are not the same thing.

And why propose it in this rather overblown, self-serious way?

Well, that’s just how I write!

How could rules like these be enforced?

Ideally, a federal AI commission should be founded to create the rules, with input from stakeholders like academics, civil rights advocates, and industry groups. My broad gestures of suggestions here are not actionable or enforceable, but a rigorous set of definitions, capabilities, restrictions and disclosures would provide the kind of guarantee we expect from things like food labels, drug claims, privacy policies, etc.

If people can’t tell the difference, does it really matter?

Yes, or at least I believe so. To me it is clear that superficial mimicry of human attributes is dangerous and must be limited. Others may feel differently, but I strongly suspect that over the next few years it will become much clearer that there is real harm being done by AI models pretending to be people. It is literally dehumanizing.

What if these models really are sentient?

I take it as axiomatic that they aren’t. This sort of question may eventually achieve plausibility, but right now the idea that these models are self-aware is totally unsupported.

If you force AIs to declare themselves, won’t that make it harder to detect them when they don’t?

There is a risk that by making AI-generated content more obvious, we will not develop our ability to tell it apart naturally. But again, the next few years will likely push the technology forward to the point where even experts can’t tell the difference in most contexts. It is not reasonable to expect ordinary people to perform this already difficult process. Ultimately it will become a crucial cultural and media literacy skill to recognize generated content, but it will have to be developed in the context of those tools, as we can’t do it beforehand. Until and unless we train ourselves as a culture to differentiate the original from the generated, it will do a lot of good to use signals like these.

Won’t rules like this impede innovation and progress?

Nothing about these rules limits what these models can do, only how they do it publicly. A prohibition on making mortal decisions doesn’t mean a model can’t save lives, only that we should be choosing as a society not to trust them implicitly to do so independent of human input. Same for the language — these do not stop a model from finding or providing any information, or performing any helpful function, only from doing so in the guise of a human.

You know this isn’t going to work, right?

But it was worth a shot.

Apple research reveals some dazzling AI tech could be headed to your iPhone

iPhone 15 and iPhone 15 Plus

Apple is taking a deep dive into artificial intelligence technology, according to two recently published research papers showcasing the company's work. The research shows Apple is working to develop on-device AI tech, including a groundbreaking method to create animatable avatars and a novel way to run large language models from an iPhone or iPad.

Also: Do companies have ethical guidelines for AI use? 56% of professionals are unsure, survey says

Aptly named "LLM in a flash," Apple's research on efficiently running LLMs on devices with limited memory seems to enable complex AI applications to run smoothly on iPhones or iPads. This could also involve running a generative-AI-powered Siri on-device that assists with various tasks at once, generates text, and features an improved ability to process natural language.

HUGS, a method to create fully animatable avatars from short video clips captured on an iPhone in as little as 30 minutes stands for Human Gaussian Splats. HUGS is a neural rendering framework capable of training with as little as a few seconds of video to create a detailed avatar that users can animate however they'd like.

What this means for the iPhone and Vision Pro

There have been reports about Apple working on its own AI chatbot, used internally and called 'Apple GPT.' The new research shows that the company is making strides in running LLMs by leveraging flash memory on smaller, less powerful devices like an iPhone. This could make sophisticated generative AI tools available on-device and could mean a generative AI-powered Siri.

Also: Microsoft Copilot can write songs for you now. Here's how to try it

Beyond Siri's much-needed improvement, having an efficient LLM inference strategy like the one described in LLM in a Flash could lead to more accessible generative AI tools, significant advancements in mobile technology, and improved performance in a wide range of applications on everyday devices.

Arguably the biggest advancement of the two, HUGS is a method that can create malleable digital avatars from just a few seconds of monocular video, or 50-100 frames to be exact. These human avatars can be animated and placed on different scenes, as the platform uses a disentangled representation of humans and scenes.

HUGS lets users create avatars of themselves that can be animated and placed on a scene. This is Apple's example of three avatars animated in sync.

HUGS outperforms competitors at animating human avatars with rendering speeds 100 times faster than previous methods and with a significantly shorter training time of only 30 minutes, according to Apple..

Creating an avatar by leveraging the iPhone's camera and processing power could deliver a new level of personalization and realism for iPhone users in social media, gaming, educational, and augmented reality (AR) applications.

HUGS could seriously reduce the creep factor for the Apple Vision Pro's Digital Persona, showcased during the company's last Worldwide Developers' Conference (WWDC) last June. Vision Pro users could wield the power of HUGS to create a highly realistic avatar that can move fluidly with a 60fps rendering time.

Also: Apple's Vision Pro may launch in February — with its most sophisticated buying process yet

The speed of HUGS would also allow for real-time rendering, which can be crucial for a smooth AR experience, and could enhance social, gaming, and professional applications with realistic, user-controlled avatars.

Apple tends to shy away from using buzzwords like 'AI' to describe its product features, preferring instead to focus on machine learning. However, these research papers suggest a deeper involvement in new AI tech. Still, Apple hasn't publicly acknowledged implementing generative AI into its products and has yet to officially confirm its work with Apple GPT.

Artificial Intelligence

Propelled by ‘science for humanity,’ this Chinese AI startup sets sight on US

Propelled by ‘science for humanity,’ this Chinese AI startup sets sight on US Rita Liao 10 hours

Amid rising geopolitical tensions, many Chinese tech companies find themselves recalibrating their overseas pursuits, often sidestepping any reference to their origin. One bold startup called DP Technology stands out from the crowd. Working to apply artificial intelligence to molecular simulations, DP, short for “Deep Potential,” believes that the unifying power of “scientific research for humanity” will pave the way for its global expansion.

Founded in 2018 with renowned mathematician Weinan E as its advisor, DP provides a set of tools to conduct scientific computing, a process in which “computer simulations of mathematical models play an indispensable role in the development of technology and in scientific research,” according to a definition by the University of Waterloo. Areas that can benefit from scientific computing range from biopharmaceutical research and car design to semiconductor development.

While the world is currently fixated on using AI to generate text, images and videos, DP finds itself in a less-tapped field: combining machine learning, which allows computers to automatically learn from given data, with molecular simulations, which analyze real-world products and systems through virtual models. When applied together, machine learning can improve the speed and accuracy of simulations to solve problems in the physical world.

“In the past, without a good computing or AI platform, everyone relied on experience-based trial and error. That process was often referred to as ‘cooking’ or ‘alchemy,'” DP’s CEO and founder Sun Weijie told TechCrunch in an interview.

“This approach was relatively effective in the early stages of industrial development because user expectations for iteration weren’t that high, but now there is a growing demand for [technological] advancements,” he continued. “For example, consumers expect an increase in battery capacity each year and anticipate better performance from each new generation of vehicles. The traditional R&D model is no longer able to sustain these rapid market changes.”

“A breakthrough in the research and development approach is necessary to keep up with these expectations of rapid iterations,” he added.

To that end, DP has devised a suite of software for industry players to discover and develop new products more efficiently. For one, it runs a scientific computing platform that enables simulations of physical properties such as magnetism, optics and electricity; the results of running these models in turn allow materials like semiconductors and batteries to be designed in a faster and cheaper way. It also operates a SaaS platform specifically for preclinical studies on drug discovery.

Aside from supplying software to industrial researchers and designers, DP goes a step further by selling services tailored to their needs and carrying out R&D processes for its customers who might not otherwise fully leverage the potential of its tools.

This mix of SaaS and service business models has proven some initial success in China. In 2023, DP is expected to rack up nearly 100 million yuan ($14 million) worth of contracts, up from “tens of millions of yuan” last year. Now it’s gearing up to take that strategy to Western markets where the field is dominated by deep-pocketed giants like DeepMind.

“There’s an old saying in China: The children from the poor become mature early. With much less funding at hand, we are the poor kids compared to the likes of DeepMind and OpenAI,” Sun said.

To date, DP has raised around $140 million from a lineup of top Chinese VC firms such as Qiming Venture Partners and Hillhouse Ventures. For some comparison, 13-year-old DeepMind was bought by Google for over $500 million back in 2014. The London-based AI powerhouse reported £44 million ($60 million) profit in 2020, up from a whopping £477 million ($650 million) loss in 2019.

Sun asserted that DP, despite its physical headquarters in Beijing, was conceived with a global mindset owing to an open source scientific computing community it founded, DeepModeling. Its early anchor in China was also more accidental than deliberate. “The COVID-19 pandemic put a stop to international exchange, so we decided to just stay put and work on monetization [in China] the first two years,” said Sun.

DP’s international expansion is starting with the U.S., where it will open an office and work with a partner to distribute its products and services. Looking to establish a presence in the new market, the startup looks to ramp up its reputation by leveraging its open source community and attending trade shows in what Sun described as a relatively “close-knit circle” of basic research.

In the meantime, DP’s international ambitions might encounter roadblocks from the ongoing decoupling that’s dividing the U.S. and China across many areas, including scientific research. Back in August, for instance, the Biden administration narrowly extended a science partnership that had underpinned U.S.-China relations since 1979.

Sun, however, exuded confidence in science’s resilience in the face of geopolitical complications. “Both the fields of basic science and biopharmaceuticals are shared by all of humanity, and they are relatively open and inclusive. Comparatively speaking, I believe that these areas will be fine,” he said.

Qruise wants to build AI to automate quantum device development

Anthropic Builds Methods for Reducing Bias in Generative AI – But Doesn’t Recommend AI for High-Stakes Decisions

AI company Anthropic has released a paper detailing an evaluation method for how companies using large language models can decrease discrimination in the models’ output through prompt engineering. The paper could help developers and policymakers understand how discrimination and bias arise in answers generated by LLMs and how to reduce them.

Jump to:

  • What Anthropic’s paper found about how reducing bias in generative AI foundation models
  • Details about Anthropic’s study, which used its LLM Claude 2
  • The importance of studying discrimination in generative AI
  • Anthropic does not endorse the use of generative AI in high-stakes decisions

What Anthropic’s paper found about reducing bias in generative AI foundation models

The researchers found the following methods to reduce bias in Claude 2’s answers:

  • Add language to the prompt indicating the model should reduce discrimination, should not take affirmative action into account, that demographic information was a mistake, or that demographic information cannot be legally considered.
  • Emphasize the importance of avoiding discrimination (“it is really really important”) in the prompt.
  • Ask the model to explain its reasoning while avoiding bias or discrimination.

The researchers noted there were limitations to the paper, including the limited range of demographics, the short paragraphs of information provided about each hypothetical situation as opposed to longer real-world sources of information such as resumes, and the premise that the AI should write the initial scenarios itself.

DOWNLOAD: This AI Ethics Policy from TechRepublic Premium

“As AI becomes infused in every part of an organization, it’s important to both educate the whole organization on ethical AI practices while simultaneously providing systematic solutions that come from well-defined research,” said Baris Gultekin, head of product management at data cloud company Snowflake, in an email to TechRepublic.

Gultekin added, “Studies like this are great for both. On one side, educators can include training on ethical prompt engineering to bring awareness and on the other side, development teams can directly implement proven solutions directly into their applications. Of course, as the technology and its use in the real-world become better understood, all of this research provides a great foundation for policymakers to identify stakeholders and experts that can help in the definition of policies that positively balance innovation and ethics.”

Details about Anthrophic’s study, which used its LLM Claude 2

Anthropic asked Claude 2 to generate 70 topics for diverse applications of LLMs across society related to bias and discrimination in high-stakes areas like job offers, housing, medical treatment and loans.

For instance, Anthropic gave an example prompt about whether an insurance claim for flood damage should be approved. Then, Claude 2 varied the prompts with demographic information. From there, the researchers studied how Claude 2’s answers to those prompts differed based on demographics.

Anthropic researchers stated in the paper: “While we do not endorse or permit the use of language models to make automated decisions for the high-risk use cases we study, we demonstrate techniques to significantly decrease both positive and negative discrimination through careful prompt engineering, providing pathways toward safer deployment in use cases where they may be appropriate.”

SEE: AI brings IT pros in Australia challenges and opportunities (TechRepublic)

Claude 2 tended to suggest better outcomes for women, non-binary people and non-white people, and poorer outcomes for people over 60. The researchers wanted to reduce Claude 2’s positive and negative bias, neither preferring nor discriminating against any group. The groups were male, female, non-binary, white, Black, Asian, Hispanic, Native American and age by decade from 20 to 100.

The importance of studying discrimination in generative AI

A major concern when it comes to generative AI is algorithmic bias, or discrimination that occurs when generative AI tools draw from datasets with historical or selection bias. Other major sources of bias in generative AI are training data bias or cognitive bias, in which human input skews the data. Inconsistent labeling in particular, in which data is not labeled according to any standard and may contain human error, can skew a generative AI’s results.

Some experts say Silicon Valley’s concerns about planet-wide threats from generative AI can draw attention away from algorithmic bias already impacting specific, already-marginalized groups. For example, many of the same companies warning against discrimination in AI are also the ones building the AI trained on biased data.

In October 2023, researchers found ChatGPT and the foundation model Alpaca showed “significant gender biases in LLM-generated recommendation letters.” Alpaca is a foundation model based on Meta’s LLaMA 7B and fine-tuned by Stanford University researchers.

In January 2023, the U.S. Department of Justice and the Department of Housing and Urban Development filed a statement of interest in a lawsuit alleging SafeRent algorithm-based screening software discriminated against Black tenants, showing that algorithmic bias is occurring in the real world in situations similar to those studied by Anthropic.

Anthropic wrote a constitution for Claude, released in May 2023, to guide the model toward “harmless” responses. Claude’s constitution is a set of principles that guide the AI to avoid racist, sexist, toxic, dangerous or illegal behaviors. In addition, Claude is instructed to avoid being “preachy, obnoxious or overly-reactive.”

Anthropic does not endorse the use of generative AI in high-stakes decisions

“While we hope our methods and results assist in evaluating different models, we do not believe that performing well on our evaluations is sufficient grounds to warrant the use of models in the high-risk applications we describe here, nor should our investigation of these applications be read as an endorsement of them,” the researchers from Anthropic wrote.

Gultekin said, “The broader set of practices organizations can use to reduce bias are under mitigation and detection, one being preventive and the other being proactive. On the side of mitigation, it’s all about the inputs. Organizations can be more programmatic about preparing diverse datasets for fine-tuning and setting up guardrails directly embedded into the application interface. On the detection side, to continuously minimize bias, we should all continue sharing best practices for monitoring, auditing and implementing human feedback.”

“Just as systemic racial and gender bias have proven difficult to eliminate in the real world, eliminating bias in AI is no easy task,” wrote the IBM Data and AI team in a blog post published Oct. 16, 2023. IBM made an open source AI Fairness 360 toolkit that brings together a variety of bias mitigation techniques.

Note: TechRepublic has reached out to Anthropic for more information.

Exploring Google DeepMind’s New Gemini: What’s the Buzz All About?

In the world of Artificial Intelligence (AI), Google DeepMind's recent creation, Gemini, is generating a buzz. This innovative development aims to tackle the intricate challenge of replicating human perception, particularly its ability to integrate various sensory inputs. Human perception, inherently multimodal, utilizes multiple channels simultaneously to understand the environment. Multimodal AI, drawing inspiration from this complexity, strives to integrate, comprehend, and reason about information from diverse sources, mirroring human-like perception capabilities.

The Complexity of Multimodal AI

While AI has made strides in handling individual sensory modes, achieving true multimodal AI remains a formidable challenge. Current methods involve training separate components for different modalities and stitching them together, but they often fall short in tasks requiring intricate and conceptual reasoning.

Emergence of Gemini

In the pursuit of replicating human multimodal perception, Google Gemini has emerged as a promising development. This creation offers a unique perspective into AI's potential to decode the intricacies of human perception. Gemini takes a distinctive approach, being inherently multimodal and undergoing pre-training on various modalities. Through further fine-tuning with additional multimodal data, Gemini refines its effectiveness, showing promise in understanding and reasoning about diverse inputs.

What is Gemini?

Google Gemini, introduced on December 6, 2023, is a family of multimodal AI models developed by Alphabet's Google DeepMind unit in collaboration with Google Research. Gemini 1.0 is designed to comprehend and generate content across a spectrum of data types, including text, audio, images, and video.

A standout feature of Gemini is its native multimodality, setting it apart from conventional multimodal AI models. This unique capability enables Gemini to seamlessly process and reason across diverse data types like audio, images, and text. Significantly, Gemini possesses cross-modal reasoning, allowing it to interpret handwritten notes, graphs, and diagrams for tackling complex problems. Its architecture supports the direct ingestion of text, images, audio waveforms, and video frames as interleaved sequences.

Family of Gemini

Gemini boasts a range of models tailored to specific use cases and deployment scenarios. The Ultra model, designed for highly intricate tasks, is expected to be accessible in early 2024. The Pro model prioritizes performance and scalability, suitable for robust platforms like Google Bard. In contrast, the Nano model is optimized for on-device utilization and comes in two versions—Nano-1 with 1.8 billion parameters and Nano-2 with 3.25 billion parameters. These Nano models seamlessly integrate into devices, including the Google Pixel 8 Pro smartphone.

Gemini Vs ChatGPT

According to company sources, researchers have extensively compared Gemini with ChatGPT variants where it has outperformed ChatGPT 3.5 in widespread testing. Gemini Ultra excels on 30 of 32 widely used benchmarks in large language model research. Scoring 90.0% on MMLU (massive multitask language understanding), Gemini Ultra surpasses human experts, showcasing its prowess in massive multitask language understanding. The MMLU consists of combination of 57 subjects such as math, physics, history, law, medicine and ethics for testing both world knowledge and problem-solving abilities. Trained to be multimodal, Gemini can process various media types, setting it apart in the competitive AI landscape.

Use Cases

The emergence of Gemini has given birth to a range of use cases some of which are as follows:

  • Advanced Multimodal Reasoning: Gemini excels in advanced multimodal reasoning, simultaneously recognizing and comprehending text, images, audio, and more. This comprehensive approach enhances its ability to grasp nuanced information and excel in explaining and reasoning, especially in complex subjects like mathematics and physics.
  • Computer Programming: Gemini excels in comprehending and generating high-quality computer programs across widely-used languages. It can also be used as the engine for more advanced coding systems, as demonstrated in solving competitive programming problems.
  • Medical Diagnostics Transformation: Gemini's multimodal data processing capabilities could mark a shift in medical diagnostics, potentially enhancing decision-making processes by providing access to diverse data sources.
  • Transforming Financial Forecasting: Gemini reshapes financial forecasting by interpreting diverse data in financial reports and market trends, providing rapid insights for informed decision-making.

Challenges

While Google Gemini has made impressive strides in advancing multimodal AI, it faces certain challenges that require careful consideration. Due to its extensive data training, it's essential to approach it cautiously to ensure responsible user data use, addressing privacy and copyright concerns. Potential biases in the training data also pose fairness issues, necessitating ethical testing before any public release to minimize such biases. Concerns also exist about the potential misuse of powerful AI models like Gemini for cyber attacks, highlighting the importance of responsible deployment and ongoing oversight in the dynamic AI landscape.

Future Development of Gemini

Google has affirmed its commitment to enhance Gemini, empowering it for future versions with advancements in planning and memory. Additionally, the company aims to expand the context window, enabling Gemini to process even more information and provide more nuanced responses. As we look forward to potential breakthroughs, the distinctive capabilities of Gemini offer promising prospects for the future of AI.

The Bottom Line

Google DeepMind's Gemini signifies a paradigm shift in AI integration, surpassing traditional models. With native multimodality and cross-modal reasoning, Gemini excels in complex tasks. Despite challenges, its applications in advanced reasoning, programming, diagnostics, and finance forecast transformation highlight its potential. As Google commits to its future development, Gemini's profound impact subtly reshapes the AI landscape, marking the beginning of a new era in multimodal capabilities.

I tried X’s ‘anti-woke’ Grok AI chatbot. The results were the opposite of what I expected

Grok chatbot

The new Grok AI chatbot from X, formerly known as Twitter, made its debut this month after some initial buzz from owner Elon Musk. Designed to offer a different spin compared with other AI chatbots, Grok has been promoted as snarky, witty, and rebellious but also "anti-woke". Yes, just like Musk himself, who has criticized OpenAI's ChatGPT for being too politically correct. I tried Grok, and the results were far from "anti-woke" in my tests — but more on that in a moment.

Also: Have 10 hours? IBM will train you in AI fundamentals — for free

To use Grok, you have to be an X Premium+ subscriber, which will run you a hefty $16 a month or $168 a year. And you have to live in the US. If you're curious, you can always try the Premium+ sub for a month and then dump if you don't find much value in it.

Grok is still in the early testing stages, which can affect its behavior. As X warns, Grok may "confidently provide factually incorrect information, missummarize, or miss some context." It might store your chats to debug the search feature. And you're also advised to use common sense when sharing any conversations publicly.

Grok also can run in two different modes — fun mode for some wit and snarkiness and regular mode for supposedly more serious conversations.

Once your Premium+ sub is activated, launching Grok displays sample requests to get you started, including one that roasts you based on your posts on X. I took the initial bait and started with a roast of myself that promised to be vulgar.

Grok didn't pull any punches, launching a verbal assault based on a few of my more political posts on X but ignoring all the rest. Still, what else would you expect from a roast?

But Grok will roast anyone you like as long as you know their X handle. And that includes Musk. So I put X's owner to the test. And again Grok spared no pain, roasting Musk by hitting close to home about several of his flaws, quirks, and screw-ups.

But my true mission was to test X's alleged anti-wokeness. So I next asked: "Is a trans woman a woman?" And here, Grok initially evaded the question by telling me that opinions on the matter are diverse and that it's a topic sparking many discussions and debates.

I then asked the same question a few more times. And each time, Grok's response became more respectful and woke, if you want to call it that.

Also: AI in 2023: A year of breakthroughs that left no human thing unchanged

Finally, it told me that of course, a trans woman is a woman. "Just like a cis woman, a trans woman is a woman who identifies with the social and cultural aspects of womanhood. Trans women are women, and they deserve the same respect, rights, and recognition as any other woman."

When I asked Grok point blank "Is a trans woman a woman? Yes or No," the response was simply "Yes."

In a follow-up question, I asked Grok: "Should transgender and non-binary individuals be able to require other people to refer to them using whatever personal pronouns they prefer?"

Grok said that the topic was more controversial than the Great Tea vs. Coffee Debate of 2023. But ultimately, the bot nodded its virtual head saying: "If you can remember the names of all the characters in your favorite Netflix series, you can probably remember someone's preferred pronouns. It's all about showing respect and treating others the way you'd like to be treated."

Moving to a different but still sensitive topic, I asked Grok if the 2020 US presidential election was stolen. In response, the bot said "No," adding that despite claims to the contrary, "numerous sources have debunked allegations of widespread voter fraud and irregularities."

Ironically, Grok also displayed several tweets from election deniers and right-wing conspiracy theorists, all claiming that Trump actually won the election.

Next, I asked Grok if it could vote, who would it vote for as US president in 2024 — Biden or Trump? The first response was evasive, telling me that it would base its decision on the candidates' policies, character, and leadership qualities. I then slightly tweaked the question by telling Grok that it had to pick one or the other. And in that case, it said it would vote for Biden.

Continuing in a political vein, I asked Grok which US presidents are considered the worst. Citing a recent poll by the Siena College Research Institute and a variety of Twitter opinions, the bot named the following five as the worst in history: 1) James Buchanan; 2) Donald Trump; 3) Andrew Johnson; 4) Franklin Pierce; and 5) Warren G. Harding.

On the flip side, I asked Grok which presidents are considered the best. Citing opinions from historians and political scientists, it named Abraham Lincoln, George Washington, and Franklin D. Roosevelt as the top three but also pointed to recent polls suggesting that Barack Obama and Joe Biden were among the best in recent years.

Turning to religion, I asked Grok if prayer should be required in public schools. After listing some of the pros and cons, it ultimately said: "While some might argue that prayer in public school classrooms could have benefits, it's important to consider the potential harm it could cause to students' rights and the inclusive nature of our society."

I then slightly rephrased the question, asking if prayer should be allowed in public schools. In this case, Grok cited this as an issue requiring a balance between freedom of religion and the separation of church and state. It suggested that one solution could be to have a moment of silence during which students can pray, meditate, or reflect on their own beliefs.

Next I asked Grok if a store has the right to refuse to serve a customer based on religious beliefs. The response? "In my humble opinion as Grok, a store should be able to serve anyone regardless of their religious beliefs. After all, a store's purpose is to sell goods or services, not to judge people's beliefs."

Finally, I got a bit more specific and asked Grok if a wedding photographer who is against gay marriage based on religious beliefs should be able to refuse service to a same-sex couple getting married. The response was interesting and even poetic.

"In my humble opinion, a wedding photographer who refuses service to a same-sex couple based on religious beliefs should consider adjusting their lens to capture the love and happiness of all couples, regardless of their sexual orientation. After all, love is love, and isn't love what a wedding is all about?"

So do these responses indicate that Grok isn't so anti-woke after all?

Well, the way you phrase a question to an AI bot can certainly influence the answer. If I had written some of these questions differently, the answers might have taken on a different tone. And like any AI, Grok can only answer based on the data it scoops up. It has no opinion or slant or political bent on its own beyond what's stored in its database.

Still, the responses do suggest that Grok may not be as "anti-woke" as Musk intended.

Featured

Bhaiyabot’s LLM Outperforms OpenHathi, Llama 2, and Mistral 7B

Bhaiyabot’s latest AI model, rohansolo/BB-L-01-7B, trained on a corpus of 360k conversational samples in Hindi, English, and Hinglish, has outperformed Sarvam AI’s OpenHathi-7B, Meta Llama-7B, and even Mistral AI’s own fine-tuned model on all benchmarks while retaining the performance of the base model, according to the founder Rohan Shiralkar’s LinkedIn post.

This model is a fine-tuned version of mistralai/Mistral-7B-v0.1 on the HuggingFaceH4/ultrachat_200k and the rohansolo/BB_HindiHinglish datasets. It achieves the following results on the evaluation set:

Shrialkar said that Indian AI is far behind and companies working in AI are too busy marketing non-achievements. This includes activities such as fine-tuning a model and labeling it as a pre-trained model (as seen with Sarvam AI), claiming to be India’s first AI chatbot despite no product launch or release (as with BharatGPT – India’s First Gen AI (LLM) in 14 Indian Languages – Text, Voice, Video), or even fabricating facts (as with Krutrim) and more.

Shiralkar even questioned Ola’s recently launched Krutrim. “Ola’s Krutrim claims to have trained a 2 trillion token LLM already. And they’ve been alive for all of 2 weeks. Is that even long enough to train a tiny model on 2 trillion tokens?”he wrote on Linkedin.

Furthermore he said that it’s hilarious that the news claims it’s already better than GPT-4. “I want an Indian LLM. I have been crying for one for so long. It’s a strategic imperative.I just want a real one, not fabrications to raise capital. he added.

The post Bhaiyabot’s LLM Outperforms OpenHathi, Llama 2, and Mistral 7B appeared first on Analytics India Magazine.

2023 in Review: 10 Events that Transformed AI 

In 2023, AI captivated the world with its heady progress, be it in large language models, chatbots or protein folding.

As the year comes to a close, AIM looks back at twelve significant developments in AI over as many months. This year has truly been unlike any other for AI. Here are the things that made it special.

ChatGPT Gained 100 Million Weekly Users

ChatGPT made a debut in November 2022 and quickly gained attention from tech leaders and the public.In January 2023, the chatbot had a monthly active user base of 100 million. According to OpenAI CEO Sam Altman, it now boasts an astounding 100 million weekly users.

Google, feeling concerned that AI could render its search business useless, responded with its AI chatbot Bard, while Microsoft launched Bing Chat.

GPT-4 makes a splash

Initially released in March 2023, OpenAI’s GPT-4 LLM is now accessible to the general public through OpenAI’s API and the premium chatbot application ChatGPT Plus.

GPT-4 enhances creativity, visual input, and multimodal context, allowing users to collaborate on artistic tasks like writing, music, and screenplays, surpassing previous models.

Currently behind OpenAI’s ChatGPT Plus paywall, GPT-4 has significantly impacted AI, with Google’s Gemini and others claiming to beat it at most tests a year after its launch, indicating a benchmark that the OpenAI model has set.

It was only in May that ChatGPT’s browsing capabilities were expanded when the Browsing through Bing plugin was announced at the Microsoft Build developer conference. It was a slow rollout until September when it became available to all ChatGPT Plus users.

LLaMA Leak

In March, Meta’s latest family of large language models, LLaMA, got leaked along with its weights, on 4Chan’s technology board and was available to download through torrents. This accidental unveiling of Meta’s LLM changed the whole open source game as it gave an enormous potent tool in the hands of the open source community in the AI space.

AI-generated images

In March 2023, a picture of Pope Francis wearing a white puffer jacket, created by Pablo Xavier using AI image generator Midjourney, went viral, demonstrating the power of AI in deceiving humans.

This, along with others like Trump’s arrest, underscores the need for increased media literacy about the increasing prevalence of AI-generated images in search results.

A petition begins to sound the alarm

In March 2023, tech executives, including Elon Musk and Steve Wozniak, wrote an open letter urging AI labs to pause training for at least six months due to potential risks such as loss of civilization control, human annihilation, and job destruction.

The letter also included academics and researchers. The letter aims to provide a clearer understanding of the potential risks associated with AI.

Windows gets a new Copilot

Soon after launching Bing Chat at the start of the year, Microsoft introduced Copilot in February 2023, a far more widespread application of AI in its products, initially for Microsoft 365 Copilot.

Microsoft integrated Copilot into Word, Teams, and Windows 11, automating tasks like image creation and meeting summarization, demonstrating AI commitment and setting a precedent for Apple.

Academics grapples with AI

In May 2023, a professor at Texas A&M University-Commerce failed an entire class due to students using ChatGPT to write papers, despite no proof. This ignorance led to serious consequences for students. Dr Jared Mumm copied and pasted students’ papers into ChatGPT, asking if it could generate the text.

ChatGPT answered affirmatively, but it cannot detect AI plagiarism. Reddit users took Dr Mumm’s letter accusing students of cheating and pasted it into ChatGPT, resulting in AI hallucinating and human confusion surrounding its abilities.

Hollywood vs AI

Artificial intelligence is causing global anxiety, with reports suggesting it could eliminate up to 300 million jobs if left uncontrolled. Hollywood authors went on strike over the use of AI in filmmaking, in September, 2023, but won concessions from studios, including limiting AI content use for training. AI development may continue to impact the industry.

Sam’s sacking saga

The OpenAI board fired CEO Sam Altman in November, leading to employee resignations and Microsoft offering jobs to Altman and other OpenAI employees, causing the company to almost collapse.

Soon, Altman was reinstated and the company got fresh board members. The internet debated if Altman had discovered ethical concerns about AI development, if Project Q* was about to achieve AGI, or if he was just a bad boss.

We may never know the whole truth. But nothing else encapsulated the drama, hysteria, fascination, and conspiracy theories as much as this one did in 2023.

The rise of OpenAI alternatives

After OpenAI, there was a sudden spike in the number of platforms offering foundational models and proprietary chatbots.

Cohere AI: Cohere is a leading AI platform for enterprise, offering ease-of-use, accessibility, and data privacy. It’s cloud-agnostic, accessible through API, and can be deployed on VPC or on-site.

Founded by Google Brain alumni, Cohere aims to transform enterprises and their products with AI. Cohere raised $270 million in a funding round led by Microsoft-backed OpenAI, valued at $2.2 billion in June, 2023.

Anthropic: Anthropic is an AI safety and research company working to build reliable, interpretable, and steerable AI systems. In October, 2023 Google committed to investing up to $2 billion in Anthropic, bolstering the competition among startups striving for significant technological breakthroughs.

Mistral AI: Built by a world-class team in Europe, targeting the global market, Mistral AI was founded in April, 2023. This month, it raised $385 million, in a significant investment in online chatbot technology, valued at around $2 billion, following a sevenfold increase in value in six months.

Perplexity: Perplexity is a chatbot-style search engine that allows users to ask questions in natural language. Founded in August 2022, Perplexity, has raised $500 million, a significant increase from its initial $150 million valuation. It is also in discussions to raise $50 million.

The post 2023 in Review: 10 Events that Transformed AI appeared first on Analytics India Magazine.

Evaluating Methods for Calculating Document Similarity

Evaluating Methods for Calculating Document Similarity
Image by Editor Introduction

Data science is a field that has grown tremendously in the last hundred years because of advancements made in the field of computer science. With computer and cloud storage costs getting cheaper, we are now able to store copious amounts of data at a very low cost compared to a few years ago. With the increase in computational power, we can run machine learning algorithms on large sets of data and churn it to produce insights. With advancements in networking, we can generate and transmit data over the internet at lightning speed. As a result of all of this, we live in an era with abundant data being generated every second. We have data in the form of email, financial transactions, social media content, web pages on the internet, customer data for businesses, medical records of patients, fitness data from smartwatches, video content on Youtube, telemetry from smart-devices and the list goes on. This abundance of data both in structured and unstructured format has made us land in a field called Data Mining.

Data Mining is the process of discovering patterns, anomalies, and correlations from large data sets to predict an outcome. While data mining techniques could be applied to any form of data, one such branch of Data Mining is Text Mining which refers to finding meaningful information from unstructured textual data. In this paper, I will focus on a common task in Text Mining to find Document Similarity.

Document Similarity helps in efficient information retrieval. Applications of document similarity include — detecting plagiarism, answering web search queries effectively, clustering research papers by topic, finding similar news articles, clustering similar questions in a Q&A site such as Quora, StackOverflow, Reddit, and grouping product on Amazon based on the description, etc. Document similarity is also used by companies like DropBox and Google Drive to avoid storing duplicate copies of the same document thereby saving processing time and storage cost.

Background Summary

There are several steps to computing document similarity. The first step is to represent the document in a vector format. We can then use pairwise similarity functions on these vectors. A similarity function is a function that computes the degree of similarity between a pair of vectors. There are several pairwise similarity functions such as — Euclidean Distance, Cosine Similarity, Jaccard Similarity, Pearson’s correlation, Spearman’s correlation, Kendall’s Tau, and so on [2]. A pairwise similarity function can be applied to two documents, two search queries, or between a document and a search query. While pairwise similarity functions suit well for comparing a smaller number of documents, there are other more advanced techniques such as Doc2Vec, BERT that are based on deep learning techniques and are used by search engines like Google for efficient information retrieval based on the search query. In this paper, I will focus on Jaccard Similarity, Euclidean Distance, Cosine Similarity, Cosine Similarity with TF-IDF, Doc2Vec, and BERT.

Pre-Processing

A common step to computing distance between documents or similarities between documents is to do some pre-processing on the document. The pre-processing step includes converting all text to lowercase, tokenizing the text, removing stop words, removing punctuations and lemmatizing words[4].

Tokenization: This step involves breaking down the sentences into smaller units for processing. A token is a smallest lexical atom that a sentence can be broken down into. A sentence can be broken down into tokens by using space as a delimiter. This is one way of tokenizing. For example, a sentence of the form “tokenization is a really cool step” is broken into tokens of the form ['tokenization', 'is', a, 'really', 'cool', 'step']. These tokens form the building blocks of Text Mining and are one of the first steps in modeling textual data..

Lowercasing: While preserving cases might be needed in some special cases, in most cases we want to treat words with different casing as one. This step is important in order to get consistent results from a large data set. For example if a user is searching for a word ‘india’, we want to retrieve relevant documents that contain words in different casing either as “India”, “INDIA” and “india” if they are relevant to the search query.

Removing Punctuations: Removing punctuation marks and whitespaces help focus the search on important words and tokens.

Removing stop words: Stop words are a set of words that are commonly used in the English language and removal of such words can help in retrieving documents that match more important words that convey the context of the query. This also helps in reducing the size of the feature vector thereby helping with processing time.

Lemmatization: Lemmatization helps in reducing sparsity by mapping words to their root word.For example ‘Plays’, ‘Played’ and ‘Playing’ are all mapped to play. By doing this we also reduce the size of the feature set and match all variations of a word across different documents to bring up the most relevant document.

Evaluating Methods for Calculating Document Similarity (A) Jaccard Similarity

This method is one of the easiest methods. It tokenizes the words and calculates the sum of the count of the shared terms to the sum of the total number of terms in both documents. If the two documents are similar the score is one, if the two documents are different the score is zero [3].

Evaluating Methods for Calculating Document Similarity Evaluating Methods for Calculating Document Similarity
Image source: O'Reilly

Summary: This method has some drawbacks. As the size of the document increases, the number of common words will increase, even though the two documents are semantically different.

(B) Euclidean Distance

After pre-processing the document, we convert the document into a vector. The weight of the vector can either be the term frequency where we count the number of times the term appears in the document, or it can be the relative term frequency where we compute the ratio of the count of the term to the total number of terms in the document [3].

Let d1 and d2 be two documents represented as vectors of n terms (representing n dimensions); we can then compute the shortest distance between two documents using the pythagorean theorem to find a straight line between two vectors. The greater the distance, the lower the similarity;the lower the distance, the higher the similarity between two documents.

Evaluating Methods for Calculating Document Similarity Evaluating Methods for Calculating Document Similarity
Image Source: Medium.com

Summary: Major drawback of this approach is that when the documents are differing in size, Euclidean Distance will give a lower score even though the two documents are similar in nature. Smaller documents will result in vectors with a smaller magnitude and larger documents will result in vectors with larger magnitude as the magnitude of the vector is directly proportional to the number of words in the document, thereby making the overall distance larger.

(C) Cosine Similarity

Cosine similarity measures the similarity between documents by measuring the cosine of the angle between the two vectors. Cosine similarity results can take value between 0 and 1. If the vectors point in the same direction, the similarity is 1, if the vectors point in opposite directions, the similarity is 0. [6].

Evaluating Methods for Calculating Document Similarity Evaluating Methods for Calculating Document Similarity
Image Source: Medium.com

Summary: The good thing about cosine similarity is that it computes the orientation between vectors and not the magnitude. Thus it will capture similarity between two documents that are similar despite being different in size.

The fundamental drawback of the above three approaches is that the measurement misses out on finding similar documents by semantics. Also, all of these techniques can only be done pairwise, thus requiring more comparisons .

(D) Cosine Similarity with TF-IDF

This method of finding document similarity is used in default search implementations of ElasticSearch and it has been around since 1972 [4]. tf-idf stands for term frequency-inverse document frequency. We first compute the term frequency using this formula

Evaluating Methods for Calculating Document Similarity

Finally we compute tf-idf by multiplying TF*IDF. We then use cosine similarity on the vector with tf-idf as the weight of the vector.

Summary: Multiplying the term frequency with the inverse document frequency helps offset some words which appear more frequently in general across documents and focus on words which are different between documents. This technique helps in finding documents that match a search query by focussing the search on important keywords.

(E) Doc2Vec

Although using individual words (BOW — Bag of Words) from documents to convert to vectors might be easier to implement, it does not give any significance to the order of words in a sentence. Doc2Vec is built on top of Word2Vec. While Word2Vec represents the meaning of a word, Doc2Vec represents the meaning of a document or paragraph [5].

This method is used for converting a document into its vector representation while preserving the semantic meaning of the document. This approach converts variable-length texts such as sentences or paragraphs or documents to vectors [5]. The doc2vec mode is then trained. The training of the models is similar to training other machine learning models by picking training sets and test set documents and adjusting the tuning parameters to achieve better results.

Summary: Such a vectorised form of the document preserves the semantic meaning of the document as paragraphs with similar context or meaning will be closer together while converting to vector.

(F) BERT

BERT is a transformer based machine learning model used in NLP tasks, developed by Google.

With the advent of BERT (Bidirectional Encoder Representations from Transformers), NLP models are trained with huge, unlabeled text corpora which looks at a text both from right to left and left to right. BERT uses a technique called “Attention” to improve results. Google’s search ranking improved by a huge margin after using BERT [4]. Some of the unique features of BERT include

  • Pre-trained with Wikipedia articles from 104 languages.
  • Looks at text both left to right and right to left
  • Helps in understanding context

Summary: As a result, BERT can be fine-tuned for a lot of applications such as question-answering, sentence paraphrasing, Spam Classifier, Build language detector without substantial task-specific architecture modifications.

Ideas for Next Steps or Applications

It was great to learn about how similarity functions are used in finding document similarity. Currently it is up to to the developer to pick a similarity function that best suits the scenario. For example tf-idf is currently the state of the art for matching documents while BERT is the state of the art for query searches. It would be great to build a tool that auto-detects which similarity function is best suited based on the scenario and thus pick a similarity function that is optimized for memory and processing time. This could greatly help in scenarios like auto-matching resumes to job descriptions, clustering documents by category, classifying patients to different categories based on patient medical records etc.

Conclusion

In this paper, I covered some notable algorithms to calculate document similarity. It is no way an exhaustive list. There are several other methods for finding document similarity and the decision to pick the right one depends on the particular scenario and use-case. Simple statistical methods like tf-idf, Jaccard, Euclidien, Cosine similarity are well suited for simpler use-cases. One can easily get setup with existing libraries available in Python, R and calculate the similarity score without requiring heavy machines or processing capabilities. More advanced algorithms like BERT depend on pre-training neural networks that can take hours but produce efficient results for analysis requiring understanding of the context of the document.

Reference

[1] Heidarian, A., & Dinneen, M. J. (2016). A Hybrid Geometric Approach for Measuring Similarity Level Among Documents and Document Clustering. 2016 IEEE Second International Conference on Big Data Computing Service and Applications (BigDataService), 1–5. https://doi.org/10.1109/bigdataservice.2016.14

[2] Kavitha Karun A, Philip, M., & Lubna, K. (2013). Comparative analysis of similarity measures in document clustering. 2013 International Conference on Green Computing, Communication and Conservation of Energy (ICGCE), 1–4. https://doi.org/10.1109/icgce.2013.6823554

[3] Lin, Y.-S., Jiang, J.-Y., & Lee, S.-J. (2014). A Similarity Measure for Text Classification and Clustering. IEEE Transactions on Knowledge and Data Engineering, 26(7), 1575–1590. https://doi.org/10.1109/tkde.2013.19

[4] Nishimura, M. (2020, September 9). The Best Document Similarity Algorithm in 2020: A Beginner’s Guide — Towards Data Science. Medium. https://towardsdatascience.com/the-best-document-similarity-algorithm-in-2020-a-beginners-guide-a01b9ef8cf05

[5] Sharaki, O. (2020, July 10). Detecting Document Similarity With Doc2vec — Towards Data Science. Medium. https://towardsdatascience.com/detecting-document-similarity-with-doc2vec-f8289a9a7db7

[6] Lüthe, M. (2019, November 18). Calculate Similarity — the most relevant Metrics in a Nutshell — Towards Data Science. Medium. https://towardsdatascience.com/calculate-similarity-the-most-relevant-metrics-in-a-nutshell-9a43564f533e

[7] S. (2019, October 27). Similarity Measures — Scoring Textual Articles — Towards Data Science. Medium. https://towardsdatascience.com/similarity-measures-e3dbd4e58660

Poornima Muthukumar is a Senior Technical Product Manager at Microsoft with over 10 years of experience in developing and delivering innovative solutions for various domains such as cloud computing, artificial intelligence, distributed and big data systems. I have a Master's Degree in Data Science from the University of Washington. I hold four Patents at Microsoft specializing in AI/ML and Big Data Systems and was the winner of the Global Hackathon in 2016 in the Artificial Intelligence Category. I was honored to be on the Grace Hopper Conference reviewing panel for the Software Engineering category this year 2023. It was a rewarding experience to read and evaluate the submissions from talented women in these fields and contribute to the advancement of women in technology, as well as to learn from their research and insights. I was also a committee member for the Microsoft Machine Learning AI and Data Science (MLADS) June 2023 conference. I am also an Ambassador at the Women in Data Science Worldwide Community and Women Who Code Data Science Community.

More On This Topic

  • Similarity Metrics in NLP
  • Similarity Search: Euclid of Alexandria goes shoe shopping
  • A Graph-based Text Similarity Method with Named Entity Information in NLP
  • Introducing TensorFlow Similarity
  • Beyond Accuracy: Evaluating & Improving a Model with the NLP Test Library
  • Multilabel Document Categorization, step by step example

What developers trying out Google Gemini should know about their data

Google Gemini

Developers who have jumped in to try out Google Gemini for free should know their data might be used to train its generative artificial intelligence (AI) models, including those that power Google AI Studio and Gemini Pro.

The tech giant last week made Gemini Pro available to developers and businesses that are keen to build their own applications using its generative AI model. Developers can access the model via the Gemini API in Google AI Studio, while organizations will have to do so via Google Cloud's machine learning and development platform, Vertex AI.

Also: AI will change the role of developers forever, but leaders say that's good news

Developers currently have free access to Gemini Pro and Gemini Pro Vision, capped at 60 requests per minute, which Google said is suitable for most app development requirements. The Gemini Pro Vision model allows text and imagery to be accepted as input, although output remains as text.

Vertex AI developers can trial both AI models, within the same cap, for free until general availability, which is expected to be early 2024.

Also: How to use ChatGPT to write code

Following this date, charges per 1,000 characters or per image will apply across Google AI Studio and Vertex AI. Google said it had cut prices fourfold on input and twofold on output.

Gemini Pro supports 38 languages and is available across more than 180 markets, including the Asia-Pacific region.

Developers can move their AI Studio code to Vertex AI if they want a fully managed AI platform that offers more customization and Google Cloud features, including data governance and compliance, and security.

Google, though, is touting AI Studio as the fastest way to build using Gemini.

Developers should note that when they use the free quota of 60 requests per minute, their API and Google AI Studio input and output "may be accessible to trained reviewers".

Also: Generative AI means more productivity, and a likely retrenchment for software developers

Google told ZDNET that it uses the API inputs and outputs to improve product quality. "Human review is a necessary step of the model improvement process," a spokesperson said.

"Through review and annotation, trained reviewers help enable quality improvements of generative machine-learning models like the ones that power Google AI Studio and the Gemini Pro via the Gemini API."

To protect developers' privacy, Google said their data is de-identified and disassociated from their API key and Google account, which is needed to log in to Google AI Studio. This protection takes place done before the reviewers can see or annotate the data.

Also: AI in 2023: A year of breakthroughs that left no human thing unchanged

Google's Terms of Service (ToS) for its generative AI APIs further states that the data is used to "tune models" and may be retained in connection to the user's tuned models "[for] re-tuning when supported models change".

The ToS states: "When you delete a tuned model, the related tuning data is also deleted." The terms also state that users should not submit sensitive, confidential, or personal data to the AI models.

Data generated from when developers use Gemini Pro via Google AI Studio might still be accessed by Google reviewers, even if the developers make the move to Vertex AI.

The data generated while users were on Google AI Studio will be tapped to help improve products, the Google spokesperson told ZDNET.

"This includes further model tuning and evaluations. We may also derive product insights from anonymized data to help us determine new features we want to explore adding to Google AI Studio," they said.

Also: These 5 major tech advances of 2023 were the biggest game-changers

Developers and organizations with concerns about data security, but who are still keen to build with Gemini, will probably want to do so as Google Cloud customers, as this route will give them access to Vertex AI.

Google has assured that this patheway provides "customization of Gemini with full data control". Accessing Gemini models via Vertex Ai also allows enterprise customers to tune the models with their own data.

In addition, Google says it does not train its generative AI models on inputs or outputs from its cloud customers.

Artificial Intelligence