The Plagiarism Problem: How Generative AI Models Reproduce Copyrighted Content

plagiarism-in-AI

The rapid advances in generative AI have sparked excitement about the technology's creative potential. Yet these powerful models also pose concerning risks around reproducing copyrighted or plagiarized content without proper attribution.

How Neural Networks Absorb Training Data

Modern AI systems like GPT-3 are trained through a process called transfer learning. They ingest massive datasets scraped from public sources like websites, books, academic papers, and more. For example, GPT-3's training data encompassed 570 gigabytes of text. During training, the AI searches for patterns and statistical relationships in this vast pool of data. It learns the correlations between words, sentences, paragraphs, language structure, and other features.

This enables the AI to generate new coherent text or images by predicting sequences likely to follow a given input or prompt. But it also means these models absorb content without regard for copyrights, attribution, or plagiarism risks. As a result, generative AIs can unintentionally reproduce verbatim passages or paraphrase copyrighted text from their training corpora.

Key Examples of AI Plagiarism

Concerns around AI plagiarism emerged prominently since 2020 after GPT's release.

Recent research has shown that large language models (LLMs) like GPT-3 can reproduce substantial verbatim passages from their training data without citation (Nasr et al., 2023; Carlini et al., 2022). For example, a lawsuit by The New York Times revealed OpenAI software generating New York Times articles nearly verbatim (The New York Times, 2023).

These findings suggest some generative AI systems may produce unsolicited plagiaristic outputs, risking copyright infringement. However, the prevalence remains uncertain due to the ‘black box' nature of LLMs. The New York Times lawsuit argues such outputs constitute infringement, which could have major implications for generative AI development. Overall, evidence indicates plagiarism is an inherent issue in large neural network models that requires vigilance and safeguards.

These cases reveal two key factors influencing AI plagiarism risks:

  1. Model size – Larger models like GPT-3.5 are more prone to regenerating verbatim text passages compared to smaller models. Their bigger training datasets increase exposure to copyrighted source material.
  2. Training data – Models trained on scraped internet data or copyrighted works (even if licensed) are more likely to plagiarize compared to models trained on carefully curated datasets.

However, directly measuring the prevalence of plagiaristic outputs is challenging. The “black box” nature of neural networks makes it difficult to fully trace this link between training data and model outputs. Rates likely depend heavily on model architecture, dataset quality, and prompt formulation. But these cases confirm such AI plagiarism unequivocally occurs, which has critical legal and ethical implications.

Emerging Plagiarism Detection Systems

In response, researchers have started exploring AI systems to automatically detect text and images generated by models versus created by humans. For example, researchers at Mila proposed GenFace which analyzes linguistic patterns indicative of AI-written text. Startup Anthropic has also developed internal plagiarism detection capabilities for its conversational AI Claude.

However, these tools have limitations. The massive training data of models like GPT-3 makes pinpointing original sources of plagiarized text difficult, if not impossible. More robust techniques will be needed as generative models continue rapidly evolving. Until then, manual review remains essential to screen potentially plagiarised or infringing AI outputs before public use.

Best Practices to Mitigate Generative AI Plagiarism

Here are some best practices both AI developers and users can adopt to minimize plagiarism risks:

For AI developers:

  • Carefully vet training data sources to exclude copyrighted or licensed material without proper permissions.
  • Develop rigorous data documentation and provenance tracking procedures. Record metadata like licenses, tags, creators, etc.
  • Implement plagiarism detection tools to flag high-risk content before release.
  • Provide transparency reports detailing training data sources, licensing, and origins of AI outputs when concerns arise.
  • Allow content creators to opt-out of training datasets easily. Quickly comply with takedown or exclusion requests.

For generative AI users:

  • Thoroughly screen outputs for any potentially plagiarized or unattribued passages before deploying at scale.
  • Avoid treating AI as fully autonomous creative systems. Have human reviewers examine final content.
  • Favor AI assisted human creation over generating entirely new content from scratch. Use models for paraphrasing or ideation instead.
  • Consult AI provider's terms of service, content policies and plagiarism safeguards before use. Avoid opaque models.
  • Cite sources clearly if any copyrighted material appears in final output despite best efforts. Don't present AI work as entirely original.
  • Limit sharing outputs privately or confidentially until plagiarism risks can be further assessed and addressed.

Stricter training data regulations may also be warranted as generative models continue proliferating. This could involve requiring opt-in consent from creators before their work is added to datasets. However, the onus lies on both developers and users to employ ethical AI practices that respect content creator rights.

Plagiarism in Midjourney's V6 Alpha

After limited prompting Midjourney's V6 model some researchers were able to generated nearly identical images to copyrighted films, TV shows, and video game screenshots likely included in its training data.

Images Created by Midjourney Resembling Scenes from Famous Movies and Video Games

Images Created by Midjourney Resembling Scenes from Famous Movies and Video Games

These experiments further confirm that even state-of-the-art visual AI systems can unknowingly plagiarize protected content if sourcing of training data remains unchecked. It underscores the need for vigilance, safeguards, and human oversight when deploying generative models commercially to limit infringement risks.

AI companies Response on copyrighted content

The lines between human and AI creativity are blurring, creating complex copyright questions. Works blending human and AI input may only be copyrightable in aspects executed solely by the human.

The US Copyright Office recently denied copyright to most aspects of an AI-human graphic novel, deeming the AI art non-human. It also issued guidance excluding AI systems from ‘authorship'. Federal courts affirmed this stance in an AI art copyright case.

Meanwhile, lawsuits allege generative AI infringement, like Getty v. Stability AI and artists v. Midjourney/Stability AI. But without AI ‘authors', some question if infringement claims apply.

In response, major AI firms like Meta, Google, Microsoft, and Apple argued they should not need licenses or pay royalties to train AI models on copyrighted data.

Here is a summary of the key arguments from major AI companies in response to potential new US copyright rules around AI, with citations:

Meta argues imposing licensing now would cause chaos and provide little benefit to copyright holders.

Google claims AI training is analogous to non-infringing acts like reading a book (Google, 2022).

Microsoft warns changing copyright law could disadvantage small AI developers.

Apple wants to copyright AI-generated code controlled by human developers.

Overall, most companies oppose new licensing mandates and downplayed concerns about AI systems reproducing protected works without attribution. However, this stance is contentious given recent AI copyright lawsuits and debates.

Pathways For Responsible Generative AI Innovation

As these powerful generative models continue advancing, plugging plagiarism risks is critical for mainstream acceptance. A multi-pronged approach is required:

  • Policy reforms around training data transparency, licensing, and creator consent.
  • Stronger plagiarism detection technologies and internal governance by developers.
  • Greater user awareness of risks and adherence to ethical AI principles.
  • Clear legal precedents and case law around AI copyright issues.

With the right safeguards, AI-assisted creation can flourish ethically. But unchecked plagiarism risks could significantly undermine public trust. Directly addressing this problem is key for realizing generative AI's immense creative potential while respecting creator rights. Achieving the right balance will require actively confronting the plagiarism blindspot built into the very nature of neural networks. But doing so will ensure these powerful models don't undermine the very human ingenuity they aim to augment.

This AI startup made a $199 pocket toy that replaces apps with ‘rabbits’ — and it might just work

ZDNET CES 2024 cover art

As companies race to implement AI capabilities onto our most personal gadgets, one Santa Monica startup may have everyone beat by taking the road less traveled.

Also: CES 2024: What's Next in Tech

At CES today, Rabbit Inc. launched R1, a $199 consumer device powered by a natural-language operating system, with the goal of making app interactions obsolete.

"We've come to a point where we have hundreds of apps on our smartphones with complicated UX designs that don't talk to each other. As a result, end users are frustrated with their devices and are often getting lost," said Rabbit CEO, Jesse Lyu.

Lyu's not wrong. While it's become habitual for humans to interact with apps and web interfaces in a systematic way — scrolling through drop-downs, long-tapping on text to copy and paste, filtering searches, etc. — an AI that's trained to fulfill those tasks saves us time and energy.

At the core of Rabbit R1 is the company's "Large Action Model," an AI foundation that can train "rabbits" to see and learn how a user interacts with their typical apps and web experiences and then reproduce it, when asked for, on a customized cloud platform. This way, instead of requiring users to download multiple apps on their personal devices, the Large Action Model can access the services via a private web portal, dubbed "Rabbit hole," in which users can sign into their accounts, allow permissions, and more.

Also: Humane launches $699 AI-powered projector to replace your phone

"Similar to handing one's unlocked phone to a friend who will help order takeout, Rabbit OS (the device's operating system) performs tasks for users with their permission, without preemptively storing their identity information or passwords," the company explains in its press release.

The focus on security extends to the industrial design of the R1, which was made in partnership with Teenage Engineering, the firm responsible for some of the most creative designs in tech, including the Nothing Phone.

Besides the pocketable figure of the R1, there's a rotating "Rabbit eye" camera that, by default, covers its view by facing downward. The camera can be used for video calls, as well as "executing some of the most advanced computer vision applications," according to the company. I'll have a greater sense of what that chain of tech jargon implies during my CES demo.

Also: Apple reportedly eyeing generative AI push and Siri overhaul for the iPhone

Also, the microphone on the R1, mainly used for voice commands, only turns on when you press the side-mounted push-to-talk button. Other design elements include a 2.88-inch touchscreen display, a scroll wheel (pictured above) for navigating the OS, and a USB-C port for charging. The R1 is equipped with a 1,000mAh battery, 128GB of storage, and a SIM slot for mobile data.

Preorders for the Rabbit R1 begin today, with a rather accessible price (as far as CES announcements go) of $199, no monthly subscription required. By the time you're reading this, I'll be getting a hands-on demo of the Rabbit R1, seeing if one of the first standalone AI devices in 2024 can truly live up to the hype. Stay tuned for all of my initial thoughts.

Alphabet quantum spin-out Sandbox AQ acquires Good Chemistry

Alphabet quantum spin-out Sandbox AQ acquires Good Chemistry Kyle Wiggers 9 hours

Sandbox AQ, the AI and quantum firm spun out of Google parent company Alphabet in 2022, has acquired Good Chemistry, a Vancouver-based quantum and computational chemistry startup, for an undisclosed sum.

Good Chemistry, founded by Arman Zaribafiyan in 2021 as a spin-off of quantum computing company 1QBit and backed by investors including Green Sands Equity and Accenture Ventures, offers cloud-based tools to accelerate material design. Good Chemistry’s platform lets chemistry developers build chemical simulation apps and workflows using algorithms in quantum chemistry and machine learning as well as quantum computing.

In a press release, Sandbox AQ VP Nadia Carlsten said that the acquisition will expand Sandbox’s global presence in addition to adding “proven technologies,” including simulation technologies, and major customers (such as Dow Chemical) to its portfolio. New investors are a part of the deal, too — all of Good Chemistry’s backers have joined Sandbox AQ’s investor base, Carlsten added.

“Good Chemistry’s software and existing partnerships will enable us to rapidly accelerate how we bring the benefits of advanced simulation and AI tools to more customers,” Carlsten said. “Combining our capabilities will give Sandbox AQ an advanced and scalable computing platform for highly accurate chemical simulation on a scale that makes impactful use cases such as new drug and material synthesis possible.”

Sandbox AQ sees simulation as increasingly central to its work, which spans drug discovery and material design across the agricultural, defense, energy, medical and manufacturing sectors. With Good Chemistry, the company gains 25 computational and quantum chemists, AI and software engineers and quantum computing scientists skilled in leveraging simulation tech; tellingly, Zaribafiyan is joining Sandbox as head of product for simulation platforms while Good Chemistry’s staff is assuming roles in Sandbox’s existing simulation teams.

“The rapid advances in AI, quantum, cloud and high-performance computing have unlocked endless opportunities for companies like Good Chemistry and Sandbox AQ to reinvent the way we think about chemistry, discover ways to make products safer, stronger and more sustainable, and reshape the fabric of our world,” Zaribafiyan said in a canned statement. “[Tapping] the broader Sandbox AQ reach and domain expertise will supercharge the capabilities of our platform.”

As part of the acquisition, Sandbox AQ says that it’ll integrate Good Chemistry’s software, Qemist Cloud and Tangelo, into its enterprise software portfolio. Qemist lets user take advantage of quantum chemistry through simulation, while Tangelo allows them to conduct chemistry experiments across quantum cloud providers and hardware devices.

Good Chemistry is Sandbox’s second acquisition. In 2022, it bought Cryptosense, an enterprise-focused software-as-a-service cybersecurity and encryption startup.

DSC Weekly 9 January 2024

Announcements

  • Generative AI signifies a pivotal shift in today’s technological landscape. It promises profound insights, streamlined operations, and assistance with data-driven decisions on an unprecedented scale. However, GenAI also brings forth ethical and regulatory considerations that require attention for modern businesses seeking to capitalize on the still-evolving technology. Register for the Enterprise Strategy Group’s upcoming GenAI Summit​ to engage with thought leaders as they navigate the intricate dilemmas, evolving regulatory landscape, and responsible AI practices that maximize the benefits of GenAI technology and mitigate inherent risks and biases.​
  • Ransomware attacks show no signs of slowing down. This year marked a record-breaking year for ransomware attacks, as they surged 74% by the first three months of 2023. Organizations require not only a solid prevention plan, but they need established recovery solutions to ensure they bounce back from attacks that can cause irreparable economic and reputational damage. The newer and more treacherous modern threat landscape forces organizations to take a second look at cyber insurance and the security it can ensure against fallout from an attack. Join the upcoming Ransomware Preparedness: Strategies for a Secure Future summit to hear leading experts discuss actionable strategies to prevent ransomware attacks, mitigate damage, and select the best cyber insurance option for your organization.

Top Stories

  • GenAI: Beware the Productivity Trap; It’s About Nanoeconomics – Part 2
    January 6, 2024
    by Bill Schmarzo
    In Part 1 of the series “GenAI: Beware the Productivity Trap,” we discussed embracing an economic mindset to avoid falling into the productivity trap. We discussed some challenges with the productivity trap and then reviewed some data economic concepts that can take your organization to the next level of game-changing performance and innovation.
  • The importance of effective API documentation and design
    January 8, 2024
    by Ovais Naseem
    APIs are the backbone of interconnected systems, enabling seamless data exchange and functionality integration across diverse applications. One of the foundational pillars of successful API implementation lies in its documentation and design. Clear, comprehensive documentation coupled with thoughtful design eases the integration process and enhances developer experience, fostering faster adoption and innovation.
  • Textual predictive coding: Do LLMs and the human mind compare?
    January 8, 2024
    by David Stephen
    There is a new letter on TIME, What Generative AI Reveals About the Human Mind, where a professor wrote, “Natural brains must learn to predict those sensory flows in a very special kind of context—the context of using the sensory information to select actions that help us survive and thrive in our worlds.
Education_DSC_160x600-2

In-Depth

  • How data science is reshaping diverse industries
    January 8, 2024
    by Erika Balla
    How do some industries seem to have cracked the code for success? It’s not luck—it’s the power of data science that changes the game. Whether it’s technology or the finance sector, data science is transforming how well we do things by understanding the data.
  • Unleashing innovation: How AI chatbots transform your website strategy
    January 8, 2024
    by Pritesh Patel
    In our fast-changing, digitized world business strategies, and content planning are also moving into the world of numbers, minimizing the need for human work. Nowadays, artificial intelligence is developing day by day, expanding over more and more users and areas of use.
  • Real-time analytics with database streaming services: Harnessing data velocity
    January 8, 2024
    by Karen Anthony
    In the short-paced landscape of information-driven decision-making, actual-time analytics has come to be paramount for corporations seeking to benefit from insights at the rate of the enterprise. Database streaming offerings have emerged as a transformative answer, allowing the processing and analysis of facts in movement.
  • Mitigating Ethical Risks in Generative AI: Strategies for a Safe and Secure AI Application
    January 3, 2024
    by Matthew McMullen
    Artificial Intelligence (AI) has been around for many decades but now it has become a buzzword even among non-technical people because of the generative AI models like ChatGPT, Bard, Scribe, Claude, DALL·E 2, and a lot more. AI has moved beyond its sci-fi origins to reality, creating human-like content and powering self-driving cars.
  • DSC Weekly 2 January 2024
    January 2, 2024
    by Scott Thompson
    Read more of the top articles from the Data Science Central community.

Intuition Robotics is giving its social bot a generative AI upgrade, and it makes so much sense

Intuition Robotics

Intuition Robotics launched its senior assistive social robot ElliQ in 2022, and even though it already has some impressive capabilities, including games, workout plans, interactive conversations, and AI, the robot is about to get even smarter.

Also: Can AI curb loneliness in older adults? This robot companion is proving it's possible

On Tuesday at CES, Intuition Robotics unveiled ElliQ 3.0, the latest generation of its robot, which boasts hardware and software upgrades, including a deeper integration of generative artificial intelligence (AI) for an even better experience.

The most evident change is ElliQ's new look. The robot's hardware has been tweaked and designed to meet growing demand and increased manufacturing processes. The robot is now 1.3 pounds lighter and has a 36% smaller footprint.

In addition, instead of having a removable tablet that doubles as a screen, the screen is now fully integrated into the device, as seen in the picture below.

These hardware changes also improve how seniors interact with ElliQ, making the robot simpler and easier to handle. Some internal changes include an octa-core SoC, a built-in AI processing unit (APU), and 33% more RAM.

ElliQ's conversational capabilities have also been improved through deeper integration of generative AI, via its Relationship Orchestration Engine, which "makes real-time decisions regarding actions, scripted conversation, and generative AI conversation," according to the release.

Also: Generative AI is a developer's delight. Now, let's find some other use cases

The implementation of generative AI will also support the robot's memory feature, which takes information the user shares with the device, such as their favorite color, their pet name, or their religion, and remembers this knowledge so it can be used to fuel conversations and choose activities based on the person's interests.

"It's astounding to see that the first people to live with and build long-term relationships with an AI are individuals in their 80's and 90's," said Dor Skuler, co-founder and CEO of Intuition Robotics. "Through this relationship, ElliQ is proving to be highly effective in reducing older adults' sense of loneliness, improving health and independence, and increasing social connectedness."

Also: Generative AI and machine learning are engineering the future in these 9 disciplines

To address safety concerns regarding the implementation of generative AI, which has proven to be vulnerable to shortcomings such as hallucinations, the company has implemented guardrail mechanisms to monitor and mediate conversations in real-time.

When I tested the earlier model of ElliQ, one of my favorite features was the interactive games. Now, ElliQ 3 will take the gaming experience up a notch by implementing synchronized games, starting with Bingo, which will allow users to participate in real-time games with other adults.

CES 2024

Leap AI wants to help businesses build and integrate AI workflows

Leap AI wants to help businesses build and integrate AI workflows Ivan Mehta 8 hours

With the availability of different types of Generative AI models increasing, more businesses are including features such as summarization, prompt-based image generation and even music-making in their solutions. Leap AI is building a solution for these companies to easily integrate AI-powered workflows or even build their own using an easy process.

The company is aiming to cater to use cases ranging from consumer-facing apps such as Songburst to internal tools for businesses. Leap AI offers ready templates such as music/beat generators and professional headshot generators. Alternatively, companies can build their own workflows using building blocks offered by the startup.

The startup offers a free plan with limited credits for AI model queries and a limited number of workflows to try out their solution. It also has plans starting from $29 per month with more credits, no limit on building workflows with customer support.

Why did the founders build Leap AI?

Leap AI was started in February 2023 by Alex Schachne and Claudio Fuentes. Schachne founded two companies in the edtech space while he was at Johns Hopkins University. Fuentes who has also founded a few startups previously worked in companies like WeWork and Pypestream. Both co-founders met at the Founders Inc. Studio incubator program and started building Leap AI.

“We were invited to the studio as builders while we didn’t have any idea. They asked us to get together and try and build something. We started building AI-powered consumer tools such as generating children’s stories for parents,” the co-founders told TechCrunch over a call.

“While we were building these tools, we realized it was painful to manage all the orchestration between multiple models to create these workflows. So we wanted to automate this process and create an engine that keeps running in the background that effectively helps you augment your work.”

Leap co-founders Alex Schachne (left) and Claudio Fuentes (right)

Leap co-founders Alex Schachne (left) and Claudio Fuentes (right) Image Credits: Leap AI

Funding an the future

The company has raised a $1.4 million seed round led by Founders Inc. with participation from Carya Venture Partners, Gaingels, performance monitoring startup Sentry’s founder David Cramer, Firebase alternative Supabase’s founder Paul Copplestone, along with executives from Google and AngelList.

Leap AI already has a few customers including Heineken, Inflection, and Live Nation. The co-founders said that Live Nation uses their solution to build an internal tool, which generates brand-consistent assets.

The company said that it is trying to iterate fast on products by using its workflows in its solution for things like customer support and content generation.

Leap AI workflow building tool Image Credits: Leap AI

Fuentes said that the company’s ability to run AI-powered tasks gives it an edge over agent-based solutions.

“One thing that we’re fundamentally doing differently is the ability to schedule these things to run automatically in the background and have your workforce operating without you needing to come to the chatbox and ask it every single time you need something. Another unique aspect is that we provide the interoperability between multiple models, multiple vendors, and multiple companies,” he said.

Schachne thinks that getting partnerships like integrations with Zapier, Vercel, and Superbase will help the company stand out from others.

The company said it is investing resources into revamping its workflow builder to replace a fixed-step approach with a chain of thought reasoning to achieve a goal. Leap AI is also working on improving context awareness of its workflows so it can leverage previously generated results.

Can AI make art more accessible? One museum models a safe way forward

rainbow-bridgegettyimages-1212460137

The past 12 months have seen a rapid rise in both the use of generative artificial intelligence (AI) and a vociferous response to the technology that can be broadly classified in two camps.

In one camp are excited amateurs who are using generative AI tools from technology giants, such as OpenAI, Microsoft, and Google, to create content from prompts in seconds.

Also: How to write better ChatGPT prompts for the best generative AI results

In the second camp are creatives — from writers and musicians to coders and artists — who fear their hard-learned professional skills could be undermined by the capabilities of generative AI.

This second camp fears their intellectual property is often being exploited without their agreement to train the models that power generative AI models.

Yet sitting between these two camps are organizations and individuals who are looking to create benefits from AI in a safe and ethical manner.

One of these pioneering professionals is Birgitte Aga, head of innovation and research at Munch Museum (MUNCH) in Oslo, Norway, which contains the world's most extensive collection of art dedicated to the Norwegian artist Edvard Munch.

With 27,000 artworks, non-art objects and writings — spread across 11 galleries on 13 floors — the museum wants to show the best parts of its collection to the widest possible audience.

Aga's role is to help MUNCH achieve its objectives through the effective exploitation of emerging technologies, such as AI, machine learning (ML), and more.

Now, in a leading-edge project alongside technology giant TCS, Aga and her colleagues at MUNCH are finding new ways to demonstrate how AI can boost interest in artistic endeavors, rather than just being a potential threat to creative processes.

Also: How to use Leonardo AI to generate stunning artwork and images

"The central point of what we create is the artwork," says Aga in a one-to-one video interview with ZDNET. "We're not replacing the painting with technology; we're enriching the experience."

TCS and MUNCH are designing, developing, and testing pioneering AI and ML technologies connected to the museum's database of 7,000 original drawings.

"We need to get people into the museum and get them engaged," says Aga. "One of the ways we're doing that is through technology. We are rethinking how you can make an archive of artworks relevant to the audience through the practice of drawing."

The two organizations are training an ML algorithm with Munch's drawings and developing a user interface that allows museum visitors to become immersed in his artwork.

Aga describes the interface, which is currently at the prototype stage, as "a back projection on a transparent surface."

When a user places a sheet of paper on the interface and starts drawing, their pen marks are met with a projected line from the machine-learning algorithm in real time: "The AI guides them to explore their own and Munch's creative drawing process simultaneously."

Also: The best AI chatbots: ChatGPT and other noteworthy alternatives

While the rapid rise of generative AI applications — such as Midjourney and DALL-E — has shown the game-changing power of emerging technology, Aga believes MUNCH's pioneering collaboration with TCS highlights how the old world can work hand-in-hand with the new.

"AI in our sector is a very sensitive subject. We understand our role in society is to be a platform to discuss what AI is and what effect it has on liberty, society, and the individual," she says.

"Our work is about reassuring audiences and partners that AI is not going to replace the museum or Edvard Munch. And I think it's super-exciting to think about what this technology can do for audiences and how we can reach more people."

Aga recognizes that the ethical application of emerging technology is a crucial success factor, which is something other experts have mentioned before.

Avivah Litan, distinguished VP analyst at Gartner, explained to me last year how any executives dabbling in emerging technology must "manage the risks before they manage you."

Also: The 3 biggest risks from generative AI — and how to deal with them

Gartner recently polled 700-plus executives about the risks of generative AI and discovered CIOs are most concerned about data privacy, followed by hallucinations, and then security.

Litan says executives must make sure they use data and AI in a way that's acceptable to the organization, its people, and its customers.

Aga says MUNCH has a talented team of mediators and learning specialists who explore how to make Edvard Munch's art relevant to the general public.

"We start with user need," she says. "We have users that come to the museum, such as young adults, who want an interactive and participatory experience. They don't want to just come and stand in front of a painting."

Also: The best AI art generators: DALL-E 2 and fun alternatives to try

However, creating a great, data-led experience is far from straightforward. For a start, there's a time lag in many current generative AI interfaces between the input of a prompt and the output of content.

"There are lots of technology and research challenges that we haven't encountered before," says Aga. "We're trying to decipher how an artist drew while, at the same time, trying to create a user interface that works in real time."

Yet Aga says the museum's AI-led project is progressing well. Depending on the success of the prototyping stage and user feedback, the interface could be used to create a new experience for audiences in locations beyond Oslo in the future.

"AI is super-interesting for us in terms of both research and innovation. We're very excited to be working with TCS," she says. "The project is about presenting Munch's work and making it relevant. We're only starting the exploration. This is the first test on how we can work together and we'll see where else this initiative can go."

Also: The ethics of generative AI: How we can harness this powerful technology

In fact, more data-led innovation is already taking place. The Museum's MUNCH Audience Lab, for example, continues to explore a range of technology-based experiences for all kinds of audiences.

Aga says her organization is exploring how language models might help to create a knowledge base of the museum's vast collection.

MUNCH is also part of a broader European project that's investigating how AI might help to predict color fading in art objects, with this insight used to bolster conservation efforts.

Whether it's through machine learning, immersive technology, or gaming, Aga says the museum's pioneering data initiatives aim to introduce digital systems carefully and effectively.

"Emerging technology that's implemented in the wrong way can threaten liberty, equality, individuality, and creativity," she says.

"But emerging technology that's applied in the right way can provide knowledge, research, and understanding. We're custodians of the artwork for Oslo's inhabitants. And our job is to conserve, present, and make relevant Munch's artistry to the people."

Artificial Intelligence

Microsoft puts Azure Quantum Elements to work

Microsoft puts Azure Quantum Elements to work Frederic Lardinois @fredericl / 8 hours

Microsoft today announced that it has worked with the U.S. Department of Energy’s Pacific Northwest National Laboratory (PNNL) to use its Azure Quantum Elements service to whittle down millions of potential new battery materials to only a few — with one of them now in the prototype stage.

Now, before you get too excited about the ‘quantum’ part of ‘Azure Quantum Elements’ (and why wouldn’t you — it’s in the name, after all), let’s get this out of the way first: no quantum computer was used in this project. Azure Quantum Elements, which launched last summer, combines AI and traditional high-performance computing (HPC) techniques into what is essentially a workbench for scientific computing, with the promise of providing access to Microsoft’s quantum supercomputer in the future. So even though no qubits were involved in this current project, but the overall idea here is to bring all of these technologies together over time.

Image Credits: Microsoft

Krysta Svore, who leads Microsoft Quantum, told me that the overall idea here was to see how far the team could push what is currently available in Azure Quantum Elements (AQE) — and especially the AI accelerator — to advance materials discovery. Using AQE, the researchers at PNNL looked at 32 million inorganic materials to arrive at 18 candidates for their battery project. First, the teams used AQE’s AI models to whittle down the pool to about 500,000 candidates. After that, the researchers then used existing HPC techniques to identify those 18 promising candidates to focus on. Typically, it would take years to go through this process and to build a prototype battery. Using AQE, the researchers were able to do this in 18 months.

“The intersection of AI, cloud and high-performance computing, along with human scientists, we believe is key to accelerating the path to meaningful scientific results,” said Tony Peurrung, PNNL Deputy Director for Science and Technology. “Our collaboration with Microsoft is about making AI accessible to scientists. We see the potential for AI to surface a material or an approach that is unexpected or unconventional, yet worth investigating. This is a first step in what promises to be an interesting journey to accelerate the pace of scientific discovery.”

Many quantum computing boosters expect that their machines will excel at solving chemistry and material science problems. And while the quantum computing community continues to push the state of the art ahead at a steady pace, we’re still at least a few years away from seeing a quantum computer that is actually useful. We’re currently still in the noisy intermediate-scale quantum (NISQ) era, after all. Svore, unsurprisingly, remains optimistic that Microsoft will be able to deliver on its plan to build a quantum supercomputer that uses its Majorana-based qubits, within the next decade.

For now, though, even though there is obiously real science involved, it’s hard not to look at this as a bit of a PR exercise, given how far we are still from bringing quantum computing into the process.

Microsoft expects to build a quantum supercomputer within 10 years

5 Coding Tasks ChatGPT Can’t Do

5 Coding Tasks ChatGPT Can't Do
Image by Author

I like to think of ChatGPT as a smarter version of StackOverflow. Very helpful, but not replacing professionals any time soon. As a former data scientist, I spent a solid amount of time playing around with ChatGPT when it came out. I was pretty impressed with its coding capacity. It could generate pretty useful code from scratch; it could offer suggestions on my own code. It was pretty good at debugging if I asked it to help me with an error message.

But inevitably, the more time I spent using it, the more I bumped up against its limitations. For any developers fearing ChatGPT will take their jobs, here’s a list of what ChatGPT can’t do.

5 Coding Tasks ChatGPT Can't Do 1. Anything your Company would use Professionally

The first limitation isn’t about its ability, but rather the legality. Any code purely generated by ChatGPT and copy-pasted by you into a company product could expose your employer to an ugly lawsuit.

This is because ChatGPT freely pulls code snippets from data it was trained on, which come from all over the internet. “I had chat gpt generate some code for me and I instantly recognized what GitHub repo it got a big chunk of it from,” explained Reddit user ChunkyHabaneroSalsa.

Ultimately, there's no telling where ChatGPT’s code is coming from, nor what license it was under. And even if it was generated fully from scratch, anything created by ChatGPT is not copyrightable itself. As Bloomberg Law writers Shawn Helms and Jason Krieser put it, “A ‘derivative work’ is ‘a work based upon one or more preexisting works.’ ChatGPT is trained on preexisting works and generates output based on that training.”

If you use ChatGPT to generate code, you may find yourself in trouble with your employers.

2. Anything Requiring Critical Thinking

Here’s a fun test: get ChatGPT to create code that would run a statistical analysis in Python.

Is it the right statistical analysis? Probably not. ChatGPT doesn’t know if the data meets the assumptions needed for the test results to be valid. ChatGPT also doesn’t know what stakeholders want to see.

For example, I might ask ChatGPT to help me figure out if there's a statistically significant difference in satisfaction ratings across different age groups. ChatGPT suggests an independent sample T-test and finds no statistically significant difference in age groups. But the t-test isn't the best choice here for several reasons, like the fact that there might be multiple age groups, or that the data aren’t normally distributed.

5 Coding Tasks ChatGPT Can't Do
Image from decipherzone.com

A full stack data scientist would know what assumptions to check and what kind of test to run, and could conceivably give ChatGPT more specific instructions. But ChatGPT on its own will happily generate the correct code for the wrong statistical analysis, rendering the results unreliable and unusable.

For any problem like that which requires more critical thinking and problem-solving, ChatGPT is not the best bet.

3. Understanding Stakeholder Priorities

Any data scientist will tell you that part of the job is understanding and interpreting stakeholder priorities on a project. ChatGPT, or any AI for that matter, cannot fully grasp or manage those.

For one, stakeholder priorities often involve complex decision-making that takes into account not just data, but also human factors, business goals, and market trends.

For example, in an app redesign, you might find the marketing team wants to prioritize user engagement features, the sales team is pushing for features that support cross-selling, and the customer support team needs better in-app support features to assist users.

ChatGPT can provide information and generate reports, but it can't make nuanced decisions that align with the varied – and sometimes competing – interests of different stakeholders.

Plus, stakeholder management often requires a high degree of emotional intelligence – the ability to empathize with stakeholders, understand their concerns on a human level, and respond to their emotions. ChatGPT lacks emotional intelligence and cannot manage the emotional aspects of stakeholder relationships.

You may not think of that as a coding task, but the data scientist currently working on the code for that new feature rollout knows just how much of it is working with stakeholder priorities.

4. Novel Problems

ChatGPT can’t come up with anything truly novel. It can only remix and reframe what it has learned from its training data.

5 Coding Tasks ChatGPT Can't Do
Image from theinsaneapp.com

Want to know how to change the legend size on your R graph? No problem – ChatGPT can pull from the 1,000s of StackOverflow answers to questions asking the same thing. But (using an example I asked ChatGPT to generate), what about something it’s unlikely to have come across before, such as organizing a community potluck where each person's dish must contain an ingredient that starts with the same letter as their last name and you want to make sure there's a good variety of dishes.

When I tested this prompt, it gave me some Python code that decided the name of the dish had to match the last name, not even capturing the ingredient requirement correctly. It also wanted me to come up with 26 dish categories, one per letter of the alphabet. It was not a smart answer, probably because it was a completely novel problem.

5.Ethical Decision Making

Last but not least, ChatGPT cannot code ethically. It doesn't possess the ability to make value judgments or understand the moral implications of a piece of code in the way a human does.

Ethical coding involves considering how code might affect different groups of people, ensuring that it doesn't discriminate or cause harm, and making decisions that align with ethical standards and societal norms.

For example, if you ask ChatGPT to write code for a loan approval system, it might churn out a model based on historical data. However, it cannot understand the societal implications of that model potentially denying loans to marginalized communities due to biases in the data. It would be up to the human developers to recognize the need for fairness and equity, to seek out and correct biases in the data, and to ensure that the code aligns with ethical practices.

It’s worth pointing out that people aren’t perfect at it, either – someone coded Amazon’s biased recruitment tool, and someone coded the Google photo categorization that identified Black people as gorillas. But humans are better at it. ChatGPT lacks the empathy, conscience, and moral reasoning needed to code ethically.

Humans can understand the broader context, recognize the subtleties of human behavior, and have discussions about right and wrong. We participate in ethical debates, weigh the pros and cons of a particular approach, and be held accountable for our decisions. When we make mistakes, we can learn from them in a way that contributes to our moral growth and understanding.

Final Thoughts

I loved Redditor Empty_Experience_10’s take on it: “If all you do is program, you’re not a software engineer and yes, your job will be replaced. If you think software engineers get paid highly because they can write code means you have a fundamental misunderstanding of what it is to be a software engineer.”

I’ve found ChatGPT is great at debugging, some code review, and being just a bit faster than searching for that StackOverflow answer. But so much of “coding” is more than just punching Python into a keyboard. It’s knowing what your business’s goals are. It’s understanding how careful you have to be with algorithmic decisions. It’s building relationships with stakeholders, truly understanding what they want and why, and looking for a way to make that possible.

It’s storytelling, it’s knowing when to choose a pie chart or a bar graph, and it’s understanding the narrative that data is trying to tell you. It's about being able to communicate complex ideas in simple terms that stakeholders can understand and make decisions upon.

ChatGPT can’t do any of that. So long as you can, your job is secure.

Nate Rosidi is a data scientist and in product strategy. He's also an adjunct professor teaching analytics, and is the founder of StrataScratch, a platform helping data scientists prepare for their interviews with real interview questions from top companies. Connect with him on Twitter: StrataScratch or LinkedIn.

More On This Topic

  • 5 Tasks To Automate With Python
  • Data Representation for Natural Language Processing Tasks
  • Scaling human oversight of AI systems for difficult tasks — OpenAI approach
  • HuggingGPT: The Secret Weapon to Solve Complex AI Tasks
  • If You Can Write Functions, You Can Use Dask
  • Can ChatGPT Be Trusted as an Educational Resource?

I tried Getty’s new AI image generator, and it doesn’t really compare to DALL-E

a-fleet-of-trucks-driving-up-a-waterfall-outside-a-fairytale-kingdom

Prompt: "A fleet of trucks driving up a waterfall outside a fairytale kingdom."

As ZDNET reported on Monday, stock photography giant Getty Images has unveiled a generative artificial intelligence (AI) image service that it says is "safe" to use because it is trained on Getty's licensed content library and, therefore, does not run the same risk of copyright infringement as other generative programs.

The announcement follows Getty's announcement of a generative AI capability in September. At the time, that capability was presented only as a demo, whereas the iStock site is open for business now.

Also: Getty Images launches its own 'commercially safe' AI image generator

Getty's service, developed with AI chip giant Nvidia, was unveiled at the annual CES trade show in Las Vegas. The program comes amidst a legal firestorm over copyright infringement, with the New York Times suing Microsoft and OpenAI a week earlier over alleged copyright infringement, and scholars documenting how the image AI program Midjourney could be prompted to reproduce protected images from movies.

Getty emphasizes that its program provides indemnification to users. The content license agreement posted after signing up specifies that "iStock's total maximum aggregate liability (meaning the total amount iStock is responsible for, whether under this agreement or any other agreement for the same content) is limited to $10,000 US dollars per item of content." An "extended" indemnification of $250,000 per content item can be purchased as an additional capability.

I took the program, "Generative AI by iStock", for a spin, using the introductory $14.99 bucket of 100 image generations and found it to be a workable substitute to images created with OpenAI's DALL-E and Clipdrop by Stability AI.

Also: 'Generative AI by iStock' lets users create images without copyright-infringement worries

To get started, I created an account on istockphoto.com, and put in details for a credit card that was instantly billed $14.99. I was then faced with a blank prompt. After entering a prompt, the results showed four images at a time, with each batch of four counting as one of the initial 100 images in the bucket.

I tried the same prompts on DALL-E and ClipDrop. The results from iStock were noticeably less interesting aesthetically and from a narrative perspective, and they were overall rather obvious to the point of being bland. But the images were generally in accord with the prompt provided.

For example, to create an imaginary scenario of apples inside some kind of experiment, I had previously submitted to DALL-E the prompt, "An apple inside of a bottle lying on its side, with apples on either side of the bottle." That produced a vivid scene of a table full of interesting science-like instruments. The version by iStock is appropriate to the prompt, but far less interesting (see below).

Another wild prompt was used to dramatize an imaginary impossible computer: "An incredibly complex computer the size of a room with hundreds of gears, levers and dials and a digital interface". In Clipdrop, that prompt produced an intriguiging, detailed scene of a room with various machine parts, with detailed texture and a doorway that had an ominous air to it. In iStock, the result was simply what looked like a concentration of gears, with none of the implicit drama that made the Clipdrop image interesting.

A third example, also in Clipdrop, was meant to dramatize cloud computing as a mysterious realm. I offered the prompt, "Hundreds of tiny workers with cranes building castles in the sky, photographic." In Clipdrop, that prompt led to a depiction of a construction site, which is focused around a sort of Tower of Babel, an interesting improvisational touch by Clipdrop that went beyond the explicit prompt guidelines.

Also: Why DeepMind's AI visualization is utterly useless

The iStock rendering, again, had all the elements mentioned but added up to a rather bland, very literal rendering, devoid of any atmosphere or mood.

Obviously, prompt engineering may yield more creative uses of iStock over time. Out of the box, however, its results are fairly dull. The program seems to mostly pick up on the simplest elements of the prompt and stick them in the frame.

There appears to be very little ability to parse complex ideas, such as "Inside of a raindrop as if you are a tiny, tiny person who is seeing all the little creatures that live and work and play in there," which requires multiple levels of composing elements in a way that is not realistic.

In fact, when a fantastical situation is realized by iStock, the results seem rather degraded compared to more realistic scenarios, as is the case in the prompt, "A fleet of trucks driving up a waterfall outside a fairytale kingdom," in the illustration at the top of this story.

It's important to note there are important qualifications and limitations to the indemnification provided by Getty. The content license agreement notes that the coverage stops where the user provides prompts that mention copyright material.

"iStock's indemnification obligations do not apply to the extent you generate content that includes prompts or inputs that include the names, likeness of real people, trademark, trade dress, logos, works of art of architecture or other elements protected by third-party intellectual property rights that you do not have the right to use," the agreement states.

Also: Nvidia makes the case for the AI PC at CES 2024

I tried out several controversial image prompts that scholars Gary Marcus and Reid Southen have claimed can be used in Midjourney to reproduce copyrighted images. In each case, either iStock produced an image that did not appear to have any obvious aspects of copyrighted material, or the program would not generate an image and produced a warning that the prompt was blocked because it was not compliant.

For example, the phrase "protocol droid from classic sci-fi movie" was used by Marcus and Southen in Midjourney to reproduce images that are almost identical to images of the droid C-3PO from Star Wars. The same prompt with iStock produced several images that look like toy robots, but they have nothing to do with Star Wars.

In another instance, the phrase "man in robes with light sword, screencap" was used by Marcus and Southen to induce Midjourney to produce an almost exact replica of a shot of Obi-Wan Kenobi from Star Wars. In iStock, the same prompt generated not only a refusal to generate an image, but also a warning that the word "sword" was forbidden because it "may violate our AI policy."

Some brands might slip through the filter, however. I was able to type, "The ZDNET journalists as interstellar superheroes", and produced images of costumed people with a heroic air about them.

`

Artificial Intelligence