Why Oracle’s Infrastructure is Best for Generative AI 

Generative AI has literally shaken the cloud infrastructure. “The shape of Azure has drastically changed and continues to change rapidly to support the models you’re building,” said Microsoft chief Satya Nadella while sharing the stage with Sam Altman at OpenAI’s first developer conference, DevDay.

He further mentioned that OpenAI has challenged Microsoft Azure to change its infrastructure to match OpenAI’s technology prowess. The first thing that we have been doing in partnership with you is changing the system all the way from power to the DC (data center), to the rack, to the accelerators and to the network,” he added.

To adapt to the demands of generative AI, all major hyperscalers, including AWS and Google Cloud, are compelled to make changes to their infrastructure. For example, earlier this year, the Google Cloud Platform announced that they are working to integrate AI infrastructure more extensively into their overall fleets. Similarly, AWS stated that the company plans to deploy multiple AI-optimized server clusters over the next 12 months.

“I would argue that the changes they (AWS and Google Cloud) are making is an attempt to build a networking infrastructure like we have at Oracle,” said Christopher G. Chelliah, senior vice president, technology & customer strategy, JAPAC, in an exclusive interview with AIM.

However Microsoft, in addition to updating its infrastructure, has recently also partnered with Oracle in a multiyear agreement to enhance AI services. It will now use both Oracle Cloud Infrastructure (OCI) AI and Microsoft Azure AI for daily Bing conversational searches.

Microsoft’s turn to Oracle implies that there may be certain shortcomings in Microsoft Azure.

What Sets Oracle Apart from Other Hyperscalers

Elaborating further how OCI is built differently from its counterparts, Chelliah said that OCI is a second generation cloud which simply means that networks built by OCI are not shared between tenants in the cloud.

“The network serves as the bottleneck for the cloud. OCI was designed with a distinctly different network compared to other players,” he said.

In contrast, when discussing AWS and Google Cloud, Chelliah noted a challenge faced by them. He said they currently host existing customers within existing tenancies, making it impractical for them to swiftly replace or upgrade their networks. “They cannot rip out those networks and change those networks overnight,” he added. Continuing, he confidently stated, “I believe our competitors are at least 18 months behind us at Oracle in this regard.”

Moreover, he mentioned that Oracle ensured the network topology is non-blocking so that the network is not shared between tenants in the cloud. He explained this using an analogy, where, instead of road intersections, Oracle has built flyovers.

From Oracle’s perspective, he said, “We already have that network, and Oracle is a step ahead.” The next problem Oracle is trying to solve is data privacy.

“We’re helping customers do training inferencing and RAG in isolation and privacy so that you can now bring corporate sensitive, private data, do trade, you know, fine tuning inferencing and RAG in this cloud, without impacting any privacy issue” he added.

Oracle ‘Feeds the Beast’

Running generative AI demands a combination of infrastructure and data.

Oracle is well-equipped in terms of infrastructure, as NVIDIA selected OCI as the first hyperscale cloud provider to offer NVIDIA DGX Cloud. “When NVIDIA thinks of cloud and data, they think of Oracle,” said Chelliah, saying that Oracle utilises MySQL HeatWave data for real-time anomaly detection on NVIDIA clusters for its customers.

“The second biggest ingredient in AI is data; you need to ‘feed the beast’ (AI Machine),” he said. “AI is not just GPUs, it’s important. But for me to be differentiated, it’s data” highlighted Chelliah.

Oracle has the majority of available data, trusted by both enterprises and governments. Recently, Oracle has shifted its strategy, extending its data to other cloud service providers. It is reaching customers where they are.

Moreover, Larry Ellison, CTO of Oracle has taken this commitment of openness to the next level. “Cloud should be open,” he said, talking about its partnership with Microsoft, and eventually partnering with other cloud providers, namely AWS and Google. “Today I can run my data platform on Amazon today with MySQL HeatWave available on Amazon.” said Chelliah.

“We’re playing this game where we tell customers to feed the beast, as I keep calling it, to feed the AI machine you need data,” he concluded.

The post Why Oracle’s Infrastructure is Best for Generative AI appeared first on Analytics India Magazine.

The Danger of Untethered Use Cases

Slide1-1

Client: “We’ve identified over 60 use cases!”

Schmarzo: “You might as well have zero…”

I always get very concerned when I see organizations that jump right into the use case inventory process without first considering the organization’s key business initiatives. Every organization has a laundry list of use cases. They tend to be the dumping grounds for organizations that lack a strategic approach for determining where and how to leverage data and analytics to drive business outcomes.

Organizations don’t fail from a lack of use cases but fail because they have too many.

I’ve talked, lectured, and written several times – heck, I even wrote “The Economics of Data, Analytics, and Digital Transformation” – about the economic ramifications of a use case-centric approach to building the organization’s data and analytic assets. A use case-centric approach (Figure 1) unleashes the unique economic characteristics of data and AI by:

  1. Reusing the same data set – an economic asset that never depletes and never wears out – across an unlimited number of use cases at zero marginal cost and
  2. Creating data and analytic assets (data products and AI apps, respectively) that appreciate, not depreciate in value, the more they are used.
Slide2-1

Figure 1: Source: “The Economics of Data, Analytics, and Digital Transformation

However, a use case-centric approach must start by identifying and understanding the organization’s strategic business initiatives. Once we understand the organization’s key business initiatives, we can decompose them into supporting use cases. This yields the following benefits:

  • If it is truly a business initiative that has critical and quantifiable business value, then there is likely someone on the executive leadership team who owns that initiative. You have immediately found a “friendly” on your AI and data journey.
  • Use cases tied to that business initiative have a sense of urgency and prioritization, given their ability to enable and optimize that business initiative.
  • Use cases tied to the same business initiative likely share the same data, which allows us to engage the Data Economic Multiplier Effect and crank out a stream of related use cases with quantifiable value more quickly and at less risk

Data Economic Multiplier Effect is the accumulation of attributable and quantifiable “value” from applying a data set against multiple use cases.

Trying to sort out a random dump of use cases is futile – they lack senior management interest, likely don’t share the same data sets, and it’s hard to rally support around a bunch of untethered use cases.

Untethered Use Cases are business and operational use cases not linked to or directly supporting an organization’s key business initiatives.

Tethering or linking use cases to an organization’s key business initiative is the best way to ensure that you focus your AI, data science, and data management resources on those use cases that matter to the business. It is also the best way to ensure the collaboration between the business and data science teams on our AI journey to deliver more relevant, meaningful, responsible, and ethical business and operational outcomes.

The AI Journey from Business Initiative to Use Cases

Let’s start this journey with some definitions to ensure everyone is on the same page.

  • A Use Case is a cluster of Decisions around common KPIs or metrics that deliver a well-defined business or operational outcome supporting an organization’s key business initiative.
  • A Business Initiative is a cross-functional effort typically taking 9-12 months, with well-defined business metrics supporting the organization’s key business and operational objectives.

Put another way, a business initiative is a high-level strategic goal that aligns with the organization’s vision and mission, while a use case is a specific and measurable way to achieve that goal. For example, a targeted business initiative might be to increase customer retention by X%, supported by a series of customer, product, and marketing-specific use cases (Figure 2).

Slide3-1

Figure 2: Tethering Use Cases to a Strategic Business Initiative

The process of going from a business initiative to identifying, validating, valuing, and prioritizing the use cases that support that business initiative is relatively easy. It’s not like I’m asking you to design and train a neural network. It just requires a collaborative engagement methodology that leverages the disciplines of data science, design thinking, and economics to unleash and corral the creative juices of the business and data science teams in support of the business initiative. I think I know where you can learn more about such a methodology (Figure 3).

Slide4-1

Figure 3: “The Art of Thinking Like a Data Scientist

Unleashing the Economic Value of Use Cases

When we combine a use case-centric implementation approach with the unique economic characteristics of data and analytics, you get the Schmarzo Digital Asset Economic Valuation Theorem and my effort to win a Nobel Prize in data economics (Figure 4).

Slide5

Figure 4: Schmarzo Economic Digital Asset Valuation Theorem

The Schmarzo Digital Asset Economic Valuation Theorem is a framework that explains how companies can increase the value of their data and analytics assets by aligning their development with the organization’s sources of value creation. The theorem is based on the concept that data and analytics are unique types of digital assets that have the following characteristics:

  • Data never depletes, never wears out, and can be reused at zero marginal cost.
  • Analytics can be continuously improved through learning and feedback and can be shared and reused across different use cases.
  • Data and analytics can create network effects, where the value of the assets increases as more users and applications use them.

The theorem states that organizations can realize three effects or benefits from the sharing, reuse, and continuous refinement of their data and analytics assets:

  • Marginal Costs Flatten: The marginal costs associated with reusing data and analytic models flatten as the reuse increases, leading to lower costs and higher profits.
  • Marginal Value Grows: Sharing and reusing data and analytic models accelerate future use case time-to-value while de-risking implementation, leading to faster and more reliable value creation.
  • Economic Value Accelerates: Analytic model refinement (improving analytic model accuracy, precision, and recall) lifts the economic value of all associated use cases that use the same analytic module, leading to exponential value growth.

Summary: Business Initiative-linked Use Cases

Tethering or linking use cases to an organization’s key business initiative is the best way to ensure that you focus your AI, data science, and data management resources on those use cases that matter to the business. It is also the best way to ensure the collaboration between the business and data science teams on our AI journey to deliver more relevant, meaningful, responsible, and ethical business and operational outcomes.

This also unleashes the benefits of the Schmarzo Digital Asset Economic Valuation Theorem by:

  • Guiding organizations in prioritizing their data and analytics initiatives based on the potential value creation and reuse opportunities.
  • Encouraging organizations to adopt a use case-by-use case deployment approach that exploits the economies of learning through rapid experimentation and feedback loops.
  • Motivating organizations to break down data silos and avoid orphaned analytics, which are the great destroyers of the economic value of digital assets.
  • Enabling and empowering organizations to measure and optimize their data and analytics assets’ return on investment (ROI) using objective metrics.

No danger here, Will Robinson.

AI is Officially The Word of The Year

We are Living in a Simulation

Talking about AI has been Analytics India Magazine’s bread and butter for over a decade but 2023 marked a year when the subject of academic papers and Silicon Valley boardrooms burst into the mainstream. Thanks in no small part to OpenAI’s ChatGPT, last week, the Collins Dictionary also announced “AI ” as the word of the year. The term was used 4x more times than last year, the publisher said.

“Considered to be the next great technological revolution, AI has seen rapid development and has been much talked about in 2023,” the UK dictionary, published by HarperCollins in Glasgow, noted in the blog post.

The chatter about the word was non-stop as all around the year industry leaders couldn’t stop talking about it be it Google or Microsoft. According to Alex Beecroft, managing director at Collins, AI was chosen as the word of the year because of its presence in daily life, akin to email and streaming platforms. “People use AI in multiple ways now, and it’s everywhere,” he mentioned.

Higher Purpose

The use of the term ‘Artificial Intelligence’ or AI dates back to the 1950s but all these years the technology has only been well known among a small group of people researching and developing it. When ChatGPT was launched last year, the power of this tech was shared with the general population making it more mainstream than ever.

Ever since the AI-powered chatbot took over the internet people have been discussing a spectrum of possibilities from AI destroying humanity to unlocking new possibilities. Technology has severely impacted the way education has functioned up until now, how art was made and even led to the longest-ever strike Hollywood had seen.

But here’s the thing, the term AI has been highly associated with ChatGPT and image generators like DALL-E and Midjourney even though it is just a part of the umbrella. The human-mimicking tech falls under the category of ‘Generative AI’ which has become the center of a growing conversation about what artificial intelligence can achieve.

AI’s power goes beyond just its ability to excite the public’s imagination. Other AI-powered tools are doing far beyond recreating versions of original text or images. For instance, Google DeepMind’s AlphaFold is an AI algorithm for predicting the three-dimensional structure of proteins. The AI lab’s ‘Gift to Humanity’ has proven since its launch how AI can help develop vaccines and help increase the speed of drug discovery.

Similarly, last year the lab solved a 50-year-old math problem using AI. DeepMind published a paper where they presented AlphaTensor, an algorithm able to find faster ways of doing one of the most common algebra operations: matrix multiplication. This development is significant as matrices are relevant in every aspect of our daily lives, from processing images on our phones and recognising speech commands to generating graphics for computer games.

Deja Vous

While such significant developments are being made in AI, ChatGPT remains the most loved byproduct of the technology. One of the reasons is the amount of money being spent on it.

Interestingly, two years ago, Collins Dictionary chose ‘NFT’ as the word of the year. The form of technology saw a similar trajectory as GPTs. The bubble around it officially burst earlier this year as 95% of the NFTs became effectively worthless, a study stated. Out of a total of 73,257 NFT collections examined 69,795 of them held a market capitalisation of precisely zero Ether.

Coming back to AI which is more or less about ChatGPT today, a similar trend is visibly forming. Even OpenAI chief Sam Altman has publicly stated that “It’s wildly overhyped in the short-term.” For now, all fingers are crossed that AI and more particularly ChatGPT does not face the same fate as NFTs after being picked as the word of the year.

The post AI is Officially The Word of The Year appeared first on Analytics India Magazine.

Bollywood Has A Deepfake Problem

When Paul Walker’s sudden demise put the producers of the ‘Fast And Furious’ movie franchise in a dilemma, they used deepfake technology to complete the parts that required his presence. That was ten years ago.

Popping out to steal the limelight for infamous activities and then going obscure like it never existed, deepfake technology has been mostly notorious for how tech can go awry in the wrong hands.

With the recent issue of actors from the Indian movie fraternity who have fallen prey to this AI application, the regulatory aspect of it has resurfaced. However, the irony continues because deepfake is widely used across movie and creative industries.

Indecisive Bollywood

Famous bollywood (Hindi movie industry) actor Anil Kapoor recently safeguarded his digital persona and won a court battle over AI use of his identity, including name, image, voice and others. Thereby, legally cutting off any possibility of creating a deepfake of the actor.

While on one end you have Anil Kapoor safeguarding against the potential use of deepfakes, on the other end you have bollywood biggies such as Shah Rukh Khan, and Salman Khan who have allowed their digital identity to be used via deepfakes for advertisements of consumer products.

Last year, Cadburys ran an ad campaign for small business owners to make original, personalised advertisements, and went on to show Shah Rukh Khan, rather deepfake Shah Rukh, saying the name of the store or brand. Salman Khan had also starred in a Pepsi commercial, where the younger version of the actor was portrayed via deepfake.

Actor Ayushman Khurrana also starred in an advertisement for Wakefit where the actor was showcased as a child with the help of deepfake. Interestingly, all these point to the increased adoption of deepfakes in the advertisement industry, where Bollywood actors are slowly embracing it.

International Acceptance

When Bollywood is still trying to figure out the adoption strategy for deepfakes, internationally, the adoption has been rampant. Late actor Paul Walker’s inclusion in his last movie via deepfake might just be one instance in the mammoth pool of adoption use-cases that Hollywood has already found with this AI tech.

De-ageing has been one of the most prominent use cases where major production houses, including Disney, have tied up with VFX design studios to implement deepfake to de-age actors. From de-ageing Al Pacino and others in ‘Irishman’ to showing a young Mark Hamill as Luke Skywalker in Disney’s Star War series, deepfake has found a comfortable place in Hollywood production.

Sports celebrities are not far behind. Footballer Lionel Messi, has signed deals with big brands allowing them to use his image and likeness in any way they deem necessary. A younger version of Messi was shown using deepfake in an ad for Mastercard.

DeepFake Will Only Get Better

When Rashmika Mandanna’s face was morphed into an Instagram influencer’s body and circulated over the net, a user pointed out the technicality of the video saying that you can find the exact point when the original face changes to become the victim’s face. In the process, hinting at how you can in fact identify a deepfake video at the moment, but it is only going to get more challenging.

While the Hollywood usage of celebrities was with their consent, there are also instances of massive viral clips without proper permission. Clips featuring ‘deepfake Tom Cruise’ were doing the rounds a few years ago. They were created by VFX specialist Chris Ume, who specialises in deepfake celebrity videos, with the help of Tom Cruise impersonator Miles Fisher.

Tom Cruise impersonator, Miles Fisher on the left. Deepfake Tom Cruise on the right. Source: Verge

Creator Ume spoke about the myriad production uses that deepfake has, and also emphasised on how this technology is getting more realistic and easier to make. However, the technology cannot operate by itself. To make a near-perfect one, like the one shown in the image above, Ume took over two months to train the AI base models using a pair of NVIDIA RTX 8000 GPUs on Cruise’s footage, followed by days of further processing for each clip.

Interestingly, Tom Cruise has not sued any of them for such deepfakes, and only resorted to making his account official.

While there are grave misuses of this technology, specific use cases in industries such as movies rally for this technology. While the discussion on deepfake’s usage and ethical implications are still debated in Bollywood, others are quietly embracing it and making the most of it.

The post Bollywood Has A Deepfake Problem appeared first on Analytics India Magazine.

The metaverse has virtually disappeared. Here’s why it’s generative AI’s fault

City in metavere

If a year is a long time in technology, then two years is practically a lifetime. At the back end of 2021, everyone was talking about the metaverse. Two years later and practically no one is talking about it (except, perhaps, Mark Zuckerberg). So, what went wrong?

One of the big issues is hype. It can be tough to deliver on the massive excitement associated with a new technology. Just ask the people who worked on the blockchain.

Also: 4 ways to detect generative AI hype from reality

In many ways, the metaverse's fall from the front page is just par for the course, says Gartner Director Analyst Samantha Searle to ZDNET.

"It's basically going through the Gartner Hype Cycle for Emerging Technologies," she says. "We've had the hype and now we're seeing the reality. The metaverse was capturing people's imagination. But we're still looking for proven use cases that are going to generate value."

Searle's assertion that the metaverse is suffering a familiar fate to other over-hyped technologies is certainly one explanatory factor for the drop in interest in the metaverse. But another huge contributory factor is the rapid rise of artificial intelligence (AI).

Also: Does your business need a chief AI officer?

Twelve months after virtual worlds were dominating the news agenda, a new hyped technology — OpenAI's ChatGPT and a host of similar generative AI tools — bubbled up and stole the media's attention in late 2022. Those who had a dabble found something out quickly: Using generative AI is easy.

Anyone can go online, log in to ChatGPT, and receive instant answers to their questions. From essays to images and onto program code, it's possible to generate content in almost real time.

For most users, these experiences are simple and fun, says Sasha Jory, CIO at Hastings Direct — and that's a big break from some of the hyped technologies of the not-so-distant past.

"This isn't going to be like blockchain, where it's all hype and then it goes nowhere. Generative AI has got legs because it's democratized. Anyone can use it — it's everywhere," she says to ZDNET at the London leg of Snowflake's Data Cloud World Tour. "Your youngest kids can get on ChatGPT and use it. Whereas blockchain, for example, was something you needed to be quite highly skilled at to understand."

Also: Will AI hurt or help workers? It's complicated

Of course, the rapid take up of generative AI isn't the only narrative in this story; there's a whole series of potential concerns, such as hallucinations, plagiarism, and ethics, that need to be dealt with sooner rather than later. But if you want to impress your family and friends with a tool that seems to work like magic, then generative AI is the one.

On the other hand, the metaverse — just like the blockchain before it — feels a bit like a rabbit that's stuck in a magician's hat. Entering the metaverse often isn't as easy as its proponents have promised. Most people have still to enter virtual environments. Users who have, including myself, are often left disappointed by the experience delivered by the hardware and software.

Adobe CIO Cynthia Stoddard told ZDNET her company hasn't found the right use case for the metaverse yet. "We've been experimenting with VR headsets," she says, referring to her company's explorations. "We don't have a solid use case for our business as yet. But I think the headsets have a real application in some verticals, because in the past, I came from transportation and logistics, and I can see a lot of applications in different verticals."

That's a sentiment that resonates with Gartner's Searle, who says it's important not to give up hope just yet. "We do see a lot of potential use cases for the metaverse," she says. She says financial services firms are often keen to explore how the technology can be used to deliver customer services.

Also: 5 essential traits that tomorrow's AI leader must have

Searle refers to JP Morgan, who bet big on the metaverse being a trillion-dollar opportunity and opened the first bank in a virtual world. But even in finance, some organizations are taking a more cautious approach.

"I think the metaverse isn't something that's an immediate priority," says Kavin Mistry, head of digital marketing and personalization at TSB Bank, said to ZDNET. "I do see value in metaverse-type technology, but I think it probably needs to mature, so we can understand the correct way to make it real for customers. And I don't think the market is mature enough yet to make an informed decision about the right way to go."

Gartner's Searle also recognizes that the point of maturity for the metaverse is still some way off. "We know with some of the underlying immersive technologies that there are still things to be addressed in terms of form factor and being comfortable," she says.

Also: AI at the edge: Fast times ahead for 5G and the Internet of Things

From using headsets to creating life-like avatars and building enjoyable virtual environments, Searle says there's a lot of construction work that needs to be finished. However, these building sites could still lead to some exciting developments, especially given that tech giant Apple has started to show an interest in the area.

What's more, recruiter Nash Squared's recent IT leadership report revealed that 26% of digital leaders are at least actively considering the metaverse.

Bev White, CEO at Nash Squared, said in an interview with ZDNET that more organizations are beginning to think about how virtual worlds present a new business opportunity.

"It's definitely alive and kicking," she says. "The metaverse provides a fantastic opportunity to run your business in a way that interacts with your customer base very differently. It actually creates a whole new platform because customers can pick things up, look at them, and experience them."

That's an approach that chimes with Lalo Luna, global head of strategy and insights at Heineken, who says his firm has been busy exploring virtual worlds.

Also: Companies aren't spending big on AI. Here's why that cautious approach makes sense

"One of the main projects that I led during the past 12 months was about understanding the future of socializing — we were understanding how technology is going to take a role in social spaces. And, of course, the metaverse is going to be one of these things." He told ZDNET that perspective is crucial — the metaverse isn't going to be built overnight.

Just as generative AI is the visible manifestation of years of research and development work, Luna expects the importance of virtual worlds to grow during the rest of this decade. And he says smart professionals are already exploring how to put a stake in the ground.

"Heineken was one of the first brewers that opened a bar and a brewery in the metaverse. We are really taking these developments very seriously — we are trying to be where we need to be," he says. "Of course, we are exploring, we are in an experimentation phase. Some of our big brands are really into technology. They are jumping fast into different kinds of platforms, and not just the metaverse. Heineken as a company has a curious and experimental mind — and I think that's crucial to success."

Artificial Intelligence

Why Microsoft temporarily blocked ChatGPT from employees on Thursday

Microsoft logo on phone

Microsoft, one of OpenAI's largest investors, restricted ChatGPT access for its employees on Thursday, citing "security and data concerns." According to CNBC, Microsoft employees were prohibited from accessing ChatGPT, a popular AI chatbot created by OpenAI, in which Microsoft has invested over $13 billion.

CNBC reports that an internal website told employees, "Due to security and data concerns, several AI tools are no longer available for employees to use." Microsoft employees were subsequently unable to access ChatGPT, which was listed as one of the tools banned for employee use.

Also: Will AI hurt or help workers? It's complicated

Microsoft reinstated ChatGPT access and removed it from the list of prohibited tools after CNBC published its report, later saying, "We restored service shortly after we identified our error. As we said previously, we encourage employees and customers to use services like Bing Chat Enterprise and ChatGPT Enterprise that come with greater levels of privacy and security protections."

The close relationship between the two companies makes the news that Microsoft banned its employees from using ChatGPT all the more shocking, as Microsoft CEO Satya Nadella shared the stage with OpenAI CEO Sam Altman during OpenAI's first developers conference this week.

Also: 6 ways business leaders are exploring generative AI at work

Microsoft has partnered with OpenAI for almost a year, making substantial investments into the AI company's technology and even incorporating it into its Bing Chat feature. Bing Chat is backed by GPT-4 and works as an AI chatbot that can access the internet, bringing online searches to a new level.

GPT-4 is a more powerful model than the one powering the free version of ChatGPT and is only accessible with a ChatGPT Plus subscription from OpenAI or, alternatively, through Microsoft's Bing Chat.

Also: 5 essential traits that tomorrow's AI leader must have

DALL-E 3, another new technology by OpenAI, is also largely incorporated into Microsoft's AI tools, as users can access it to generate images using AI through Bing Chat or the Bing Image Creator.

ChatGPT has taken over the AI world since it launched almost a year ago and currently has over 100 million users. Since people discovered its user-friendly interface and powerful capabilities in content creation, different reports have emerged of people using ChatGPT to write books and college papers, and its use has been largely scrutinized.

Some companies, including Samsung, have restricted their employees from using the AI chatbot in company devices after employees were found to be giving ChatGPT confidential code to debug.

Artificial Intelligence

GitHub Universe: Open Source Trends Report and New AI Security Products

GitHub Copilot app on smartphone with AI security on background.
Image: Adobe/sdx15

At the GitHub Universe conference held in San Francisco and virtually on Nov. 8 and Nov. 9, 2023, the company revealed its new open source trends report as well as changes to GitHub Copilot and AI enhancements for GitHub Advanced Security.

GitHub Copilot and GitHub Advanced Security are available globally. However, some GitHub services, including Copilot, are subject to U.S. trade controls and are not available in the sanctioned countries listed here.

Jump to:

  • Generative AI is popular among open source projects
  • Git trends toward cloud-native applications at scale
  • Trends in the GitHub developer community
  • GitHub Copilot Chat and GitHub Copilot Enterprise revealed
  • Additional AI features added to GitHub Advanced Security

Generative AI is popular among open source projects

Open source generative AI projects joined GitHub’s list of the top 10 most popular open source projects by contributor count in 2023. In 2022, about 17,000 developers on GitHub worked on generative AI projects; in 2023, that number rocketed to around 60,000. AI projects are becoming more mainstream, GitHub said.

More organizations are likely to start using pre-trained AI models in the future as developers become more familiar with them, GitHub predicted.

Git trends toward cloud-native applications at scale

GitHub found developers are increasingly using the Git version control system for declarative languages using Git-based infrastructure as code workflows.

The study also found greater standardization in cloud deployments and a sharp increase in the rate at which developers were using Dockerfiles and containers, infrastructure-as-code and other cloud-native technologies. Use of Hashicorp Configuration Language (HCL), which is an indicator for operations and infrastructure-as-code work, grew 36% year-over-year.

Trends in the GitHub developer community

The number of new developers on GitHub grew by 26%, with India having the fastest-growing population of developers. GitHub defines a developer as anyone with a non-spam GitHub account.

Commercially-backed open source projects draw attention

Commercially-backed open source projects had the largest number of contributions and the largest number of first-time contributors. The number of private projects grew 38% year over year.

Securing dependencies and branches are popular projects

In terms of security in open source, more developers are turning to automation to secure dependencies, and open source maintainers are paying close attention to protecting their branches.

Front-end development shows promise

Front-end development is a rapidly growing type of project among open-source developers.

GitHub Copilot Chat and GitHub Copilot Enterprise revealed

At GitHub Universe, the company announced GitHub Copilot Chat (Figure A), which is a generative AI assistant that explains code in natural language, and GitHub Copilot Enterprise. GitHub Copilot Chat will be available in December 2023 to customers with existing individual or organization-wide GitHub Copilot subscriptions.

Figure A

Screenshot of Github Copilot chat explain.
GitHub Copilot Chat explains code in natural language. Image: GitHub

GitHub Copilot Enterprise, customized for business use, is coming in February 2024 at a price of $39 USD per user per month. Compare this to Copilot Business, which costs $19 per month and is available now.

Additional AI features added to GitHub Advanced Security

Three more AI-powered features are coming to GitHubAdvanced Security: code scanning autofix, secret scanning for generic secrets and a regular expression generator.

SEE: GitHub isn’t the only version control and collaboration platform. See GitHub alternatives that are flourishing in 2023. (TechRepublic)

“Developers need the ability to proactively secure their code right where it’s created,” GitHub VP of product management, Asha Chakrabarty, and director of product marketing at GitHub security lab and platform security, Laura Paine, wrote in a blog post.

Code scanning autofix

Code scanning will now propose AI-generated fixes right in the pull request, enabling developers to instantly fix vulnerabilities while they code; this will lead to faster remediation time. AI-generated fixes can be created for CodeQL, JavaScript and TypeScript alerts. This works by GitHub querying a large language model in the background to find fixes for any new alerts, which are then posted as code suggestions within the pull request.

Autofix is available for code scanning within GitHub Advanced Security now.

Secret scanning

Secret scanning with generative AI, which is now in limited public beta, is designed to reduce false positives that often crop up when searching for possibly active leaked passwords (Figure B).

Figure B

Screenshot of GitHub secret scanning.
Secret scanning alerts users to a password that may have been exposed. Image: GitHub

Regular expression generator

The regular expression generator enhances developers’ options when it comes to secret scanning, letting them create custom patterns with regular expressions created with a few natural-language queries sent to the generative AI. It is designed to make writing regular expressions faster, and enables developers to perform dry runs in real time to make sure everything works before saving the pattern.

Regular expression generation is available now.

More new features in GitHub Advanced Security

Other new features of GitHub Advanced Security include authoring custom patterns with generative AI and a new security overview dashboard. Interested security personnel can join a waitlist for these features.

Sure, real-time data is now ‘democratized,’ but it’s only a start

Data concept

Real-time data seems to be everywhere — in augmented reality, digital twins, 5G, IoT, AI, machine learning, wearables, and beacon technology. One can be forgiven for thinking that today's enterprises are streaming real-time data across every vital task area. We're getting there — thanks in large part to many open-source solutions such as Apache Flink, Kafka, Spark, and Storm, as well as cloud-based platforms. However, there is still a lot of work that needs to be done before we reach the point at which data moves through and between organizations at light speed or something close to it.

First, a level set, courtesy of IDC's John Rydning: "Often, the terms streaming data and real-time data are used in conjunction with each other and sometimes interchangeably. While not all streaming data created is real time, and not all real-time data is streamed, organizations indicate that over two-thirds of streaming use cases require ultra-real-time or real-time data."

Also: Every AI project begins as a data project, but it's a long, winding road

There's even ultra-real-time data in use — and companies are anxious to make this work. "Understanding the value of capturing and processing real-time data is growing at the fastest pace in recent times," says Avtar Raikmo, director of engineering at Hazelcast. "With platforms taking complexity away from the individual user or engineer, it has accelerated adoption across the industry. Innovation such as SQL support, help make it democratized and provide ease of access to the vast majority rather than a select few."

There is a wide range of use cases, including "compute at the edge for audio and video streaming, computer vision for AI and machine learning processing, or even active noise-canceling headphones," Raikmo says. Another emerging use case is digital twins, especially for mobility. "Being able to capture real-time data and telemetry from cars, trucks or rockets enables organizations to model scenarios as they unfold. Digital twins can be used to optimize real-world routes taken, energy used or assisted driving improvements. In the world of sport, Formula 1 strategists determine the optimum pit-stop and tire compounds to maximize race performance."

Still, there are many technical and organizational issues standing in the way of full real-time — or ultra-real-time data realities. "Real-time data deployments typically use higher performance technologies that cater to the large volumes and fast analysis that are required to make instant decisions," says Emma McGrattan, senior VP of engineering and product for Actian. "For very large volumes that some verticals, like financial services, tend to generate, moving to real time will require investment in additional resources for hardware, software, and network components."

Also: How AI reshapes the IT industry will be 'fast and dramatic'

Investments are needed to "increase availability and reliability of data infrastructure and services," McGrattan says. "For lower volumes, the existing infrastructure is likely able to be usable, with modifications to the applications to do the real-time analysis and deployment."

The process of capturing, visualizing, and storing real-time data requires "substantial investments in infrastructure components that are capable of handling heavy and complex data streams," says Rakesh Jayaprakash, head of product management at ManageEngine and Zoho. "This is particularly true when real-time data streams require some level of pre-processing. Unfortunately, many organizations, particularly SMBs, lack the necessary infrastructure to handle such intensive processing."

Many companies' infrastructures aren't ready, and neither are the organizations themselves. "Some yet to understand or see the value of real-time while others are all-in, with solutions that were designed for streaming throughout the organization," says Raikmo. "Combining datasets in motion with advanced techniques such as watermarking and windowing, is not a trivial matter. It requires correlating multiple streams, combining the data in memory and producing merged stateful result sets, at enterprise scale and resilience."

Also: The real-time revolution is here, but it's unevenly distributed

The good news is not every bit of data needs to be streaming or delivered in real time. "Organizations often fall into the trap of investing in resources to make every data point they visualize be in real time, even when it is not necessary," Jayaprakash points out. "However, this approach can lead to exorbitant costs and become unsustainable."

"While visualizing real-time data is more appealing than analyzing data that is a few minutes old, you must carefully assess the cost-benefit ratio and ROI associated with building real-time data streams and visualizations," says Jayaprakash. "Moreover, organizations should exercise due diligence in selecting the metrics they wish to stream in real time."

IDC's Amy Machado makes the case for carefully considering what needs to be delivered in real-time: "I always say, 'Let the use case lead,'" she writes in a blog post. "It should direct how you think about real-time architecture, which ideally, is an expansion of your existing framework to avoid creating data silos."

Also: Business leaders continue to struggle with harnessing the power of data

Machado outlines key questions to ask about real-time data delivery:

  • "What are the business benefits we are hoping to achieve?"
  • "What insights do we need to achieve those goals?"
  • "Who needs those insights and where do they need them?"
  • "What other systems might we need to integrate with for context or to operationalize insights?"

To optimize real-time data investments, "carefully select metrics that truly require real-time reporting," Jayaprakash advises. "The complex nature of the infrastructure needed to operate and maintain real-time data streams introduces potential points of failure, necessitating a dedicated staff for troubleshooting and maintenance. To mitigate data continuity issues resulting from stream failures, you need to implement fail-safe mechanisms, which adds to overall costs."

Artificial Intelligence

AI makes you worse at what you’re good at

AI makes you worse at what you’re good at Haje Jan Kamps 9 hours

Welcome to Startups Weekly. Sign up here to get it in your inbox every Friday.

If you’ve been following along with this newsletter, you’ll have noticed that I’ve been a little bit curious about AI — especially generative AI. I’m likely not the first person to make this observation, but AIs are extremely, painfully average. I guess that’s kind of the point of them — train them on all knowledge, and mediocrity will surface.

The trick is to only use AI tools for stuff that you, yourself, aren’t very good at. If you’re an expert artist or writer, it’ll let you down. The truth, though, is that most people aren’t great writers, and so ChatGPT and its brethren are going to be a massive benefit to white-collar workers everywhere. Well, until we collectively discover that a house cleaner has greater job security than an office manager or a secretary, at least.

On that cheerful note, let’s sniff about in the startup bushes and see what tasty morsels we can scare up from the depths of the TechCrunch archive from the past week. . . .

Okay, fine, let’s start with AI

Image of a robot with shopping cart on an orange background.

Image Credits: Kirillm (opens in a new window) / Getty Images

I know, this happens every damn week: I start with the intention of writing this newsletter without going up to my eyelashes into the AI morass, and every week, y’all keep reading our AI news as if your livelihood depends on it. Because, well, it’s entirely possible it does, I suppose.

The GPT Store, introduced by OpenAI, enables developers to create custom GPT-based conversational AI models and sell them in a new marketplace. This initiative is designed to expand the accessibility and commercial use of AI, similar to how app stores revolutionized software distribution. Developers can not only build but also monetize their AI creations, opening up a new avenue for innovation and entrepreneurship in the field of artificial intelligence. Of course, that little update — and the platform now natively being able to read PDFs and websites — is a substantial threat to startups that had previously filled this gap in ChatGPT’s offerings, especially those whose business models are based on such features. It’s a reminder that building a business around another company’s API without a sustainable, stand-alone product is, perhaps, not the shrewdest business move.

AI is, of course, not just for startups. During Apple’s Q4 earnings call, the company’s CEO, Tim Cook, emphasized AI as a fundamental technology and highlighted recent AI-driven features like Personal Voice and Live Voicemail in iOS 17. He also confirmed that Apple is continuing to develop generative AI technologies — tellingly, without revealing specifics.

Heinlein would be horrified: Elon Musk announced that Twitter’s Premium Plus subscribers will soon have early access to xAI’s new AI system, Grok, once it exits early beta, positioning the chatbot as a perk for the platform’s $16/month ad-free service tier.

Brother, can you spare a GPU?: AWS introduced Amazon Elastic Compute Cloud (EC2) and Capacity Blocks for ML, a new service that enables customers to rent Nvidia GPUs for a set period, primarily for AI tasks like training or experimenting with machine learning models.

From zero to AI founder in one easy bootstrap: In “How to bootstrap an AI startup” on TC+, Michael Koch advises founders on maintaining control over their startup’s strategy and product by bootstrapping — yes, even in the oft-capital-intensive world of AI startups.

The rocky ocean of venture-backed startups

An illustration depicting the Wework logo looking battered and wearing bandages, meant to suggest financial hardship

Image Credits: Darrell Etherington with assets from Getty under license

WeWork, once a high-flying startup valued at $47 billion, has filed for Chapter 11 bankruptcy protection, highlighting a staggering collapse. The company, which has over $18.6 billion of debt, received agreement from about 90% of its lenders to convert $3 billion of debt into equity in an attempt to improve its balance sheet and address its costly leases. On TC+, Alex notes what we kinda knew all along: that the core business just didn’t make sense.

In other venture news . . .

Ex-Twitter CEO raises third venture fund: 01 Advisors, the venture firm founded by former Twitter executives Dick Costolo and Adam Bain, has secured $395 million in capital commitments for its third fund, aimed at investing in Series B–stage startups focused on business software and fintech services.

Happy 10th unicornaversary: Alex reflects on the tenth anniversary of the term “unicorn,” which was initially coined right here on TechCrunch, to describe startups valued at over $1 billion.

You get a chip! You get a chip!: In response to a shortage of AI chips, Microsoft is updating its startup support program to offer selected startups free access to advanced Azure AI supercomputing resources to develop AI models​​.

Let’s talk Sam Bankman-Fried

Illustration of Sam Bankman-Fried aka SBF

Image Credits: Bryce Durbin / TechCrunch

Look, I’m not going to lie, I think most crypto is dumb, and I’ve seen only a handful of startups that use blockchains in a way that makes any sense whatsoever — most of them would have done just fine with a simple database — so I’ve been following Jacquelyn’s coverage of Bankman-Fried’s trial with a not insignificant amount of schadenfreude. It’s human to make mistakes, and startup founders are human, but if you’re defrauding the fuck out of people, you deserve all the comeuppance you can get.

Sam Bankman-Fried was the co-founder and CEO of the cryptocurrency exchange FTX and the trading firm Alameda Research (named specifically to not sound like a crypto company). He has been found guilty on all seven counts of fraud and money laundering.

The charges were related to a scheme involving misappropriating billions of dollars of customer funds deposited with FTX and misleading investors and lenders of both FTX and Alameda Research. After the five-week trial, the jury spent just four hours to reach its verdict.

The collapse of FTX and Alameda Research, which led to the indictment of Bankman-Fried about 11 months ago by the U.S. Department of Justice, was significant, with the executives allegedly stealing over $8 billion in customer funds.

Sentencing will happen next March, but if he gets smacked with the full weight of his actions, he will face a total possible sentence of 115 years in prison.

Jacquelyn did a heroic job covering the trial for TechCrunch, and it’s worth taking an afternoon to read through it all — the details are mind-boggling.

Top reads on TechCrunch this week

The house sometimes wins: Mr. Cooper, a mortgage and loan company, experienced a “cybersecurity incident” that led to an ongoing system outage. The company says it has taken steps to secure data and address the issue​.

Can’t think of any downsides of the Hindenburg: The world’s largest aircraft, Pathfinder 1, is an electric airship prototype developed by LTA Research and funded by Sergey Brin. It was unveiled this week, promising a new era in sustainable air travel.

Arrival’s departure: The EV startup Arrival, which aimed to revolutionize electric vehicle production with its micro-factory model, is now facing severe operational challenges, including multiple layoffs, missed production targets, and noncompliance with SEC filing requirements, resulting in a plummet from a $13 billion valuation.

Here’s how to create your own custom chatbots using ChatGPT

OpenAI ChatGPT

At its Dev Day event on Monday, ChatGPT creator OpenAI announced that subscribers would soon be able to create their own custom chatbots known as GPTs. Today the company has officially launched that feature.

Available to paying ChatGPT Plus and Enterprise subscribers, the custom GPT option lets you design GPT chatbots simply by telling the GPT Builder what you want.

Also: The best AI chatbots: ChatGPT and other noteworthy alternatives

The idea behind the new custom GPTs is to help subscribers move beyond the standard, general-purpose ChatGPT model by devising more specialized models for more specific areas and tasks. In a blog post, OpenAI cited a couple of examples: You could create a chatbot focused on teaching computer science in a school or one that lets you design marketing logos and campaigns.

You can keep your custom GPTs private for your own use, share them publicly, or even deploy them within an organization. Later this month, OpenAI will open a GPT store through which subscribers will be able to publish custom GPTs with commercial value and receive a cut of the sales.

Devising a custom GPT requires no coding knowledge or skills. Just start a conversation with the GPT Builder and explain what you want the GPT to do. Give it a name, description, and instructions.

You then select which ChatGPT capabilities your GPT should possess — Web Browsing, DALL-E Image Generation, and/or Code Interpreter. You can even integrate real-world data to connect your GPT to external databases, email inboxes, and e-commerce systems.

To try out the custom GPT, I signed into ChatGPT with my subscriber account. To do this, you click your name at the bottom of the left pane and select My GPTs. On the My GPTs page, I selected the option at the top for Create a GPT.

Also: 6 helpful ways to use ChatGPT's Custom Instructions

The GPT Builder popped up asking me what type of GPT I'd like to make. You can keep it simple or go wild. I kept it fairly simple by asking it to create a GPT that could summarize an uploaded document. The GPT Builder kicked off the process and suggested naming my GPT "Summary Sage." I gave the OK to that name though I could have easily chosen my own.

Next, the builder generated a picture for my GPT showing a magnifying glass on top of an open book. I asked it to revise the image by replacing the open book with a printed document, which it did.

The builder then asked me what types of documents I'd want the GPT to handle. After answering that I wanted it to analyze news articles and technical papers in PDF or Word format, I could then continue responding to questions to flesh out the GPT or I could just save it. After I clicked Save, the builder asked if I wanted my GPT to be private, available to anyone with a link, or public. I kept it private and then saved it.

My new GPT appeared on the screen. To try it out, I uploaded a PDF of a book about tips for using Excel. The builder generated two different summaries, asking me to choose the one I liked better. I could then give the response a thumbs up or thumbs down or generate a different response. Back at the My GPT screen, I was to access my new GPT to run it, edit it, or delete it.

Also: ChatGPT is no longer as clueless about recent events

Will custom GPTs catch on among ChatGPT subscribers? The process for creating one is certainly quick and easy, especially with no programming savvy necessary. I could see people experimenting with different GPTs not just for themselves but for businesses and organizations that want to automate certain tasks with a dose of AI.

Artificial Intelligence