Databricks Raises $500 Million in Pre-IPO Round

Databricks Raises $500 Million in Pre-IPO Round September 14, 2023 by Alex Woodie

(T. Schneider/Shutterstock)

Databricks hasn’t gone public yet–the market just hasn’t been right this year for that–but it did the next best thing today when it raised $500 million in a growth round of funding. The Series I round included investments by GPU giant Nvidia and valued the Spark backer at $43 billion.

Databricks has been building market momentum for years, first as an efficient way to run Apache Spark big data workloads in the cloud, and now as a way to build and scale artificial intelligence (AI) applications.

The San Francisco company hasn’t hid its desire to eventually go public, but the combination of factors like rampant inflation in 2022 and bank runs in early 2023, among others, have put a damper on the market for initial public offerings (IPOs).

“The markets are closed,” Databricks CEO Ali Ghodsi told Bloomberg, which broke the news on the new funding round last month. “If they had not been closed, we would have already been public.”

There have been rumors of cash flow issues at the San Francisco company since it spent $1.3 billion to purchase MosaicML, a GenAI company, earlier this year. Databricks envisions MosaicML, which develops own language models that customers can train themselves for a fraction of the cost it takes to train a model like GPT-3.5, serving as an AI “factory” that can churn out models en masse for customers.

Ali Ghodsi is the CEO and co-founder of Databricks

The Series I was led by T. Rowe Price Associates, and included a number of existing investors, including Andreeson Horowitz and others. New investors participating in the round included include Nvidia, which develops the GPUs that are so essential to training AI models, as well as Capital One Ventures.

Enterprise data is “a goldmine for generative AI,” said Jensen Huang, founder and CEO of Nvidia, in a press release. “Databricks is doing incredible work with Nvidia technology to accelerate data processing and generative AI models.”

Databricks released some additional figures about itself. For instance, it said that it crossed the $1.5 billion revenue run rate. That’s notable because it comes just a few months after it said it crossed the $1 billion annual run rate.

Crossed $1.5 billion revenue run rate at over 50% revenue year-over-year growth [with the second quarter representing the strongest quarterly incremental revenue growth in Databricks’ history. The company has more than 10,000 customers, including more than 300 that are spending at the rate of at least $1 million per year with Databricks.

A $500 million investment sounds like a lot, but it’s nowhere close to being Databricks’ biggest round. Its previous round cleared $1.6 billion in August 2021, while the Series G round just six months before that brought in $1 billion in funding.

The new round brings Databricks’ valuation to $43 billion. That’s up from $38 billion in August 2021. The company says the Series I establishes the company’s price per share at $73.50. Echoes of a Wall Street opening bell can’t be far.

Related

About the author: Alex Woodie

Alex Woodie has written about IT as a technology journalist for more than a decade. He brings extensive experience from the IBM midrange marketplace, including topics such as servers, ERP applications, programming, databases, security, high availability, storage, business intelligence, cloud, and mobile enablement. He resides in the San Diego area.

The 5 Best AI Tools For Maximizing Productivity

The 5 Best AI Tools For Maximizing Data Science Productivity

Efficiency and productivity are essential when it comes to data science and processing the massive datasets involved. As these datasets rapidly balloon in size and complexity, the tools we use to manage and analyze them must not only keep pace—but also propel us forward; none more so than AI.

While most apps and tools primarily focus on aspects like data analysis, transcription, and IT operations, the AI landscape has touched our everyday workflows, too. You can now merge PDF files, rearrange spreadsheets and do just about any mundane task in a matter of seconds. So, when you consider the following tools, think broader than just data science.

In this article, we look at the following 5 varied AI tools for maximizing the productivity of data scientists:

  • Assembly AI
  • DataRobot
  • H2O.ai
  • Hugging Face
  • BigPanda

Assembly AI

Assembly AI Assembly AI Key Points

  • Transcription and Speech Recognition
  • Customizable API solutions
  • Accuracy and scalability

✅ Pros

  • Exceptional accuracy even in challenging audio conditions
  • Flexible and robust API for developers
  • Scalable for both small and large enterprises

❌ Cons

  • Requires integration work
  • Might be overkill for simple transcription tasks

Assembly.ai is hailed as one of the leading solutions in transcription and speech recognition, focusing on delivering highly accurate transcriptions even in noisy environments. Their API provides developers with a customizable and flexible framework, ensuring seamless integration with different platforms and tools.

What makes Assembly.ai stand out is its commitment to scalability, ensuring that organizations of all sizes can benefit from its functionalities. It boasts a hybrid system that combines deep learning capabilities with traditional speech recognition techniques, making it suitable for both real-time and batch-processing tasks.

Beyond transcription, their suite offers a range of audio processing tools, including keyword spotting and speaker diarization. As companies increasingly rely on voice data for insights and analytics, Assembly.ai's role becomes even more pivotal. Their commitment to continuous development promises enhancements in both speed and precision.

DataRobot

DataRobot DataRobot Key Points

  • Cloud-based Automated Machine Learning (AutoML) Tool
  • Model interpretability and deployment
  • User-friendly interface

✅ Pros

  • Simplifies the machine learning process
  • Robust model interpretability features
  • Suitable for both novices and experts

❌ Cons

  • Can be costly for smaller companies
  • Advanced users might seek more customization options

DataRobot has risen to prominence by pioneering today's automated machine learning, or AutoML, space.

Its platform allows data professionals to swiftly build, tune, and deploy predictive models without the intricate details associated with manual modeling. Users can get recommendations on the best models to use by simply uploading a dataset, with the platform automatically handling feature engineering and hyperparameter tuning.

Its cloud-native architecture helps ensure that models can be deployed anywhere you need them with ease. Furthermore, with a focus on collaboration, teams can share insights, models, and findings, amplifying productivity.

Beyond its automation capabilities, DataRobot emphasizes model interpretability. This ensures that the models produced aren't just ‘black boxes and that their workings can be understood and explained. With its user-friendly interface, even those with minimal machine learning experience can harness the power of sophisticated algorithms for their data projects.

H2O.ai

H2O.ai H2O.ai Key Points

  • Open-source AI platform
  • Support for a wide range of algorithms
  • Scalability and integration capabilities

✅ Pros

  • Cost-effective due to its open-source nature
  • Broad algorithm support
  • High scalability and compatibility

❌ Cons

  • Might have a steeper learning curve for beginners
  • Less user-friendly than some competitors

H2O.ai offers a comprehensive, open-source platform catering to various AI and machine learning needs. It supports a broad spectrum of algorithms, from deep learning to generalized linear models. Data scientists can access and experiment with these without requiring licenses or extra excessive costs.

The true strength of H2O.ai lies in its scalability, catering to tasks from small datasets on personal computers to big data analytics on enterprise-level clusters. Its platform integrates seamlessly with popular data platforms like Hadoop and Spark, ensuring a cohesive workflow in any environment.

Moreover, they offer courses and resources to assist users, ensuring even those new to the field can get started swiftly. Their continuous innovation, based on user feedback, ensures that they're consistently addressing the evolving needs of the data science community.

Hugging Face

Hugging Face Hugging Face Key Points

  • Leading platform for Natural Language Processing (NLP)
  • Expansive model library
  • Active community and frequent updates

✅ Pros

  • Comprehensive resources for NLP tasks
  • Strong community support and contributions
  • Frequent updates and a growing model library

❌ Cons

  • Primarily focused on NLP, limiting versatility
  • Might be overwhelming for beginners

Hugging Face has established itself as the go-to platform for natural language processing (NLP) tasks. Their Transformers library is a repository of state-of-the-art models in NLP, making cutting-edge technology accessible to developers and data scientists alike. From chatbots to sentiment analysis, their tools cover a wide array of applications.

The continuous contributions from an active community ensure that Hugging Face remains at the forefront of advancements in NLP. They also provide ample resources, including various pre-trained models, making it easier for users to bootstrap their NLP-powered projects.
On top of all this, their community-first approach means frequent updates, ensuring users are always equipped with the latest innovations in NLP and LLM technology.

BigPanda

BigPanda BigPanda Key Points

  • AI-driven IT operations
  • Centralized event management
  • Real-time insights and analytics

✅ Pros

  • Streamlines IT operations with AI
  • Provides centralized event management
  • Comprehensive insights and analytics

❌ Cons

  • Primarily caters to large-scale IT operations
  • May require initial setup and integration efforts

BigPanda provides a platform that bolsters IT operations using artificial intelligence. It efficiently consolidates IT alerts into high-level incidents, enabling teams to identify and address critical issues faster. By centralizing event management, BigPanda offers a holistic view of the operational landscape, preventing the chaos of scattered notifications.

The platform also provides real-time insights, allowing teams to understand root causes and correlations quickly. With its analytics, teams can prioritize tasks and preemptively address potential issues. BigPanda seamlessly integrates with a plethora of IT systems, making it a central hub for all operational needs.

The tools we choose to use can ultimately make or break our productivity, especially in the complicated intersection of data science and IT operations.

From the meticulous transcription abilities of Assembly.ai to the IT operational wizardry of BigPanda, advancements in AI tools are shaping the future of how researchers in data science work and manage datasets.

Whether you're diving deep into Natural Language Processing with HuggingFace or seeking to streamline your machine learning processes with DataRobot and H2O.ai, the market for innovative AI-driven solutions is rich with a vast selection of options tailored to various needs.

Choosing the right tool for your data science needs hinges on recognizing your own specific requirements, potential budget constraints, and possible integration capabilities. As AI tools continue to improve, staying informed and adaptable at all times is vital.

Nahla Davies is a software developer and tech writer. Before devoting her work full time to technical writing, she managed—among other intriguing things—to serve as a lead programmer at an Inc. 5,000 experiential branding organization whose clients include Samsung, Time Warner, Netflix, and Sony.

More On This Topic

  • 7 AI-Powered Tools to Enhance Productivity for Data Scientists
  • Top 6 Tools to Improve Your Productivity on Snowflake
  • The Seven Best ELT Tools for Data Warehouses
  • 8 Best Python Image Manipulation Tools
  • 7 Best Tools for Machine Learning Experiment Tracking
  • A List of 7 Best Data Modeling Tools for 2023

Mind the trust gap: Data concerns prompt customer caution over generative AI

ailockgettyimages-1437761046

Generative artificial intelligence (AI) is being championed as essential for organizations to ensure their market relevance, but some remain hesitant to take the plunge over concerns about data and trust.

These issues are especially pertinent for businesses that operate in sectors with stringent data governance rules and large customer bases, pushing them to bide their time in adopting generative AI tools despite their touted benefits.

Also: Businesses need pricing clarity as generative AI services hit the market

The ability to generate sales reports via a prompt, for instance, instead of having to manually fiddle with spreadsheets, offers an interesting potential for generative AI tools such as Salesforce's Einstein Copilot, said Sarwar Faruque, head of development operations at Jollibee Foods Corporation. The Philippine restaurant chain operator uses Salesforce's Heroku to build its applications and Mulesoft as the middleware to connect its applications, including ERP and order management systems.

Jollibee has 15,000 employees and operates almost 4,000 stores worldwide across 34 countries. Its applications run predominantly on the cloud, so it does not maintain its own data centers, with the exception of a small intranet.

Faruque also sees potential for AI to be used in manufacturing, where it can drive efficiencies in its production pipeline and assembly. For instance, AI can help monitor food quality and forecast demand.

His interest in the potential use of AI, however, remains limited to backend operations. Faruque is adamant about keeping generative AI away from customer interactions and customer-facing operations — for now, at least.

With the technology still in its infancy, there still is a lot that needs to be understood and worked through, he noted.

"We see the output [and responses] it generates, but we don't really understand how [it got to the answer]," he said. "There's this black box…it needs to be demystified. I want to know how it works, how it arrived at its response, and whether this answer is repeatable [every time the question is asked]."

Also: Why companies must use AI to think differently, and not simply to cut costs

Currently, this is not the case, he said, adding that the risk of hallucination also is a concern. And in the absence of a security incident, little is known about whether there are any inherent cybersecurity issues that need to be resolved, he noted.

"Right now, there's just a lot of marketing [hype]," Faruque said, adding that it was not enough to simply talk about "trust" without providing details about what exactly that meant.

He urged AI vendors to explain how their large language models are formed, what data they consume, and what exactly they do to generate responses. "They need to stop acting like it's magic [when] there's a code running it and there's science behind it," he said. "Help us understand it [because] we don't like adopting a technology that we don't have a solid understanding of."

He underscored the need for accountability and transparency, alongside guarantees that customers' data used to train AI models will not be made public. This is critical, especially for organizations that need to comply with data privacy regulations in their local jurisdiction.

Also: Measuring trust: Why every AI model needs a FICO score

Until these issues are ironed out, he said he is not willing to put his own customers' data at risk.

Trust also is something Singapore's Ministry of Trade and Industry (MTI) takes seriously, specifically, in terms of data privacy and security. Ten government agencies sit under the ministry, including EDB and the Singapore Tourism Board.

In particular, the ministry's data must be retained in Singapore, and this is proving to be a big hurdle in ensuring data security and governance, said MTI's ministry family CIO Sharon Ng. It means any AI and large language models it uses should be hosted in its own environment, even those run by US vendors such as Salesforce's Einstein Copilot platform.

Like Faruque, Ng also stressed the need for transparency, in particular the details of how the security layer operates, including what kind of encryption is used and whether data is retained, she noted.

Also: How trusted generative AI can improve the connected customer experience

Her team currently is exploring how generative AI tools, including Salesforce's, can benefit the ministry, which remains open to using different AI and large language models that are available in the market. This would be less costly than building its own models and would shorten the time to market, she said.

The use of any AI model, however, still would be subject to trust and security considerations, she noted. MTI currently is running generative AI pilots that aim to improve operational efficiencies and ease work tasks across its agencies.

For Singapore telco M1, delivering better customer service is the clear KPI for generative AI. Like MTI and Jollibee, though, data compliance and trust are critical, said Jan Morgenthal, chief digital officer of M1. The telco currently is running proof-of-concepts to assess how generative AI can enhance interactions its chatbot has with customers and whether it can support additional languages, other than English.

This means working with vendors to figure out the parameters and understand where the large language and AI models are deployed, Morgenthal said. Similar to MTI and Jollibee, M1 also has to comply with regulations that require some of its data, including those hosted on cloud platforms, to reside in its local market.

Also: Ahead of AI, this other technology wave is sweeping in fast

This necessitates the training of AI models to be carried out in M1's network environment, he said.

The Singapore telco also needs to be careful about the data used to train the models and responses generated, which should be tested and validated, he said. These not only need to be checked against guidelines stipulated by the vendor, such as Salesforce's Trust Layer, but also against the guardrails that M1's parent company Keppel has in place.

Addressing the generative AI trust gap

Such efforts will prove critical amid falling trust in the use of AI.

Both organizations and consumers now are less open to the use of AI than they were before, according to a Salesforce survey released last month. Some 73% of business buyers and 51% of consumers are receptive to the technology being used to improve their experiences, a drop from 82% and 65%, respectively, in 2022.

And while 76% of customers trust businesses to make honest claims about their products and services, a lower 57% trust them to use AI ethically. Another 68% believe AI advancements have made it more important for companies to be trustworthy.

The trust gap is a significant issue and concern for organizations, said Tim Dillon, founder and director of Tech Research Asia, pointing to the backlash Zoom experienced when it changed its Terms of Service, giving it the right to use its users' video, audio, and chat data to train its AI models.

Also: AI, trust, and data security are key issues for finance firms and their customers

Generative AI vendors would want to avoid a similar scenario, Dillon said in an interview with ZDNET, on the sidelines of Dreamforce 2023 held in San Francisco this week. Market players such as Salesforce and Microsoft have made efforts to plug the trust gap, which he noted was a positive step forward.

Apart from addressing trust issues, organizations planning to adopt generative AI also should look at implementing change management, noted Phil Hassey, CEO and founder of research firm CapioIT.

This is an area that often is left out of the discussion, Hassey told ZDNET. Organizations have to figure out the cost involved and skillsets they need to acquire and roles that have to be reskilled, as a result of rolling out generative AI.

A proper change management strategy is key to ensuring a smooth transition and retaining talent, he said.

Based in Singapore, Eileen Yu reported for ZDNET from Dreamforce 2023 in San Francisco, at the invitation of Salesforce.com.

Artificial Intelligence

Microsoft open sources EvoDiff, a novel protein-generating AI

Microsoft open sources EvoDiff, a novel protein-generating AI Kyle Wiggers 8 hours

Proteins, the natural molecules that carry out key cellular functions within the body, are the building blocks of all diseases. Characterizing proteins can reveal the mechanisms of a disease, including ways to slow it or potentially reverse it, while creating proteins can lead to entirely new classes of drugs and therapeutics.

But the current process for designing proteins in the lab is costly — both from a computational and human resource standpoint. It entails coming up with a protein structure that could plausibly perform a specific task inside the body, then finding a protein sequence — the sequence of amino acids that make up a protein — likely to “fold” into that structure. (Proteins must correctly fold into three-dimensional shapes to carry out their intended function.)

It doesn’t necessarily have to be this complicated.

This week, Microsoft introduced a general-purpose framework, EvoDiff, that the company claims can generate “high-fidelity,” “diverse” proteins given a protein sequence. Different from other protein-generating frameworks, EvoDiff doesn’t require any structural information about the target protein, cutting out what’s typically the most laborious step.

Available in open source, EvoDiff could be used to create enzymes for new therapeutics and drug delivery methods as well as new enzymes for industrial chemical reactions, Microsoft senior researcher Kevin Yang says.

“We envision that EvoDiff will expand capabilities in protein engineering beyond the structure-function paradigm towards programmable, sequence-first design,” Yang, one of the co-creators of EvoDiff, told TechCrunch in an email interview. “With EvoDiff, we’re demonstrating that we may not actually need structure, but rather that ‘protein sequence is all you need’ to controllably design new proteins.”

Core to the EvoDiff framework is a 640-parameter model trained on data from all different species and functional classes of proteins. (“Parameters” are the parts of an AI model learned from training data and essentially define the skill of the model on a problem — in this case generating proteins.) The data to train the model was sourced from the OpenFold data set for sequence alignments and UniRef50, a subset of data from UniProt, the database of protein sequence and functional information maintained by the UniProt consortium.

EvoDiff is a diffusion model, similar in architecture to many modern image-generating models such as Stable Diffusion and DALL-E 2. EvoDiff learns how to gradually subtract noise from a starting protein made almost entirely of noise, moving it closer — slowly, step by step — to a protein sequence.

Microsoft EvoDiff

The process by which EvoDiff generates proteins.

Diffusion models have been increasingly applied to domains outside of image generation, from conjuring up designs for novel proteins, like EvoDiff, to creating music and even synthesizing speech.

“If there’s one thing to take away [from EvoDiff], I think it’d be this idea that we can — and should — do protein generation over sequence because of the generality, scale and modularity that we’re able to achieve,” Microsoft senior researcher Ava Amini, another co-contributor on EvoDiff, said via email. “Our diffusion framework gives us the ability to do that and also to control how we design these proteins to meet specific functional goals.”

To Amini’s point, EvoDiff can not only create new proteins but fill in the “gaps” in an existing protein design, so to speak. Provided a part of a protein that binds to another protein, the model can generate a protein amino acid sequence around that part that meets a set of criteria, for example.

Because EvoDiff designs proteins in the “sequence space” rather than the structure of proteins, it can also synthesize “disordered proteins” that don’t end up folding into a final three-dimensional structure. Like normal functioning proteins, disordered proteins play important roles in biology and disease, like enhancing or decreasing other protein activity.

Now, it should be noted that the research behind EvoDiff hasn’t been peer reviewed — at least not yet. Sarah Alamdari a data scientist at Microsoft who contributed to the project, admits that there’s “a lot more scaling work” to be done before the framework can be used commercially.

“This is just a 640-million-parameter model, and we may see improved generation quality if we scale up to billions of parameters,” Alamdari said via email. “While we demonstrated some coarse-grained strategies, to achieve even more fine-grained control, we would want to condition EvoDiff on text, chemical information or other ways to specify the desired function.”

As a next step, the EvoDiff team plans to test the proteins that the model generated in the lab to determine whether they’re viable. If they turn out to be, they’ll begin work on the next generation of the framework.

A Comprehensive Review of Blockchain in AI

AI and Blockchain have emerged as two of the most groundbreaking technical innovations in recent times.

  • Artificial Intelligence (AI): Enables machines and computers to emulate human thinking and decision-making processes.
  • Blockchain: A distributed and immutable ledger that securely stores data and information in a decentralized and trusted manner.

Recently, scientists have delved into exploring potential applications of these technologies across various sectors. In this article, we'll provide a brief overview of how blockchain can be integrated with AI, a concept that might be coined as “decentralized AI”. Let's dive in.

Decentralized AI: An Introduction to Blockchain in AI

In the past decade or so, blockchain has been one of the most hyped innovations, and it started to gain momentum when it found its application in other fields. Ever since its inception in 2008, it continued to emerge as a disruptive technology that had the potential to revolutionize the way we store or exchange data or information, and revolutionize the way we trace & track transactions or automate them.

One of the most talked about points of blockchain is that every blockchain transaction is signed cryptographically, and the mining nodes that hold a replica of the entire ledger of chained block of all transactions verifies each such transaction that results in the creation of synchronized, secure, and shared timestamped records that are impossible to alter. Resultantly, blockchain can be an effective option to eliminate the requirement for a central authority to verify & govern the transactions & interactions between users on the network.

Moving along, the technical industry has been producing and generating a huge amount of data thanks to technical innovations like IoT devices, smartphones, social media, and web applications that have contributed significantly in the rise of AI because to perform effectively & efficiently, AI systems often utilize a large amount of data using deep learning and machine learning practices to perform different analytics.

Even today, a vast chunk of machine learning and deep learning techniques for AI models rely on a centralized model that trains a group of servers that run or train a specific model against training data, and then verifies the learning using validation or training dataset. The high requirement to effectively train an AI model is the reason why major tech organizations and development teams often store a large amount of data to train their models for the best possible results and performance.

Most AI models and practices today are centralized, and although centralization has brought a lot of success to the AI industry, there is a major drawback with centralized data storage for AI models. When the entire data is stored in a centralized manner, the possibility of data tampering, or data corruption increases as centralized data storage is always a subject to malware and cybersecurity attacks. Furthermore, when dealing with a large amount of data, it is a challenging task to verify the authenticity & provenance of the data source is not guaranteed which may result in wrong training of the model that can further result in unwanted, inaccurate, and even dangerous outcomes.

The challenges with data storage for AI models is the major reason behind the use of blockchain in AI and the development of decentralized AI. The primary aim of decentralized AI is to enable a process and perform decision making or analytics using a digitally signed, secured, and trusted shared data that has been stored & transacted on the blockchain network in a decentralized or distributed manner without using external Third-Party resources.

AI models have the reputation of often working with a large amount of data, and scientists have already predicted blockchain to be the future of data storage. Furthermore, blockchain have smart contracts that allow users to program the blockchain network to govern transactions amongst the participants involved in generating or accessing the data, or decision-making. Autonomous applications and machines based on blockchain smart contracts can learn and adapt to changes as time progresses, and they can also make accurate and trusted decisions, outcomes verified and validated by the mining nodes of the blockchain network.

How Blockchain can Transform Artificial Intelligence?

Several shortcomings of the artificial intelligence and blockchain industry can be addressed efficiently by combining both the technical systems. Blockchain acts as a distributed ledger that stores and transmits data in a cryptographically signed method that is agreed and verified by the mining nodes of the network. Blockchain networks store data with high resilience & integrity that makes it almost impossible to tamper with the data which is the major reason why the outcome of machine learning algorithms when they make decisions using blockchain smart contracts cannot be disputed, and can be trusted. The use of blockchain networks with AI technologies can help in creating decentralized, immutable, and secure systems for highly sensitive data that can be collected, processed, and utilized by AI-powered applications. The security and safety offered by the use of blockchain in AI can have revolutionary applications across industries, especially the more sensitive ones like healthcare & hospitals, finance, defense, and more.

Moving along, some of the prominent benefits of integrating AI and blockchain are listed below.

  • Enhanced Data Security

A major reason behind blockchain’s immense popularity is that it offers a highly safe & secure method to store information on the web. Blockchains offer an alternative to store sensitive and critical information on disks, which is by storing digitally signed data that can be accessed only by using private keys. Hence, using blockchain to store data for AI algorithms can allow AI models to work with sensitive data, thus resulting in more accurate & trusted information.

  • Collective Decision Making

In a technical ecosystem, the involved applications or tools must work in coordination with each other to achieve the goal with maximum efficiency. Blockchain systems offer decentralized and distributed solutions for decision making algorithms that can replace the requirement for a central authority. Eliminating the central authority will allow the robots to discuss the problem internally, vote on any issue, and resolve the matter with majority until a conclusion is agreed upon.

  • Enhanced Trust on Robotic Decisions

Blockchain stores the data in a highly secure way that cannot be altered with which ensures the quality of the data throughout the development of the training process. As a result, the model will train on highly accurate data that will ultimately help in increasing the accuracy of the mode.

  • Higher Efficiency

One of the major reasons why business processes that often involve multiple users like multiple shareholders or stakeholders, governmental organizations, and business firms are often inefficient is because of numerous authorization of business transactions. Using blockchain and smart contracts will enable DAOs or Decentralized Autonomous Agents that will validate data or asset transfers amongst different stakeholders automatically, efficiently, and quickly.

Taxonomy of Blockchain in AI

In this section, we will be talking about some of the key concepts used in the application of blockchain technologies for AI applications that are mentioned in the figure below.

Decentralized AI Applications

Current AI applications generally operate in an autonomous manner to execute informed decisions using different planning, searching, optimizing, learning, knowledge recovery and management strategies. However, decentralizing AI applications is a difficult and challenging task for numerous reasons.

  • Autonomic Computing

One of the major goals of AI applications is to enable partially or fully autonomous operations where numerous intelligence agents or small computer programs will perceive & analyze their local environments, preserve their internal states, and execute specified actions accordingly.

  • Optimization

One of the major features of AI applications is their potential to make the most effective & efficient decisions by filtering a set of ideal solutions amongst all the possible solutions, and its possible because of optimization of AI algorithms and models. Optimization techniques aim to find the best solution to a problem by operating in a constrained or unconstrained environment depending upon the system level and application level objectives. Decentralized optimization will result in better efficiency & boosted performance.

  • Planning

AI applications make use of planning strategies when collaborating with other applications & systems to solve complex problems in new or challenging environments. Planning strategies play an important role in maintaining the resilience & efficiency of AI models. Using blockchain for planning strategies can result in devising more immutable and critical strategies used for mission critical systems and strategic applications.

  • Knowledge Discovery and Knowledge Management

AI applications have a reputation of working with a large amount of data, and their reliance on centralized data processing systems. With the use of decentralization, the knowledge discovery and knowledge management processes will be able to provide personalized knowledge patterns that considers the needs of all the stakeholders involved.

  • Learning

At the heart of Ai applications sits the learning algorithms that enable the knowledge discovery & automation processes. There are different kinds of learning algorithms like supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, ensemble, deep learning models, and much more that solve different machine learning problems. The use of decentralized learning models can result in highly autonomous learning systems that support local intelligence across different verticals in AI systems.

Decentralized AI Operations

AI models and algorithms often train, test and validate on a large amount of data to make better, and more versatile decisions. However, using centralized data storage solutions like data centers, clouds, and clusters act as a major hurdle in developing highly secure AI applications that preserve the privacy of its users. Here are some of the top blockchain implementations that can be adopted by numerous AI applications.

  • Decentralized Storage

Centralized data storage solutions are highly susceptible when it comes to security and privacy as these data storage solutions involve a user’s personal and sensitive data along with their locations, health records, activities, and financial information. Blockchain offers decentralized and cryptographically secure storage solutions across the participating applications & networks. Decentralized data storage solutions use nodes, and each node in the network keeps a client-centric encrypted copy of the database to ensure data availability for clients. Clients are free to use and mine their data as per their needs and requirements.

Two of the most common storage techniques used in decentralized data storage solutions are Sharding and Swarming. Sharding is the process in which you create logical partitions of the databases known as “Shards” where each partition is assigned a unique key that can be used to access the partition. On the other hand, Swarming is a method that uses “Swarms” to enable parallel data access from multiple nodes in the network to reduce the latency in AI applications, and thus resulting in more efficient & smooth performance. The shards are grouped together resulting in the formation of a collected storage that is supported in the network by a group of nodes in the form of swarms.

The use of decentralized storage solutions can result in enhanced reliability & scalability of storage because of multiparty geographical distributions offered by the decentralized storage solutions. Some of the emerging decentralized storage solutions include Storj, Swarm, Sia, FileCoin, IPFS, and more.

  • Data Management

One of the major requirements of developing an AI application is to manage data in a way that highly accurate, relevant, and complete datasets can be collected from reliable and trusted data sources. Conventionally, AI applications and algorithms have run centralized data management methods like data segmentation, data filtration, and content-aware data storage that are executed across all the nodes in the network. When compared against decentralized data storage offered by blockchain networks, centralized data management fares poorly because not only will the rate of data duplication be high even when only minor changes are made to the data, but the need to transfer similar datasets repeatedly will also be high.

Decentralized data management methods on the other hand have been designed to be deployed at the node levels in the network considering the spatial and temporal attributes in the data. Furthermore, to maintain the provenance and security of the data, decentralized management schemes can put the metadata on the blockchain.

Blockchain-types for AI Applications

The Blockchain technology can be grouped into two categories: Permissioned where only the authorized users can access the blockchain applications in cloud-based, consortium, or private settings, and Permissionless where anyone can publicly access the systems using the internet.

  • Public Blockchains

Public blockchain belongs to the permissionless category of blockchain networks, where users have the freedom to download the blockchain code on their systems, modify the code, and use the code as per their own needs and requirements. Furthermore, public blockchains are often open-source for read & write operations, and easily accessible. Because public blockchains are accessible by everyone, these systems make use of complex protocols for safety, and the identity & transactional privacy information of the users on the network is managed using pseudonymous and anonymous data on the network. For data and asset transfer, each public blockchain network uses native tokens also known as value pointers or cryptocurrencies.

  • Private Blockchains

Unlike public blockchains, private blockchain networks are permissioned systems that are managed by a single organization, and they are designed as permissionless systems where the users or participants are always known within the network, and they have the pre-approval for read and write operations on the network. Private blockchains often offer higher efficiency because the identity of the visitors is known, and they are pre-approved participants of the network to eliminate the need for complex algorithms and mathematical operations to validate any transaction on the network. Additionally, private blockchain networks can transfer any kind of assets, values, or indigenous data within the network.

Just like in public blockchain networks, the approval of a transaction and asset transfers in the private blockchain network is done by multi party consensus algorithms or voting that not only enable faster transactions but also consumes low energy. Astonishingly, the average transaction approval time on a private blockchain network is under a second.

  • Consortium Blockchain Networks

Consortium Blockchains, also known as Federated Blockchains are operated by a group of organizations where the groups are generally formed on the basis of mutual interest shared by these organizations. Consortium blockchain networks are generally offered by government organizations & bodies, banks, and some private blockchain companies as well.

Just like their private blockchain counterparts, the Consortium blockchain network operates as permissioned systems although a few users on the network have both read and write privileges on the network. Generally, all the users on the Consortium blockchain network have read access, but only a handful of individuals can write data on the network.

Decentralized Infrastructure for AI Applications

Blockchain architectures were traditionally designed by developers as linear infrastructure using a combination of hashing strategies, and linked lists data structures. However, recently, developers have been working on nonlinear infrastructures using queuing information, and graph theory to handle big data, and cater the requirements of real-time AI-based applications.

Blockchain-enabled AI Applications

Decentralized Data Storage and Data Management with AI

Using Blockchain with AI has allowed developers to work on developing stable systems that support the interaction of different technical innovations, and thus providing a platform for secure and safe data management, data transfer, and data storage. The below figure demonstrates the combined features of blockchain and AI technologies for the medical industry that includes different stages like analytics, diagnosis, validation of medical discoveries & reports, and critical decision making.

In recent years, handling a large amount of data, increasing the computing power of algorithms & models exponentially, and growing user acceptance of connected systems and applications have been the top priorities in the AI and ML industry. As artificial neural networks often require a large amount of data and computing power for training purposes, it is essential to create powerful data centers to acquire large datasets. During an audit process, blockchain networks can be used to store the data & the query information while achieving a higher level of security and privacy. Furthermore, the integration of AI and Blockchain technologies will provide a strong consensus mechanism that is immutable, robust, decentralized.

Decentralized Infrastructure for AI

The introduction of the Blockchain network infrastructure added three new characteristics to the traditional distributed architectures: decentralized and shared control of data and assets, native asset exchanges, and immutable audit trails. When the blockchain infrastructure was combined with AI technologies, the infrastructure provided users with new data models, and offered shared control of AI models & training data while adding to the trustworthiness of the data. To produce better and more efficient data models, AI models need access to a large amount of data that is provided by blockchain networks.

Decentralized networks like IPFS and Ethereum can handle data storage, and huge computational resources respectively, therefore providing tamper-free records with a high level of privacy. Open-source decentralized AI platforms like ChainIntel aim to get rid of the monopolization of AI services by big companies.

Decentralized AI Applications

Collective decision making, and decentralized intelligence can have numerous applications. For example, the figure below demonstrates the features & benefits of combining Blockchain with IoT and AI technologies to increase the yield in farming fields. IoT sensors can monitor soil’s nutrients levels, and capture images that can help in monitoring the growth of crops over time. AI can make use of the data received from IoT sensors to provide predictive analysis that allows the farmers to monitor different conditions. The use of blockchain ensures that every user on the network has access to the transactions that helps in reducing the time spent on logistics.

The above image demonstrates blockchain-based systems used for unmanned automated intelligent exploration of the ocean beds.

The above image demonstrates the use of Blockchain and AI for financial and banking purposes, and how blockchain and AI can improve the efficiency, safety & security of the financial system.

Conclusion

In this article, we have talked about the application and use cases of blockchain in AI. The article gives an overview of decentralized storage, and how blockchain can be the key to solving several issues with AI. Moving along, we also discussed the taxonomy of blockchain in AI, and the related technologies, and the comparison of blockchain implementations in terms of blockchain types & infrastructure, decentralized AI operations, and protocols. Finally, we discuss the various applications of blockchain in AI.

To sum things up, it would be safe to say that the implementation of blockchain in AI has the potential to address and solve existing issues in the AI industry related to user privacy, secured oracles, smart contract security, consensus protocols, standardization, and governance.

After London, OpenAI Now Expands to Dublin

OpenAI London Office

OpenAI is in expansion mode. The company has opened a new office in Dublin, Ireland. The objective is to collaborate with the Irish government and support their national AI strategy and work with industries and startups.

This is OpenAI’s second presence in Europe after it opened the London office in June this year where the company will focus on conducting new research and strengthen its engineering capabilities. This will also put the research lab close to policymakers, given that the European Union is on the roads to regulate AI. The United Kingdom’s corporate tax policy is lower compared to the United States.

The expansion of the nonprofit-lab-turned-startup comes after Altman took a tour around the globe, meeting several country leaders on his stops. OpenAI now globally has 3 offices, with their headquarters in San Francisco, United States.

“We believe that companies such as OpenAI operating in Ireland can help build on our foundation to support emerging AI research and innovation, and ensure our workforce is well prepared,” said Simon Coveney, Minister for Enterprise, Trade, and Employment.

Jason Kwon, chief strategy officer at OpenAI, told Reuters that the company wants to grow deliberately and not too rapidly because they want to make sure that the “culture of the company is established” first in the new office before they scale up. The motive to open this office in Dublin was not to benefit from lower taxes, as OpenAI is yet to be profitable, he said.

The Microsoft-backed company’s next stop might be Japan as Sam Altman, OpenAI’s CEO met the Japanese Prime Minister, Fumio Kishda in April this year to discuss expanding services in their country.

The post After London, OpenAI Now Expands to Dublin appeared first on Analytics India Magazine.

Pursue A Master’s In Data Science With The 3rd Best Online Program

Sponsored Content

Pursue A Master’s In Data Science With The 3rd Best Online Program

Data science teams need general industry experts who understand data science and technical specialists who can make it happen. Bay Path University will provide you with a career path in data science, regardless of your background and experience. We were one of the first institutions to develop two tracks to complete the Master of Science (MS) in Applied Data Science degree, which is right for you?

Generalist Track — This track prepares students to be well-rounded, collaborative, and skilled data scientists and analysts regardless of their background or area of expertise. Coursework in this track provides the foundation needed for breaking into the fast-growing field of data science.

Specialist Track — This track prepares students to take on more technical roles on data science teams, such as data modeler, data mining engineer, or data warehouse architect.

"I gained practical experience applying a diverse set of data science skills and was able to build a diverse portfolio."

—Aspen Gulley G'23

Our MS in Applied Data Science Degree Program Provides:

  • Small class settings, led by an extraordinary team of faculty who teach and mentor students throughout the program
  • Hands-on application using essential programming languages such as Python, SAS, R, and SQL
  • A project-based curriculum teaching students to solve real-world business challenges, using both "small" and "big" data and cutting-edge practices in statistical modeling, machine learning, and data mining
  • A project-oriented capstone that will harness the skills gained throughout the program
  • Flexibility for working professionals with convenient one and two-year schedules

LEARN MORE

More On This Topic

  • Maximize Your Value With The 3rd Best Online Master’s In Data Science…
  • Advance your Career with the 3rd Best Online Master's in Data Science…
  • Read This Before You Apply to a Business Analytics Master's Program
  • Join UC's Information Session for the Master's in Business Analytics…
  • Online Master’s in Data Science from Northwestern
  • Northwestern Online Master's in Data Science

Collaborative visual knowledge graph modeling at the system level

Collaborative visual knowledge graph modeling at the system level
Image by Chaitawat Pawapoowadon from Pixabay

The best way to model business and consumer dynamics is collaboratively, with stakeholders all in the same virtual room contributing. Of course, this has been happening asynchronously for some time now, but the potential exists for more real-time interaction.

Modelers don’t work in a vacuum, of course. The iterations between a modeler who develops a straw man and the stakeholders who know the domain is critical.

No wonder shared online whiteboarding has become popular. But how do you capture and reuse the data that underpins the model without disrupting the ongoing collaboration? Front- and back-end designers need to work together and solve this key technical challenge.

Interactive graph modeling

Synergy Codes (SC) is a custom data visualization product development firm that has been working more on graph-related projects. For internal project management and human resources purposes, they’ve built a knowledge graph-based recommendation app to match skills to particular projects.

Collaborative visual knowledge graph modeling at the system level
Synergy Codes, 2023

Lately, SC has been working toward adding an interactive capability. For now, the company makes it possible to optical character recognition (OCR) scan a node-and-edge diagram on a shared screen and make the diagram editable. Obviously more needs to be done to enable seamless back-and-forth between shared visual diagramming and a knowledge graph stored in standard form in a graph database. SC supports Neo4j and is considering support for Mongo DB or Postgres AgensGraph.

Knowledge graphing has become an effective and scalable way to model things at the data layer and make them interoperable at a system level. Graphs–models expressed in terms of any-to-any edges and nodes–are the mothers of all the data modeling children, which means interoperation between disparate elements of the model is built in.

A data modeler working in graphs can bring together tables (in relational databases, for example), documents (in document databases) from different contexts and images and other multimedia (annotated with the help of standard semantic metadata) from other contexts into a single, cohesive graph.

The best graph databases make system-level analysis and management using thousands of heterogeneous sources possible.

How data capture and reuse of visual models happens

Back-end knowledge modelers have been thinking and designing visually since the 2000s, but tooling potential has been haphazard and underfunded.

Here’s how a front-end design process that ties into the full power of back-end connectivity can help. User experience (UX) designers are in the habit of interviewing and learning from end users as a part of their research effort. Now the challenge for companies such as SC is to collaborate also with domain modelers who manage and scale the data connectivity, to help them work visually and expedite their efforts.

Ultimately, the goal of knowledge graph modelers is to create extensible, flexible data contexts that can be shared and reused as simply as possible.

Ethernet pioneer Bob Metcalfe coined the term “network effect.” He describes the connected contexts in a knowledge graph and a less centralized architecture as a means of delivering the “network effect to the Nth power”.

Consider for a minute how knowledge graph modeling can pay off “to the Nth power” in the way that Metcalfe alludes to:

  1. A knowledge graph makes data the organization wants to keep findable, accessible, interoperable, and reusable (FAIR) in as many ways as there are contexts.
  2. It also provides a means of sharing description logic–how things fit and interoperate together across contexts–so that this logic can be reused over and over again by calling it from the graph. Rather than trapping this information in apps, the graph makes zero-copy code and data possible.
  3. This zero-copy capability makes it possible for interactions that create, replace, update, and delete (CRUD) data and description logic to happen just once, for every user and every use.

How digital twins and agents fit in

User and modeler experience designers can harness the power of the zero-copy capability by designing with data in terms of a digital twin and agent paradigm. A twin in this scenario is a knowledge subgraph for a context and function. An agent becomes the messenger relaying the needed information back and forth between users and the graph, as well as modelers and the graph.

All of a sudden, the designer starts to have an impact, not only with a user and one app but at the system level.

The result: Visual model-driven development

When UX + MX designers can help knowledge, engineers, modelers, and architects build systems visually, they scale the impact they can have on innovation. When it’s visual, a higher percentage of the workforce can get involved in the business modeling effort. When it’s visual, even the C-suite can see the benefits of how processes are expressed at the data layer.

Moreover, when the knowledge graph is at the center of the effort, the modeling changes only have to take place once, because conceptual, logical, and physical models are all unified. When adding the interactive visual component with an extensible knowledge graph-oriented architecture, the whole effort boosts interaction and workforce efficiency.

Nvidia Releasing Open-Source Optimized Tensor RT-LLM Runtime with Commercial Foundational AI Models to Follow Later This Year

Nvidia Releasing Open-Source Optimized Tensor RT-LLM Runtime with Commercial Foundational AI Models to Follow Later This Year September 14, 2023 by Agam Shah

(thodonal88/Shutterstock)

Nvidia's large-language models will become generally available later this year, the company confirmed.

Organizations widely rely on Nvidia's graphics processors to write AI applications. The company has also created proprietary pre-trained models similar to OpenAI's GPT-4 and Google's PaLM-2.

Customers can use their own corpus of data, embed it in Nvidia's pre-trained large language models, and build their own AI applications. The foundational models cover text, speech, images, and other forms of data.

Nvidia has three foundational models. The most publicized is NeMo, which includes Megatron, in which customers can build ChatGPT-style chatbots. NeMo also has TTS, which converts text to human speech.

The second model, BioNemo, is a large-language model targeted at the biotech industry. Nvidia's third AI model is Picasso, which can manipulate images and videos. Customers will be able to engage Nvidia's foundational models through software and services products from the company and its partners.

"We'll be offering our foundational model services a little later this year," said Dave Salvator, a product marketing director at Nvidia, during a conference call. Nvidia's spokeswoman was not specific on availability dates for particular models.

The NeMo and BioNeMo services are currently in early access to customers via the AI Enterprise software and will likely be the first ones available commercially. Picasso is still further out from release, and services around the model may not become available as quickly.

"We are currently working with select customers, and others interested can sign up to get notified for when the service opens up more broadly," an Nvidia spokeswoman said.

The models will run best on Nvidia's GPUs, which are in short supply. The company is working to meet the demand, said Nvidia CFO Colette Kress at the recent Citi Global Technology Conference this week.

The GPU shortage creates a barrier to adoption, but customers can access Nvidia's software and services through the company's DGX Cloud or through Amazon Web Services, Google Cloud, Microsoft Azure, or Oracle Cloud, which have H100 installations.

Nvidia's foundational models are important ingredients in the company's concept of an "AI factory," in which customers do not have to worry about coding or hardware. An AI factory can take in raw data and churn it through GPUs and LLMs. The output is actionable data for companies.

The LLMs will be part of the AI Enterprise software suite, which includes frameworks, foundation models, and other AI technologies. The technology stack also includes tools like Tao, which is a no-code AI programming environment, and NeMo Guardrails, which can analyze and redirect output to provide more reliability on responses.

Nvidia is relying on its partners to sell and help companies deploy AI models such as NeMo to its accelerated computing platform.

Some Nvidia partners include software companies Snowflake and VMware and AI service providers Huggingface. Nvidia has also partnered with consulting company Deloitte for larger deployments. Nvidia has already announced it will bring its NeMo LLM to Snowflake Data Cloud, on which top organizations deposit data. Snowflake Data Cloud users will be able to generate AI-related insights and create AI applications by connecting their data to NeMo and Nvidia's GPUs.

The partnership with VMware brings the AI Enterprise software to VMware Private Cloud. VMware's vSphere and Cloud Foundation platforms provide administrative and management tools for AI deployments in virtual machines across Nvidia's hardware in the cloud. The deployments can also extend to non-Nvidia CPUs.

Nvidia is about 80% a software company, and its software platform is the operating system for AI, said Manuvir Das, vice president for enterprise computing at the company during Goldman Sachs' Communacopia+Technology conference.

Last year, people were still wondering how AI would help, but this year, "customers come to see us now as they already know what the use case is," Das said. The barrier to entry for AI remains high, and the challenge has been in the development of foundational models such as NeMo, GPT-4, or Meta's Llama 2.

"You have to find all the data, the right data, you have to curate it. You have to go through this whole training process before you get a usable model," Das said.

But after millions in investments for development and training, the models are now becoming available to customers.

"Now they're ready to use. You start from there, you finetune with your own data, and you use the model," Das said.

Nvidia has projected a $150 billion market opportunity for the AI Enterprise software stack, which is half that of the $300 billion hardware opportunity, which includes GPUs and systems. The company's CEO, Jensen Huang, has previously talked about AI computing being a radical shift from the old style of computing reliant on CPUs.

Open-Source Tensor-RT LLM

Nvidia separately announced Tensor-RT LLM, which improves the inferencing performance of foundational models on its GPUs. The runtime can extract the best inferencing performance of a wide range of models such as Bloom, Falcon, and Meta's latest Llama models.

A heavy-duty H100 is considered the best for training models but may be overkill for inferencing when factoring in the power and performance of the GPU. Nvidia has the lower-power L40s and L4 GPUs for inferencing but is making the H100 viable for inference if the GPUs are not busy.

Nvidia low power L4 GPU

The Tensor-RT LLM is specially optimized for low-level inferencing on H100 to reduce idle time and keep the GPU occupied at close to 100%, said Ian Buck, vice president of hyperscale and HPC at Nvidia.

Buck said that the combination of Hopper and Tensor-RT LLM software improved inference performance by eight times compared to the A100 GPU.

"As people develop new large language models, these kernels can be reused to continue to optimize and improve performance and build new models. As the community implements new techniques, we will continue to place them … into this open-source repository," Buck continued.

Tensor RT-LLM has a new kind of scheduler for the GPU, which is called inflight batching. The scheduler allows work to enter and exit the GPU independently of other tasks.

"In the past, batching came in as work requests. The batch was scheduled onto a GPU or processor, and then when that entire batch was completed … the next batch would come in. Unfortunately, in high variability workloads, that would be the longest workload … and we often see GPUs and other things be underutilized," Buck said.

With Tensor-RT LLM and in-flight batching, work can enter and leave the batch independently and asynchronously to keep the GPU 100% occupied.

"This all happens automatically inside the Tensor RT-LLM runtime system and it dramatically improves H100 efficiency," Buck said.

The runtime is in early access now and will likely be released next month.

Related

Conversational AI Company Uniphore Leverages Red Box Acquisition for New Data Collection Tool

3D illustration of conversation bubbles and a chat bot merging together.
Image: VectorMine/Adobe Stock

In February 2023, conversational AI company Uniphore acquired voice recording firm Red Box. Today, the result of that partnership has emerged: U-Capture, which is an AI-powered metadata solution for data collected across the enterprise. U-Capture is available beginning on Tuesday in North America, Europe, the Middle East, Australia, New Zealand and the Asia Pacific region.

U-Capture will be delivered as part of Uniphore’s existing X platform and provides AI-ready data in real time for automation and analysis of conversations. U-Capture is remarkable because it’s a practical use for an AI assistant combined with existing data recording tools, but security is still a concern for Uniphore, which bills its platform as AI-native. Uniphore gained a lot of buzz in 2022 for a $400 million funding round and support from Cisco CEO John Chambers.

Jump to:

  • U-Capture builds on Red Box data recording capabilities
  • Part of Red Box’s appeal: Open architecture
  • How Uniphore addresses security and data sharing concerns

U-Capture builds on Red Box data recording capabilities

One problem Uniphore wanted to solve through the Red Box acquisition was that a lot of call center data was unstructured. It was multimodal: some might be video, some might be audio, some might be text.

U-Capture collects data from call center conversations or other business audio or video records such as meetings. Then, U-Capture applies automation and connected applications to understand and reformat the content to provide coaching, analysis or tailored responses (Figure A).

Figure A

Screenshot of conversational AI analytics dashboard in U-Capture.
Conversational AI analytics dashboard in U-Capture. Image: Uniphore

The tool also applies so-called emotion AI, which can read sentiment and tailor its suggestions and coaching to that sentiment. For example, the AI system might let an irate customer skip through a welcome message on a phone tree to get them to a human agent faster, and then coach that agent on how to respond.

SEE: AI can enhance, not necessarily replace, jobs, Microsoft says. (TechRepublic)

“We re-engineered the Red Box capability on our platform, which is now called and sold as U-Capture,” said Umesh Sachdev, chief executive officer of Uniphore, in an interview with TechRepublic.

“As opposed to a data company that just bought an AI platform, we are an AI company that through an acquisition, got ourselves a data platform,” said Sachdev.

Part of Red Box’s appeal: Open architecture

Some companies that sold contact center software aren’t allowed by their vendors to take their own audio from the call center and use it for analytics; Uniphore wanted to open up the ways in which that data could be used, while still keeping the data secure.

“Red Box was the third largest player in the segment of call center, audio streaming, recording and capture, and Red Box was the only player in the segment globally, which on its website advertised that it’s an open architecture,” said Sachdev. “So it was the very ethos that we were looking for to take control of our own destiny. We don’t have to be reliant on a third-party data provider to give us the data.

“With Red Box, we’ve got access to some of the world’s most open format data on recording, capture and streaming in the call center. We got access to a terrific leadership team, and the talent that they had built,” Sachdev said.

Other new features coming to the X platform today include natural language generation for summarizing chats and intent training for business analysis.

How Uniphore addresses security and data sharing concerns

Umesh noted that a public AI such as ChatGPT isn’t particularly useful for a highly regulated company in finance or telecommunications; some of Uniphore’s largest customers are in those fields. So, keeping the data inside the enterprise is both more secure and more practical than feeding it out to a generative AI that trains on the entire internet. But there are benefits to keeping some of the data open, Umesh said.

“Our customers contractually allow us to learn and improve our models from the data that’s flowing through our system, as long as it’s anonymized and aggregated at all points,” Sachdev said. “At no point does the AI model know the source of this data. It can never trace it back to an individual customer or their record … it’s about a series of trends or patterns.”

Uniphore builds data redaction into its monitoring software, meaning that personal data such as credit card numbers are obscured.

Data coming out of customers’ call centers is open in that it moves through an API framework. Both Uniphore and the customer company are free to use that data for whatever analytics they may wish to run on it, whether that be through Uniphore or a third-party AI vendor.

“Companies are looking at more and more applications of generative AI, and we are going to see an explosion of the number of applications within the enterprise in different departments through generative AI,” said Umesh. “Which means these companies need a data layer which is open, which is fast and seamless [and] which is secure and reliable.”

Subscribe to the Daily Tech Insider AU Newsletter

Stay up to date on the latest in technology with Daily Tech Insider Australian Edition. We bring you news on industry-leading companies, products, and people, as well as highlighted articles, downloads, and top resources. You’ll receive primers on hot tech topics that are most relevant to AU markets that will help you stay ahead of the game.

Delivered Thursdays Sign up today