What Does Closed-Door Meeting With AI Industry Leaders Mean for Business?

Some of the United States’ top tech executives and generative AI development leaders met with senators last Wednesday in a closed-door, bipartisan meeting about possible federal regulations for generative artificial intelligence. Elon Musk, Sam Altman, Mark Zuckerberg, Sundar Pichai and Bill Gates were some of the tech leaders in attendance, according to reporting from the Associated Press. TechRepublic spoke to business leaders about what to expect next in terms of government regulation of generative artificial intelligence and how to remain flexible in a changing landscape.

Jump to:

  • AI summit included tech leaders and stakeholders
  • U.S. regulation of generative AI is still developing
  • Generative AI’s impact on cybersecurity for businesses
  • Considerations for business leaders working with generative AI

AI summit included tech leaders and stakeholders

Each participant had three minutes to speak, followed by a group discussion led by Senate Majority Leader Chuck Schumer and Republican Sen. Mike Rounds of South Dakota. The goal of the meeting was to explore how federal regulations might respond to the benefits and challenges of rapidly-developing generative AI technology.

Musk and former Google CEO Eric Schmidt discussed concerns about generative AI posing existential threats to humanity, according to the Associated Press’ sources inside the room. Gates considered solving problems of hunger with AI, while Zuckerberg was concerned with open source vs. closed source AI models. IBM CEO Arvind Krishna pushed back against the idea of AI licenses. CNN reported that NVIDIA CEO Jensen Huang was also present.

All of the forum attendees raised their hands in support of the government regulating generative AI, CNN reported. While no specific federal agency was named as the owner of the task of regulating generative AI, the National Institute of Standards and Technology was suggested by several attendees.

The fact that the meeting, which included civil rights and labor group representatives, was skewed toward tech moguls was dissatisfying to some senators. Sen. Josh Hawley, R-Mo., who supports licensing for certain high-risk AI systems, called the meeting a “giant cocktail party for big tech.”

“There was a lot of care to make sure the room was a balanced conversation, or as balanced as it could be,” Deborah Raji, a researcher at the University of California, Berkeley who specialized in algorithmic bias and attended the meeting, told the AP.(Note: TechRepublic contacted Senator Schumer’s office for a comment about this AI summit, and we have not received a reply by the time of publication.)

U.S. regulation of generative AI is still developing

So far, the U.S. federal government has issued suggestions for AI makers, including watermarking AI-generated content and putting guardrails against bias in place. Companies including Meta, Microsoft and OpenAI have attached their names to the White House’s list of voluntary AI safety commitments.

Many states have bills or legislation in place or in progress related to a variety of applications of generative AI. Hawaii has passed a resolution that “urges Congress to begin a discussion considering the benefits and risks of artificial intelligence technologies.”

Questions of copyright

Copyright is also a factor being considered when it comes to legal rules around AI. AI-generated images cannot be copyrighted, the U.S. Copyright Office determined in February, although parts of stories created with AI art generators can be.

Raul Martynek, chief executive officer of data center solutions maker DataBank, emphasized that copyright and privacy are “two very clear problems stemming from generative AI that legislation could mitigate.” Generative AI consumes massive amounts of energy and information about people and copyrighted works.

“Given that states from California to New York to Texas are forging ahead with state privacy legislation in the absence of unified federal action, we may soon see the U.S. Congress act to bring the U.S. on par with other jurisdictions that have more comprehensive privacy legislation,” said Martynek.

SEE: The European Union’s AI Act bans certain high-risk practices such as using AI for facial recognition. (TechRepublic)

He brought up the case of Barry Diller, chairman and senior executive of media conglomerate IAC, who suggested companies using AI content should share revenue with publishers.

“I can see privacy and copyright as the two issues that could be regulated first when it ultimately happens,” Martynek said.

Ongoing AI policy discussions

In May 2023, the Biden-Harris administration created a roadmap for federal investments in AI development, made a request for public input on the topic of AI risks and benefits, and produced a report on the problems and advantages of AI in education.

“Can Congress work to maximize AI’s benefits, while protecting the American people—and all of humanity— from its novel risks?,” Schumer wrote in June.

“The policymakers must ensure vendors realize if their service can be used for a darker purpose and likely provide the legal path for accountability,” said Rob T. Lee, a technical consultant to the U.S. government and chief curriculum director and faculty lead at the SANS Institute, in an email to TechRepublic. “Trying to ban or control the development of services could hinder innovation.”He compared artificial intelligence to biotech or pharmaceuticals, which are industries that could be harmful or beneficial depending on how they are used. “The key is not stifling innovation while ensuring ‘accountability’ can be created,” Lee said.

Generative AI’s impact on cybersecurity for businesses

Generative AI will impact cybersecurity in three main ways, Lee suggested:

  • Data integrity problems.
  • Conventional crimes such as theft or tax evasion.
  • Vulnerability exploits such as ransomware.

“Even if policymakers get involved more — all of the above will still occur,” he said.

“The value of AI is overstated and not well understood, but it is also attracting a lot of investment from both good actors and bad actors,” Blair Cohen, founder and president of identity verification firm AuthenticID, said in an email to TechRepublic. “There is a lot of discussion over regulating AI, but I am sure the bad actors won’t follow those regulations.”

On the other hand, Cohen said, AI and machine learning may also be critical to protecting against malicious uses of the hundreds or thousands of digital attack vectors open today.

Business leaders should keep up-to-date with cybersecurity in order to protect against both artificial intelligence and traditional digital threats. Lee noted that the speed of the development of generative AI products creates its own dangers.

“The data integrity side of AI will be a challenge, and vendors will be rushing to get products to market (and) not putting appropriate security controls in place,” Lee said.

Policymakers might learn from corporate self-regulation

With large companies self-regulating some of their uses of generative AI, the tech industry and governments will learn from each other.

“So far, the U.S. has taken a very collaborative approach to generative AI legislation by bringing in the experts to workshop needed policies and even simply learn more about generative AI, its risk and capabilities,” said Dan Lohrmann, field chief information security officer at digital solutions provider Presidio, in an email to TechRepublic. “With companies now experimenting with regulation, we are likely to see legislators pull from their successes and failures when it comes time to develop a formal policy.”

Considerations for business leaders working with generative AI

Regulation of generative AI will move “reasonably slowly” while policymakers learn about what generative AI can do, Lee said.

Others agree that the process will be gradual. “The regulatory landscape will evolve gradually as policymakers gain more insights and expertise in this area,” predicted Cohen.

64% of Americans want generative AI to be regulated

In a survey published in May 2023, global customer experience and digital solutions provider TELUS International found that 64% of Americans want generative AI algorithms to be regulated by the government. 40% of Americans do not believe companies using generative AI in their platforms are doing enough to stop bias and false information.

Businesses can benefit from transparency

“Importantly, business leaders should be transparent and communicate their AI policies publicly

and clearly, as well as share the limitations, potential biases and unintended consequences of

their AI systems,” said Siobhan Hanna, vice president and managing director of AI and machine learning at TELUS International, in an email to TechRepublic.

Hanna also suggested that business leaders should have human oversight over AI algorithms, be sure that the information conveyed by generative AI is appropriate for all audiences and address ethical problems through third-party audits.

“Business leaders should have clear standards with quantitative metrics in place measuring the accuracy, completeness, reliability, relevance and timeliness of its data and its algorithms’ performance,” Hanna said.

How businesses can be flexible in the face of uncertainty

It is “incredibly challenging” for businesses to keep up with changing regulations, said Lohrmann. Companies should consider using GDPR requirements as a benchmark for their policies around AI if they handle personal data at all, he said. No matter what regulations apply, guidance and norms around AI should be clearly defined.

“Keeping in mind that there is no widely accepted standard in regulating AI, organizations need to invest in creating an oversight team that will evaluate a company’s AI projects not just around already existing regulations, but also against company policies, values and social responsibility goals,” Lohrmann said.

When decisions are finalized, “Regulators will likely emphasize data privacy and security in generative AI, which includes protecting sensitive data used by AI models and safeguarding against potential misuse,” Cohen said.

Subscribe to the Innovation Insider Newsletter

Catch up on the latest tech innovations that are changing the world, including IoT, 5G, the latest about phones, security, smart cities, AI, robotics, and more.

Delivered Tuesdays and Fridays Sign up today

Generative AI can be the academic assistant an underserved student needs

Robot drawing on a chalkboard

From essay writing to standardized test prep and scores, navigating the higher education path involves complex twists and turns that can put students from lower socioeconomic backgrounds at a disadvantage. AI tools could provide meaningful aid, but ethical questions still loom.

Although tutoring services, consultants, or essay mentors cond be beneficial, they often come with a steep cost, while educational advancement is a process that can be difficult for students who can't afford these extra resources.

As a first-generation applicant, I had minimal family guidance when looking at schools and programs or when filling out applications. When applying for college, I noticed this familiar pattern negatively impacting me and many other students in a lower economic bracket. These disadvantages also influenced college enrollment.

A Brookings Institute report found that 89% of students from well-off families go to college, 64% of students from middle-class families, and only 51% of students from low-income families.

AI, however, has the unexpected possibility to level out the playing field. ChatGPT, for example, is a free, all-encompassing resource able to chat in real-time to answer all my questions. Other educational tools also incorporate generative AI to advance students' education.

How AI chatbots can help in education

Google is a highly effective internet tool, but users still have to comb through results to find a piece of information they can then string together for a final answer.

In summarizing the benefits of AI, Sid Nag, a Gartner research analyst, tells ZDNET that "generative AI technology democratizes the whole aspect of knowledge and information access."

"In the past, knowledge was obtained through extensive research through reading a lot of different textbooks and going to libraries, and getting online with paid subscriptions."

Answer your questions in conversations

If I searched for the next SAT test date, I'd spend roughly five minutes sifting through articles, choosing one, and navigating its contents to retrieve the simple answer.

Meanwhile, internet-connected AI chatbots like ChatGPT or Bing Chat conduct the scanning process for the user, providing an automatically straightforward answer plus additional resources.

As ZDNET previously reported, a study that compared Google's responses to those of ChatGPT found that ChatGPT's outputs outperformed Google's responses in both intermediate and advanced questions in both thoughtfulness and context.

Also: ChatGPT or Google: Which gives the best answers?

"The internet has been transformative for particularly adept, self-driven learners, who can become experts in many areas of human knowledge by judicious, directed consumption," Tom Lippincott, director of digital humanities and assistant research professor at John Hopkins University, tells ZDNET.

"But that process is itself a major challenge, not everyone learns the same way or has the time to find, curate, and digest what they need."

The natural language processing (NLP) capabilities AI chatbot possess makes the system more capable of understanding questions and able to answer in a conversational manner.

The personal statement

The personal statement is arguably one of the most important components of a college application, meant to showcase both the student's writing skills and personality beyond GPA and test scores.

Usually, the personal statement is the first of many supplemental essays, each unique to every school, ultra-specific in the prompt, and also weighs heavily on the applicant's portfolio. For the top 250 colleges, these essays generally account for 25% of their overall application, according to CollegeVine.

Due to time and importance, many students seek costly outside services. A Google search of "college essay assistance" revealed an oversaturation of services. One such service, PrepMaven, costs $79 to $349 per hour, with a minimum $510 package. With a PrepMaven subscription, students are entitled to an initial consultation and multiple essay revision cycles, according to the website.

Conversely, ChatGPT and other AI writing assistants have the ability to provide the same ideation services and grammar-specific essay guidance — for free.

Homework help

Outside of the application process itself, AI tools can also further academics once you've committed to a classroom, whether it be grasping new material, sourcing a specific concept, or summarizing a complex reading.

Sierra President, a sophomore at the University of North Carolina at Chapel Hill, has previously written about the ethics of college students using ChatGPT. She says to ZDNET, "I think that there are aspects of AI chatbots that can be helpful to students, such as in cases of gaining quick information on things like paper planning and studying help."

Additionally, AI chatbots can aid in actual learning, including, writing, studying, math, coding, advice, research, and more.

Tutoring services

For students who can't rely on their parent(s) as an academic assistant for homework and assignments, chatbots also have the potential to serve as a 24/7, live tutor.

A 2021 SWNS digital survey asked 2,000 American parents with school-aged children about their ability to help their children with homework and 56% of parents said they feel helpless helping their kids with homework. Moreover, two-thirds of the parents surveyed said they would turn to Google to try to find ways to help.

Now, even parents can take advantage of tools like ChatGPT for a more efficient, thorough medium for solving even the trickiest homework answers. Similarly, students of all ages could use AI's efficient answers as a tool to better grasp the material.

Also: 5 ways ChatGPT can make parenting easier this summer

"Anyone who has become proficient in a complex domain of human knowledge can probably remember how important it was to hear the same idea presented in several different ways to make it 'click': the ability of these models to rephrase, simplify, and interactively explain is perfect for this," says Lippincott.

An AI chatbot allows all students with an internet connection, regardless of background, access to a powerful and mostly accurate learning tool.

Other educational tools that employ generative AI

Aside from AI chatbots, generative AI could play an integral role in other learning platforms, which could improve students' learning and education accessibility. Even prior to the AI chatbot "boom," many free education tools leveraged AI to develop helpful and educational content.

Quizlet, a free learning platform, provides users with free study sets and virtual custom-made flashcards, as well as millions of study sets previously created by other Quizlet users.

Over six years ago, Quizlet implemented AI to introduce its "Learn" mode and generated features such as "example sentences" for vocabulary learning and multiple-choice questions.

Then in 2020, it partnered with OpenAI (before its claim to fame ChatGPT) as part of the beta for GPT-3. More recently, Quizlet's adaptive AI tutor called Q-chat, powered by OpenAI's ChatGPT API, is available in beta for free. The tool also has premium features for $39.99 annually.

Quizlet CEO Lex Bayer cited a study that found the outcomes of one-to-one tutoring can be as much as two standard deviations better than those of classroom instruction. He tells ZDNET that although tutoring is the best way for students to learn, it's expensive and hard to scale.

"But now with these amazing technologies that have come through large language models, we can make this a reality for students."

Similarly, the popular language learning platform Duolingo embraced AI before the technology's rise to popularity, allowing users to learn 40+ languages with fun exercises.

Duolingo's AI works with human experts to create and personalize the user's curriculum and recently employed GPT-4 to create more lessons, extend content length, and provide better suggestions.

Duolingo is free, and users are able to learn a new language at their own pace while bypassing expensive tutors or group classes. This access provides equal opportunities to students who might not otherwise have it, Bozena Pajak, Duolingo's VP of learning and curriculum, tells ZDNET.

In the US, Duolingo's GPT-4 integration helps students learning English as a second language develop and sharpen their verbal and writing skills. Pajak says "teaching people English, that's something that we know just helps improve people's prospects," adding that the company has been focusing on English.

In March, the company announced Duolingo Max with AI-powered, tutor-like "roleplay" and "explain my answer" features for $168 annually or $30 a month.

The ethics surrounding AI in education

Whenever AI is discussed in the educational field, there's apprehension regarding how it could negatively affect learning in students, including cheating and the spread of misinformation.

Because the new technology raises so many questions, it's helpful to look back at history. When the printing press, calculator, or even the internet, were first developed and popularized, the public had a similar reaction: fear.

Misinformation

While obtaining a "generative AI education" so to speak, students need to be aware that AI chatbots have been guilty of outputting misinformation.

Trained on vast amounts of data, generative AI models use vast amounts of pre-existing content to then create new outputs like text and images. Since the chatbot makes inferences on data it's trained on to understand what you are saying, and how to respond to it, there might be a disconnect, which causes incorrect outputs.

These false outputs are referred to as hallucinations, which often result in plausible but incorrect answers. In turn, that output can spread misinformation and make it easy for someone to misinterpret as misinterpret incorrectly as truth.

However, having just been a student myself, I think that if students are warned about misinformation, they can be careful about taking AI chatbot output as a fact.

"In general, I would say treat it as a conversation with an unreliable but knowledgeable person: identify factual claims that can be checked directly, try to have a clear goal in mind that you can evaluate in itself, and ask the model how it arrived at it until you're satisfied," says Lippincott.

AI educational policies not yet set

Although AI has the potential to revolutionize how students complete their work, schools have not set clear guidelines for proper usage.

In fact, some of the biggest public school districts in the country including New York City, Seattle, and Los Angeles have blocked access to ChatGPT on school networks and devices. Though some districts like New York City have lifted the block, AI policies still remain vague or too restrictive.

President emphasized this as a concern for schools using AI platforms, which causes "students to not understand what they can use and to what extent."

Standards and structures in the classroom regarding generative AI could help students take appropriate advantage of nuanced tech without compromising work integrity or "cheating" against peers who don't use AI.

Unfair advantage and cheating

As a student, President also adds that she feels AI gave students who used the technology an unfair advantage over those who opted to do work without AI assistance.

A possible workaround for this issue is for teachers to create assignments and testing where students' critical thinking is challenged as that is something AI models (tools) are not yet capable of doing.

Even if AI is a tool used in these types of "critical thinking"-based assignments, it won't be a sure way to an "A" as what's really being tested is higher level synthesis and application rather than data or information output.

Lippincott says the most immediate concerns to education are "misinformation and cheating."

"We want students to come away with a better understanding of the world, which is undermined by models generating falsehoods — and we want to be able to endorse students' achievements, which is undermined by models doing the work for them," he says.

A solution to prevent cheating is to keep AI as simply a learning tool.

"The cheating aspect is a bit more immediate and actionable: I don't see any way around needing to have careful, controlled evaluations where students have to demonstrate their understanding in class," says Lippincott.

In addition to being a hot topic now, AI has the potential to play a bigger role in our future. With its rapid popularity, Bayer notes the power of preparing students with a greater understanding of these tools, their limitations, as well as how and when to use them.

Since the launch of ChatGPT, AI has grown tremendously with no sign of slowing down. According to a study from Grandview Research, AI is expected to have an annual growth of 37.3% between 2023 and 2030.

As a recent first-generation graduate, AI would have been an invaluable tool in both my application process and studies. I hope that students of all socioeconomic backgrounds will look into AI's one-to-one tutoring aspects, language learning assistance, ideation, and more.

Artificial Intelligence

AI and Blockchain Integration for Preserving Privacy

With the widespread attention, and potential applications of blockchain and artificial intelligence technologies, the privacy protection techniques that arise as a direct result of integration of the two technologies is gaining notable significance. These privacy protection techniques not only protect the privacy of individuals, but they also guarantee the dependability and security of the data.

In this article, we will be talking about how the collaboration between AI and blockchain gives birth to numerous privacy protection techniques, and their application in different verticals including de-identification, data encryption, k-anonymity, and multi-tier distributed ledger methods. Furthermore, we will also try to analyze the deficiencies along with their actual cause, and offer solutions accordingly.

Blockchain, Artificial Intelligence, and their Integration

The blockchain network was first introduced to the world when in 2008 Nakamoto introduced Bitcoin, a cryptocurrency built on the blockchain network. Ever since its introduction, blockchain has gained a lot of popularity, especially in the past few years. The value at which Bitcoin is trading today, and it crossing the Trillion-dollar market cap mark indicates that blockchain has the potential to generate substantial revenue and profits for the industry.

Blockchain technology can be categorized primarily on the basis of the level of accessibility and control they offer, with Public, Private, and Federated being the three main types of blockchain technologies. Popular cryptocurrencies and blockchain architectures like Bitcoin and Ethereum are public blockchain offerings as they are decentralized in nature, and they allow nodes to enter or exit the network freely, and thus promotes maximum decentralization.

The following figure depicts the structure of Ethereum as it utilizes a linked list to establish connections between different blocks. The header of the block stores the hash address of the preceding block in order to establish a linkage between the two successive blocks.

The development, and implementation of the blockchain technology is followed with legitimate security and privacy concerns in various fields that cannot be neglected. For example, a data breach in the financial industry can result in heavy losses, while a breach in military or healthcare systems can be disastrous. To prevent these scenarios, protection of data, user assets, and identity information has been a major focus of the blockchain security research community, as to ensure the development of the blockchain technology, it is essential to maintain its security.

Ethereum is a decentralized blockchain platform that upholds a shared ledger of information collaboratively using multiple nodes. Each node in the Ethereum network makes use of the EVM or Ethereum Vector Machine to compile smart contracts, and facilitate the communication between nodes that occur via a P2P or peer-to-peer network. Each node on the Ethereum network is provided with unique functions, and permissions, although all the nodes can be used for gathering transactions, and engaging in block mining. Furthermore, it is worth noting that when compared to Bitcoin, Ethereum displays faster block generation speeds with a lead of nearly 15 seconds. It means that crypto miners have a better chance at acquiring rewards quicker while the interval time for verifying transactions is reduced significantly.

On the other hand, AI or Artificial Intelligence is a branch in modern science that focuses on developing machines that are capable of decision-making, and can simulate autonomous thinking comparable to a human’s ability. Artificial Intelligence is a very vast branch in itself with numerous subfields including deep learning, computer vision, natural language processing, and more. NLP in particular has been a subfield that has been focussed heavily in the past few years that has resulted in the development of some top-notch LLMs like GPT and BERT. NLP is headed towards near perfection, and the final step of NLP is processing text transformations that can make computers understandable, and recent models like ChatGPT built on GPT-4 indicated that the research is headed towards the right direction.

Another subfield that is quite popular amongst AI developers is deep learning, an AI technique that works by imitating the structure of neurons. In a conventional deep learning framework, the external input information is processed layer by layer by training hierarchical network structures, and it is then passed on to a hidden layer for final representation. Deep learning frameworks can be classified into two categories: Supervised learning, and Unsupervised learning.

The above image depicts the architecture of deep learning perceptron, and as it can be seen in the image, a deep learning framework employs a multiple-level neural network architecture to learn the features in the data. The neural network consists of three types of layers including the hidden layer, the input payer, and the output layer. Each perceptron layer in the framework is connected to the next layer in order to form a deep learning framework.

Finally, we have the integration of blockchain and artificial intelligence technologies as these two technologies are being applied across different industries and domains with an increase in the concern regarding cybersecurity, data security, and privacy protection. Applications that aim to integrate blockchain and artificial intelligence manifest the integration in the following aspects.

  • Utilizing blockchain technology to record and store the training data, input and output of the models, and parameters, ensuring accountability, and transparency in model audits.
  • Using blockchain frameworks to deploy AI models to achieve decentralization services among models, and enhancing the scalability and stability of the system.
  • Providing secure access to external AI data and models using decentralized systems, and enabling blockchain networks to acquire external information that is reliable.
  • Using blockchain-based token designs and incentive mechanisms to establish connections and trust-worthy interactions between users and AI model developers.

Privacy Protection Through the Integration of Blockchain and AI Technologies

In the current scenario, data trust systems have certain limitations that compromise the reliability of the data transmission. To challenge these limitations, blockchain technologies can be deployed to establish a dependable and secure data sharing & storage solution that offers privacy protection, and enhances data security. Some of the applications of blockchain in AI privacy protection are mentioned in the following table.

By enhancing the implementation & integration of these technologies, the protective capacity & security of current data trust systems can be boosted significantly.

Data Encryption

Traditionally, data sharing and data storing methods have been vulnerable to security threats because they are dependent on centralized servers that makes them an easily identifiable target for attackers. The vulnerability of these methods gives rise to serious complications such as data tampering, and data leaks, and given the current security requirements, encryption methods alone are not sufficient to ensure the safety & security of the data, which is the main reason behind the emergence of privacy protection technologies based on the integration of artificial intelligence & blockchain.

Let’s have a look at a blockchain-based privacy preserving federated learning scheme that aims to improve the Multi-Krum technique, and combine it with homomorphic encryption to achieve ciphertext-level model filtering and model aggregation that can verify local models while maintaining privacy protection. The Paillier homomorphic encryption technique is used in this method to encrypt model updates, and thus providing additional privacy protection. The Paillier algorithm works as depicted.

De-Identification

De-Identification is a method that is commonly used to anonymize personal identification information of a user in the data by separating the data from the data identifiers, and thus reducing the risk of data tracking. There exists a decentralized AI framework built on permissioned blockchain technology that uses the above mentioned approach. The AI framework essentially separates the personal identification information from non-personal information effectively, and then stores the hash values of the personal identification information in the blockchain network. The proposed AI framework can be utilized in the medical industry to share medical records & information of a patient without revealing his/her true identity. As depicted in the following image, the proposed AI framework uses two independent blockchain for data requests with one blockchain network storing the patient's information along with data access permissions whereas the second blockchain network captures audit traces of any requests or queries made by requesters. As a result, patients still have complete authority and control over their medical records & sensitive information while enabling secure & safe data sharing within multiple entities on the network.

Multi-Layered Distributed Ledger

A multi-layered distributed ledger is a data storage system with decentralization property and multiple hierarchical layers that are designed to maximize efficiency, and secure the data sharing process along with enhanced privacy protection. DeepLinQ is a blockchain-based multi-layered decentralized distributed ledger that addresses a user’s concern regarding data privacy & data sharing by enabling privacy-protected data privacy. DeepLinQ archives the promised data privacy by employing various techniques like on-demand querying, access control, proxy reservation, and smart contracts to leverage blockchain network’s characteristics including consensus mechanism, complete decentralization, and anonymity to protect data privacy.

K-Anonymity

The K-Anonymity method is a privacy protection method that aims to target & group individuals in a dataset in a way that every group has at least K individuals with identical attribute values, and therefore protecting the identity & privacy of individual users. The K-Anonymity method has been the basis of a proposed reliable transactional model that facilitates transactions between energy nodes, and electric vehicles. In this model, the K-Anonymity method serves two functions: first, it hides the location of the EVs by constructing a unified request using K-Anonymity techniques that conceal or hide the location of the owner of the car; second, the K-Anonymity method conceals user identifiers so that attackers are not left with the option to link users to their electric vehicles.

Evaluation and Situation Analysis

In this section, we will be talking about comprehensive analysis and evaluation of ten privacy protection systems using the fusion of blockchain and AI technologies that have been proposed in recent years. The evaluation focuses on five major characteristics of these proposed methods including: authority management, data protection, access control, scalability and network security, and also discusses the strengths, weaknesses, and potential areas of improvement. It's the unique features resulting from the integration of AI and blockchain technologies that have paved ways for new ideas, and solutions for enhanced privacy protection. For reference, the image below shows different evaluation metrics employed to derive the analytical results for the combined application of the blockchain and AI technologies.

Authority Management

Access control is a security & privacy technology that is used to restrict a user’s access to authorized resources on the basis of pre-defined rules, set of instructions, policies, safeguarding data integrity, and system security. There exists an intelligent privacy parking management system that makes use of a Role-Based Access Control or RBAC model to manage permissions. In the framework, each user is assigned one or more roles, and are then classified according to roles that allows the system to control attribute access permissions. Users on the network can make use of their blockchain address to verify their identity, and get attribute authorization access.

Access Control

Access control is one of the key fundamentals of privacy protection, restricting access based on group membership & user identity to ensure that it is only the authorized users who can access specific resources that they are allowed to access, and thus protecting the system from unwanted to forced access. To ensure effective and efficient access control, the framework needs to consider multiple factors including authorization, user authentication, and access policies.

Digital Identity Technology is an emerging approach for IoT applications that can provide safe & secure access control, and ensure data & device privacy. The method proposes to use a series of access control policies that are based on cryptographic primitives, and digital identity technology or DIT to protect the security of communications between entities such as drones, cloud servers, and Ground Station Servers (GSS). Once the registration of the entity is completed, credentials are stored in the memory. The table included below summarizes the types of defects in the framework.

Data Protection

Data protection is used to refer to measures including data encryption, access control, security auditing, and data backup to ensure that the data of a user is not accessed illegally, tampered with, or leaked. When it comes to data processing, technologies like data masking, anonymization, data isolation, and data encryption can be used to protect data from unauthorized access, and leakage. Furthermore, encryption technologies such as homomorphic encryption, differential privacy protection, digital signature algorithms, asymmetric encryption algorithms, and hash algorithms, can prevent unauthorized & illegal access by non-authorized users and ensure data confidentiality.

Network Security

Network security is a broad field that encompasses different aspects including ensuring data confidentiality & integrity, preventing network attacks, and protecting the system from network viruses & malicious software. To ensure the safety, reliability, and security of the system, a series of secure network architectures and protocols, and security measures need to be adopted. Furthermore, analyzing and assessing various network threats and coming up with corresponding defense mechanisms and security strategies are essential to improve the reliability & security of the system.

Scalability

Scalability refers to a system’s ability to handle larger amounts of data or an increasing number of users. When designing a scalable system, developers must consider system performance, data storage, node management, transmission, and several other factors. Furthermore, when ensuring the scalability of a framework or a system, developers must take into account the system security to prevent data breaches, data leaks, and other security risks.

Developers have designed a system in compliance with European General Data Protection Rules or GDPR by storing privacy-related information, and artwork metadata in a distributed file system that exists off the chain. Artwork metadata and digital tokens are stored in OrbitDB, a database storage system that uses multiple nodes to store the data, and thus ensures data security & privacy. The off-chain distributed system disperses data storage, and thus improves the scalability of the system.

Situation Analysis

The amalgamation of AI and blockchain technologies has resulted in developing a system that focuses heavily on protecting the privacy, identity, and data of the users. Although AI data privacy systems still face some challenges like network security, data protection, scalability, and access control, it is crucial to consider and weigh these issues on the basis of practical considerations during the design phase comprehensively. As the technology develops and progresses further, the applications expand, the privacy protection systems built using AI & blockchain will draw more attention in the upcoming future. On the basis of research findings, technical approaches, and application scenarios, they can be classified into three categories.

  • Privacy protection method application in the IoT or Internet of Things industry by utilizing both blockchain and AI technology.
  • Privacy protection method application in smart contract and services that make use of both blockchain and AI technology.
  • Large-scale data analysis methods that offer privacy protection by utilizing both blockchain and AI technology.

The technologies belonging to the first category focus on the implementation of AI and blockchain technologies for privacy protection in the IoT industry. These methods use AI techniques to analyze high volumes of data while taking advantage of decentralized & immutable features of the blockchain network to ensure authenticity and security of the data.

The technologies falling in the second category focus on fusing AI & Blockchain technologies for enhanced privacy protection by making use of blockchain’s smart contract & services. These methods combine data analysis and data processing with AI and use blockchain technology alongside to reduce dependency on trusted third parties, and record transactions.

Finally, the technologies falling in the third category focus on harnessing the power of AI and blockchain technology to achieve enhanced privacy protection in large-scale data analytics. These methods aim to exploit blockchain’s decentralization, and immutability properties that ensure the authenticity & security of data while AI techniques ensure the accuracy of data analysis.

Conclusion

In this article, we have talked about how AI and Blockchain technologies can be used in sync with each other to enhance the applications of privacy protection technologies by talking about their related methodologies, and evaluating the five primary characteristics of these privacy protection technologies. Furthermore, we have also talked about the existing limitations of the current systems. There are certain challenges in the field of privacy protection technologies built upon blockchain and AI that still need to be addressed like how to strike a balance between data sharing, and privacy preservation. The research on how to effectively merge the capabilities of AI and Blockchain techniques is going on, and here are several other ways that can be used to integrate other techniques.

  • Edge Computing

Edge computing aims to achieve decentralization by leveraging the power of edge & IoT devices to process private & sensitive user data. Because AI processing makes it mandatory to use substantial computing resources, using edge computing methods can enable the distribution of computational tasks to edge devices for processing instead of migrating the data to cloud services, or data servers. Since the data is processed much nearer the edge device itself, the latency time is reduced significantly, and so is the network congestion that enhances the speed & performance of the system.

  • Multi-chain Mechanisms

Multi-chain mechanisms have the potential to resolve single-chain blockchain storage, and performance issues, therefore boosting the scalability of the system. The integration of multi-chain mechanisms facilitates distinct attributes & privacy-levels based data classification, therefore improving storage capabilities and security of privacy protection systems.

ChatGPT-supported Bing Chat is now available in Microsoft Launcher

How to use the new Bing

Microsoft unveiled its Launcher application in 2017 to enable users to bring the Microsoft interface to their Android phones' home screen experience. Now, Launcher is getting an update that will incorporate generative AI.

Also: Everything we're expecting at Amazon's Devices and Services event this week

On Friday, the tech giant announced Bing Chat's integration into Microsoft Launcher's home screen experience for Android.

This means Launcher users on Android will no longer have to download and open the Bing app to start a conversation with Bing. Instead, all the users need to do is swipe down to access the Launcher's search functionality, where a Bing Chat icon will appear.

Once you tap the Bing Chat icon, you're taken to the Bing Chat interface, where you can ask questions and write prompts like you would with it regularly.

This feature is especially convenient if you use the image search feature. If you want to identify something, such as a flower you saw on your walk, you can quickly open the Bing Chat search interface and upload the image there.

Also: Back to school? How ChatGPT can help you with your essay writing

Microsoft also announced a "Continue on Your Phone" feature, which allows users to continue a Bing Chat conversation that is occurring on their desktop on their phone by simply scanning a QR code.

To access the QR code, all you have to do is click on the "Continue on phone" option in the upper-right corner of the last chat response. After scanning, you will be able to pick up right where you left off on your smartphone.

Artificial Intelligence

Ensemble Learning Techniques: A Walkthrough with Random Forests in Python

Machine learning models have become an integral component of decision-making across multiple industries, yet they often encounter difficulty when dealing with noisy or diverse data sets. That is where Ensemble Learning comes into play.

This article will demystify ensemble learning and introduce you to its powerful random forest algorithm. No matter if you are a data scientist looking to hone your toolkit or a developer looking for practical insights into building robust machine learning models, this piece is meant for everyone!

By the end of this article, you will gain a thorough knowledge of Ensemble Learning and how Random Forests in Python work. So whether you are an experienced data scientist or simply curious to expand your machine-learning abilities, join us on this adventure and advance your machine-learning expertise!

1. What Is Ensemble Learning?

Ensemble learning is a machine learning approach in which predictions from multiple weak models are combined with each other to get stronger predictions. The concept behind ensemble learning is decreasing the bias and errors from single models by leveraging the predictive power of each model.

To have a better example let's take a life example imagine that you have seen an animal and you do not know what species this animal belongs to. So instead of asking one expert, you ask ten experts and you will take the vote of the majority of them. This is known as hard voting.

Hard voting is when we take into account the class predictions for each classifier and then classify an input based on the maximum votes to a particular class. On the other hand, soft voting is when we take into account the probability predictions for each class by each classifier and then classify an input to the class with maximum probability based on the average probability (averaged over the classifier's probabilities) for that class.

2. When to Use Ensemble Learning

Ensemble learning is always used to improve the model performance which includes improving the classification accuracy and decreasing the mean absolute error for regression models. In addition to this ensemble learners always yield a more stable model. Ensemble learners work at their best when the models are not correlated then every model can learn something unique and work on improving the overall performance.

3. Ensemble Learning Strategies

Although ensemble learning can be applied in many ways, however when it comes to applying it to practice there are three strategies that have gained a lot of popularity due to their easy implementation and usage. These three strategies are:

  1. Bagging: Bagging which is short for bootstrap aggregation is an ensemble learning strategy in which the models are trained using random samples of the data set.
  2. Stacking: Stacking which is short for stacked generalization is an ensemble learning strategy in which we train a model to combine multiple models trained on our data.
  3. Boosting: Boosting is an ensemble learning technique that focuses on selecting the misclassified data to train the models on.

Let's dive deeper into each of these strategies and see how we can use Python to train these models on our dataset.

4. Bagging Ensemble Learning

Bagging takes random samples of data, and uses learning algorithms and the mean to find bagging probabilities; also known as bootstrap aggregating; it aggregates results from multiple models to get one broad outcome.

This approach involves:

  1. Splitting the original dataset into multiple subsets with replacement.
  2. Develop base models for each of these subsets.
  3. Running all models concurrently before running all predictions through to obtain final predictions.

Scikit-learn provides us with the ability to implement both a BaggingClassifier and BaggingRegressor. A BaggingMetaEstimator identifies random subsets of an original dataset to fit each base model, then aggregates individual base model predictions?—?either through voting or averaging?—?into a final prediction by aggregating individual base model predictions into an aggregate prediction using voting or averaging. This method reduces variance by randomizing their construction process.

Let's take an example in which we use the bagging estimator using scikit learn:

from sklearn.ensemble import BaggingClassifier  from sklearn.tree import DecisionTreeClassifier  bagging = BaggingClassifier(base_estimator=DecisionTreeClassifier(),n_estimators=10, max_samples=0.5, max_features=0.5)

The bagging classifier takes into consideration several parameters:

  • base_estimator: The base model used in the bagging approach. Here we use the decision tree classifier.
  • n_estimators: The number of estimators we will use in the bagging approach.
  • max_samples: The number of samples that will be drawn from the training set for each base estimator.
  • max_features: The number of features that will be used to train each base estimator.

Now we will fit this classifier on the training set and score it.

bagging.fit(X_train, y_train)  bagging.score(X_test,y_test)

We can do the same for regression tasks, the difference will be that we will be using regression estimators instead.

from sklearn.ensemble import BaggingRegressor  bagging = BaggingRegressor(DecisionTreeRegressor())  bagging.fit(X_train, y_train)  model.score(X_test,y_test)

5. Stacking Ensemble Learning

Stacking is a technique for combining multiple estimators in order to minimize their biases and produce accurate predictions. Predictions from each estimator are then combined and fed into an ultimate prediction meta-model trained through cross-validation; stacking can be applied to both classification and regression problems.

Ensemble Learning Techniques: A Walkthrough with Random Forests in Python
Stacking ensemble learning

Stacking occurs in the following steps:

  1. Split the data into a training and validation set
  2. Divide the training set into K folds
  3. Train a base model on k-1 folds and make predictions on the k-th fold
  4. Repeat until you have a prediction for each fold
  5. Fit the base model on the whole training set
  6. Use the model to make predictions on the test set
  7. Repeat steps 3–6 for other base models
  8. Use predictions from the test set as features of a new model (the meta model)
  9. Make final predictions on the test set using the meta-model

In this example below, we begin by creating two base classifiers (RandomForestClassifier and GradientBoostingClassifier) and one meta-classifier (LogisticRegression) and use K-fold cross-validation to use predictions from these classifiers on training data (iris dataset) for input features for our meta-classifier (LogisticRegression).

After using K-fold cross-validation to make predictions from the base classifiers on test data sets as input features for our meta-classifier, predictions on test sets using both sets together and evaluate their accuracy against their stacked ensemble counterparts.

# Load the dataset  data = load_iris()  X, y = data.data, data.target    # Split the data into training and testing sets  X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)    # Define base classifiers  base_classifiers = [     RandomForestClassifier(n_estimators=100, random_state=42),     GradientBoostingClassifier(n_estimators=100, random_state=42)  ]    # Define a meta-classifier  meta_classifier = LogisticRegression()    # Create an array to hold the predictions from base classifiers  base_classifier_predictions = np.zeros((len(X_train), len(base_classifiers)))    # Perform stacking using K-fold cross-validation  kf = KFold(n_splits=5, shuffle=True, random_state=42)  for train_index, val_index in kf.split(X_train):     train_fold, val_fold = X_train[train_index], X_train[val_index]     train_target, val_target = y_train[train_index], y_train[val_index]       for i, clf in enumerate(base_classifiers):         cloned_clf = clone(clf)         cloned_clf.fit(train_fold, train_target)         base_classifier_predictions[val_index, i] = cloned_clf.predict(val_fold)    # Train the meta-classifier on base classifier predictions  meta_classifier.fit(base_classifier_predictions, y_train)    # Make predictions using the stacked ensemble  stacked_predictions = np.zeros((len(X_test), len(base_classifiers)))  for i, clf in enumerate(base_classifiers):     stacked_predictions[:, i] = clf.predict(X_test)    # Make final predictions using the meta-classifier  final_predictions = meta_classifier.predict(stacked_predictions)    # Evaluate the stacked ensemble's performance  accuracy = accuracy_score(y_test, final_predictions)  print(f"Stacked Ensemble Accuracy: {accuracy:.2f}")

6. Boosting Ensemble Learning

Boosting is a machine learning ensemble technique that reduces bias and variance by turning weak learners into strong learners. These weak learners are applied sequentially to the dataset; firstly by creating an initial model and fitting it to the training set. Once errors from the first model have been identified, another model is designed to correct them.

There are popular algorithms and implementations for boosting ensemble learning techniques. Let's explore the most famous ones.

6.1. AdaBoost

AdaBoost is an effective ensemble learning technique, that employs weak learners sequentially for training purposes. Each iteration prioritizes incorrect predictions while decreasing weight assigned to correctly predicted instances; this strategic emphasis on challenging observations compels AdaBoost to become increasingly accurate over time, with its ultimate prediction determined by aggregating majority votes or weighted sum of its weak learners.

AdaBoost is a versatile algorithm suitable for both regression and classification tasks, but here we focus on its application to classification problems using Scikit-learn. Let’s look at how we can use it for classification tasks in the example below:

from sklearn.ensemble import AdaBoostClassifier  model = AdaBoostClassifier(n_estimators=100)  model.fit(X_train, y_train)  model.score(X_test,y_test)

In this example, we used the AdaBoostClassifier from scikit learn and set the n_estimators to 100. The default learn is a decision tree and you can change it. In addition to this, the parameters of the decision tree can be tuned.

2. EXtreme Gradient Boosting (XGBoost)

eXtreme Gradient Boosting or is more popularly known as XGBoost, is one of the best implementations of boosting ensemble learners due to its parallel computations which makes it very optimized to run on a single computer. XGBoost is available to use through the xgboost package developed by the machine learning community.

import xgboost as xgb  params = {"objective":"binary:logistic",'colsample_bytree': 0.3,'learning_rate': 0.1,                 'max_depth': 5, 'alpha': 10}  model = xgb.XGBClassifier(**params)  model.fit(X_train, y_train)  model.fit(X_train, y_train)  model.score(X_test,y_test)

3. LightGBM

LightGBM is another gradient-boosting algorithm that is based on tree learning. However, it is unlike other tree-based algorithms in that it uses leaf-wise tree growth which makes it converge faster.

Ensemble Learning Techniques: A Walkthrough with Random Forests in Python
Leaf-wise tree growth / Image by LightGBM

In the example below we will apply LightGBM to a binary classification problem:

import lightgbm as lgb  lgb_train = lgb.Dataset(X_train, y_train)  lgb_eval = lgb.Dataset(X_test, y_test, reference=lgb_train)  params = {'boosting_type': 'gbdt',               'objective': 'binary',               'num_leaves': 40,               'learning_rate': 0.1,               'feature_fraction': 0.9               }  gbm = lgb.train(params,     lgb_train,     num_boost_round=200,     valid_sets=[lgb_train, lgb_eval],     valid_names=['train','valid'],    )

Ensemble learning and random forests are powerful machine learning models that are always used by machine learning practitioners and data scientists. In this article, we covered the basic intuition behind them, when to use them, and finally, we covered the most popular algorithms of them and how to use them in Python.

References

  • A Gentle Introduction to Ensemble Learning Algorithms
  • A Comprehensive Guide to Ensemble Learning: What Exactly Do You Need to Know
  • A Comprehensive Guide to Ensemble Learning (with Python codes)

Youssef Rafaat is a computer vision researcher & data scientist. His research focuses on developing real-time computer vision algorithms for healthcare applications. He also worked as a data scientist for more than 3 years in the marketing, finance, and healthcare domain.

More On This Topic

  • Decision Trees vs Random Forests, Explained
  • When Would Ensemble Techniques be a Good Choice?
  • Classification Metrics Walkthrough: Logistic Regression with Accuracy,…
  • Hyperparameter Tuning Using Grid Search and Random Search in Python
  • Microsoft Explores Three Key Mysteries of Ensemble Learning
  • A Comprehensive Guide to Ensemble Learning – Exactly What You Need to Know

Everything we’re expecting at Amazon’s Devices and Services event this week

Amazon boxes with the Amazon logo stacked on top of each other

Amazon is constantly looking for new ways to innovate through hardware and software, from Echo devices to new Alexa capabilities. On Wednesday, September 20, the company will hold its annual Devices and Services event at its new headquarters in Arlington, VA. The event is an opportunity to unveil new products and improvements made to its platforms.

Also: Amazon's Echo Show 5 made me a smart display believer (and my daughter, too)

Last year, we saw the launch of the Kindle Scribe, a new generation of Echo Dot and Fire TV Cube, the addition of spatial audio to the Echo Studio, a Fire TV Pro Remote, and the eero built-in capability added to some Echo devices. Though we can't say for certain what is on the docket this year, we've got a pretty good idea of what could be.

Amazon

Python in Excel: This Will Change Data Science Forever

Python in Excel: This Will Change Data Science Forever
Image by Author

As a data scientist working in industry, the past year has felt like a rollercoaster ride of new tech breakthroughs and AI innovations.

Tools like ChatGPT, Notable, Pandas AI, and the Code Interpreter have saved me considerable amounts of time in performing tasks like writing, research, programming, and data analysis.

And just when I thought things couldn’t get any better, Microsoft and Anaconda announced the integration of Python into Excel!

You can now write Python code to analyze data, build machine learning models, and create visualizations within Excel spreadsheets.

So…Why the Hype Around a Python-Excel Integration?

The ability to write Python code within Excel will open new doors for data scientists and analysts.

When I got my first data science job, I assumed I’d be doing most of my work in Jupyter Notebooks. To my surprise, I ended up having to learn to use Excel on my first day of the job, since upper management, stakeholders, and clients preferred to interpret results from spreadsheets.

In fact, I’ve even created Tableau dashboards in the past to present results to clients, only to end up rebuilding the charts in Excel since they were more familiar with the platform.

And this isn’t unique to my organization. As of 2023, over a million companies and 1.5 billion people around the world use Excel.

Many data practitioners, like myself, find themselves constantly switching between Python IDEs and Excel spreadsheets. We use the former to build machine learning models and analyze data, and the latter to present our findings.

A Python-Excel integration will help data scientists and analysts streamline our workflows, by allowing us to perform data analysis, modeling, and presentation within a single platform.

Still not convinced?

Let’s explore some potential use cases of this combination.

Ways Data Scientists Can Use Python in Excel

Here are some ways in which data scientists can combine the functionality of spreadsheets with Python’s vast array of libraries:

1. Data Pre-Processing

If there is one part of my job I would gladly outsource, it is data preparation. This is a cumbersome task that becomes extremely time-consuming when using native Excel functions.

With the new Python-Excel integration, users can now import libraries like Pandas directly into Excel, and perform advanced filtering and data aggregation directly within Excel spreadsheets.

You can simply type “=PY” into a cell in a spreadsheet and highlight the data you want to analyze with Python, and a Pandas dataframe will be created for you. You can proceed to group and manipulate this data as you would in a Jupyter Notebook.

Here is an example of how you can create a Pandas dataframe in Excel:

Python in Excel: This Will Change Data Science Forever
Source: Microsoft

2. Machine Learning

While Excel offers basic tools like linear regression and trendline fitting in charts, most machine-learning use cases require more complex modeling techniques that go beyond the native capabilities of Excel.

With this Python-Excel integration, users can now build and train advanced statistical models within Excel using libraries like Scikit-Learn. The model outcomes can be visualized and presented in Excel, bridging the gap between modeling and decision-making in a single platform.

Here is an image showcasing just how simple it is to build a decision tree classifier in Excel with Python:

Python in Excel: This Will Change Data Science Forever
Source: Microsoft

3. Data Analysis

The process of analyzing data in Excel can be painstaking — when working with multiple files at once, users need to copy and paste data manually, drag formulas across cells, and combine data manually.

For example, if I have five sheets of monthly sales data that looks like this:

XXXXX

If I wanted to find products with more than 100 units sold in the span of a month, I’d first have to manually copy data from all sheets and paste it below the data in the first sheet. Then, I’d have to change the date format and create a pivot table.

Finally, I’d have to add a filter to find the products that match my criteria.

Every time I get new sales data in a different file or sheet, I need to copy and paste it manually.

This process becomes increasingly difficult and error-prone as the amount of data increases.

Instead, the entire analysis can be streamlined in Python using the following lines of code:

# 1. Merge the data  df_merged = pd.concat([df_jan, df_feb], ignore_index=True)    # 2. Convert the date format  df_merged['Date'] = pd.to_datetime(df_merged['Date']).dt.strftime('%Y-%m-%d')    # 3. Compute the total units sold for each product  grouped_data = df_merged.groupby('Product').agg({'Units Sold': 'sum'}).reset_index()    # 4. Identify products that sold more than 100 units  products_over_100 = grouped_data[grouped_data['Units Sold'] > 100]    products_over_100

Every time new data comes in, I just need to change one line of code and re-run the program to get the desired result. With a Python-Excel integration, I get to maximize efficiency while overseeing the entire data analysis workflow within a single platform.

4. Data Visualization

Although Excel itself offers a multitude of visualization options, the tool is still somewhat limited in the types of charts you can build. Charts like violin plots, heatmaps, and pair plots aren’t readily available in Excel, making it difficult for data scientists to represent complex statistical relationships.

The ability to run Python code will allow Excel users to use libraries like Matplotlib and Seaborn to create more complex, highly customizable charts.

Python in Excel: This Will Change Data Science Forever
Source: Microsoft How Can You Use Python in Excel?

At the time of writing this article, the Python-Excel feature is only available via the Microsoft 365 Insider Program. You need to sign up and choose the Beta Channel Insider level to access this feature, since it hasn’t been rolled out to the public yet.

Once you join the 365 Insider program, you will find a Python section in the Formulas tab. You just need to click on “Insert Python.” You can click on it to start writing your own Python code.

Alternatively, you can just type =PY into any cell to get started.

Python in Excel: This Will Change Data Science Forever
Source: Anaconda A Python-Excel Integration Will Democratize Data Science

With the release of ChatGPT, along with plugins such as the Code Interpreter and Notable, many tasks that once required strong technical expertise have become easier to perform.

This is especially true for data scientists and analysts — you can now upload CSV files to ChatGPT, and it will clean, analyze, and build models on your datasets.

In my opinion, the Python-Excel integration brings us one step closer to the democratization of data science and analytics.

In fields like marketing and finance, industry experts who work solely in Excel will now be able to execute Python code to analyze their data without even having to download a programming IDE.

The ability to work with data in an interface they are familiar with, coupled with ChatGPT’s proficiency in writing code, will allow non-programmers to perform data science workflows and solve problems with Python code.

If you are an Excel user who doesn’t know how to code, this is a great opportunity for you to learn Python programming within an interface that you are already comfortable with.
Natassha Selvaraj is a self-taught data scientist with a passion for writing. You can connect with her on LinkedIn.

More On This Topic

  • Forget ChatGPT, This New AI Assistant Is Leagues Ahead and Will Change the…
  • Automate Microsoft Excel and Word Using Python
  • Do You Read Excel Files with Python? There is a 1000x Faster Way
  • Free Microsoft Excel for Beginners Course
  • Essential Math for Data Science: Basis and Change of Basis
  • Relax! Data Scientists will not go extinct in 10 years, but the role will…

Procurement management platform Levelpath raises $30M

Procurement management platform Levelpath raises $30M Kyle Wiggers 8 hours

Back in 2014, Stan Garber and Alex Yakubovich set out to reinvent the request for proposals (RFP) process with the launch of Scout RFP, which provides a cloud-based sourcing solution designed to help organizations source faster — and, ideally, easier. Scout RFP was acquired by Workday in 2019, and Garber and Yakubovich decided to stay on under Workday’s management following the purchase. But while at Workday, the pair experienced major challenges with business procurement.

“After Workday acquired Scout RFP, we began experiencing everyday pain points — chasing down the right person to approve a happy hour budget, finding which vendors we could purchase swag from or even getting an NDA spun up to sign,” Yakubovich told TechCruch via email. “Then, of course, there are the endless emails and Slacks associated with these various activities. The time-consuming nature of the tasks created barriers to getting our work done, and we realized just how much data was siloed and the time inefficiencies it was causing across the organization.”

These blockers drove Garber and Yakubovich to found Levelpath, a software-as-a-service platform to manage various enterprise procurement services. In an apparent sign that they’re on the right path (no pun intended), Levelpath today announced that it raised $30 million in a Series A round led by Redpoint with participation from Menlo Ventures, which follows an unannounced $14.5 million seed round led by Benchmark with participation from NewView Capital and World Innovation Lab and brings the startup’s total raised to $44.5 million.

It’s not a secret that enterprises struggle with procurement. According to a Harvard Business Review study from 2020, approximately 60% of business leaders say that a lack of transparency between their finance and procurement functions represents a risk to their business. Data quality and governance are frequently cited as the biggest roadblocks for procurement teams, which often struggle to gain visibility into procurement processes.

“Companies are looking to save money and focus more on the procurement process,” Yakubovich, who serves as Levelpath’s CEO, told TechCrunch. “It’s our hypothesis that creating an enjoyable experience will help maximize adoption and, in turn, company-wide efficiencies and immediate return on investment.”

Now, it should be noted that lots of startups are angling to corner the market for procurement software, which Fortune Business Insights valued at $6.15 billion in 2021.

Zip is one of the larger players in the sector, having recently raised $100 million at a $1.5 billion valuation. Fintech startup Ramp expanded into procurement just a few months ago. And then there’s small-time, more specialized vendors like Focal Point, Keelvar and Tropic.

So what sets Levelpath apart?

Yakubovich claims it’s the platform’s mobile-first (yes, really) interface, which he describes rather subjectively as “next-gen” and “easy-to-use.” While Levelpath can be used by smaller firms, Yakubovich says, it’s intended for enterprises managing hundreds or thousands of vendors and employees — offering tools customized for each company’s approval workflows.

Levelpath

Levelpath’s procurement management interface, which is mobile-centric.

“Often, what seems like a simple procurement request, such as purchasing a software seat, turns into an endless trail of phone calls and emails — trying to garner approval from all the right people,” Yakubovich said. “For example, suppose a marketing executive is purchasing swag for an event. In that case, a filter their company may have on is one that flags any purchases over 5,000 to the department head for approval. If this person’s purchase is only 3,000, they may proceed to their list of approved vendors and begin ordering. Or, if they’re signing a sponsorship contract, it may connect this person with their department head and the legal team to complete their specific approval process … Levelpath funnels their responses to the right procurement leaders.”

AI plays a differentiating role, too, according to Yakubovich. Algorithms built into the Levelpath platform provide “actionable insights” to reduce instances of vendor redundancy. And Levelpath’s building an AI model that understands the purchasing and workflow habits of employees and adjusts the procurement experience based on this. For example, if someone in an organization wanted to purchase software, Levelpath would run through the enrichment data and let the user know of software vendors that offer similar, potentially cheaper products that meet their typical criteria.

The goal is to help companies decide where to consolidate and restructure their services, Yakubovich says. “We’re the first platform to build with the end-user experience as the guiding light for our entire product roadmap,” he added. “Our mission is to make procurement delightful.”

That’s an ambitious mission. But Levelpath claims to have dozens of enterprise customers already, including Ace Hardware, Qualtrics and Innovacare.

With a staff of 26 employees — a number Yakubovich expects will double next year — Levelpath plans to commence a broader go-to-market strategy in 2024 while putting its latest funding tranche toward product development and research.

“Companies are looking to save money and focus more on the procurement process,” Yakubovich said. “It’s never been a better time to invest in this space.”

This Indian AI Healthcare Model Outperformed GPT-4 and MedPaLM

Recently, ChatGPT helped diagnose a child’s disease after a string of doctors failed to identify the condition. When use cases for large language models in healthcare are slowly finding its way, an Indian health-tech startup has created a model that has not only outperformed OpenAI’s GPT-4 and Google’s MedPaLM on USMLE (US Medical Licensing Examination), but is available for free and has already helped over 2000 people since its launch.

Indian ‘August AI’

AI Model’s performance in USMLE. Source: GetBeyondHealth

August AI, a large language model created by Bangalore-based health AI startup Beyond, aims to democratise access to high-quality health information. “August has been developed by a team of engineers, data scientists, and doctors, and aims to bridge the gap between doctors and patients,” said Samarth Sharma, part of the founding team at Beyond, in an interview with AIM. “Most importantly, August has been trained on proprietary health data that Beyond has generated over the years.” The model scored 94.8% on USMLE, whereas, MedPaLM scored 85%, and GPT-4 87.8%.

August is multimodal, that accepts input in various formats such as audio, text and PDF Lab, and delivers the output in text format. It currently aids people in handling physical and mental health concerns and simplifying complex health procedures. The platform is currently free for use on Whatsapp.

“Whatsapp is already on 2 billion phones around the world and people are comfortable using it, so it made sense to bring August to users through this platform,” said Sharma. “We want August to be accessible to everyone, regardless of socioeconomic status and technical sophistication.” Though the company is working on creating an app for August, they believe that Whatsapp is a great way to access the platform, and their key focus is to enable natural conversation on Whatsapp.

Fusion of LLMs

August is built on a combination of LLMs from multiple providers. “Over the past year, we have experimented with, disassembled, and fine-tuned various large language models to arrive at the core engine behind August’s health AI,” said Sharma. “We’re currently using a combination of BERT, LlaMA, GPT-4, to name a few, and most recently have experimented with LLama-2.

August has completed 1500 health consultations over the last year to understand the best ways to guide a person, and even fine-tuned the model around those conversations. “Utilising a proprietary, highly-tuned version of ensemble refinement, August provides high-quality answers to health-related questions, said Sharma, who has optimised this technique for improving token efficiency and minimising hallucinations.

Manoeuvring AI Healthcare

It has been noted in the past where GPT-4 has inclined towards providing societal biases in clinical decisions. With the medical field being a critical use case, any misinformation can have grave consequences, and August AI is critically working towards tackling it.

“Understanding the health of people from different ethnicities is still something that the overall healthcare ecosystem is not good at, and that bias translates to LLMs as well,” said Sharma. The company has been focusing on India and has been collecting India-specific data that has been integrated to August AI.

“August does not do any diagnosis. We don’t think health AIs are there yet,” said Sharma. “August is highly tuned to answer only health-information questions. The ensemble refinement deals with hallucinations and we’re using retrieval augmented generation to ensure it is drawing information from reliable sources when someone asks for specific information like exercise or products”

A Long Path Ahead

August AI has an elaborate roadmap where they wish to scale the product through direct outreach to users and partnerships with distribution platforms that can benefit from having August as part of their ecosystem. The company even wishes to integrate booking of doctor appointments in the future.

Sharma believes that August AI is much more than just medical competency which is Google MedPaLM’s focus. “We’re doing much better than Google MedPaLM on the critical elements of empathy, holding a conversation and proactively checking in on people. While Google is building for the US, August’s focus on India and its empathetic conversation will be key differentiators for us,” said Sharma.

The post This Indian AI Healthcare Model Outperformed GPT-4 and MedPaLM appeared first on Analytics India Magazine.

How to Identify Missing Data in Time-Series Datasets

Time-series data, collected nearly every second from a multiplicity of sources, is often subjected to several data quality issues, among which missing data.

In the context of sequential data, missing information can arise due to several reasons, namely errors occurring on acquisition systems (e.g. malfunction sensors), errors during the transmission process (e.g., faulty network connections), or errors during data collection (e.g., human error during data logging). These situations often generate sporadic and explicit missing values in our datasets, corresponding to small gaps in the stream of collected data.

Additionally, missing information can also arise naturally due to the characteristics of the domain itself, creating larger gaps in the data. For instance, a feature that stops being collected for a certain period of time, generating non-explicit missing data.

Regardless of the underlying cause, having missing data in our time-series sequences is highly prejudicial for forecasting and predictive modeling and may have serious consequences for both individuals (e.g., misguided risk assessment) and business outcomes (e.g., biased business decisions, loss of revenue and opportunities).

When preparing the data for modeling approaches, an important step is therefore being able to identify these patterns of unknown information, as they will help us decide on the best approach to handle the data efficiently and improve its consistency, either through some form of alignment correction, data interpolation, data imputation, or in some cases, casewise deletion (i.e., omit cases with missing values for a feature used in a particular analysis).

For that reason, performing a thorough exploratory data analysis and data profiling is indispensable not only to understand the data characteristics but also to make informed decisions on how to best prepare the data for analysis.

In this hands-on tutorial, we’ll explore how ydata-profiling can help us sort out these issues with the features recently introduced in the new release. We’ll be using the U.S. Pollution Dataset, available in Kaggle (License DbCL v1.0), that details information regarding NO2, O3, SO2, and CO pollutants across U.S. states.

Hands-on Tutorial: Profiling the U.S. Pollution Dataset

To kickstart our tutorial, we first need to install the latest version of ydata-profiling:

pip install ydata-profiling==4.5.1

Then, we can load the data, remove unnecessary features, and focus on what we aim to investigate. For the purpose of this example, we will focus on the particular behavior of air pollutants’ measurements taken at the station of Arizona, Maricopa, Scottsdale:

import pandas as pd     data = pd.read_csv("data/pollution_us_2000_2016.csv")  data = data.drop('Unnamed: 0', axis = 1) # dropping unnecessary index      # Select data from Arizona, Maricopa, Scottsdale (Site Num: 3003)  data_scottsdale = data[data['Site Num'] == 3003].reset_index(drop=True)

Now, we’re ready to start profiling our dataset! Recall that, to use the time-series profiling, we need to pass the parameter tsmode=True so that ydata-profiling can identify time-dependent features:

# Change 'Data Local' to datetime  data_scottsdale['Date Local'] = pd.to_datetime(data_scottsdale['Date Local'])     # Create the Profile Report  profile_scottsdale = ProfileReport(data_scottsdale, tsmode=True, sortby="Date Local")  profile_scottsdale.to_file('profile_scottsdale.html')

Time-Series Overview

The output report will be familiar to what we already know, but with an improved experience and new summary statistics for time-series data:

XXXXX

Immediately from the overview, we can get an overall understanding of this dataset by looking at the presented summary statistics:

  • It contains 14 different time-series, each with 8674 recorded values;
  • The dataset reports on 10 years of data from January 2000 to December 2010;
  • The average period of time sequences is 11 hours and (nearly) 7 minutes. This means that on average, we have measures being taken every 11 hours.

We can also get an overview plot of all series in data, either in their original or scaled values: we can easily grasp the overall variation of the sequences, as well as the components (NO2, O3, SO2, CO) and characteristics (Mean, 1st Max Value, 1st Max Hour, AQI) being measured.

Inspecting Missing Data

After having an overall view of the data, we can focus on the specifics of each time sequence.

In the latest release of ydata-profiling, the profiling report was substantially improved with dedicated analysis for time-series data, namely reporting on the “Time Series” and “Gap Analysis”' metrics. The identification of trends and missing patterns is extremely facilitated by these new features, where specific summary statistics and detailed visualizations are now available.

Something that stands out immediately is the flaky pattern that all time series present, where certain “jumps” seem to occur between consecutive measurements. This indicates the presence of missing data (“gaps” of missing information) that should be studied more closely. Let’s take a look at the S02 Mean as an example.

XXXXX

XXXXX

When investigating the details given in the Gap Analysis, we get an informative description of the characteristics of the identified gaps. Overall, there are 25 gaps in the time-series, with a minimum length of 4 days, a maximum of 32 weeks, and an average of 10 weeks.

From the visualization presented, we note somewhat “random” gaps represented by thinner stripes, and larger gaps which seem to follow a repetitive pattern. This indicates that we seem to have two different patterns of missing data in our dataset.

Smaller gaps correspond to sporadic events generating missing data, most likely occurring due to errors in the acquisition process, and can often be easily interpolated or deleted from the dataset. In turn, larger gaps are more complex and need to be analyzed in more detail, as they may reveal an underlying pattern that needs to be addressed more thoroughly.

In our example, if we were to investigate the larger gaps, we would in fact discover that they reflect a seasonal pattern:

df = data_scottsdale.copy()  for year in df["Date Local"].dt.year.unique():      for month in range(1,13):          if ((df["Date Local"].dt.year == year) & (df["Date Local"].dt.month ==month)).sum() == 0:              print(f'Year {year} is missing month {month}.')
# Year 2000 is missing month 4.  # Year 2000 is missing month 5.  # Year 2000 is missing month 6.  # Year 2000 is missing month 7.  # Year 2000 is missing month 8.  # (...)  # Year 2007 is missing month 5.  # Year 2007 is missing month 6.  # Year 2007 is missing month 7.  # Year 2007 is missing month 8.  # (...)  # Year 2010 is missing month 5.  # Year 2010 is missing month 6.  # Year 2010 is missing month 7.  # Year 2010 is missing month 8.

As suspected, the time-series presents some large information gaps that seem to be repetitive, even seasonal: in most years, the data was not collected between May to August (months 5 to 8). This may have occurred due to unpredictable reasons, or known business decisions, for example, related to cutting costs, or simply related to seasonal variations of pollutants associated with weather patterns, temperature, humidity, and atmospheric conditions.

Based on these findings, we could then investigate why this happened, if something should be done to prevent it in the future, and how to handle the data we currently have.

Final Thoughts: Impute, Delete, Realign?

Throughout this tutorial, we’ve seen how important it is to understand the patterns of missing data in time-series and how an effective profiling can reveal the mysteries behind gaps of missing information. From telecom, healthcare, energy, and finance, all sectors collecting time-series data will face missing data at some point and will need to decide the best way to handle and extract all possible knowledge from them.

With a comprehensive data profiling, we can make an informed and efficient decision depending on the data characteristics at hand:

  • Gaps of information can be caused by sporadic events that derive from errors in acquisition, transmission, and collection. We can fix the issue to prevent it from happening again and interpolate or impute the missing gaps, depending on the length of the gap;
  • Gaps of information can also represent seasonal or repeated patterns. We may choose to restructure our pipeline to start collecting the missing information or replace the missing gaps with external information from other distributed systems. We can also identify if the process of retrieval was unsuccessful (maybe a fat-finger query on the data engineering side, we all have those days!).

I hope this tutorial has shed some light on how to identify and characterize missing data in your time-series data appropriately and I can’t wait to see what you’ll find in your own gap analysis! Drop me a line in the comments for any questions or suggestions or find me at the Data-Centric AI Community!
Fabiana Clemente is cofounder and CDO of YData, combining data understanding, causality, and privacy as her main fields of work and research, with the mission of making data actionable for organizations. As an enthusiastic data practitioner she hosts the podcast When Machine Learning Meets Privacy and is a guest speaker on the Datacast and Privacy Please podcasts. She also speaks at conferences such as ODSC and PyData.

More On This Topic

  • Handling Missing Values in Time-series with SQL
  • 4 Factors to Identify Machine Learning Solvable Problems
  • KDnuggets News, June 29: 20 Basic Linux Commands for Data Science…
  • Multidimensional multi-sensor time-series data analysis framework
  • Teaching AI to Classify Time-series Patterns with Synthetic Data
  • Market Data and News: A Time Series Analysis