5 Simple Steps Series: Master Python, SQL, Scikit-learn, PyTorch & Google Cloud

5 Simple Steps Series: Master Python, SQL, Scikit-learn, PyTorch & Google Cloud
New to the world of data science and machine learning? Welcome to your ultimate guide and starting point, whether you’re looking to break into the industry, learn something new or sharpen your current skills. The Back to Basics: Getting Started in 5 Steps series is all you need and is designed to turn complex concepts into simple straightforward knowledge.

As part of KDnuggets 30 year journey within the data science, machine learning and artificial intelligence space, the team have come together to curate a variety of articles for you to soak up all the knowledge you can.

When starting something new, it’s always hard to get started. The KDnuggets team are taking that weight off your shoulder with our Back to Basics: Getting Started in 5 Steps series, which includes:

  • Python Data Structures
  • SQL
  • Scikit-learn
  • PyTorch
  • Google Cloud Platform

So let’s get right into it…

Python Data Structures in 5 Steps

This tutorial covers Python's foundational data structures — lists, tuples, dictionaries, and sets. Learn their characteristics, use cases, and practical examples, all in 5 steps.
When it comes to learning how to program, regardless of the particular programming language you use for this task, you find that there are a few major topics of your newly chosen discipline to which most of what you are being exposed could be categorized.

A few of these, in general order of grokking, are syntax (the vocabulary of the language); commands (putting the vocabulary together into useful ways); flow control (how we guide the order of command execution); algorithms (the steps we take to solve specific problems… how did this become such a confounding word?); and, finally, data structures (the virtual storage depots that we use for data manipulation during the execution of algorithms (which are, again… a series of steps).

Learn the 5 steps: Getting Started with Python Data Structures in 5 Steps

SQL in 5 Steps

This comprehensive SQL tutorial covers everything from setting up your SQL environment to mastering advanced concepts like joins and subqueries, and optimizing query performance. With step-by-step examples, this guide is perfect for beginners looking to enhance their data management skills.

When it comes to managing and manipulating data in relational databases, Structured Query Language (SQL) is the biggest name in the game. SQL is a major domain-specific language which serves as the cornerstone for database management and provides a standardized way to interact with databases.

With data being the driving force behind decision-making and innovation, SQL remains an essential technology demanding top-level attention from data analysts, developers, and data scientists.

Learn the 5 steps: Getting Started with SQL in 5 Steps

Scikit-learn in 5 Steps

This tutorial offers a comprehensive hands-on walkthrough of machine learning with Scikit-learn. Readers will learn key concepts and techniques including data preprocessing, model training and evaluation, hyperparameter tuning, and compiling ensemble models for enhanced performance.

When learning about how to use Scikit-learn, we must obviously have an existing understanding of the underlying concepts of machine learning, as Scikit-learn is nothing more than a practical tool for implementing machine learning principles and related tasks. Machine learning is a subset of artificial intelligence that enables computers to learn and improve from experience without being explicitly programmed. The algorithms use training data to make predictions or decisions by uncovering patterns and insights.

Learn the 5 steps: Getting Started with Scikit-learn in 5 Steps

PyTorch in 5 Steps

This tutorial provides an in-depth introduction to machine learning using PyTorch and its high-level wrapper, PyTorch Lightning. The article covers essential steps from installation to advanced topics, offering a hands-on approach to building and training neural networks, and emphasizing the benefits of using Lightning.

PyTorch is a popular open-source machine learning framework based on Python and optimized for GPU-accelerated computing. Originally developed by Meta AI in 2016 and now part of the Linux Foundation, PyTorch has quickly become one of the most widely used frameworks for deep learning research and applications.

Unlike some other frameworks like TensorFlow, PyTorch uses dynamic computation graphs which allow for greater flexibility and debugging capabilities.

Learn the 5 steps: Getting Started with PyTorch in 5 Steps

Google Cloud Platform in 5 Steps

Explore the essentials of Google Cloud Platform for data science and ML, from account setup to model deployment, with hands-on project examples.

This article aims to provide a step-by-step overview of getting started with Google Cloud Platform (GCP) for data science and machine learning. We'll give an overview of GCP and its key capabilities for analytics, walk through account setup, explore essential services like BigQuery and Cloud Storage, build a sample data project, and use GCP for machine learning.

Whether you're new to GCP or looking for a quick refresher, read on to learn the basics and hit the ground running with Google Cloud.

Learn the 5 steps: Getting Started with Google Cloud Platform in 5 Steps

That’s a Wrap

This Back to Basics: Getting Started in 5 Steps series will have enlightened you on the foundational tools used in data science. You will have familiarised yourself with the basics of Python, SQL, machine learning with Scikit-learn and PyTorch, but also ventured into the Google Cloud Platform.

The path to data mastery does not end here, it is an ongoing journey which requires you to continuously learn new skills and the tools acquired to be proficient.

Keep an eye on KDnuggets for more insights, advanced guides, and the support of a community that’s just as passionate about data science as you are.

Nisha Arya is a Data Scientist and Freelance Technical Writer. She is particularly interested in providing Data Science career advice or tutorials and theory based knowledge around Data Science. She also wishes to explore the different ways Artificial Intelligence is/can benefit the longevity of human life. A keen learner, seeking to broaden her tech knowledge and writing skills, whilst helping guide others.

More On This Topic

  • Getting Started with Scikit-learn in 5 Steps
  • Getting Started with Google Cloud Platform in 5 Steps
  • Simplifying Decision Tree Interpretability with Python & Scikit-learn
  • Multilabel Classification: An Introduction with Python's Scikit-Learn
  • How to Speed up Scikit-Learn Model Training
  • The Best Machine Learning Frameworks & Extensions for Scikit-learn

Brave’s Leo AI assistant is now available to desktop users

Brave’s Leo AI assistant is now available to desktop users Ivan Mehta 8 hours

Brave, a company building an alternative web browser, is releasing its AI-powered assistant, Leo, to all desktop users. The company is also releasing a $15 per month paid version called Leo Premium with features like access to faster and better large language models (LLMs) and higher-rate limits.

Leo has been under testing for a few months. Brave made it available to its Nightly version users in August. Starting today, it will be available to all Brave desktop browser users with version 1.60.

Users can access the Leo assistant by clicking the Leo icon in the sidebar to start conversing or just by typing a question in the address bar and clicking the Leo icon to get a direct answer.

Brave Leo

Image Credits: Brave

Brave’s AI-powered assistant can handle context-aware requests such as summarizing webpages or videos, translating text, and rewriting phrases.

Leo is based on Llama 2 and Anthropic’s Claude LLMs. While free users get the basic version of these models, paying users will get access to models like Llama 2 70B, Code Llama 70B, and Anthropic Claude Instant. These models enable faster and more accurate responses.

Brave said that all requests to Leo use an anonymous server as a proxy, so they can’t be linked back to a particular IP. Additionally, the company specified that responses are immediately discarded after generation, and not stored on any server or used to train models. Brave noted that all subscriptions are validated by unlinkable tokens, so the company can’t know about your activity or your email.

Browser-based AI assistants and features are now becoming more common. Some browsers including Opera and Microsoft Edge have also introduced AI-assistant in the sidebar. Other upstarts such as SigmaOS and Browser Company’s Arc have experimented with different formats and features for AI-based features. As more browsers are adopting AI, they will have to innovate beyond summarization and rewriting features to help companies monetize.

Brave, which laid off 9% of its staff in October, is currently focusing on generating more revenue. In April, the company Search stopped using Bing’s index to start depending on its own indexing solution. In May, Brave launched its own search API for clients, with prices starting from $3 per 1,000 queries.

Now You Can Reserve NVIDIA GPUs on AWS

Now You Can Reserve NVIDIA GPUs on AWS

After a month of struggle, AWS has now introduced a solution to address the surging demand for GPU compute capacity for machine learning workloads. The company announced the general availability of Amazon Elastic Compute Cloud (EC2) Capacity Blocks for ML, offering an innovative consumption model for customers, powered by NVIDIA H100 GPUs.

With this new offering, customers can access highly sought-after GPU compute capacity on a flexible, short-term basis, removing the need for long-term commitments.

“This is an innovative new way to schedule GPU instances where you can reserve the number of instances you need for a future date for just the amount of time you require,” said Channy Yun.

EC2 Capacity Blocks allow customers to reserve GPU capacity for durations ranging from one to 14 days, with advanced scheduling up to eight weeks. These capacity blocks are deployed in EC2 UltraClusters with low-latency, high-throughput connectivity, offering the flexibility to scale up to hundreds of GPUs. This new solution is ideal for training and fine-tuning ML models, short experimentation runs, and handling temporary surges in inference demand for product launches.

David Brown, the vice president of compute and networking at AWS, highlighted the importance of this innovation in democratising access to generative AI capabilities. AWS and NVIDIA have collaborated for over a decade to deliver scalable GPU solutions, and the introduction of Amazon EC2 Capacity Blocks is a significant step in broadening access to GPU capacity for generative AI applications.

The EC2 Capacity Blocks are available for reservation in the AWS US East (Ohio) Region, with plans for expansion to additional AWS Regions and Local Zones. This development is a game-changer for startups and organisations looking to harness the power of generative AI without making long-term capital commitments.

The launch of EC2 Capacity Blocks has received positive feedback from industry leaders and organisations, such as Amplify Partners, Canva, Leonardo.Ai, and OctoML. These stakeholders believe that this new solution will provide the predictability and timely access to GPU compute capacity necessary to drive innovation and meet customer demands in today’s supply-constrained environment.

The post Now You Can Reserve NVIDIA GPUs on AWS appeared first on Analytics India Magazine.

An Honest Comparison of Open Source Vector Databases

An Honest Comparison of Open Source Vector Databases
Image frm DALL-E 3

Vector databases offer a wide range of benefits, particularly in generative artificial intelligence (AI), and more specifically, large language models (LLMs). These benefits can range from advanced indexing to accurate similarity searches, helping to deliver powerful, state-of-the-art projects,

In this article, we will provide an honest comparison of three open-source vector databases that have established an impressive reputation—Chroma, Milvus, and Weaviate. We will explore their use cases, key features, performance metrics, supported programming languages, and more to provide a comprehensive and unbiased overview of each database.

What Are Vector Databases?

In its most simplistic definition, a vector database stores information as vectors (vector embeddings), which are a numerical version of a data object.

As such, vector embeddings are a powerful method of indexing and searching across very large and unstructured or semi-unstructured datasets. These datasets can consist of text, images, or sensor data and a vector database orders this information into a manageable format.

Vector databases work using high-dimensional vectors which can contain hundreds of different dimensions, each linked to a specific property of a data object. Thus creating an unrivaled level of complexity.

Not to be confused with a vector index or a vector search library, a vector database is a complete management solution to store and filter metadata in a way that is:

  • Is completely scalable
  • Can be easily backed up
  • Enables dynamic data changes
  • Provides a high level of security

The Benefits of Using Open Source Vector Databases

Open source vector databases provide numerous benefits over licensed alternatives, such as:

  • They are a flexible solution that can be easily modified to suit specific needs, unlike licensed options which are typically designed for a particular project.
  • Open source vector databases are supported by a large community of developers who are ready to assist with any issues or provide advice on how projects could be improved.
  • An open-source solution is budget-friendly with no licensing fees, subscription fees, or any unexpected costs during the project.
  • Due to the transparent nature of open-source vector databases, developers can work more effectively, understanding every component and how the database was built.
  • Open source products are constantly being improved and evolving with changes in technology as they are backed by active communities.

Open Source Vector Databases Comparison: Chroma Vs. Milvus Vs. Weaviate

Now that we have an understanding of what a vector database is and the benefits of an open-source solution, let’s consider some of the most popular options on the market. We will focus on the strengths, features, and uses of Chroma, Milvus, and Weaviate, before moving on to a direct head-to-head comparison to determine the best option for your needs.

1. Chroma

Chroma is designed to assist developers and businesses of all sizes with creating LLM applications, providing all the resources necessary to build sophisticated projects. Chroma ensures a project is highly scalable and works in an optimal way so that high-dimensional vectors can be stored, searched for, and retrieved quickly.

It has grown in popularity due to its reputation as being an extremely flexible solution, with a wide range of deployment options. In addition, Chroma can be deployed directly on the cloud or it can be run on-site, making it a viable option for any business, regardless of its IT infrastructure.

Use Cases

Multiple data types and formats are also supported by Chroma, making it suitable for almost any application. However, one of Chroma’s key strengths is its support for audio data, making it a top choice for audio-based search engines, music recommendation applications, and other sound-based projects.

2. Milvus

Milvus has gained a strong reputation in the world of ML and data science, boasting impressive capabilities in terms of vector indexing and querying. Utilizing powerful algorithms, Milvus offers lightning-fast processing and data retrieval speeds and GPU support, even when working with very large datasets. Milvus can also be integrated with other popular frameworks such as PyTorch and TensorFlow, allowing it to be added to existing ML workflows.

Use Cases

Milvus is renowned for its capabilities in similarity search and analytics, with extensive support for multiple programming languages. This flexibility means developers aren't limited to backend operations and can even perform tasks typically reserved for server-side languages on the front end. For example, you could generate PDFs with JavaScript while leveraging real-time data from Milvus. This opens up new avenues for application development, especially for educational content and apps focusing on accessibility.

This open-source vector database can be used across a wide range of industries and in a large number of applications. Another prominent example involves eCommerce, where Milvus can power accurate recommendation systems to suggest products based on a customer’s preferences and buying habits.

It’s also suitable for image/ video analysis projects, assisting with image similarity searches, object recognition, and content-based image retrieval. Another key use case is natural language processing (NLP), providing document clustering and semantic search capabilities, as well as providing the backbone to question and answer systems.

3. Weaviate

The third open source vector database in our honest comparison is Weaviate, which is available in both a self-hosted and fully-managed solution. Countless businesses are using Weaviate to handle and manage large datasets due to its excellent level of performance, its simplicity, and its highly scalable nature.

Capable of managing a range of data types, Weaviate is very flexible and can store both vectors and data objects which makes it ideal for applications that need a range of search techniques (E.G. vector searches and keyword searches).

Use Cases

In terms of its use, Weaviate is perfect for projects like Data classification in enterprise resource planning software or applications that involve:

  • Similarity searches
  • Semantic searches
  • Image searches
  • eCommerce product searches
  • Recommendation engines
  • Cybersecurity threat analysis and detection
  • Anomaly detection
  • Automated data harmonization

Now we have a brief understanding of what each vector database can offer, let’s consider the finer details that set each open source solution apart in our handy comparison table.

Comparison Table

Chroma Milvus Weaviate
Open Source Status Yes — Apache-2.0 license Yes — Apache-2.0 license Yes — BSD-3-Clause license
Publication Date February 2023 October 2019 January 2021
Use Cases Suitable for a wide range of applications, with support for multiple data types and formats.

Specializes in Audio-based search projects and image/video retrieval.

Suitable for a wide range of applications, with support for a plethora of data types and formats.

Perfect for eCommerce recommendation systems, natural language processing, and image/video-based analysis

Suitable for a wide range of applications, with support for multiple data types and formats.

Ideal for Data classification in enterprise resource planning software.

Key Features Impressive ease of use.

Development, testing, and production environments all use the same API on a Jupyter Notebook.

Powerful search, filter, and density estimation functionality.

Uses both in-memory and persistent storage to provide high-speed query and insert performance.

Provides automatic data partitioning, load balancing, and fault tolerance for large-scale vector data handling.

Supports a variety of vector similarity search algorithms.

Offers a GraphQL-based API, providing flexibility and efficiency when interacting with the knowledge graph.

Supports real-time data updates, to ensure the knowledge graph remains up-to-date with the latest changes.

Its schema inference feature automates the process of defining data structures.

Supported Programming Languages Python or JavaScript Python, Java, C++, and Go Python, Javascript, and Go
Community and Industry Recognition Strong community with a Discord channel available to answer live queries. Active community on GitHub, Slack, Reddit, and Twitter.

Over 1000 enterprise users.

Extensive documentation.

Dedicated forum and active Slack, Twitter, and LinkedIn communities. Plus regular Podcasts and newsletters.

Extensive documentation.

Performance Metrics N/A https://milvus.io/docs/benchmark.md https://weaviate.io/developers/weaviate/benchmarks/ann
GitHub Stars 9k 23.5k 7.8k

Conclusion

Each open-source vector database in our honest comparison guide is powerful, scalable, and completely free. This can make choosing the perfect solution a little difficult but the process can be made easier by knowing the exact project you are working on and the level of support required.

Chroma is the newest solution and is not as well backed as the other two in terms of community support, however, its ease of use and flexibility make it a great option, especially for projects that involve audio search.

Milvus has the highest GitHub Star rating and strong community support, with an impressive number of enterprise businesses trusting this vector database to meet their needs. Therefore, Milvus is a good choice for natural language processing and image/ video analysis projects.

Finally, Weaviate offers self-hosted and fully managed solutions, with extensive documentation and support available. A key use case is data classification in enterprise resource planning software, but this solution is perfect for a range of projects.

Nahla Davies is a software developer and tech writer. Before devoting her work full time to technical writing, she managed—among other intriguing things—to serve as a lead programmer at an Inc. 5,000 experiential branding organization whose clients include Samsung, Time Warner, Netflix, and Sony.

More On This Topic

  • Python Vector Databases and Vector Indexes: Architecting LLM Apps
  • Qdrant: Open-Source Vector Search Engine with Managed Cloud Platform
  • A Comprehensive Guide to Pinecone Vector Databases
  • What are Vector Databases and Why Are They Important for LLMs?
  • Closed Source VS Open Source Image Annotation
  • A Critical Comparison of Machine Learning Platforms in an Evolving Market

UK AI Safety Summit: Global Powers Make ‘Landmark’ Pledge to AI Safety

Representatives from 28 countries and tech companies convened on the historic site of Bletchley Park in the U.K. for a landmark two-day summit held Nov. 1-2, 2023, focusing on the safety and regulation of artificial intelligence. Day one of the AI Safety Summit culminated in the signing of the “landmark” Bletchley Declaration on AI Safety, which commits 28 participating countries — including the U.K., U.S. and China — to jointly manage and mitigate risks from artificial intelligence while ensuring safe and responsible development and deployment.

Jump to:

  • What is the Bletchley Declaration on AI safety?
  • What is the AI Safety Summit?
  • Who is attending the AI Safety Summit?
  • What are experts’ reactions to the AI Safety Summit?
  • Why is AI safety important?
  • How has the UK invested in AI?

What is the Bletchley Declaration on AI safety?

The Bletchley Declaration states that developers of advanced and potentially dangerous AI technologies shoulder a significant responsibility for ensuring their systems are safe through rigorous testing protocols and safety measures to prevent misuse and accidents.

It also emphasizes the need for common ground in understanding AI risks and fostering international research partnerships in AI safety while recognizing that there is “potential for serious, even catastrophic, harm, either deliberate or unintentional, stemming from the most significant capabilities of these AI models.”

U.K. Prime Minister Rishi Sunak called the signing of the declaration “a landmark achievement that sees the world’s greatest AI powers agree on the urgency behind understanding the risks of AI.”

In a written statement, Sunak said: “Under the UK’s leadership, more than twenty five countries at the AI Safety Summit have stated a shared responsibility to address AI risks and take forward vital international collaboration on frontier AI safety and research.

“The UK is once again leading the world at the forefront of this new technological frontier by kickstarting this conversation, which will see us work together to make AI safe and realize all its benefits for generations to come.”

What is the AI Safety Summit?

The AI Safety Summit is a major conference taking place on November 1 and 2, 2023 in Buckinghamshire, U.K. It will bring together international governments, technology companies and academia to consider the risks of AI “at the frontier of development” and discuss how these risks can be mitigated through a united, global effort.

The inaugural day of the AI Safety Summit saw a series of talks from business leaders and academics aimed at promoting a deeper understanding of what the U.K. government has dubbed “frontier AI” — advanced artificial intelligence systems that could pose as-yet unknown risks to society.

This included a number of roundtable discussions with “key developers,” including OpenAI, Anthropic and U.K.-based Google DeepMind, that centered on how risk thresholds, effective safety assessments and robust governance and accountability mechanisms can be defined.

SEE: ChatGPT Cheat Sheet: Complete Guide for 2023 (TechRepublic)

The first day of the summit also featured a virtual address by King Charles III, who labeled AI one of humanity’s “greatest technological leaps” and highlighted the technology’s potential in transforming healthcare and various other aspects of life. The British Monarch called for robust international coordination and collaboration to ensure AI remains a secure and beneficial technology.

Day two of the summit will feature a press conference and closing remarks from Prime Minister Sunak.

Who is attending the AI Safety Summit?

Representatives from the Alan Turing Institute, Stanford University, the Organisation for Economic Co-operation and Development and the Ada Lovelace Institute are among the attendees at the AI Safety Summit, alongside tech companies including Google, Microsoft, IBM, Meta and AWS, as well as leaders such as SpaceX boss Elon Musk. Also in attendance is U.S. Vice President Kamala Harris.

What are experts’ reactions to the AI Safety Summit?

Poppy Gustafsson, chief executive officer of AI cybersecurity company Darktrace, told PA Media she had been concerned that discussions would focus too much on “hypothetical risks of the future” — like killer robots — but that the discussions were more “measured” in reality.

Rajesh Ganesan, president of Zoho-owned ManageEngine, commented in an email statement that, “While some may be disappointed if the summit falls short of establishing a global regulatory body,” the fact that global leaders were discussing AI regulation was a positive step forward.

“Gaining international agreement on the mechanisms for managing the risks posed by AI is a significant milestone — greater collaboration will be paramount to balancing the benefits of AI and limiting its damaging capacity,” Ganesan said in a statement.

“It’s clear that regulation and security practices will remain critical to the safe adoption of AI and must keep pace with its rapid advancements. This is something that the EU’s AI Act and the G7 Code of Conduct agreements could drive and provide a framework for.”

Ganesan added: “We need to prioritize ongoing education and give people the skills to use generative AI systems securely and safely. Failing to make AI adoption about the people who use and benefit from it risks dangerous and suboptimal outcomes.”

Why is AI safety important?

There is currently no comprehensive set of regulations governing the use of artificial intelligence, though the European Union has drafted a framework that aims to establish rules for the technology in the 28-nation bloc.

The potential misuse of AI, either maliciously or via human or machine error, remains a key concern. The summit heard that cybersecurity vulnerabilities, biotechnological dangers and the spread of disinformation represented some of the most significant threats posted by AI, while issues with algorithmic bias and data privacy were also highlighted.

U.K. Technology Secretary Michelle Donelan emphasized the importance of the Bletchley Declaration as a first step in ensuring the safe development of AI. She also stated that international cooperation was essential to building public trust in AI technologies, adding that “no single country can face down the challenges and risks posed by AI alone.”

She noted: “Today’s landmark Declaration marks the start of a new global effort to build public trust by ensuring the technology’s safe development.”

How has the UK invested in AI?

On the eve of the UK AI Safety Summit, the UK government announced £118 million ($143 million) funding to boost AI skills funding in the United Kingdom. The funding will target research centers, scholarships and visa schemes and aims to encourage young people to study AI and data science fields.

Meanwhile, £21 million ($25.5 million) has been earmarked for equipping the U.K.’s National Health Service with AI-powered diagnostic technology and imaging technology, such as X-rays and CT scans.

Repeated Data Leaks Cast Doubt on India Stack

India achieved significant success with its Digital Public Goods (DPG) and is now sharing these open-source technologies with the rest of the world, particularly nations in the global south. While addressing the G20 digital economy ministers’ meeting in Bengaluru earlier this year, Indian Prime Minister Narendra Modi said that India is ready to share its experiences with the world.

“We offered our CoWIN (Covid Vaccine Intelligence Work) platform for global good during the COVID pandemic. We have now created an online global public digital goods repository — the India Stack. This is to ensure that no one is left behind,” he said.

India Stack is a set of open APIs (Application Programming Interfaces) and DPG that includes Aadhaar, Unified Payment Interface (UPI) and DigiLocker, among others. Several countries, including Antigua and Barbuda in the Caribbean, Trinidad and Tobago, Sierra Leone in Africa, Suriname in South America, Armenia in Eastern Europe, and Papua New Guinea in Southeast Asia, have expressed interest in India Stack.

Frequent data breaches are a great concern

The India Stack has had a significant impact on financial inclusion, economic development, and innovation in India. However, recently, there has been a big controversy around it related to data leaks.

Last month, a hacker put on sale a massive database of Personal Identifiable Information (PII) of 815 million Indians on the dark web, which includes information like Aadhaar and passport details, as well as individuals’ names, phone numbers, and addresses.

When Resecurity, a US-based cybersecurity firm that reported the breach, reached out to this hacker, they were willing to sell the complete dataset for Aadhaar and Indian passports for USD 80,000. The alleged data breach, which is being seen as the biggest data breach in India’s history, raises serious questions about the security and reliability of India’s digital public infrastructure.

Interestingly in 2018, a Chandigarh-based newspaper reported that around a billion Indian’s PII were sold online for a few dollars. They also claimed that there is a software available for purchase on the internet that can create counterfeit Aadhaar cards.

In June of this year, Malayala Manorama reported that a bot operating on the Telegram messaging platform was responsible for disclosing the personal information of Indian citizens who had registered on the CoWIN portal for vaccination purposes.

The bot allegedly exposed sensitive details such as individuals’ names, Aadhaar numbers, and passport numbers when provided with their phone numbers.

Similarly, in 2020, a flaw in DigiLocker allowed hackers to access over 3.8 crore accounts without passwords. The same was reported by security researcher Ashish Ghalot, who found the flaw while analysing the authentication mechanism.

The frequent data breaches and the government’s indifferent attitude toward them have raised significant concerns. From civil rights activists, and technology lawyers to opposition leaders, many have expressed concerns about the recurrent data breaches and have questioned the government’s efforts to address the issue.

With alleged breaches occurring every few months, the Indian government should prioritise strengthening these technologies before promoting them to the global south, ensuring their robustness before sharing them with the world.

Apar Gupta, advocate and founding director of the Internet Freedom Foundation, in a LinkedIn post also raised similar concerns. “Seriously, what Digital Public Infrastructure is being built in India? How can we possibly offer a model for the democratic world?” he asked while sharing his thoughts on the recent data leak.

Should not lose public trust

According to the latest findings revealed by cybersecurity firm Surfshark, India has been ranked as the 7th most breached country in the world in the second quarter of 2023. These data breaches can potentially be extremely harmful and may lead to identity theft and banking fraud.

The frequent data breaches are already creating a sense of distrust among Indian citizens. This will lead to identity theft and other costs amounting to billions, according to Mishi Choudhary, technology lawyer, online civil rights activist and the founder of SFLC.in.

“With the apparent digital divide and exclusion from certain services due to these digital public infrastructures, there is already a sense of certain public distrust. Such breaches further lead to security and privacy concerns,” Vaishnavi Sharma, research associate at the Dialogue, a public-policy think tank, told AIM.

Moreover, the current situation highlights potential vulnerabilities within various government establishments, particularly in the realm of faster cyber threat detection and response, according to Mr. Kiran Vangaveti, founder & CEO at BluSapphire.

“Relying on the speculation that the breach originated from the government entity Indian Council of Medical Research (ICMR) raises crucial questions about accountability within government institutions, an aspect that was not adequately addressed in the Digital Personal Data Protection (DPDP) Act,” Vangaveti told AIM.

Better safeguarding of citizen’s data

Hence, before promoting these technologies to the democratic world, the Ministry of Electronics & IT (MeitY), Indian Computer Emergency Response Team (CERT-In) and the other parties involved should ensure that these platforms are more secure and that the data of its citizens are better protected.

“The benefits that India’s digital public infrastructure brings can barely be matched anywhere else in the world. Deployments have been fast-moving. However, focus towards ensuring robust and secure data protection practices is the need of the hour,” Choudhary told AIM.

Furthermore, data of millions of Indian citizens have been collected by the government in the absence of a concrete data protection law in the country. Alongside the relevant data protection laws, it would be important to establish institutional policies and technological protections to make systems more secure and privacy-friendly.

“While efforts are made to curtail data breaches ex post facto, it would be necessary to establish clear and transparent ex-ante protocols for data transfers between systems and methods of transfers, and delineate clear roles and responsibilities of stakeholders engaged not only in the sense of different governments (national, state, and district) but also horizontally between different departments who have access to these data,” Sharma said.

There must be a clear and transparent identification and assessment of risks and harms that apply across the processes, systems, and stakeholders including third-party players, Sharma adds. “This further includes capacity-building, conducting security training for employees, and emphasising the potential harm that data breaches could pose.”

If they cannot guarantee data security, the government should stop collecting PII for every interaction, according to Choudhary. Vangaveti also points out that a majority of Indians are not well-versed in privacy matters. Hence it is even more imperative that the government takes significant steps towards data protection of its citizens.

“Their interactions with government, quasi-government entities, financial institutions, and service providers often involve sharing physical copies of Aadhaar/passport without clear accountability for the storage, use, and protection of this sensitive data,” Vangaveti said.

He hopes this incident sparks a meaningful discourse on the government’s and its institutions’ responsibility in safeguarding the personal data of its citizens.

The post Repeated Data Leaks Cast Doubt on India Stack appeared first on Analytics India Magazine.

Factory wants to use AI to automate the software dev lifecycle

Factory wants to use AI to automate the software dev lifecycle Kyle Wiggers 8 hours

Developer velocity, the speed at which an organization ships code, is often impacted by necessary but lengthy processes like code review, writing documentation and testing. Inefficiencies threaten to make theses processes even longer. According to one source, developers waste 17.3 hours per week due to technical debt and bad — i.e. nonfunctional — code.

Machine learning Ph.D. Matan Grinberg and Eno Reyes, previously a data scientist at Hugging Face and Microsoft, thought that there had to be a better way.

During a Hackathon in San Francisco, Grinberg and Reyes built a platform that could autonomously solve simple coding problems — a platform that they later came to believe had commercial potential. After the hackathon, the pair expanded the platform to handle more software development tasks and founded a company, Factory, to monetize what they’d built.

“Factory’s mission is to bring autonomy to software engineering,” Grinberg told TechCrunch in an email interview. “More concretely, Factory helps large engineering organizations automate parts of their software development lifecycle via autonomous, AI-powered systems.”

Factory’s systems — which Grinberg calls “Droids,” a term Lucasfilm might have a problem with — are built to juggle various repetitive, mundane but normally time-consuming software engineering tasks. For example, Factory has “Droids” for reviewing code, refactoring or restructuring code and even generating new code from prompts a la GitHub Copilot.

Grinberg explains: “The review Droid leaves insightful code reviews and provides context for human reviewers on every change to the codebase. The documentation Droid generates and continually updates documentation as needed. The test Droid writes tests and maintains test coverage percentage as new code is merged. The knowledge Droid lives in your communication platform (e.g. Slack) and answers deeper questions about the engineering system. And the project Droid helps plan and design requirements based on customer support tickets and feature requests.”

All of Factory’s Droids are built on what Grinberg refers to as the “Droid core”: an engine that ingests and processes a company’s engineering system data to build a knowledge base, and an algorithm that pulls insights from the knowledge base to solve various engineering problems. A third Droid core component, Reflection Engine, acts as a filter for the third-party AI models that Factory leverages, enabling the company to implement its own safeguards, security best practices and so on on top of those models.

“The enterprise angle here is that this is a software suite that allows engineering organizations to output better product faster, while also improving engineering morale by lightening the load of tedious tasks like code review, docs and testing,” Grinberg said,. “Additionally, due to the autonomous nature of the Droids, little is required by way of user education and onboarding.”

Now, if Factory can consistently, reliably automate all those dev tasks, the platform would pay for itself indeed. According to a 2019 survey by Tidelift and The New Stack, developers spend 35% of their time managing code, including testing and responding to security issues — and less than a third of their time actually coding.

But the question is, can it?

Even the best AI models today aren’t above making catastrophic mistakes. And generative coding tools can introduce insecure code, with one Stanford study suggesting that software engineers who use code-generating AI are more likely to cause security vulnerabilities in the apps they develop.

Grinberg was upfront about the fact that Factory didn’t have the capital to train all of its models in house — and thus is at the mercy of third-party limitations. But, he asserts, the Factory platform is still delivering value while relying on third-party vendors for some AI muscle.

“Our approach is building these AI systems and reasoning architectures, making use of cutting-edge … models and establishing relationships with customers to deliver value now,” Grinberg said. “As an early startup, it’s a losing battle to train [large] models. Compared to incumbents, you have no monetary advantage, no chip access advantage, no data advantage and (almost certainly) no technical advantage.”

Factory’s long-term play is to train more of its own AI models to build an “end-to-end” engineering AI system — and to differentiate these models by soliciting engineering training data from its early customers, Grinberg said.

“As time goes on, we’ll have more capital, the chip shortage will clear up and we’ll have direct access (with permission) to a treasure trove of data (i.e. the historical timeline of entire engineering organizations),” he continued. “We’ll build Droids to be robust, fully autonomous — with minimal required human interaction — and tailored to customers’ needs from day one.”

Is that an overly optimistic view? Perhaps. The market for AI startups grows more competitive by the day.

But to Grinberg’s credit, Factory’s already working with a core group of around 15 companies. Grinberg wouldn’t name names, save the clients — which have used Factor’s platform to author thousands of code reviews and hundreds of thousands of lines of code to date — range in size from “seed stage” to “public.”

And Factory recently closed a $5 million seed round co-led by Sequoia and Lux with participation from SV Angel, BoxGroup, DataBricks CEO Ali Ghodsi, Hugging Face co-founder Clem Delangue and others. Grinberg says that the new capital will be put toward expanding Factory’s six-person team and platform capabilities.

“The major challenges in this AI code generation industry are trust and differentiation,” he said. “Every VP of engineering wants to improve their organization’s output with AI. What stands in the way of this is the unreliable nature of many AI tools, and the reticence of large, labyrinthine organizations to trust this new, futuristic sounding technology … Factory is building a world where software engineering itself is an accessible, scalable commodity.”

When Rajeev Chandrasekhar Met Elon Musk

When Rajeev Chandrasekhar Met Elon Musk

Rajeev Chandrasekhar, Minister of State for Electronic and Information Technology of India, posted on X, that he bumped into Elon Musk at the AI Safety Summit at Bletchley Park in UK.

Look who i bumped into at #AISafetySummit at Bletchley Park, UK.@elonmusk shared that his son with @shivon has a middle name "Chandrasekhar" – named after 1983 Nobel physicist Prof S Chandrasekhar pic.twitter.com/S8v0rUcl8P

— Rajeev Chandrasekhar 🇮🇳 (@Rajeev_GoI) November 2, 2023

In the post, Chandrasekhar wrote that Musk’s son with Shivon has the same middle name “Chandrasekhar”, which is after the 1983 Nobel physicist professor S Chandrasekhar.

Rajeev Chandrasekhar is representing India at this two-day summit, where the £80 million ($100 million) funding initiative is a collaborative effort between Britain, Canada, and the Bill and Melinda Gates Foundation. The objective is to promote “safe and responsible” programming. The UK’s AI for Development Programme will contribute £38 million to this collaboration, demonstrating the UK’s commitment to investing in partnerships that leverage cutting-edge technology to address global challenges.

Chandrasekhar emphasised India’s significant progress in digitising its economy over the past eight years. During the global AI summit hosted by Britain, he stated, “India has rapidly digitised its economy in the last eight years, moving from about 4.5 percent of the total GDP to a target of 20 percent by 2025-26. Presently, we are at about 11 percent, but the digital and innovation economies are growing 2.5 to 3 times faster than the non-digital part of the GDP.”

Chandrasekhar further conveyed Prime Minister Narendra Modi’s vision for a collaborative approach to shaping the future of technology and innovation, emphasising that the institutional framework should be sustained, strategic, and involve multiple nations, rather than just one or two countries. He underlined the vital role of artificial intelligence as a catalyst for the burgeoning digital economy, growth, and governance.

The summit brought together international digital ministers, technology sector leaders, top academics, and civil society representatives to discuss shared risks associated with emerging AI technology and potential mitigations.

Chandrasekhar highlighted the Indian government’s commitment to international collaboration, stating, “International conversations between countries, such as this Global AI summit, are extremely important as we shape the future of technology, especially in a year where technology presents unprecedented opportunities in the history of mankind.”

The post When Rajeev Chandrasekhar Met Elon Musk appeared first on Analytics India Magazine.

Building Data Pipelines to Create Apps with Large Language Models

Building Data Pipelines to Create Apps with Large Language Models
Image from DALL-E 3

Enterprises currently pursue two approaches for LLM powered Apps — Fine tuning and Retrieval Augmented Generation (RAG). At a very high level, RAG takes an input and retrieves a set of relevant/supporting documents given a source (e.g., company wiki). The documents are concatenated as context with the original input prompt and fed to the LLM model which produces the final response. RAG seems to be the most popular approach to get LLMs to market especially in real-time processing scenarios. The LLM architecture to support that most of the time includes building an effective data pipeline.

In this post, we’ll explore different stages in the LLM data pipeline to help developers implement production-grade systems which work with their data. Follow along to learn how to ingest, prepare, enrich, and serve data to power GenAI apps.

What are the Different Stages of an LLM pipeline?

These are the different stages of an LLM pipeline:

Data ingestion of unstructured data

Vectorization with enrichment (with metadata)

Vector indexing (with real-time syncing)

AI Query Processor

Natural Language User interaction (with Chat or APIs)

Building Data Pipelines to Create Apps with Large Language Models

Data Ingestion of unstructured data

The first step is gathering the right data to help with the business goals. If you are building a consumer facing chatbot then you have to pay special attention to what data is going to be used. The sources of data could range from a company portal (e.g. Sharepoint, Confluent, Document storage) to internal APIs. Ideally you want to have a push mechanism from these sources to the index so that your LLM app is up to date for your end consumer.

Organizations should implement data governance policies and protocols when extracting text data for LLM in context training. Organizations can start by auditing document data sources to catalog sensitivity levels, licensing terms and origin. Identify restricted data that needs redaction or exclusion from datasets.

These data sources should also be assessed for quality — diversity, size, noise levels, redundancy. Lower quality datasets will dilute the responses from LLM apps. You might even need an early document classification mechanism to help with the right kind of storage later in the pipeline.

Adhering to data governance guardrails, even in fast-paced LLM development, reduces risks. Establishing governance upfront mitigates many issues down the line and enables scalable, robust extraction of text data for context learning.

Pulling messages via Slack, Telegram, or Discord APIs gives access for real-time data is what helps RAG but raw conversational data contains noise — typos, encoding issues, weird characters. Filtering out messages real-time with offensive content or sensitive personal details that could be PII is an important part of data cleansing.

Vectorization with metadata

Metadata like author, date, and conversation context further enriches data. This embedding of external knowledge into vectors helps with smarter and targeted retrieval.

Some of the metadata related to documents could lie in the portal or in the document’s metadata itself, however if the document is attached to a business object( e.g. Case, Customer , Employee information) then you would have to fetch that information from a relational database. If there are security concerns around data access, this is a place where you can add security metadata which also helps with the retrieval stage later in the pipeline.

A critical step here is to convert text and images into vector representations using the LLM’s embedding models. For documents, you need to do chunking first, then you do encoding preferably using on-prem zero shot embedding models.

Vector indexing

Vector representations have to be stored somewhere. This is where vector databases or vector indexes are used to efficiently store and index this information as embeddings.

This becomes your “LLM source of truth” and this has to be in sync with your data sources and documents. Real-time indexing becomes important if your LLM app is servicing customers or generating business related information. You want to avoid your LLM app being out of sync with your data sources.

Fast retrieval with a query processor

When you have millions of enterprise documents, getting the right content based on the user query becomes challenging.

This is where the early stages of pipeline starts adding value : Cleansing and Data enrichment via metadata addition and most importantly Data indexing. This in-context addition helps with making prompt engineering stronger.

User interaction

In a traditional pipelining environment, you push the data to a data warehouse and the analytics tool will pull the reports from the warehouse. In an LLM pipeline, an end user interface is usually a chat interface which at the simplest level takes a user query and responds to the query.

Summary

The challenge with this new type of pipeline is not just getting a prototype but getting this running in production. This is where an enterprise grade monitoring solution to track your pipelines and vector stores becomes important. The ability to get business data from both structured and unstructured data sources becomes an important architectural decision. LLMs represent the state-of-the-art in natural language processing and building enterprise grade data pipelines for LLM powered apps keeps you at the forefront.

Here is access to a source available real-time stream processing framework.

Anup Surendran is a VP of Product and Product Marketing who specializes in bringing AI products to market. He has worked with startups that have had two successful exits (to SAP and Kroll) and enjoys teaching others about how AI products can improve productivity within an organization.

More On This Topic

  • How to Create Stunning Web Apps for your Data Science Projects
  • Top Open Source Large Language Models
  • More Free Courses on Large Language Models
  • Learn About Large Language Models
  • Introducing Healthcare-Specific Large Language Models from John Snow Labs
  • What are Large Language Models and How Do They Work?

“Leadership is Extremely Lonely,” says Narayana Murthy

Narayana Murthy Lonely

“Leadership is extremely lonely,” said Narayana Murthy, the Infosys co-founder on the second episode of 3one4 Capital podcast, The Record, with T V Mohandas Pai. This should have been the first episode instead, as it gives in so much more context to why Murthy urged youth to work 12-hours a day.

He said while leadership is very lonely, the memories of the joys, and the challenges that you faced in the organisation is what really gives you company. He explained further, saying how taking decisions required him oftentimes to not sleep at night and sit alone to think about it.

“I enjoyed those successes. I wept for those mistakes. I think in some sense, the experiences that you have gone through within the organisation is your only companion. There’s nobody else.”

Pai further asked Murthy why he has been obsessed with Infosys. “It’s 24/7 in your life,” highlighting how he spends most of his time working for Infosys. He asked about Murthy’s work-life balance.

Citing Mahatma Gandhi and Nelson Mandela, Murthy said that they were not successful in family life, and were attached to their passion – i.e. to build a great nation.

“When you are thinking about it all your waking hours, when you are spending all the disposable hours with that institution, it is quite possible that you would not be judged a successful family person,” said Murthy, saying that he was fortunate enough to have wife like Sudha Murthy, who understood Murthy’s vision.

“When we founded Infosys, I will look after the children. I won’t worry you about their issues. You go ahead and look after your child, Infosys,” recalled Murthy.

This clarification was absolutely necessary after Murthy, in the last episode, urged the youth to work 12 hours in a day. Murthy’s comment has been intensively discussed, and oftentimes scrutinised for urging people to work for around 70 hours in a week. There is always so much to learn from such leaders.

The post “Leadership is Extremely Lonely,” says Narayana Murthy appeared first on Analytics India Magazine.