Microsoft releases its internal generative AI red teaming tool to the public

Abstract tech colorful image

Despite the advanced capabilities of generative AI (gen AI) models, we have seen many instances of them going rogue, hallucinating, or having loopholes malicious actors can exploit. To help mitigate that issue, Microsoft is unveiling a tool that can help identify risks in generative AI systems.

On Thursday, Microsoft released its Python Risk Identification Toolkit for generative AI (PyRIT), a tool Microsoft's AI Red Team has been using to check for risks in its gen AI systems, including Copilot.

Also: How renaissance technologists are connecting the dots between AI and business

In the past year, Microsoft red-teamed more than 60 high-value gen AI systems, through which it learned that the red-teaming process differs vastly for these systems from classical AI or traditional software, according to the blog post.

The process looks different because Microsoft has to consider the usual security risks, in addition to responsible AI risks, such as ensuring harmful content cannot be intentionally generated, or that the models don't output disinformation.

Additionally, gen AI models vary widely in architecture, and there are deviations in outcomes that can be produced from the same input, making it difficult to find one streamlined process that fits all models.

Also: Want to work in AI? How to pivot your career in 5 steps

As a result, manually probing for all of these different risks ends up being a time-consuming, tedious, and slow process. Microsoft shares that automation can help red teams by identifying risky areas that require more attention and automating routine tasks, and that's where PyRIT comes in.

The toolkit, "battle-tested by the Microsoft AI team," sends a malicious prompt to the generative AI system, and once it receives a response, its scoring agent gives the system a score, which is used to send a new prompt based on previous scoring feedback.

Microsoft says that PyRIT's biggest advantage is that it has helped Microsoft's red team efforts be more efficient, significantly shortening the amount of time a task would take.

Also: How tech professionals can survive and thrive at work in the time of AI

"For instance, in one of our red teaming exercises on a Copilot system, we were able to pick a harm category, generate several thousand malicious prompts, and use PyRIT's scoring engine to evaluate the output from the Copilot system all in the matter of hours instead of weeks," said Microsoft in the release.

The toolkit is available for access today and includes a list of demos to help familiarize users with the tool. Microsoft is also hosting a webinar on PyRIT that demonstrates how to use it in red teaming generative AI systems, which you can register for through Microsoft's website.

Artificial Intelligence

Chrome gets a built-in AI writing tool powered by Gemini

Chrome gets a built-in AI writing tool powered by Gemini Frederic Lardinois @fredericl / 9 hours

Google Chrome is getting a new AI writing generator today. At its core, this Gemin-powered tool is essentially the existing “Help me write” feature from Gmail, but extended to the entire web and powered by one of Google’s latest Gemini AI models. The company first announced this new tool in January and it remains in its ‘experimental’ phase, meaning you must explicitly enable it.

To get started, head to the Chrome settings menu and look for the ‘Experimental AI’ page. From there, you can easily enable the new writing feature, as well as Google’s new automatic tab organizer (which I haven’t found particularly useful or smart so far) and the new Chrome theme manager). For now, the AI writer is only available in English on Windows, Mac and Linux. After that, right-click on any text field and select ‘Help me write.” You can use this to write something completely now Gemini can also rewrite existing text.

Image Credits: Google

If you’re subscribed to Gemini Advanced, this new tool will not give you access to an enhanced writing model, a Google spokesperson told us. It’s very much meant for short-form content like emails or support requests and a bigger model may not even be of much help there anyway.

One nifty feature here is that the tool will take into account the site you are on when it makes its recommendations “The tool will understand the context of the webpage you’re on to suggest relevant content,” Google engineering director Adriana Porter Felt writes in today’s announcement. “For example, if you’re writing a review for a pair of running shoes, Chrome will pull out key features from the product page that support your recommendation so it’s more valuable to potential shoppers.”

As with the ‘Help me write” feature in Gmail, it’s easy enough to change the length and tone of the results, too.

It’s important to note that the text, content and the URL of the page you are using the service on will be sent to Google under its existing privacy policy. Google explicitly notes that this information “is used to improve this feature, which includes generative model research and machine learning technologies,” which includes a review process with humans in the loop. Caveat scriptor.

Image Credits: Google

Google’s Duet AI can now write your emails for you

Chrome now has a new AI writing tool to help you write almost anything online

"Help me write" in Chrome

If you've ever struggled to find the right words online, a new AI-powered feature for the Google Chrome browser is here to help.

Late last year, Google began testing a "Help me write" feature for Chrome that promised to help users write just about any kind of text online — from reviews to customer service requests to marketplace listings.

Also: Need Google Chrome to load pages faster? Enable this feature to speed it up

Starting today, that feature is rolling out to all Chrome users.

To use it, just right-click in any open text field. After plugging in a simple prompt describing what you want written, you'll be asked to choose the length and tone. From there, the AI tool, which is powered by the company's Gemini models, works its magic to make you sound better.

"Help me write" will understand the context of the website you're on and suggest relevant content. For example, Google says, if you want to write a review about a pair of running shoes, the tool will pull out key features from the product page to make your review more valuable to other shoppers.

Also: How to write better ChatGPT prompts in 5 steps

If you don't like what was generated, there's a retry button that will take another shot. There's also an edit button to fine-tune your prompt or change the tone or length. Once you get your response, you can rate it with a thumbs-up or a thumbs-down to help the AI learn.

Already have something composed? You can get help rewriting existing text by highlighting it and then right-clicking.

The feature was already available for Gmail, Google Docs, and other Google Workspace products. And while you can accomplish the same thing with ChatGPT, this option will be more accessible to the common user.

Also: How ChatGPT (and other AI chatbots) can help you write an essay

In one example offered by Google, a user submits a prompt of "moving to a smaller place selling air fryer for 50 bucks." The AI finesses that text into a more appealing response of "I'm moving to a smaller place and won't have any room for my air fryer. It's in good condition and works great. I'm selling it for $50. Please contact me if you're interested."

To turn on this feature, click the three-dot menu on the top right of your Chrome window and head to settings. From there, you'll see an "Experimental AI" page. Clicking that presents several options you can enable, including "Help me write." Once it's turned on, right-click in any open text field and choose "Help me write."

Artificial Intelligence

Women in AI: Krystal Kauffman, research fellow at the Distributed AI Research Institute

Women in AI: Krystal Kauffman, research fellow at the Distributed AI Research Institute Kyle Wiggers 8 hours

To give AI-focused women academics and others their well-deserved — and overdue — time in the spotlight, TechCrunch is launching a series of interviews focusing on remarkable women who’ve contributed to the AI revolution. We’ll publish several pieces throughout the year as the AI boom continues, highlighting key work that often goes unrecognized. Read more profiles here.

Krystal Kauffman worked as an organizer on political and issue campaigns for a decade before pursuing a degree in geology. Then, she turned to gig work, which lead her to Turkopticon, a nonprofit organization dedicated to fighting for the rights of gig workers — specifically those using Amazon’s Mechanical Turk (AMT) platform.

Now the lead organizer at Turkopticon, Kauffman recently started as a research fellow with the Distributed AI Research Institute (DAIR) Institute, working alongside others to build — in her words — “a community of workers united in righting the wrongs of the big-tech marketplace platforms.”

Q&A

Briefly, how did you get your start in AI? What attracted you to the field?

In 2015, I became ill, and couldn’t work outside of my home. While doctors were trying to sort things out, I found the AMT platform. For the next two years, I was able to support myself doing data work in which I completed tasks that helped program AI, build LLMs and so on. During my time working on AMT, I became very passionate about solving issues with the platform and taking on the ethics of data work in general.

What work are you most proud of (in the AI field)?

When I first started data work nine years ago, very few people knew that there was a global workforce quietly programming smart devices, developing AI and building data sets from their homes. Over the last several years, I’ve spoken out about this workforce and the ethical challenges that come with data work through interviews, conference panels, articles, forums, aiding legislators, speaking engagements, workshops and social media. It’s an honor to be in a position in which I can help educate the general public, congressional leaders and labor advocates about this workforce and all that comes with it.

How do you navigate the challenges of the male-dominated tech industry, and, by extension, the male-dominated AI industry?

I consider myself very fortunate because I have a great support system that includes my colleagues and mentors. I choose to surround myself with people who want to see female and non-binary folks succeed. My mentors are women and I also seek advice from supportive men. One thing that has to continue, however, is speaking up about inequity and moving the conversation forward to change it.

What advice would you give to women seeking to enter the AI field?

I would tell any woman wanting to enter the AI field to go for it! Finding a good mentor or mentors is so important. Look to the many strong women and non-binary folks in the field for guidance when needed. Forge relationships with supportive men. Lastly, don’t be afraid to speak up. Great ideas come from confronting some of the hardest questions!

What are some of the most pressing issues facing AI as it evolves?

One of the most pressing issues facing the evolution of AI is accessibility. Who has access to the tools? Who’s providing the data and maintaining the system? Who’s benefiting from AI? What populations are being left behind and how do we change that? How are the workers behind the system being treated?

The other issue I would raise here would be bias. How do we create systems completely free from bias?

What are some issues AI users should be aware of?

I would always tell users to look at how the workers training AI are being treated. That’s an indicator of so many things.

What is the best way to responsibly build AI?

It’s imperative that we involve underrepresented populations in the creation of AI. The people who will be impacted by the tech should always have a seat at the table. Similarly, the creation of AI legislation has to involve data workers. They are the foundation of these systems and to have the discussion without them would be irresponsible.

How can investors better push for responsible AI?

I will just say what I have been saying: Nothing is set in stone. We do not have to accept what is being presented to us. The only way things improve is to speak up and act. Look for other organizations pushing for responsible AI. Challenge working conditions, challenge implementation, usage, etc. Challenge anything that feels unfair or irresponsible.

Apollo Health & Lifestyle Revamps Finance Systems with Oracle Cloud ERP

Retail healthcare company, Apollo Health & Lifestyle Limited,has migrated its core financial systems to Oracle Fusion Cloud Enterprise Resource Planning (ERP) to optimise its financial operations and increase productivity, the company announced in a statement.

With Oracle Cloud ERP, Apollo Health & Lifestyle will be able to eliminate manual processes and embrace continuous innovation to improve speed and accuracy in reporting, align financial and operational planning, and gain insights to drive better decisions.

A subsidiary of Apollo Hospitals Enterprise Limited, Apollo Health & Lifestyle manages an extensive network of standardized primary healthcare models, including Multispecialty Clinics, Diabetes Management Clinics, Dental Clinics, and Diagnostic Centers under various brand names.

Oracle Cloud ERP is set to increase productivity, reduce costs, and improve controls for Apollo Health & Lifestyle. “By integrating Oracle Cloud ERP into our operations, we aim to fortify our financial infrastructure, ensuring agility and efficiency across their business verticals,” said Sriram Iyer, CEO, Apollo Health & Lifestyle Ltd.

The incorporation of Oracle Fusion Cloud Enterprise Performance Management (EPM) within the ERP system will further enhance the speed and accuracy of financial reporting, reducing the time required for book closures and facilitating better decision-making by senior leaders.

The decision to select Oracle Cloud ERP was made in September 2022, with the cloud-based solution offering touchless operations, predictive insights, and embedded collaboration features.

“With Oracle Cloud ERP, we can simplify finance processes, increase business agility, and improve business decisions so employees can focus on providing the best care to our customers,” said Ashish Maheshwari, CFO, Apollo Health & Lifestyle Ltd.

The implementation project was executed by PwC, a long-term member of the Oracle PartnerNetwork (OPN). Hirak Kayal, Partner at PwC, highlighted the need for an efficient and scalable back-office operation to meet the growing demand in India’s healthcare delivery market.

The post Apollo Health & Lifestyle Revamps Finance Systems with Oracle Cloud ERP appeared first on Analytics India Magazine.

5 Airflow Alternatives for Data Orchestration

5 Airflow Alternatives for Data Orchestration
Image by Author

Data orchestration has become a critical component of modern data engineering, allowing teams to streamline and automate their data workflows. While Apache Airflow is a widely used tool known for its flexibility and strong community support. However, there are several other alternatives that offer unique features and benefits.

In this blog post, we will discuss five alternatives to manage workflows: Prefect, Dagster, Luigi, Mage AI, and Kedro. These tools can be used for any field, not just limited to data engineering. By understanding these tools, you'll be able to choose the one that best suits your data and machine learning workflow needs.

1. Prefect

Prefect is an open-source tool for building and managing workflows, providing observability and triaging capabilities. You can build interactive workflow applications using a few lines of Python code.

5 Airflow Alternatives for Data Orchestration

Prefect offers a hybrid execution model that allows workflows to run in the cloud or on-premises, providing users with greater control over their data operations. Its intuitive UI and rich API enable easy monitoring and troubleshooting of data workflows.

2. Dagster

Dagster is a powerful, open-source data pipeline orchestrator that simplifies the development, maintenance, and observation of data assets throughout their entire lifecycle. Built for cloud-native environments, Dagster offers integrated data lineage, observability, and a user-friendly development environment, making it a popular choice for data engineers, data scientists, and machine learning engineers.

5 Airflow Alternatives for Data Orchestration

Dagster is an open-source data orchestration system that allows users to define their data assets as Python functions. Once defined, Dagster manages and executes these functions based on a user-defined schedule or in response to specific events. Dagster can be used at every stage of the data development lifecycle, from local development and unit testing to integration testing, staging environments, and production.

3. Luigi

Luigi, developed by Spotify, is a Python-based framework for building complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization, and more, focusing on reliability and scalability.

5 Airflow Alternatives for Data Orchestration

Luigi is a powerful tool that excels in managing task dependencies, ensuring that tasks are executed in the correct order and only if their dependencies are met. It is particularly suitable for workflows that involve a mix of Hadoop jobs, Python scripts, and other batch processes.

Luigi provides an infrastructure that supports various operations, including recommendations, toplists, A/B test analysis, external reports, internal dashboards, etc.

4. Mage AI

Mage AI is a newer entrant in the data orchestration space, offering a hybrid framework for transforming and integrating data, combining the flexibility of notebooks with the rigor of modular code. It is designed to streamline the process of extracting, transforming, and loading data, enabling users to work with data in a more efficient and user-friendly manner.

5 Airflow Alternatives for Data Orchestration

Mage AI provides a simple developer experience, supports multiple programming languages, and enables collaborative development. Its built-in monitoring, alerting, and observability features make it well-suited for large-scale, complex data pipelines. Mage AI also supports dbt for building, running, and managing dbt models.

5. Kedro

Kedro is a Python framework that provides a standardized way to build data and machine learning pipelines. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.

5 Airflow Alternatives for Data Orchestration

Kedro provides a standardized project template, data connectors, pipeline abstraction, coding standards, and flexible deployment options, which simplify the process of building, testing, and deploying data science projects. By using Kedro, data scientists can ensure a consistent and organized project structure, easily manage data and model versioning, automate pipeline dependencies, and deploy projects on various platforms.

Conclusion

While Apache Airflow continues to be a popular tool for data orchestration, the alternatives presented here offer a range of features and benefits that may better suit certain projects or team preferences. Whether you prioritize simplicity, code-centric design, or the integration of machine learning workflows, there is likely an alternative that meets your needs. By exploring these options, teams can find the right tool to enhance their data operations and drive more value from their data initiatives.

If you are new to the field of Data Engineering, consider taking the Data Engineering Professional Course to become job-ready and start earning $300K/Yr.

Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master's degree in Technology Management and a bachelor's degree in Telecommunication Engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.

More On This Topic

  • Workflow Orchestration with Prefect and Coiled
  • Build a synthetic data pipeline using Gretel and Apache Airflow
  • Pandas not enough? Here are a few good alternatives to processing…
  • The Top 5 Alternatives to GitHub for Data Science Projects
  • GitHub Copilot Open Source Alternatives
  • Top 5 Free Alternatives to GPT-4

How renaissance technologists are connecting the dots between AI and business

Group of young adults, photographed from above, on various painted tarmac surface, at sunrise.

The so-called "subject matter expert" has been a staple for most companies since the invention of fire. Lately, however, with many simple and not-so-simple tasks being handed to AI, leaders are realizing that someone with a blend of technology and business acumen needs to be steering things in the right — and, hopefully, profitable — direction. Call it the renaissance technologist.

Also: AI has caused a renaissance of tech industry R&D, says Meta's chief AI scientist

An analysis of companies posting the most AI jobs finds they "exhibited a higher demand for AI professionals combining technical expertise with leadership, innovation, and problem-solving skills, underscoring the importance of these competencies in the AI field," according to a report out of the Organization for Economic Cooperation and Development (OECD). The researchers looked across 32,000 unique skills within leading economies.

The renaissance technologist — be they inside or outside the IT department — is an important track for career development, especially among people designing, developing, and deploying AI-based systems. This role will ensure that AI is delivering what the business needs, while suppressing bias, addressing errors, and maintaining security. After all, you don't want to be doing the wrong thing at subsecond speed.

The following are some of the elements that renaissance technologists need to surface in their businesses:

  • What business are we in, what do we want to be in, and how will AI help us get there?
  • What is the potential return on investment for this effort?
  • How will AI open new perspectives and ideas to people?
  • How will AI boost people's career potential?
  • What are competitors missing with the AI opportunity?
  • Is the system taking in data relevant to the desired outcome?
  • Is the system producing insights relevant to the direction the business wants to take?

Technology managers and professionals are increasingly being called upon to answer such questions. "Beyond basic technical skills and AI management knowledge, broader subject matter expertise will remain essential across the various domains that make up the business world," says Ted Lango, senior vice president at Intradiem.

Also: How tech professionals can survive and thrive at work in the time of AI

The emerging renaissance technologist "will need expertise in fostering an environment where AI tools and human workers complement each other to enhance productivity and employee well-being," Lango urges. "Basic skills related to emotional intelligence and problem-solving, and the ability to communicate clearly, think critically, and work collaboratively will remain as important as ever."

Simply stated: less coding and development work, and more focus on the business. We've heard this prediction many times over the past couple of decades, but AI seems to be effectively picking up the slack of heads-down programming.

"The time spent doing non-value-added work to create software — like searching through internal knowledge systems — will be sped up through AI, so engineers can focus on more high-value tasks," says Jonny LeRoy, chief technology officer at Grainger. For example, he relates, "Grainger's most recent internal hackathon was centered on using AI to gain productivity. The winning team created a question-answering agent that retrieves relevant internal Grainger information from various systems including Jira, Confluence, and GitHub."

Also: Soon, every employee will be both AI builder and AI consumer

It's important that technology teams understand the business context "so they can frame the problem they're trying to solve with AI," says LeRoy. "At Grainger, we encourage 'ride-alongs' with our sellers or customer service agents, so team members deepen their understanding of our customers, their wants and needs."

There's no question that technology professionals and engineers are using new AI tools to write code, debug, create APIs, and define test cases faster," says Ilya Goldin, principal data scientist at Phenom. "But AI also frees up professionals to engage with their work in more meaningful ways through opportunities for creativity and innovation."

Renaissance technologists, then, will lead the way with "idea generation, driven by the knowledge, skills, and abilities of creative personnel to produce novel solutions," Goldin illustrates. "Generative AI can help with idea generation, finding highly complex problems that tech organizations can work to solve. It can also help with content generation in text, images, video, and code."

Also: AI will have a big impact on jobs this year. Here's why that could be good news

Going forward, success in the technology career arena requires "interdisciplinary learning," says Lango. "Courses in AI and machine learning are fundamental, of course, but studies in psychology, human behavior, and ethics are also really relevant to the evolution of technology work."

Training programs "should focus on developing a holistic understanding of how technology impacts human behavior and organizational culture," Lango continues.

At Grainger, prompt engineer is a relatively new emerging job title, but AI is opening up a multitude of potential career paths. "We're trying to do multiple things," says LeRoy. "We have also hired great talent often using broader titles like machine learning scientist, and we're teasing apart the ML-specific skills from the more engineering-oriented ones to allow focus for experts and fungibility for generalists. I expect to see more roles dedicated to quality, safety, and governance, which will require a fascinating set of skills to manage the emerging risks around bias, confabulation, and quality."

Also: Agile Intelligence: AI gives tech and business collaboration a much-needed boost

Across today's companies, additional titles emerging include "AI ethics officer, human-tech integration specialist, and employee experience officer," says Lango. "These roles focus on ensuring the ethical usage of AI, managing the integration of AI into workplaces, and enhancing employee experiences in increasingly tech-driven work environments."

Renaissance technologists will help deliver breakthrough innovation, which "is a cognitively demanding task that human intelligence struggles to meet," Goldin points out. "It takes strategic thinking to identify when creative solutions are needed, and to anticipate the consequences of a particular solution. With AI handling more of the idea and content generation, tech professionals can spend more time applying their abilities to elements such as problem identification and idea evaluation."

Artificial Intelligence

How HARMAN is Solving Healthcare Problems with GenAI

Back in October 2023, Samsung-owned HARMAN entered the generative AI race in the healthcare space with HealthGPT – a private LLM built on TII’s open-source model Falcon 7B. The latest version, HealthGPT Chat, is built on Llama 2.

Healthcare is one of the most inherently complicated fields for generative AI to be harnessed given the vast amount of unstructured and sensitive data.

“The primary motivation behind developing a private LLM like HealthGPT stems from enterprises’ concerns about data privacy and security. This issue arises particularly when using public LLMs, as they require transferring sensitive data to external entities, leading to potential uncertainties regarding how this data might be used,” Dr Jai Ganesh, chief product officer, told AIM, in an exclusive interaction.

Ganesh further explained that another challenge with public models is their structural limitations, such as token restrictions and a lack of control over the entire processing chain. This lack of control can be detrimental, as any issues with public models can directly impact business performance.

Understanding the severity of these issues, HARMAN entered the space of private models, coinciding with the open-sourcing moment of several foundational models around late February last year. It was around the same time when Meta’s LLaMA 1 was leaked online. Following this, models like Falcon 7B by TII and Llama 2 were released. However, the company chose to tap into the potential of open-source models instead of building an LLM in-house.

“We chose not to develop a foundational model ourselves due to the immense cost and resource investment required—estimated to be around $30 to $40 million. Instead, we focused on utilising existing open-source foundation models,” he commented.

Dearth of Data

“We identified healthcare as the most relevant domain for our model due to the sector’s inefficiencies, the vast amount of unstructured data, and the challenges decision-makers face in deriving insights from this data,” said Ganesh.

So given the scarcity of training data, the team found a rich source of such publicly reported data in clinical trial studies related to cancer, immune diseases, and heart diseases, leveraging this data to train HealthGPT.

However, the problem is not confined to the amount of data available, but also its poor quality. Models trained on skewed datasets can lead to biased outcomes. “We have mechanisms in place to identify and correct bias in the data. This is not limited to the data preparation stage; the output is also scrutinised for bias,” he added.

Rigorous testing frameworks, involving extensive querying, are integral to maintaining the integrity and privacy of processed data.

Tackling Hallucination Responsibly

It is not new that LLMs are prone to spewing wrong information. But in critical areas like life science, where accuracy and reliability of diagnosis, such hallucinations can cost you serious threats. To overcome these risks, HARMAN uses a combination of automated mechanisms and human-in-loop intervention.

HARMAN has a multifaceted approach to managing hallucinations. Firstly, HealthGPT employs guardrails to regulate the level of hallucination. “Initially, the accuracy of the models was around 74%, but with continuous refinement, it has improved significantly, reaching over 85-90%,” said Ganesh.

Secondly, the model’s interface provides control over settings like temperature and token number, allowing users to check the extent of hallucinatory outputs. Higher temperatures increase hallucination, while lower settings reduce it. Thirdly, human oversight is involved with the medical doctor of the team validating AI-generated results. This is supplemented by a feedback mechanism for a continuous loop of user input that helps in refining the model. Furthermore, RAG feature adds references to answers for better information credibility. Finally, the system includes a benchmarking section that compares the model’s performance with other studies and models.

However, responsible AI is at the heart of HealthGPT’s success. “One of the reasons why we chose to go for private LLM is because it ensures end-user control. Unlike models hosted on unknown cloud instances, HealthGPT model operates within the user’s own Virtual Private Cloud (VPC) and cloud environment,” explained Ganesh, emphasising how the approach improves security and privacy as the model is fine-tuned on the user’s data within their controlled environment.

Pre-fine-tuning checks are also implemented to detect anomalies, with a focus on privacy through automated mechanisms for handling Personally Identifiable Information (PII) and Protected Health Information (PHI).

Harman’s approach in the healthcare domain, as well as others, is characterised by a human-centric philosophy. This approach involves understanding problems holistically and placing the human, or decision-maker, at the center of solutions. This philosophy is fundamental to Harman’s interactions with its global customer base, which varies from companies in the proof-of-concept (POC) stage to those in more advanced stages of implementation.

What Next

Although not yet deployed for live customers, HealthGPT has found compelling proof-of-concept (POC) stories showcasing its potential applications in diverse domains. One POC demonstrates the use of HealthGPT for personalised data analysis across sectors, allowing customisation for individual needs, such as drug discovery in pharmaceuticals.

Another story involves using the model for extracting insights from medical instrument data, emphasising its capacity to handle large-scale, structured information. A third user story features a pharmaceutical company leveraging the model to bolster drug discovery by integrating clinical trial data and information from sources like PubMed. IQVIA, Roche, Aetrex are among the major clients that HARMAN serves.

“Currently, we are experimenting Mistral AI’s Mixtral 7B for future iterations. The aim is to constantly push the boundaries of auto foundation models,” said Ganesh.

The company’s generative AI strategy involves integrating more diverse data sources with a focus on ensuring data quality and compliance with regulations like HIPAA. This necessitates extensive training for developers handling healthcare data. Alongside this, Ganesh and his team are working towards introducing multimodal features.

But no matter how strong the generative AI product it comes up with, HARMAN’s strategy will always be characterised by a human-centric philosophy. This approach involves understanding problems holistically and placing the human, or decision-maker, at the center of solutions.

Moreover, there are plans in place to expand HealthGPT beyond healthcare into domains like manufacturing and IT management, the other core focus areas of HARMAN, aligning with HARMAN’s strategic plans.

The post How HARMAN is Solving Healthcare Problems with GenAI appeared first on Analytics India Magazine.

Former OpenAI Researcher Andrej Karpathy Unveils Tokenisation Tutorial, Decodes Google’s Gemma 

Shocking: Andrej Karpathy leaves OpenAI

Former OpenAI researcher Andrej Karpathy has released a new tutorial on LLM tokenisation. Karpathy, introducing the course, mentions, ‘In this lecture, we build from scratch the Tokenizer used in the GPT series from OpenAI.

Tokenizers are a completely separate stage of the LLM pipeline: they have their own training set, training algorithm (Byte Pair Encoding), and after training implement two functions: encode() from strings to tokens, and decode() back from tokens to strings.

“We will see that a lot of weird behaviors and problems of LLMs actually trace back to tokenization. We’ll go through a number of these issues, discuss why tokenization is at fault, and why someone out there ideally finds a way to delete this stage entirely,” said Karpathy.

Furthermore, he has released a new repository on GitHub named ‘minbpe.’ It contains minimal, clean code for the Byte Pair Encoding (BPE) algorithm, commonly used in LLM tokenization. The repository can be found at: https://github.com/karpathy/minbpe.

Karpathy recently departed from OpenAI. He confirmed his departure in a post on X, saying that it is solely due to his intention to focus on personal projects. During his time at OpenAI, he actively contributed to the development of an AI assistant, collaborating closely with the company’s research head, Bob McGrew.

Decodes Google’s Gemma

Moreover, Karpathy also analysed Google’s recently released open source model Gemma’s tokenizer. Karpathy, in his pursuit of understanding the Gemma tokenizer, decoded the model protobuf in Python and presented a detailed comparison with the Llama 2 tokenizer.

Seeing as I published my Tokenizer video yesterday, I thought it could be fun to take a deepdive into the Gemma tokenizer.
First, the Gemma technical report [pdf]: https://t.co/AgxjBJfh0T
says: "We use a subset of the SentencePiece tokenizer (Kudo and Richardson, 2018) of… https://t.co/5mCrXTTvs3

— Andrej Karpathy (@karpathy) February 21, 2024

Key observations from the comparison include a substantial increase in vocabulary size from 32K to 256K tokens. Additionally, Gemma’s departure from the Llama 2 tokenizer is marked by the “add_dummy_prefix” setting being set to False, aligning with GPT practices and promoting consistency with minimal preprocessing.

Noteworthy aspects of the Gemma tokenizer include its model_prefix, representing the path of the training dataset, which hints at a substantial training corpus of approximately 51GB. The presence of numerous user-defined symbols, including special tokens like newline sequences and HTML elements, adds complexity to Gemma’s tokenization process.

In summary, Gemma’s tokenizer shares a foundational similarity with the Llama 2 tokenizer but distinguishes itself by its larger vocabulary size, inclusion of more special tokens, and a departure in the functional approach with the “add_dummy_prefix” setting. This comprehensive exploration by Karpathy sheds light on the nuances of Gemma’s tokenization methodology.

The post Former OpenAI Researcher Andrej Karpathy Unveils Tokenisation Tutorial, Decodes Google’s Gemma appeared first on Analytics India Magazine.

7 Free Kaggle Micro-Courses for Data Science Beginners

7 Free Kaggle Micro-Courses for Data Science Beginners
Image by Author

Do you remember that one data science course you signed up for but never got around to finishing? Well, you’re not alone.

Most data science beginners enroll in one or more courses: free or paid. But because data science courses typically cover a wide range of topics—from programming to data analysis, visualization, and more—it takes several weeks to work through them. And even if they start strong, most learners start feeling overwhelmed after the first few modules and fail to make progress. Enter Kaggle (micro)courses.

The series of micro-courses from Kaggle are a good alternative if you find longer courses harder to get through. They are great resources to learn data science skills—Python, pandas, machine learning, and more—without feeling overwhelmed. The courses are designed such that they take only a few hours to finish, and include tutorial and practice components. Now let's go over some beginner-friendly courses and what they cover.

1. Python

Python is one of the most widely used languages in data science. Besides helping you in your data career, Python is also helpful if you want to break into software engineering at some point. The Python course on Kaggle will help you learn the following:

  • Python basics (syntax and variables)
  • Functions
  • Booleans and conditionals
  • Lists, loops, and list comprehensions
  • Strings and dictionaries
  • Working with external libraries

If you feel like you need an even simpler intro to programming before diving into Python, you can check out the intro to programming course.

Because the subsequent courses on Pandas and data visualization require you to be comfortable with the contents of this course, you should not skip the Python course if you are new to programming with Python.

Link: Learn Python

2. Pandas

Once you’re familiar with basic Python you can learn pandas, a powerful data analysis and manipulation library.

Through a series of short lessons and hands-on coding exercise, the pandas will help you learn to perform the following operations on pandas dataframes:

  • Creating, reading, and writing
  • Indexing, selecting, and assigning
  • Renaming and combining
  • Summary functions and maps
  • Grouping and sorting
  • Data types and missing values

Link: Learn Pandas

3. Data Visualization

Now that you know how to analyze data with Python and pandas, it's time to build on that by learning how to visualize your data.

The Data Visualization course covers the fundamentals of creating helpful plots and charts using the Python library Seaborn. The course covers the following:

  • Line charts
  • Bar charts and heat maps
  • Scatterplots
  • Histograms and density plots
  • Choosing plot types

You also need to work on a final project to apply what you learned.

Link: Learn Data Visualization

4. Intro to SQL

SQL is the single most essential data science skill that you can learn. To understand why SQL is super important for data science, read "Why SQL is the Language to Learn for Data Science" by KDnuggets contributor Nate Rosidi.

The Intro to SQL course will teach you how to you query data ets with SQL using the BigQuery Python client and covers SQL fundamentals, filtering, and writing readable SQL queries:

  • Getting started with SQL and BigQuery
  • Select, from, and where
  • Group by, having, and count
  • Order by
  • As and with
  • Joining data

Link: Learn Intro to SQL

5. Advanced SQL

Now that you are comfortable with SQL basics, you can take the Advanced SQL course to develop your SQL skills further. This course builds on the intro to SQL course and covers the following topics on combining data from multiple tables and performing more complex operations:

  • Joins and unions
  • Analytic functions
  • Nested and repeated data
  • Writing efficient queries

Link: Learn Advanced SQL

6. Intro to Machine Learning

If you’ve already worked your way through the above courses, you should be comfortable with programming and data analysis with Python and SQL. You’re now ready to get started with machine learning.

The Intro to Machine Learning course covers:

  • How ML models work
  • Basic data exploration
  • Model validation
  • Underfitting and overfitting
  • Random forests

You can also make a submission to a beginner-friendly Kaggle competition.

Link: Learn Intro to Machine Learning

7. Intermediate Machine Learning

The Intermediate Machine Learning course builds on the Intro to Machine Learning course and teaches you how to handle missing values, categorical variables, and avoid the tricky problem of data leakage when training machine learning models.

The topic covered include:

  • Missing values
  • Categorical variables
  • ML pipelines
  • Cross validation
  • XGBoost
  • Data leakage

Link: Intermediate Machine Learning

Wrapping Up

I hope you found this round-up of courses helpful.

As mentioned, they’re all free. And it only takes a few hours to learn an essential data science skill. So you can start out on your data science journey one micro-course at a time. Happy learning!

Bala Priya C is a developer and technical writer from India. She likes working at the intersection of math, programming, data science, and content creation. Her areas of interest and expertise include DevOps, data science, and natural language processing. She enjoys reading, writing, coding, and coffee! Currently, she's working on learning and sharing her knowledge with the developer community by authoring tutorials, how-to guides, opinion pieces, and more.

More On This Topic

  • Micro, Macro & Weighted Averages of F1 Score, Clearly Explained
  • 3 Free Machine Learning Courses for Beginners
  • KDnuggets News, December 14: 3 Free Machine Learning Courses for…
  • KDnuggets News, October 5: Top Free Git GUI Clients for Beginners •…
  • Free Data Engineering Course for Beginners
  • Free MLOps Crash Course for Beginners