Get an AI Education in 2024 for Just $20

Businessman holding a projection of a digital brain.
Image: StackCommerce

Every business these days could benefit from artificial intelligence and machine learning. If your business isn’t, it may be time to invest in some education. With The Ultimate Artificial Intelligence & Machine Learning E-Degree Bundle, you can develop the skills you need to bring your business into 2024. It’s on sale for $19.99 (reg. $80).

This e-degree program includes six modules from Eduonix Learning Solutions (4.0/5-star instructor rating). Eduonix’s professional team of trainers has more than a decade of experience and will help you learn the vital skills you need to succeed in the AI and machine learning space.

At the start of the program, you’ll get an introduction to Python for AI and machine learning, ironing down the basics of how algorithms work. As you progress, you’ll explore additional tools for AI and machine learning, like R, and progress into data science. There are courses dedicated to regression analysis, supervised learning, unsupervised learning and neural networks to give you a comprehensive understanding of how AI and machine learning algorithms work.

To fortify your learning, the bundle includes quizzes, real-life test cases and exams, so you truly feel like you’re getting a degree education that will guide you in 2024 and beyond.

Help bring your business into the future. Right now, you can get The Ultimate Artificial Intelligence & Machine Learning E-Degree Bundle for 75% off $80 at just $19.99 for a limited time.

Prices and availability are subject to change.

Nabla raises another $24 million for its AI assistant for doctors that automatically writes clinical notes

Nabla raises another $24 million for its AI assistant for doctors that automatically writes clinical notes Romain Dillet @romaindillet / 8 hours

Paris-based startup Nabla just announced that it has raised a $24 million Series B funding round led by Cathay Innovation, with participation from ZEBOX Ventures — the corporate VC fund of CMA CGM. This funding round comes just a few months after Nabla signed a large-scale partnership with Permanente Medical Group, a division of U.S. healthcare giant Kaiser Permanente.

Nabla has been working on an AI copilot for doctors and other medical staff. The best way to describe it is that it’s a silent work partner that sits in the corner of the room, takes notes and writes medical reports for you.

The startup was originally founded by Alexandre Lebrun, Delphine Groll and Martin Raison. Lebrun, Nabla’s CEO, was the CEO of Wit.ai, an AI assistant startup that was acquired by Facebook. He then became the head of engineering of Facebook’s AI research lab FAIR.

A few weeks ago, I saw a live demo of Nabla with a real doctor and a fake patient pretending that they had back pains. When a physician starts a consultation, they hit the start button in Nabla’s interface and forget about their computer.

In addition to the physical examination part, a consultation also includes a long discussion with a bunch of questions about what brings you here and your medical history. At the end of the consultation, there might also be recommendations and prescriptions.

Nabla uses speech-to-text technology to turn the conversation into a written transcript. It works with both in-person consultations and telehealth appointments.

After the patient has left, the doctor hits the stop button. Nabla then uses a large language model refined with medical data and health-related conversations to identify the important data points in the consultation — medical vitals, drug names, pathologies, etc.

Nabla generates a thorough medical report in a minute or two with a summary of the consultation, prescriptions and follow-up appointment letters.

These reports can be customized to the doctor’s needs with a personalized format for your notes. For instance, you can add instructions to make the note more concise or more verbose. Or you can ask to generate notes that follow the Subjective, Objective, Assessment and Plan (SOAP) note pattern that is widely used in the U.S.

During the demo that I saw, I was extremely surprised by the effectiveness of Nabla in general. Even though we were in a crowded room and Nabla was running on a laptop a couple of meters away from the demo presenters, the tool was able to generate an accurate transcript and a useful report.

With Nabla Copilot, as the name suggests, the startup isn’t trying to take the human out of the medical loop. Physicians still have a final say as they can edit reports before they are filed in their electronic health record system (EHR).

Instead, the company thinks it can help doctors save time on admin work so that they can spend more time focusing on patients.

“What we know is the near future is we don’t want to try to replace doctors. You’ve seen companies — like Babylon in the U.K. — burning $1 billion trying to do chatbots and trying to automate things right away and remove doctors from the loop. And we’ve decided a long time ago with Nabla Copilot that [doctors] are the pilots and we work by their side,” Nabla co-founder and CEO Alexandre Lebrun said.

“It’s a little bit like automation for autonomous vehicles. We are still at level two today. We will start level three very soon with clinical assurance support. Then level four is clinical decision support, but with FDA approval, because you make decisions that you cannot really explain,” he added.

At some point, you could even imagine a level five of autonomous healthcare, which would mean removing physicians from the room. But Lebrun is still very cautious on this front.

“For some situations in some markets, like in some countries where they don’t have any access to healthcare, it would be a relevant thing,” he said. Over the long term, he sees the diagnostic process as a “pattern matching problem” that could be solved with AI. Doctors would focus on empathy, surgery procedures and critical decisions.

While Nabla is based in France, most of the company’s customers are in the U.S. following a rollout across Permanente Medical Group. Nabla isn’t just a work in progress, it is actively used every day by thousands of doctors.

Nabla’s privacy model

Nabla is currently available as a web app or a Google Chrome extension. The company is well aware that it is handling sensitive data. That’s why it doesn’t store audio or medical notes on its servers, unless both the doctor and the patient give their consent.

Nabla focuses on data processing instead of data storing. After a consultation, the audio file is discarded and the transcript is stored in the EHR that doctors are already using for their patient files.

In more technical terms, when a physician starts a recording, the audio is transcribed in real-time using a fine-tuned speech-to-text API. The company uses a combination of an off-the-shelf speech-to-text API from Microsoft Azure and its own speech-to-text model (a refined model based on the open-source Whisper model).

“When you have just a normal speech-to-text algorithm, they may or may not be good on medical data. But we have a fine-tuned one. And, as you probably have seen, the text is very light at first, and then it becomes dark. And when it becomes dark, it means that we verified it with our own model and we corrected it with medication names or medical conditions,” Nabla ML engineer Grégoire Retourné said during the demo that I saw.

The transcript is first pseudonymized, meaning that personally identifiable information are replaced with variables. Pseudonymized transcripts are processed by a large language model. Historically, Nabla has been using GPT-3 and then GPT-4 as its main large language model. As an enterprise customer, Nabla can tell OpenAI that it can’t store its data and train its large language model on those consultations.

But Nabla has also been playing with a fine-tuned version of Llama 2. “In the future, we envision using more and more narrow models as opposed to general models,” Lebrun said.

Once the LLM has processed the transcript, Nabla de-pseudonymizes the output. Doctors can see the note, which is stored on the computer in the local web browser storage file. Notes can be exported to EHRs.

However, doctors can give their approval and ask for the patient consent to share medical notes with Nabla so that they can be used to correct transcription errors. And given that Nabla is on track to process more than 3 million consultations per year in three languages, chances are Nabla will improve really quickly thanks to real world data.

Image Credits: Romain Dillet / TechCrunch

Legacy of The Titan of Programming

The developer community is mourning the demise of Niklaus Emil Wirth, an 89- year-old pioneer of programming languages who passed away on the 1st of January. He is celebrated as the creator of Pascal programming language but that was only one in a list of languages he contributed to.

Niklaus Wirth was the chief designer of Euler, PL360, ALGOL W, Pascal, Modula, Modula-2, Oberon, Oberon-2, and Oberon-07. His expertise extended to operating systems as well, where he significantly contributed to the design and creation of Medos-2 (1983, for the Lilith workstation) and Oberon (1987, for the Ceres workstation). Additionally, he was instrumental in the development of Lola, a digital hardware design and simulation system, in 1995.

Recognising his contributions, Wirth received the prestigious ACM Turing Award in 1984 for his advancements in programming languages and was named an ACM Fellow in 1994.

His legacy

He was 29 when he completed his doctorate in Berkeley under Harry Huskey on the generalisation of the Algol 60 programming language. “When I was 22 or 25, I realised that actually my calling was to teach, like my father,” he said in an interview.

Niklaus Wirth’s academic path included assistant professorships at Stanford University and the University of Zurich before returning to ETH Zurich in 1968 as a Professor of Computer Science, a role he held until 1999. He also enhanced his expertise with two stints at Xerox’s Palo Alto Research Center (PARC) in 1976–1977 and 1984–1985, focusing on advanced research and collaborations.

Anecdotes about his wit and passion for teaching were shared by his students on HackerNews who recounted, “Undergraduate students were all in awe of him, and he seemed to have a good relationship with his graduate students.”

His PhD students Martin Odersky who is known for creating the Scala programming language; William Marshall McKeeman, Peter U. Schulthess; Edouard Marmier; Rudolf Schild, and Michael Franz, all of whom have significant contributions in computer science.

A Role Model for Developers

Most developers today posted on social media that Pascal was one of the first programming languages they’ve learnt. Before Wirth’s contribution, many programming languages lacked a clear structure, often leading to what is known as “spaghetti code” – code with complex and tangled control structures.

Wirth’s approach advocated for a more organised and modular code structure, which is evident in his design of Pascal. This approach enforced a disciplined approach to programming, with clearly defined beginning and end points for loops and conditional statements, reducing complexity and increasing readability.

With each language, he worked on being as precise and his greatest legacy is his focus on engineering vs. science in programming language design. He was not concerned with adding the latest and greatest features, but on what is proven, and what has a well understood means of efficient implementation. He explained, “The driving force for all these years was to be able to teach programming in a scientific, mathematical manner, devoid of computer jargon and idiosyncrasies and based on proper abstractions.”

His first language, Euler, simplified programming syntax, making it more accessible to users. With PL360, a system programming language, he balanced low-level hardware access with structured programming principles, bridging a gap between hardware interaction and readable code design.

In ALGOL W, Wirth expanded upon the foundations of ALGOL 60 by introducing new data types, enhancing the language’s versatility and setting the stage for the development of Pascal in 1970.

Pascal itself was a landmark achievement in programming, emphasising readability, data structuring, and efficient code organisation, which had a lasting impact on educational programming and became a model for subsequent languages.

Wirth later introduced Modula and Modula-2, which were significant for introducing the concept of modules. This innovation in language design transformed how software systems were structured, allowing for better organisation and modularity in large-scale software development. This influenced languages like Ada and Java, which adopted similar approaches to organisation and structure.

When asked about what valuable features of Pascal and Modula-2 were brought into C, he said, “Most of the statement structures and data structures appear in them, and, I believe, also the important concept of static typing of all constants, variables, and procedures.”

His expertise extended to operating systems as well, where he significantly contributed to the design and creation of Medos-2 for Lilith workstation and Oberon for Ceres workstation. With his understanding of hardware, these systems were designed to work seamlessly with the hardware they were built for. Additionally, Wirth was instrumental in the development of Lola, a digital hardware design and simulation system, in 1995.

The Wirth’s Law, an adage he wrote about in a paper states that software becomes slower more rapidly than hardware becomes faster. This highlights the tendency of software applications to grow larger, slower, and more resource-intensive over time, a phenomenon known as software bloat.

It occurs due to factors like adding new features, poor code optimization, and inefficient programming practices. As hardware advances, developers often add more features to software, contributing to this bloat, a phenomenon true

He also gained a reputation for his wit, Europeans tend to pronounce his name properly, as Nih-klaus Virt, while Americans usually mangle it into something like Nickels Worth. He joked about it saying, “Europeans call me by the name while Americans call me by the value.”

Niklaus Wirth’s contributions have improved programming as we know it. His philosophy of simplicity and precision in programming language design, his emphasis on structured and efficient code, and his dedication to teaching have shaped the foundations of modern programming.

The post Legacy of The Titan of Programming appeared first on Analytics India Magazine.

Enroll in a 4-year Computer Science Degree Program For Free

Enroll in a 4-year Computer Science Degree Program For Free
Image by Author

Have you ever wanted to study computer science but didn't want to pay the high cost of college tuition? Well, you're in luck! There is an incredible open-source curriculum called OSSU (Open Source Society University) that allows you to enroll in the equivalent of a 4-year computer science degree program entirely for free.

What is the OSSU Computer Science Degree Program?

ossu/computer-science provides a complete education in computer science concepts fundamental to all computing disciplines. The curriculum is designed according to the degree requirements of undergraduate computer science majors at leading universities. It uses high-quality courses from platforms like edX, Coursera, and Udacity taught by professors from schools like MIT, Harvard, and Princeton.

Enroll in a 4-year Computer Science Degree Program For Free
Image from ossu/computer-science

The coursework covers everything from programming languages, algorithms, and data structures to operating systems, computer architecture, and software engineering. After completing the core CS requirements, you can choose advanced electives to specialize in areas like software testing, game theory, linear algebra, and more.

The best parts are that all the course materials are freely available online, and you can complete the program at your own pace. While it is possible to finish in 2 years if you study around 20 hrs/week, you can adjust based on your schedule. You also join a worldwide community of independent learners who support each other.

Curriculum

Prerequisites

Computer science curriculum has prerequisite expectations at different stages:

  1. Core CS: Students must have prior high school level math background including algebra, geometry and pre-calculus.
  2. Advanced CS: Students can only choose advanced electives after finishing all required courses in the Core CS section first.
  3. Advanced Systems: Any student wanting to specialize in the Advanced Systems electives must have previously taken at least one basic physics course in high school or college.

Intro CS

The Intro CS section has beginner-level courses to help students new to computer science understand if it is the right fit for them. It covers introductory programming to teach basic coding concepts and introductory computer science to provide students with an understanding of computation's role in solving problems.

Core CS

The Core CS section has all the coursework equivalent to the first 3 years of a university computer science degree. It builds a strong base across essential areas like:

  • Core programming: Covers languages, testing, design patterns, architecture, etc.
  • Core math: Builds mathematical maturity required for data structures, algorithms, etc.
  • CS tools: Introduces commonly used tools for tasks like version control, shell scripting, etc.
  • Core systems: Deals with OS, networking, compilation, and computer architecture.
  • Core theory: Fundamental theoretical concepts like algorithms, NP-completeness, etc.
  • Core security: Secure coding, cryptography, and vulnerabilities.
  • Core applications: Databases, machine learning, computer graphics, etc.
  • Core ethics: Explores ethical implications of technology in society.

Advanced CS

After completing all required Core CS courses, students should choose additional Advanced CS courses based on their interests and intended field.

  • Advanced Programming: Covers topics like debugging, parallel computing, UML, software architecture, compilers, functional programming with Haskell, and more.
  • Advanced Systems: Goes deeper into digital logic, computer organization, pipelining, parallel processing, virtualization, and other lower-level computing concepts.
  • Advanced Theory: Includes formal language theory, Turing machines, computability, concurrency models, computational geometry, logic, and game theory.
  • Advanced Information Security: Provides more specialized security knowledge like compliance, digital forensics, secure development lifecycle, and verification.
  • Advanced Math: Includes linear algebra, numerical methods, formal logic, probability theory, and other important mathematical foundations for computer science.

Final Project

The final project requires students to apply all their learnings by building something useful. This provides tangible proof of their knowledge and skills for showcasing to potential employers.

Creating a project not only looks great on a resume but also validates and reinforces your knowledge. You can build something new from scratch or contribute to an existing open-source project in need of help.

For more guidance, there are structured course specializations oriented around projects you may pursue. Topics include full-stack development, data science, robotics, and beyond. With the core foundations, you can now identify series matching your interests.

When your final project is complete, submit information on it to the OSSU PROJECTS via a pull request. Also, add the OSSU badge to your project README. Then, use community channels to announce your creation to fellow students.

The evaluation is done by getting feedback from peers and showcasing abilities rather than traditional grading. It also lets OSSU evaluate how well its curriculum prepares independent learners for real achievements.

Conclusion

The OSSU computer science degree program offers a fantastic opportunity for those who are interested in studying computer science without the burden of high tuition fees. With a well-structured curriculum that covers all the fundamental concepts of computer science, you can gain a comprehensive education in this field. The program's flexibility allows you to learn at your own pace, making it an accessible option for anyone, regardless of their schedule. So why wait? Start your computer science education today for FREE with OSSU.

Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master's degree in Technology Management and a bachelor's degree in Telecommunication Engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.

More On This Topic

  • Don't Miss Out! Enroll in FREE Courses Before 2023 Ends
  • 8 Best Data Science Courses to Enroll in 2022 For Steep Career Advancement
  • How our Obsession with Algorithms Broke Computer Vision: And how…
  • Launch your career with a Northwestern data science degree
  • 7 reasons you should get a formal degree in Data Science
  • Which is Best: Data Science Bootcamp vs Degree vs Online Course

$2 Bn Design Startup InVision Shuts Down Amidst Industry Rivalry

In a startling turn of events, design startup InVision, once valued at an impressive $2 billion, has announced its impending shutdown at the end of this year, as revealed in a company blog post on Thursday.

The news follows a period of decline, during which the business experienced a significant drop in revenue and ultimately sold its core business line to competitor Miro last fall.

Once celebrated as a market leader in collaborative design software, InVision struggled to maintain its foothold after Figma, a rival firm, gained widespread popularity, successfully diverting its customer base. Previous reports highlight InVision’s challenges, noting a 50% reduction in revenue to $50 million in 2022.

The startup, backed by substantial investments totaling more than $350 million from notable contributors like Goldman Sachs and Spark Capital, now finds itself in the company of other failed ventures that faltered when investors hesitated to inject additional funds.

Read: Conversational AI Startup Coqui Shuts Down

Convoy, Veev, and OliveAI, despite collectively raising around $2.5 billion from venture investors, also succumbed to the difficulties of sustaining their operations.

The demise of InVision reflects the harsh reality faced by startups in the current economic landscape. As borrowing costs rise, cash reserves dwindle, and venture capitalists grow more cautious, companies like InVision find themselves with fewer opportunities for financial recovery.

The downfall of InVision can be traced back to a pivotal moment when its founder and former CEO, Clark Valberg, recognised the rising threat from competitors. Despite early success and rapid growth, InVision faced a formidable challenge from Figma, a startup that ultimately eclipsed its market share.

Valberg’s attempts to salvage the situation, including an unsuccessful attempt to sell the company, highlight the difficulties startups encounter in securing second or third chances in an increasingly competitive and unforgiving business environment.

InVision, once boasting a valuation of $2 billion and support from prominent investors, now stands as a cautionary tale in an era where market dynamics can swiftly shift, leaving even well-established startups vulnerable to the forces of change.

The post $2 Bn Design Startup InVision Shuts Down Amidst Industry Rivalry appeared first on Analytics India Magazine.

Data Cleaning in SQL: How To Prepare Messy Data for Analysis

Data Cleaning in SQL: How To Prepare Messy Data for Analysis
Image generated with Segmind SSD-1B model

Excited to start analyzing data using SQL? Well, you may have to wait just a bit. But why?

Data in database tables can often be messy. Your data may contain missing values, duplicate records, outliers, inconsistent data entries, and more. So cleaning the data before you can analyze it using SQL is super important.

When you're learning SQL, you can spin up database tables, alter them, update and delete records as you like. But in practice, this is almost never the case. You may not have permission to alter tables, update and delete records. But you’ll have read access to the database and will be able to run a bunch of SELECT queries.

In this tutorial, we’ll spin up a database table, populate it with records, and see how we can clean the data with SQL. Let's start coding!

Creating a Database Table with Records

For this tutorial, let’s create an employees table like so:

-- Create the employees table  CREATE TABLE employees (  	employee_id INT PRIMARY KEY,  	employee_name VARCHAR(50),  	salary DECIMAL(10, 2),  	hire_date VARCHAR(20),  	department VARCHAR(50)  );

Next, let’s insert some fictional sample records into the table:

-- Insert 20 sample records   INSERT INTO employees (employee_id, employee_name, salary, hire_date, department) VALUES  (1, 'Amy West', 60000.00, '2021-01-15', 'HR'),  (2, 'Ivy Lee', 75000.50, '2020-05-22', 'Sales'),  (3, 'joe smith', 80000.75, '2019-08-10', 'Marketing'),   (4, 'John White', 90000.00, '2020-11-05', 'Finance'),  (5, 'Jane Hill', 55000.25, '2022-02-28', 'IT'),  (6, 'Dave West', 72000.00, '2020-03-12', 'Marketing'),  (7, 'Fanny Lee', 85000.50, '2018-06-25', 'Sales'),  (8, 'Amy Smith', 95000.25, '2019-11-30', 'Finance'),  (9, 'Ivy Hill', 62000.75, '2021-07-18', 'IT'),  (10, 'Joe White', 78000.00, '2022-04-05', 'Marketing'),  (11, 'John Lee', 68000.50, '2018-12-10', 'HR'),  (12, 'Jane West', 89000.25, '2017-09-15', 'Sales'),  (13, 'Dave Smith', 60000.75, '2022-01-08', NULL),  (14, 'Fanny White', 72000.00, '2019-04-22', 'IT'),  (15, 'Amy Hill', 84000.50, '2020-08-17', 'Marketing'),  (16, 'Ivy West', 92000.25, '2021-02-03', 'Finance'),  (17, 'Joe Lee', 58000.75, '2018-05-28', 'IT'),  (18, 'John Smith', 77000.00, '2019-10-10', 'HR'),  (19, 'Jane Hill', 81000.50, '2022-03-15', 'Sales'),  (20, 'Dave White', 70000.25, '2017-12-20', 'Marketing');

If you can tell, I’ve used a small set of first and last names to sample from and construct the name field for the records. You can be more creative with the records, though.

Note: All the queries in this tutorial are for MySQL. But you’re free to use the RDBMS of your choice.

1. Missing Values

Missing values in data records are always a problem. So you have to handle them accordingly.

A naive approach is to drop all the records that contain missing values for one or more fields. However, you should not do this unless you’re sure there is no other better way of handling missing values.

In the employees table, we see that there is a NULL value in the ‘department’ column (see row of employee_id 13) indicating that the field is missing:

SELECT * FROM employees;

Data Cleaning in SQL: How To Prepare Messy Data for Analysis

You can use the COALESCE() function to use the ‘Unknown’ string for the NULL value:

SELECT  	employee_id,  	employee_name,  	salary,  	hire_date,  	COALESCE(department, 'Unknown') AS department  FROM employees;

Running the above query should give you the following result:

Data Cleaning in SQL: How To Prepare Messy Data for Analysis 2. Duplicate Records

Duplicate records in a database table can distort the results of analysis. We’ve chosen the employee_id as the primary key in our database table. So we’ll not have any repeating employee records in the employee_data table.

You can still the SELECT DISTINCT statement:

SELECT DISTINCT * FROM employees;

As expected, the result set contains all the 20 records:

Data Cleaning in SQL: How To Prepare Messy Data for Analysis 3. Data Type Conversion

If you notice, the ‘hire_date’ column is currently VARCHAR and not a date type. To make it easier when working with dates, it’s helpful to use the STR_TO_DATE() function like so:

SELECT  	employee_id,  	employee_name,  	salary,  	STR_TO_DATE(hire_date, '%Y-%m-%d') AS hire_date,  	department  FROM employees;

Here, we’ve only selected the ‘hire_date’ column amongst others and haven’t performed any operations on the date values. So the query output should be the same as that of the previous query.

But if you want to perform operations such as adding an offset date to the values, this function can be helpful.

4. Outliers

Outliers in one or more numeric fields can skew analysis. So we should check for and remove outliers so as to filter out the data that is not relevant.

But deciding which values constitute outliers requires domain knowledge and data using knowledge of both the domain and historical data.

In our example, let's say we know that the ‘salary’ column has an upper limit of 100000. So any entry in the ‘salary’ column can be at most 100000. And entries greater than this value are outliers.

We can check for such records by running the following query:

SELECT *  FROM employees  WHERE salary > 100000;

As seen, all entries in the ‘salary’ column are valid. So the result set is empty:

Data Cleaning in SQL: How To Prepare Messy Data for Analysis 5. Inconsistent Data Entry

Inconsistent data entries and formatting are quite common especially in date and string columns.

In the employees table, we see that the record corresponding to employee ‘bob johnson’ is not in the title case.

But for consistency let's select all the names formatted in the title case. You have to use the CONCAT() function in conjunction with UPPER() and SUBSTRING() like so:

SELECT  	employee_id,  	CONCAT(      	UPPER(SUBSTRING(employee_name, 1, 1)), -- Capitalize the first letter of the first name      	LOWER(SUBSTRING(employee_name, 2, LOCATE(' ', employee_name) - 2)), -- Make the rest of the first name lowercase      	' ',      	UPPER(SUBSTRING(employee_name, LOCATE(' ', employee_name) + 1, 1)), -- Capitalize the first letter of the last name      	LOWER(SUBSTRING(employee_name, LOCATE(' ', employee_name) + 2)) -- Make the rest of the last name lowercase  	) AS employee_name_title_case,  	salary,  	hire_date,  	department  FROM employees;

Data Cleaning in SQL: How To Prepare Messy Data for Analysis 6. Validating Ranges

When talking about outliers, we mentioned how we’d like the upper limit on the ‘salary’ column to be 100000 and considered any salary entry above 100000 to be an outlier.

But it's also true that you don't want any negative values in the ‘salary’ column. So you can run the following query to validate that all employee records contain values between 0 and 100000:

SELECT  	employee_id,  	employee_name,  	salary,  	hire_date,  	department  FROM employees  WHERE salary < 0 OR salary > 100000;

As seen, the result set is empty:

Data Cleaning in SQL: How To Prepare Messy Data for Analysis 7. Deriving New Columns

Deriving new columns is not essentially a data cleaning step. However, in practice, you may need to use existing columns to derive new columns that are more helpful in analysis.

For example, the employees table contains a ‘hire_date’ column. A more helpful field is, perhaps, a ‘years_of_service’ column that indicates how long an employee has been with the company.

The following query finds the difference between the current year and the year value in ‘hire_date’ to compute the ‘years_of_service’:

SELECT  	employee_id,  	employee_name,  	salary,  	hire_date,  	department,  	YEAR(CURDATE()) - YEAR(hire_date) AS years_of_service  FROM employees;

You should see the following output:

Data Cleaning in SQL: How To Prepare Messy Data for Analysis

As with other queries we’ve run, this does not modify the original table. To add new columns to the original table, you need to have permissions to ALTER the database table.

Wrapping Up

I hope you understand how relevant data cleaning tasks can improve data quality and facilitate more relevant analysis. You’ve learned how to check for missing values, duplicate records, inconsistent formatting, outliers, and more.

Try spinning up your own relational database table and run some queries to perform common data cleaning tasks. Next, learn about SQL for data visualization.

Bala Priya C is a developer and technical writer from India. She likes working at the intersection of math, programming, data science, and content creation. Her areas of interest and expertise include DevOps, data science, and natural language processing. She enjoys reading, writing, coding, and coffee! Currently, she's working on learning and sharing her knowledge with the developer community by authoring tutorials, how-to guides, opinion pieces, and more.

More On This Topic

  • Messy Data is Beautiful
  • SQL for Data Visualization: How to Prepare Data for Charts and Graphs
  • Prepare Behavioral Questions for Data Science Interviews
  • Prepare Your Data for Effective Tableau & Power BI Dashboards
  • How to Prepare for a Data Science Interview
  • A Faster Way to Prepare Time-Series Data with the AI & Analytics Engine

Is This The End of Cell Towers?

SpaceX recently launched six of the 21 Starlink satellites which is equipped with “Direct-to-Cell” capabilities – also the end of cell towers, or cell sites, as we know it.

In other words, the direct-to-cell capability transforms Starlink satellites into high-tech cell towers orbiting the Earth. Equipped with eNodeB modems, these satellites can communicate directly with mobile phones, offering a seamless connection without bulky antennas.

The six @Starlink satellites on this mission with Direct to Cell capability will further global connectivity and help to eliminate dead zones → https://t.co/FgiJ7LOYdK pic.twitter.com/zFy7SrpsYs

— SpaceX (@SpaceX) January 3, 2024

“This will allow for mobile phone connectivity anywhere on Earth,” said Elon Musk, saying that this only supports ~7Mb per beam and the beams are very big, so while this is a greate solution for locations with no cellular connectivity, it is not meaningfully comopetitive with existing terrestrial cellular networks.

While it might not disrupt the existing cellular towers, there are changes that Starlink would be a go-to provider in the future for many when internet gets shut down, and cell service gets disrupted – “the only way humanity will be able to stay connected and informed is by using the everything app “X”, on a @Starlink connected X phone (and Grok of course),” said AI enthusiast Alex Covo.

Read: Are You Living in Musk’s Simulation?

This development also comes in the backdrop of SpaceX and T-Mobile partnership, which happened two years ago, which has already made remarkable strides into advancing Direct-to-Cell technology. The collaboration has facilitated early testing through T-Mobile’s customer base and international network.

Starlink in India

Starlink’s potential entry into the Indian market is marked by significant developments, for instance reports indicated that Starlink submitted additional documents addressing concerns related to data storage and transfer norms. This followed earlier hints from government officials in September 2023, suggesting a potential approval of Starlink’s Global Mobile Personal Communication by Satellite (GMPCS) license after receiving satisfactory compliance answers.

Starlink’s trial demonstrations in August 2023 showcased its technology to Indian authorities, while Elon Musk hinted at an early launch in India, emphasizing potential benefits for rural areas. Last year, Starlink also formally applied for a GMPCS license, marking a crucial step in offering voice and data services.

Along with that, Vodafone Idea officially denied any ongoing talks or collaboration with Elon Musk’s Starlink, a satellite internet company, dismissing recent reports that fueled stock surges. The denial followed heightened speculation and a surge in Vodafone Idea’s stock price, prompting the Bombay Stock Exchange to seek clarification from the company.

Despite these advancements, ongoing concerns include pricing and affordability for Indian users, regulatory hurdles, and competition from domestic players like Reliance Jio, contributing to the market’s complexity.

Starlink alternatives

Amazon’s Project Kuiper also entered the scene as a direct competitor to SpaceX’s Starlink. Notably, Project Kuiper achieved successful tests for an optical mesh network in low-Earth orbit, aiming to deliver high-speed, low-latency internet access using advanced technologies such as space lasers.

While Starlink currently leads with over 2 million active customers and government contracts, Project Kuiper distinguishes itself through innovations like a constellation of 3,276 satellites and the application of generative AI for efficient constellation management.

In India, Reliance Jio had earlier unveiled JioSpaceFiber, India’s inaugural satellite-based giga-fibre service at the India Mobile Congress. This initiative aimed to provide high-speed broadband connectivity to previously inaccessible regions, aligning with India’s Digital India initiative to foster digital inclusivity, offering reliable, low-latency, high-speed internet services even in remote areas.

OneWeb, with the backing of Bharati Airtel, had also deployed 34 satellites, reaching a total of 428 in orbit, with plans to launch an additional 228. The goal was to establish a global Low Earth Orbit (LEO) network in 2022. CEO Neil Masterson emphasized partnerships with Hughes, Marlink, and Field Solutions Holdings to enhance connectivity, especially in hard-to-reach areas.

In India, Starlink’s potential entry, along with competitors like Project Kuiper, JioSpaceFiber, and OneWeb, reflects a dynamic landscape. The race among these satellite internet providers suggests a transformative shift in how the world stays connected, with Starlink at the forefront of reshaping the future of communication.

While ISRO is diligently working on its new updates, India anticipates positive initiatives from the space agency that could offer support to companies navigating the challenges of ongoing technological advancements. The hope is for collaborative efforts that will help businesses stay abreast of the rapidly evolving landscape.

The post Is This The End of Cell Towers? appeared first on Analytics India Magazine.

How This Coimbatore-based AI Lab is Helping Developers Build LLMs

“We believe in the Zoho School of Thought, inspired by Sridhar Vembu, that you don’t need to build a company from a city like Bengaluru,” said Vishnu Subramanian, founder and chief of Jarvislabs.ai in an exclusive interview with AIM. Situated on the outskirts of Coimbatore, Jarvislabs.ai is silently playing a crucial role in providing AI infrastructure (GPUs on rent) for budding entrepreneurs and developers at an affordable price, alongside helping them advance AI research.

Jarvislabs.ai is one of the few companies in India that provides GPUs on rent for training, fine-tuning, and deploying AI models. Moreover, utilising services from hyperscalers like Microsoft Azure, Google Cloud, and AWS can be a costly affair for those who simply want to experiment. “If you compare with a hyper scalar, I would say we are like 60 to 70% cheaper” shared Subramanian.

“We wanted to make it easy and affordable for a small company in Chennai or Bengaluru to deploy a model natively, which can be expensive, especially when considering platforms like AWS or Google Cloud Platform. Even for older GPUs, they often charge around 6.7 cents, whereas we provide a modern GPU at a comparable price.” added Subramanian.

Unlocking Affordable GPU Prices

Subramanian clarified that their strategy to keep GPU prices low involves not pursuing NVIDIA’s H100s. Instead, they provide alternatives that offer similar capabilities. “The reason we are able to keep the costs low is we are not going behind the fancy H100s which are super expensive today. But you can do a lot of really amazing stuff with the Quadro or the next generation GPUs.” said Subramanian.

“We have A100s right now. We are not directly investing in H100; instead, we plan to partner with different hyperscalers or similar platforms to incorporate H100s into our platform. Currently, the GPUs we host include Quadro RTX 5000 and 6000. These are the latest generations and are as good as the older generation A100, but they come at a much cheaper price.” he added.

“For example, RTX 6000 Ada comes at around $1.14, providing similar performance to A100, for which you would end up spending around $3,” he explained.

While Subramanian didn’t disclose the exact number of GPUs acquired by Jarvislabs.ai, he did shed light on the strategy behind securing them. “There is a national distributor named RP Tech in India. They are the sole company responsible for importing GPUs into India. Fortunately, we have partnered with them, and from the start, they have been generous enough to assist us in obtaining GPUs at a very competitive price compared to the open market,” he said, saying that the prices are usually higher on ecommerce websites like Amazon, compared to buying it directly from distributors.

Subramanian also mentioned that they’ve barely spent a penny on marketing. Word of mouth has always helped them, as they have been offering a decent and good product. “People have been kind enough to introduce us to their friends, and word of mouth continues,” he said.

This is quite true as we came across Jarvislabs.ai when AIM was talking to one of the local developers, where he, in passing, mentioned about this AI lab, and how it gives GPUs-on-rent for building AI models and advancing AI research.

Current Customers

Subramanian said that their target customers are anyone ‘who is not a Fortune 500 company.’

Some of the notable customers of Jarvislabs.ai include Zoho and upGrad.”Some of the top companies in India, like Zoho, have been using LLMs for different workloads, probably to help students learn programming code, like a copilot or something similar. I’m not sure exactly what they’re doing, but I would assume that they could be fine-tuning some of the models.” said Subramanian.

Universities like CMU (Carnegie Mellon University) from the US have been using Jarvislabs.ai’s services to train large language models. Additionally, Weights and Biases are also leveraging its services.

“Our primary emphasis is on the global market, and until at least mid-2023, there wasn’t much usage from Indian customers. It was predominantly outside India, in the US and Europe. However, we see a slow change now; Indian customers are also starting to adopt more and more AI,” said Subramanian.

Not Just a Hardware Company

Subramanian emphasised that Jarvislabs.ai is not merely a hardware company but rather a comprehensive solution provider saying “How to train your model and how to deploy it are much more challenging engineering problems to solve. We are packaging this entire thing for our customers.”

Additionally, as a bootstrapped company, Jarvislabs.ai cannot afford to offer substantial discounts to its customers, unlike VC-backed firms that often provide significant discounts as they aim to capture a larger market share.

To compensate for this, Jarvislabs.ai is providing additional value to its customers at the same price. “We’re not just focusing on renting GPUs as raw material. We are trying to help. Let’s say, for example, you are a small company with 40-50 people. You may not have the DevOps team or the AI team to create an optimsed environment. So, we also simplify things for them. We have automated all these steps. With just a few clicks, all of this is done for you.”

He shared that, at the end of the day, as a company, it also has to make money, and it’s not sustainable to buy GPUs at a very high price. “NVIDIA has been increasing prices year-on-year with every release. For example, the RTX 6000 used to cost us around three lakhs. The current generation costs around six to 6.5 lakhs per GPU, but the cost at which we give it to the customer has not increased a lot,” he said.

Subrmanian said that he intends Jarvislabs.ai to be like an app store which will host multiple frameworks. “We are working on a new version of the product, and we hope to release it this month. This version will incorporate a lot more frameworks. When we started, there were only three popular frameworks: PyTorch, TensorFlow, and Fast.ai,” he concluded.

The post How This Coimbatore-based AI Lab is Helping Developers Build LLMs appeared first on Analytics India Magazine.

What to Expect from Indian IT in 2024

For Indian IT, the year 2023 unfolded as a narrative of paradoxes and challenges. While luminaries like TCS, Infosys, HCLTech, Wipro, and Tech Mahindra celebrated record-high deal wins, the overarching ambiguity surrounding revenue growth cast a shadow over the sector’s trajectory for the entire year.

Industry pundits, reflecting on the incidents of 2023, echoed sentiments of a sector grappling with its weakest performance since the recession of 2008. In the midst of this uncertainty, the beacon of hope for the Indian IT sector emanates from the realm of generative AI.

Can 2024 be the year of AI deals?

Pareekh Jain, CEO of EIIRTrend, predicts that around 2% of the revenue in the coming year would be through generative AI directly. He also highlighted how AI was one of the positive things in 2023 for Indian IT, and this would influence a lot of partnerships and deals in 2024.

On the contrary, the $1.5 billion Infosys AI deal that got cancelled would also play a major impact in the industry. “A lot of companies would realise that moving too fast and investing in technologies would not be ideal. Thus, they would move with a lot of caution,” said Jain. But he believes that deals would definitely be a big focus this year.

As industry analysts predict a transformation from Q3 onwards, a nuanced narrative of deal divergence and recalibration unfolds. Large deals are anticipated to derive impetus from vendor consolidation, steering away from conventional digital transformation paradigms.

As generative AI emerges as a focal point of interest, its potential impact on revenue is forecasted to crystallise in the third quarter of 2024. Jain said that though in 2023 AI was not the focus of deal in Indian IT, it was one of the biggest factors of the deals. “Generative AI could also be a deal breaker in 2024,” said Jain, highlighting that it would be the focus.

That said, direct generative AI focused deals would definitely be the focus. “More than 50% of the deals would be directly influenced by generative AI,” said Jain.

For example, HCLTech announced a slew of new deals, where it will be involved in digital and cloud transformation, alongside generative AI initiatives. TCS also predicts that Indian IT firms would be a major focus globally and expects a lot of new deals in the year, along with 12 million more net jobs by 2025.

LTIMindtree has expressed its commitment to integrating generative AI into its products and solutions, revealing that they have participated in more than 100 discussions and currently have over 20 active engagements with clients.

Wipro has doubled the number of customers as compared to the last quarter, and would focus on large deals in 2024. HCLTech is working with a handful of customers with generative AI projects, while LTIMindtree is engaged in over 20 clients for generative AI.

Infosys, along with a few other parties including Elon Musk, AWS and others donated USD 1 billion to OpenAI. TCS and Wipro also announced generative AI capabilities in partnership with Google Cloud.

A lot in the pipeline

Companies have actively discussed various ongoing experiments and proof of concepts (PoCs) in their pipelines. In the previous month, Accenture disclosed an industry-leading generative AI pipeline valued at $450 million in new bookings, a significant increase from the $300 million reported for the entire fiscal year of FY23.

This along with a $3billion investment in AI for three years. The company anticipates a shift from general AI experimentation to more proof of concepts and pilot projects in 2024.

Similarly, TCS reported having over 250 generative AI opportunities in its pipeline, Infosys is engaged in over 50 active generative AI projects, and HCLTech is involved in more than 140 generative AI PoCs at various stages.

However, the caveat is that these PoCs are still in the pre-production stage and are far from generating immediate revenue. Analysts believe that only 10% of the PoCs will move to production. Considering this, a rapid recovery in the IT services sector in 2024 seems unlikely.

Analysts from Kotak note that despite high expectations for a recovery in discretionary spending in 2024, enterprises across most sectors are still focused on cost reduction. Many have set cost-saving targets extending into 2024, and the reprioritisation of spending toward areas of investment is not yet complete.

2023 continued?

During the earnings call, CP Gurnani, former managing director of Tech Mahindra said, “Tech Mahindra is now working in about 60 customer locations on actually using generative AI to enhance operations, innovation, and productivity.”

On the other hand, Indian IT firms have also largely stopped hiring freshers. This trend might continue this year, along with rampant possible layoffs.

In the initial half of fiscal year 2024, the top-tier IT firms collectively shed 39,000 employees, signalling a strategic recalibration in response to evolving market dynamics. Delays in lateral hires and adjustments in compensation strategies underscored a pragmatic approach to talent acquisition.

Overall, the Indian IT majors have trained close to seven lakh employees in generative AI, in partnership with companies like Google, Microsoft, Oracle and NVIDIA. The monetary impact of the same would be on the forefront of discussions in 2024.

The post What to Expect from Indian IT in 2024 appeared first on Analytics India Magazine.

GPT-4 Drives WHOOP’s Personalised Fitness Coaching 

WHOOP, known for its fitness-tracking wearables, has introduced WHOOP Coach, a personalised health and fitness coach driven by OpenAI’s GPT-4.

The LLM-powered coach can provide answers to a wide range of fitness and health-related queries. For instance, it can address questions like “What was my lowest resting heart rate ever?” or “What weekly workout schedule would help me reach my goal?” — all while offering personalised guidance based on each individual’s unique body and goals.

The WHOOP engineering team experimented with integrating GPT-4 into their companion app, fine-tuning it with anonymized member data and proprietary algorithms. The result is an AI-powered coach that offers personalised, relevant, and conversational insights based on each user’s unique body and goals.

Introduced in September, WHOOP coach stands out as the first wearable to provide highly individualised performance coaching on demand “WHOOP Coach leverages GPT-4 to essentially serve as a search engine for your body,” said Jaime Waydo, chief technology officer at WHOOP.

Members can now access WHOOP Coach at any time, receiving tailored guidance and answers to specific queries related to their fitness and health. This ranges from inquiries about resting heart rates to recommendations for weekly workout schedules. The capability to process thousands of individual data points enables Coach to act as a personalised search engine for one’s body.

“Now, we can offer on-demand, personalised health and fitness coaching within seconds. This is the first of its kind, and it will transform our members’ relationship with their data, as well as their access to information in the health and wellness space.” said Will Ahmed, WHOOP founder and chief.

The post GPT-4 Drives WHOOP’s Personalised Fitness Coaching appeared first on Analytics India Magazine.