Intel Soon to be on Par with NVIDIA

Intel Soon to be on Par with NVIDIA

Everyone knows why NVIDIA is on the top of the market in generative AI when there are competitors like Intel and AMD who are also making strides. Now the behemoth Intel is going all in into the AI hardware segment, and it might have just cracked it.

Intel is planning to onboard another version of AI accelerator superchip, Falcon Shores 2, by 2026. “We have a simplified roadmap as we bring together our GPU and our accelerators into a single offering,” CEO Pat Gelsinger said.

The recently released Intel Xeon Max 9480 combines 56 cores and is not a standard DDR5 memory, but a 64 GB HBM2e, which is on par with the ones used in GPUs and AI accelerators. Interestingly, Intel believes that these GPUs would be mostly used for inference based tasks, and not actually training AI models.

More leaks about the upcoming 14th Gen Meteor Lake processor also suggest that the CPU might have a DDR5 memory, which is also similar to Apple’s M2 chip design. Moreover, it is also expected that AI will play a major role in the Meteor Lake CPUs. Much is awaited at the upcoming Intel Innovation 2023 event on September 19.

All roads lead to AI

While Intel navigates its strategic adjustments, it’s noteworthy that NVIDIA has also taken a substantial leap by venturing into the CPU market with the GH200 supercomputer. This expansion into CPUs complements NVIDIA’s existing prowess in GPUs and AI technologies, while venturing into the CPU market.

Furthermore, Intel’s Falcon Shores chips were originally conceived as a fusion of CPU and GPU cores, representing the company’s inaugural venture into the ‘XPU’ architecture for high-performance computing. Nonetheless, a few months ago, Intel astounded the industry by opting for a GPU-only approach and deferring the chip’s release until 2025. The company’s voyage into the realm of AI and GPUs has encountered a series of twists and turns.

This comes after Intel has already been providing Gaudi2 AI chips for training models. Interestingly, Gaudi2 works 2.4 times faster than the NVIDIA A100, and is almost coming close to the H100 Hopper GPU.

On the flip side, Intel’s decision to decelerate its GPU release cadence could potentially place it at a disadvantage against more advanced architectures like NVIDIA Grace Superchips and AMD’s Instinct MI300, both slated for launch in 2023. This strategic choice may hinder Intel’s competitiveness in the HPC market.

Intel is convinced with two AI markets

One that deals with the infrastructure, for which the company has the Habana Labs Gaudi. The other is for inference, which according to Intel can be adequately done on a CPU like Xeon.

This seemed like an almost good approach until NVIDIA jumped onto the same wagon. At NVIDIA’s recent financial call, Jensen Huang said that the company plans to introduce L40S, a GPU that is specifically designed for fine-tuning and inference. Given that people are already using NVIDIA H100s for training, shifting to Intel’s Xeon processors might be a big leap to make.

Amid these strategic shifts, in May Intel had announced a strategic collaboration with the Boston Consulting Group (BCG) to facilitate generative AI. This partnership aimed to harness Intel’s AI hardware and software to craft tailor-made generative AI solutions for enterprises, all while ensuring the sanctity of data privacy and security.

Cut to August, Anthropic announced its partnership with BCG for bringing responsible generative AI for enterprise clients. But after that, NVIDIA and Microsoft also made an investment in Anthropic, which makes all of this a little confusing. This proves that AI is for everyone to take. This might be a hint that the Google-backed startup might be leveraging Intel supercomputers for building generative AI, which is quite rarely heard given the NVIDIA GPU dominance.

Intel GPUs, NVIDIA CPUs

You read that right. All of these strategic shifts were interpreted as Intel’s departure from direct competition with AMD’s Instinct MI300 and NVIDIA’s Grace Hopper processors, both of which boast a combined CPU+GPU design. NVIDIA went into the CPU business in March and it left people wondering, what else does the GPU giant want to take on Intel with.

On the other hand, Intel has shed light on the reasoning behind this strategic reconfiguration. While the initial plan for Falcon Shores permitted flexible CPU/GPU configurations, Intel emphasised the significance of enabling customers to utilise various CPUs, including those from rivals like AMD and NVIDIA. But given the announcements around Gaudi2, a potential Gaudi3, and NVIDIA venturing into CPUs, Intel might be able to take a bigger slice of the GPU market soon.

The rivalry between Intel and NVIDIA, encompassing both CPUs and GPUs, is poised to intensify, potentially reshaping the landscape of AI and HPC. Intel is the best CPU to buy, and NVIDIA is the best GPU to buy. This is widely believed. But it might take a turn soon given that the conversation about computing has almost shifted around AI.

The post Intel Soon to be on Par with NVIDIA appeared first on Analytics India Magazine.

How to Get a Job in Data Science as a Student

How to Get a Job in Data Science as a Student
Illustration by Author. Source: Flaticon

Data Science is a challenging field and just showing that you have a certification or a university degree is not enough to obtain a job in industry. The recruiters need to understand what value you can add to the company compared to the other candidates.

Since a qualification is not enough by itself to land a job in data science, the best moment in getting started to do different experiences is when you are a student. You still are young, have more time than you think and energies that can be exploited in increasing your possibilities in getting your first data science job.

In this article, I want to show five different ways to demonstrate your skills and even earn money. Let’s get started!

Build a Portfolio of Data Science Projects

When I started to apply for my first job in data science while I was a student, I didn’t have any experience in that field obviously. The projects I did during my master degree helped me in demonstrating my skills to the company.

To show your abilities in coding, the best way is to share your projects on GitHub. You surely should not only focus your attention on python scripts/jupyter notebooks, but also a READ.me file that explains well how the project is organised and the results obtained.

For example, one of my personal projects that have attracted the interest of managers was Topic Modeling with BERTopic, in which I trained a BERTopic model to identify topics from e-commerce clothing reviews and interpreted the results through interactive data visualisations. It helped to demonstrate that I was able to solve a NLP problem.

Write Data Science Articles

While I was a student, I started to write data science articles in different publications, like Towards AI and Towards Data Science, allowing me to reach a lot of readers from different countries, put into practice what I have learned during my university courses, go deeper in unknown topics that I was interested to explore, receive job proposals and, even, earn money.

The best teacher is the feedback from the community, that can seem scary at first when you start, especially the negative comments, but it really helps to improve your critical thinking and open your mind to other possible solutions in solving a problem.

You can also create your own blog from scratch, instead of publishing on a publication, but you should take into consideration how much time it will take to build the website. Anyway, it has become easier thanks to Chat-GPT and other AI tools based on Large Language Models.

Create Youtube Videos

Articles can be a good way to demonstrate your abilities in communication, which is an important skill for a data scientist, but also recording Youtube Videos can prove what you are made of.

It can be done in combination with the articles, but even creating videos by itself can be enough. There are a lot of known DS influencers that are known for their great capacity in building video content, like StatQuest with Josh Starmer, Data Professor and Patrick Loeber.

Work as Freelance Data Scientist

Education content is not the only way to break into data science. Another possibility to gain experience is working as a freelance data scientist. It’s a good alternative to the traditional 9-to-5 job: it can guarantee freedom and flexibility, especially when you are a student and you also need time to attend and study the university courses.

You can get started from existing platforms, like Upwork and Fiverr, where you need to create your profile and apply for data science freelance gigs published by clients. You can also reach or be reached by clients in LinkedIn, which is considered the best social network for connecting with people belonging to the data science community.

Participate to Data Science Competitions

Another way to improve your chances of getting a job is by entering the data science competitions. I would highlight that it’s not so important that you do a lot of competitions, it’s better that you focus on real-world projects that can be valuable for the company that is interviewing you. Quality is always more important than quantity.

The most known and popular platform for hosting competitions is Kaggle. You have surely consulted this website while searching data science topics, datasets and other stuff. Other competitions can be found in DrivenData, TopCoder and DevPost.

Final Thoughts

I hope that this article has inspired you to take action, allowing you to improve your resume. It’s also a good way to put into practice the knowledge you have acquired during your studies, overcome the imposter syndrome and become more flexible.

Surely, you shouldn’t do it only for improving your CV, but also for the other good points I listed before that can allow you to become more confident of your capabilities while doing different experiences outside the traditional job. Thanks for reading! Have a nice day
Eugenia Anello is currently a research fellow at the Department of Information Engineering of the University of Padova, Italy. Her research project is focused on Continual Learning combined with Anomaly Detection.

More On This Topic

  • 5 Supporting Skills That Can Help You Get a Data Science Job
  • How to Get Your First Job in Data Science without Any Work Experience
  • 5 Tips to Get Your First Data Scientist Job
  • How to Get a Job as a Data Engineer
  • How to Get a Job as a Data Scientist
  • Are you satisfied in your job? Take our Data Community Job Satisfaction…

This $30 Guide Will Help You Become a ChatGPT Expert

Person using a laptop with ChatGPT webpage on display.
Image: StackCommerce

ChatGPT has taken the business world by storm, helping people ideate content, automate work and much more. But ChatGPT really is only as good as you are at giving it the right information. And that’s why The Ultimate ChatGPT Prompts Guide is such a valuable resource.

ChatGPT is all about saving you time, so it defeats the purpose if you’re typing in prompts over and over and not quite satisfied with what the system is outputting. The prompts guide helps you with that.

This comprehensive guide features more than 3,000 highly-optimized prompts covering a wide range of topics in a user-friendly layout, with easy-to-follow instructions. Whether you need help with marketing, research, ideation, business planning, SEO, content creation, social media ads, problem-solving and much more, this guide will help you unlock the full potential of ChatGPT for your business or research.

With these prompts, you’ll be able to save time while generating fresh ideas, refining your strategies, streamlining your workflows and becoming a ChatGPT adherent. There are virtually countless applications, whether you’re a marketer, researcher, salesperson, entrepreneur, data analyst, content creator or practically anything else.

In addition to the prompts, you’ll also get a multi-chapter guide on ChatGPT best practices and a complete guide on how to use prompts to fully master ChatGPT. Before you know it, you’ll be cruising with ChatGPT like a pro and saving all kinds of time.

Tap into another level of productivity with The Ultimate ChatGPT Prompts Guide. Right now, you can get a lifetime subscription for 74% off $119 at just $29.99.

Prices and availability are subject to change.

Subscribe to the Innovation Insider Newsletter

Catch up on the latest tech innovations that are changing the world, including IoT, 5G, the latest about phones, security, smart cities, AI, robotics, and more.

Delivered Tuesdays and Fridays Sign up today

Unveiling the AI Potential: FinBert beats ChatGPT in Financial Text Analytics

There are several aspects in a study conducted by JPMorgan and Queen’s University where FinBert, an LLM fine-tuned on financial domain data, specific terminology, language structures, and concepts demonstrated its superiority over ChatGPT in the context of financial text analytics.

Gpt-3.5-turbo and GPT-4 with 8k tokens were compared with FinBert, and tested on Arithmetic Reasoning, News Classification Sentiment Analysis and Named Entity Recognition.

FinBert outperforms ChatGPT in sentiment analysis tasks related to financial texts. Sentiment analysis in finance requires understanding nuanced expressions and the impact of news on investors.

FinBert, being specifically designed for the financial domain, does not require extensive adaptation or fine-tuning to perform well in these tasks. In Few-shot Learning: ChatGPT’s performance on various tasks required more extensive prompts.

FinBert outperforms ChatGPT and even competes with human experts in arithmetic reasoning. This indicates that FinBert is a highly specialised model for financial tasks, while ChatGPT, might not reach the level of expertise demonstrated by FinBert in the financial domain.

In tasks such as financial named entity recognition (NER) and sentiment analysis, where a deep well of domain-specific knowledge is essential, ChatGPT and GPT-4 struggles. Their inability to grasp the intricacies of financial terminologies becomes evident.

Comparative analysis puts both the models against fine-tuned models tailored for the financial sector, like FinBert and FinQANet. The outcome underscores the fact that, while these LLMs hold potential, they are not yet on par with their specialised counterparts.

This study paves the way for further enhancements.

The gap between these state-of-the-art generative language models and domain-specific proficiency remains, but it also presents a promising opportunity for refinement.

Rajiv Shah, a machine learning engineer at Hugging Face said on linkedin, “A domain-specific model like FinBERT is more accurate for finance tasks than GPT-4”.

The post Unveiling the AI Potential: FinBert beats ChatGPT in Financial Text Analytics appeared first on Analytics India Magazine.

AWS Receives Cloud Service Provider Empanelment From MeitY

Amazon Web Services (AWS) India has announced that it has received cloud service provider (CSP) empanelment from India’s Ministry of Electronic and Information Technology (MeitY) for cloud services provided using the AWS Asia Pacific (Hyderabad) Region.

Operational since November 2022, as part of AWS’s total investment of INR 1,36,500 crores (US $16.4 billion) in cloud infrastructure in India by 2030, the AWS Asia-Pacific (Hyderabad) Region is the second AWS Region in India to be fully empaneled by MeitY.

In 2017, AWS India became the first global CSP in India to receive full empanelment for its cloud service offerings after the AWS Asia-Pacific (Mumbai) Region completed MeitY’s STQC (Standardization Testing and Quality Certification) audit.

AWS Regions are comprised of Availability Zones (AZs) that place infrastructure in separate and distinct geographic locations. AZs are located far enough from each other to support customers’ business continuity, and near enough to provide low latency for high-availability applications that use multiple data centres.

Public sector organisations and financial services institutions across India, including banking and payment providers, can achieve greater operational resiliency from the AWS Asia-Pacific (Hyderabad) Region – which consists of three AZs – by using its cloud infrastructure for higher availability, disaster recovery, data backup, while meeting security and regulation requirements, and enjoying lower latency performance, AWS said.

Moreover, government organisations can improve e-governance standards and enable on-demand digital services for citizens and businesses across India. Similarly, financial services organisations can benefit from agile, efficient, and security-compliant cloud solutions at scale, providing consumers faster and secure digital banking, insurance, and payment innovations.

“With both AWS Regions in India empaneled by MeitY, AWS is providing customers more choice to access resilient, secure, and low-latency cloud infrastructure, while offering more ability for AWS Partners to develop innovative solutions and address customer needs,” “Shalini Kapoor, director and chief technologist for the public sector with AWS India Private Limited (AWS India), said.

The post AWS Receives Cloud Service Provider Empanelment From MeitY appeared first on Analytics India Magazine.

Demystifying Machine Learning

Demystifying Machine Learning
Image by Author Tradition vs. Transformation: A Look Back and Forward

Traditionally, computers used to follow an explicit set of instructions. For instance, if you wanted the computer to perform a simple task of adding two numbers, you had to spell out every step. However, as our data became more complex, this manual approach of giving instructions for each situation became inadequate.

This is where Machine Learning emerged as a game changer. We wanted computers to learn from examples just like we learn from our experiences. Imagine teaching a child how to ride a bicycle by showing it a few times and then letting him fall, figure it out, and learn on his own. That's the idea behind Machine Learning. This innovation has not only transformed industries but has become an indispensable necessity in today's world.

Learning the Basics

Now that we have a basic understanding of the term ”Machine learning“, let us familiarize ourselves with some fundamental terms:

Data

Data is the lifeblood of Machine learning. It refers to the information that a computer uses to learn. This information can be numbers, pictures, or anything else that a computer can understand. This is further divided into 2 categories:

  • Training Data: This data refers to the examples that we use to teach the computer.
  • Testing Data: After learning, we test the performance of the computer using some new, unseen data referred to as the test data.

Label and Features

Imagine that you are teaching a kid how to differentiate between different animals. The name of the animals (dog, cat, etc) would be the labels while the characteristics of these animals (number of legs, fur, etc) that help you recognize them are the features.

Models

It is the outcome of the Machine Learning process. It is the mathematical representation of the patterns and relationships within the data. It's like making a map after exploring a new place.

Types of Machine Learning

There are four main types of Machine Learning:

Supervised Machine Learning

It is also referred to as guided learning. We provide the labeled dataset to our Machine Learning algorithm where the correct output is already known. Based on these examples it learns the hidden patterns in the data and can predict or correctly classify the new data. The common categories within supervised learning are:

  • Classification: Sorting things into separate distinct categories for example classifying pictures as cats or dogs, emails as spam or not spam, etc.
  • Regression: It involves predicting numerical values for example price of the house, your GPA, or the number of sales based on certain features.

Unsupervised Machine Learning

Here the computer is provided the unlabelled data without prior hints and it explores the hidden patterns on its own. Just consider that you are handed a box of puzzle pieces with no picture and your task is to group similar pictures to form a complete picture. Clustering is the most common type of unsupervised learning where similar data points are grouped into a group. For example, we can employ clustering to group similar kinds of social media posts and users can follow the sub-topics of their interest.

Semi-Supervised Machine Learning

Semi-supervised learning contains a mix of labeled and unlabelled datasets where the labeled dataset acts as the guiding point in identifying the patterns in data. For example, you give a chef a list of the main ingredients to use but do not provide the complete recipe. So although they don’t have the recipe some hints that might help them to get started.

Reinforcement Learning

Reinforcement learning is also called learning by doing. It interacts with the environment and gets a reward as a penalty for its actions. With time, it learns to maximize the reward and perform well. Imagine that you are training a puppy and you give positive feedback by rewarding him when he behaves well and negative feedback in the form of withholding rewards. Over time, the puppy learns the actions that lead to rewards and also the ones that don’t

High-Level Machine Learning Process

Machine Learning, much like the art of cooking, possesses the magical ability to transform raw, disparate elements into profound insights. Just as a skilled chef adeptly combines various ingredients to craft a delicious dish. These are the 6 basic steps used to perform a Machine Learning Task:

Demystifying Machine Learning
Image by Author

1. Data Collection

Data is an important resource and its quality matters a lot. Diverse, more relevant data yields better results. You can think of it as the Chef gathering various ingredients from different markets.

2. Data Preprocessing

Most of our data is not in the desired form. Like washing, chopping, and preparing ingredients before cooking, data preprocessing involves cleaning and organizing data for the learning process. Some common issues that you might face are missing data, outliers, incorrect format, etc.

3. Choosing an Algorithm

Similar to selecting the recipe for a specific dish, you choose an algorithm based on the problem that you are trying to solve. This choice may also be influenced by the type of data that you have.

4. Training the Model

Think of it as the cooking process where we wait unless the flavors come together. Similarly, we let the model learn from the training data. An important concept of learning rate also comes into play here that determines how big of a step your model takes during each iteration of training. If you add too much salt or spice at once, the dish could become overpowering. Conversely, if you add too little, the flavors might not develop fully. The learning rate finds the perfect balance for gradual flavor enhancement.

5. Testing & Evaluation

Once the learning process wraps up, we put it to the test using special test data, much like tasting a dish and examining its appearance before sharing it with others. Common evaluation metrics include accuracy, precision, recall, and F1 score, depending on the problem at hand.

6. Tuning and Iteration

Adjusting the seasoning or ingredients to perfect the dish, you fine-tune your models by introducing more variables, choosing a different learning algorithm, and adjusting parameters or the learning rate.

Wrapping Up

As we wrap up our exploration of the basics of Machine learning, remember that it's all about empowering the computers to learn and make decisions with minimal human intervention. Stay curious and keep an eye out for our next articles, where we'll dive deeper into the various types of machine learning algorithms. Here are some beginner-friendly resources for you to explore further:

  • Introduction to Machine Learning with Python
  • Machine Learning For Absolute Beginners
  • Machine Learning — Coursera
  • Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow

Kanwal Mehreen is an aspiring software developer with a keen interest in data science and applications of AI in medicine. Kanwal was selected as the Google Generation Scholar 2022 for the APAC region. Kanwal loves to share technical knowledge by writing articles on trending topics, and is passionate about improving the representation of women in tech industry.

More On This Topic

  • Demystifying AI: The prejudices of Artificial Intelligence (and human…
  • Demystifying Bad Science
  • Explainable AI: 10 Python Libraries for Demystifying Your Model's Decisions
  • 5 Machine Learning Skills Every Machine Learning Engineer Should Know in…
  • KDnuggets News, December 14: 3 Free Machine Learning Courses for Beginners…
  • A Solid Plan for Learning Data Science, Machine Learning, and Deep Learning

Snowflake Sees Great Potential in Indian Public Sector

Snowflake sees great business potential in the government and public sector in India. “I think in the next five to ten years, that will be one of our biggest opportunities,” Sanjay K Deshmukh, senior regional vice president, ASEAN, and India, Snowflake, told AIM at the Bangalore leg of its Data Cloud World Tour (DCWT).

The Data Cloud company headquartered in Montana, US, wants to tap into the public sector to expand its presence in the country further.

“We have our service available in India. It’s running in the Indian data centres of AWS and Azure so we qualify in that sense.”

Deshmukh adds that Snowflake is looking to address two primary use cases. Firstly, to facilitate collaboration between government agencies that currently operate in silos. Often, agencies like the land department, transport, or infrastructure work independently, making policy decisions without a comprehensive data understanding.

“We’re working on multi-ministry and multi-agency initiatives to foster collaboration. This will lead to more informed policy decisions and improved citizen services,” he said.

Secondly, the focus for Snowflake is on government agencies that provide citizen services. “These agencies are under pressure to modernise their services and actively seek ways to improve citizen interactions. This is a common thread among the agencies we are engaged with, and it presents a significant opportunity for us in the government sector.”

The post Snowflake Sees Great Potential in Indian Public Sector appeared first on Analytics India Magazine.

NVIDIA Expands Cloud Business with Investments, Partnerships

Hugging Face is steadily growing to be the one stop solution for AI models, after their partnership with NVIDIA, even more so. Recently, they announced that Hugging Face users will have access to NVIDIA DGX Cloud AI supercomputing to train and fine tune their AI models.

With this partnership, Hugging Face users get access to SOTA GPUs and infrastructure needed to rapidly train and finetune foundation models at scale and drive a new wave of enterprise LLM development.

Hugging Face has divided the costs of building specific models parameters on DGX, tokens and datasets from the range of $32,902 to $18,461,354 making the process more efficient. “We hope to spur a new wave of experimentation and learning in AI – exploration that simply wasn’t feasible before,” said Julien Chaumond.

NVIDIA CEO Jensen Huang also acknowledged the immense potential of Hugging Face as an AI enabler and said, “I think there are 50,000 companies with 2 million users and there’s some 275,000 models and 50,000 datasets. Just about everybody who creates an AI model and wants to share it with the community puts it up on Hugging Face.”

NVIDIA, now a stakeholder at hugging face, is further driving their own product through them. They are likely to reap the benefits in revenue this year and next because of massive customer investments.

While Hugging Face got a big boost from this partnership, NVIDIA has another plan. They’re growing the user base of their DGX cloud by inviting the open source community and accelerating AI innovation at scale.

Is NVIDIA Eating Cloud?

NVIDIA has positioned themselves smartly between the cloud service providers and their customers.

In March, NVIDIA announced that the company is partnering with leading cloud service providers to host DGX Cloud infrastructure, starting with Oracle Cloud Infrastructure (OCI). The company also said that Microsoft Azure, Google Cloud will Host DGX Cloud soon. Because of the partnership, users get to quickly access GPU servers and DGX Cloud without having to make commitments to multiple cloud vendors.

While NVIDIA insists this partnership is a shared success, it is clearly at odds with other cloud providers. The use of DGX as a “single software platform” from NVIDIA allows companies to streamline the operation of its new AI software across various cloud providers and within its own data centres, enhancing efficiency.

Additionally, NVIDIA’s DGX cloud servers are built by engineers who leverage their knowledge of the company’s chips and are in a better position to fine-tune DGX Cloud servers. As expected, this surpasses their performance in comparison to other AI-centric servers rented out by cloud providers, as confirmed by individuals closely acquainted with the service.

Interestingly, the similar partnership offer by NVIDIA was also made to AWS, but it refused. Joshua Bernstein, a former manager at AWS and Google Cloud, commenting on this said, “It puts NVIDIA’s brand front and centre over a cloud providers’ brand.” AWS is the biggest player in the cloud sector and is already a formidable competitor with their EC2 P5 service.

NVIDIA for Startups

Nvidia is slowly building up its foothold in the AI startup and enterprise ecosystem. The startups themselves might be slow in arrival but NVIDIA is making sure to enable them.

Last year, NVIDIA open-sourced certain parts of their GPU software for Linux after criticism of not being open-source friendly.

They also said that they’re making code run more efficiently across different types of processors like CPUs, GPUs, and AI accelerators. To support the open source projects, they have a team of engineers for support and build of their services. This inevitably helps startups and enterprises to build on top of their services as a comprehensive solution.

Apart from their DGX Cloud capabilities, NVIDIA has a host of cloud products and services at a reasonable cost for enterprises. They unveiled a suite of cloud-based services tailored to develop generative AI models specialised for specific tasks within their domains, such as medical imaging.

These services fall under the umbrella of NVIDIA AI Foundations and encompass two distinct offerings: NVIDIA NeMo, dedicated to language models, and NVIDIA Picasso, geared towards generating image, video, and 3D content. Once these models are prepared for deployment, businesses have the flexibility to run them either within NVIDIA’s cloud infrastructure or on other platforms of their choice.

Source: NVIDIA NeMo

NVIDIA is also heavily investing in AI startups. They’ve invested in 20+ companies this year apart from Hugging Face. The most recent was yesterday, when Databricks announced it received funding from NVIDIA and Capital One. These agreements enable Nvidia to keep AI companies loyal to their products. Majority of these startups and enterprises are already Nvidia patrons, from which the company stands to regain its investments.

The post NVIDIA Expands Cloud Business with Investments, Partnerships appeared first on Analytics India Magazine.

PyCharm vs. Spyder: Choosing the Right Python IDE

PyCharm vs Spyder: Choosing the Right Python IDE

Python is immensely popular among developers and data scientists due to its simplicity, versatility, and robustness, making it one of the most used programming languages in 2023. With around 147,000 packages, the Python ecosystem continues to evolve with better tools, plugins, and community support.

When we talk about Python development, Integrated Development Environments (IDEs) take center stage, allowing developers to enhance their coding experience. Two popular IDEs for Python development are PyCharm and Spyder. This article briefly compares Python vs. Spyder to help developers make an informed choice.

A Brief Look Into Pycharm & Spyder

Before comparing PyCharm vs. Spyder to determine the best IDE for Python development, it’s essential to understand what these tools entail.

PyCharm: Python IDE for Professional Developers

PyCharm Dashboard UI

PyCharm is a product by JetBrains that offers a feature-rich integrated development environment for Python. The IDE has two editions – PyCharm Community and PyCharm Professional. The former is a free, open-source version, while the latter is a paid version for full-stack development. Both versions support several features, including code completion, code analysis, debugging tools, and integration with various version control systems. The professional edition further includes frameworks for web development and data science.

Spyder: Python IDE for Scientists, Engineers & Data Analysts

Spyder dashboard UI

Spyder, or Scientific Python Development Environment, is an open-source IDE primarily focusing on data science and scientific computing in Python. It’s part of the Anaconda distribution, a popular package manager and distribution platform for Python. Spyder provides comprehensive tools for advanced data analysis, visualization, and scientific development. It features automatic code completion, code analysis, and vertical/horizontal screen splits with a multi-language editor pane that developers can use for creating and modifying source files. Moreover, developers can extend Spyder’s functionality with powerful plugins.

Pycharm vs. Spyder Comparison – Who Wins?

Pycharm vs. Spyder Comparison - Who Wins?

Several similarities and differences exist between these two IDEs. Below, we compare them against various dimensions, including code editing and navigation features, debugging capability, support for integrated tools, customizability, performance, usability, community support, and pricing.

Code Editing & Navigation

Both PyCharm and Spyder offer powerful code editing and navigation features, making it easy for developers to write and understand code across files. While Spyder provides similar code completion and navigation ability, it is less robust than PyCharm's code editing features, which offer context-based recommendations for faster development. For instance, developers get code completion suggestions (sorted by priority) based on other developers' work in a similar scenario.

PyCharm leads this category with its advanced code analysis and completion capabilities.

Debugger

PyCharm’s professional version has a Javascript-based debugger that supports various debugging modes, including remote debugging. It also provides a visual debugger with breakpoints, variable inspection, and step-by-step execution.

Spyder includes a PDB debugger. PDB is a source debugging library for Python that lets developers set conditional breakpoints and inspect stack frames. Its variable explorer is particularly helpful for checking variable states at several breakpoints.

While Spyder’s debugging capabilities are robust, PyCharm’s visual debugger is better as it helps in more complex debugging scenarios.

Integrated Tools

PyCharm has extensive integration with third-party tools and services. For instance, it has built-in support for version control systems like Git, SVN, Perforce, etc. The professional edition supports web development frameworks, such as Django, Flask, Angular, etc., making it an excellent choice for full-stack development.

Spyder, primarily a data science and scientific computing utility, comes with numerous libraries and tools, such as NumPy, SciPy, Matplotlib, and Jupyter Notebooks. Also, it shares all libraries that come with the Anaconda distribution. However, Spyder only supports Git for version control.

Overall, PyCharm overtakes Spyder in this category since the former offers integration with diverse tools through plugins.

Customization

PyCharm offers a high level of visual customization, allowing developers to tailor the IDE according to their workflow and preferences. They can change font type and color, code style, configure keyboard shortcuts, etc.

Spyder is relatively less customizable compared to PyCharm. The most a user can do is change the user interface’s (UI’s) theme using a few options among light and dark styles.

Again, PyCharm takes the win in the customization category.

Performance

While performance can vary depending on the size and complexity of the projects, Spyder is relatively faster than PyCharm. Since PyCharm has many plugins installed by default, it consumes more system resources than Spyder.

As such, Spyder's lightweight architecture can make it a better choice for data scientists who work on large datasets and complex data analysis.

Spyder is the clear winner in the performance category.

Usability & Learning Curve

PyCharm has many customization options for its user interface (UI). Developers benefit from an intuitive navigation system with a clean layout. However, its extensive feature set means it has a steep learning curve, especially for beginners.

In contrast, Spyder's interface is much more straightforward. Like R, it has a variable navigation pane, a console, a plot visualization section, and a code editor, all on a single screen. The simplified view is best for data scientists who want a holistic view of model results with diagnostic charts and data frames. Also, Spyder's integration with Jupyter Notebooks makes data exploration and visualization easier for those new to data science.

Overall, Spyder is ideal for beginners, while PyCharm is more suited to experienced Python developers.

Pricing

PyCharm has a free and paid version. The free community version is suitable for individual developers and teams working on a small scale. The paid version, the Professional Edition, comes in two variants – for organizations and individuals. The organization version costs US 24.90 monthly, while the individual one costs USD 9.90 monthly.

In contrast, Spyder is open-source and entirely free to use. It comes as part of the Anaconda distribution, which is also open-source and free.

In terms of cost, Sypder is a clear winner. However, in Python development, it is up to the practitioners and organizations to choose based on their business requirements.

Community Support

Both PyCharm and Spyder have active communities that provide extensive support to users. PyCharm benefits from JetBrains' strong reputation and rich experience in building Python development tools. As such, developers can utilize its large user community and get help from a dedicated support team. They also have access to many tutorials, help guides, and plugins.

Spyder leverages the Anaconda community for user support. With an active data science community, Spyder benefits from the frequent contributions of data scientists who provide help through forums and online resources, data science tutorials, frameworks, and computation libraries.

Again, it is up to the practitioners and organizations to choose a community that aligns with their task or business requirements.

PyCharm vs. Spyder: Ideal Use Cases

PyCharm vs. Spyder: Ideal Use Cases

Choosing between PyCharm and Spyder can be challenging. It’s helpful to consider some of their use cases so practitioners can decide which IDE is better for their task.

PyCharm is ideal for full-stack developers as the IDE features several web and mobile app development tools and supports end-to-end testing. It’s best for working on large-scale projects requiring extensive collaboration across several domains.

Spyder, in contrast, is suitable for data scientists, researchers, and statisticians. Its lightweight architecture allows users to perform exploratory data analysis and run simple ML models for experimentation. Instructors can use this IDE to teach students the art of data storytelling and empower them to train machine learning models efficiently.

PyCharm vs. Spyder: The Final Choice

The choice between PyCharm and Spyder ultimately depends on user needs, as both IDEs offer robust features for specific use cases.

PyCharm is best for experienced professionals who can benefit from its advanced web development tools, making it an excellent choice for building web and mobile apps. Users wishing to learn data science or work on related projects should go for Spyder.

To read more interesting technology-related content, navigate through Unite.ai‘s extensive catalog of insightful resources to amplify your knowledge.

Will AGI Be Built in China?  

Chinese artificial intelligence startup Baichuan Intelligent Technology recently introduced two open-source AI-powered large language models called Baichuan 2-7B and Baichuan 2-13B.

Interestingly, one thing that caught everyone’s eye was that it performed better than ChatGPT on AGIEval – a benchmark created by Microsoft Research. ChatGPT’s score on AGIEval was 46.13 while Baichuan 2-13 B’s was 48.17. The word soon spread that Baichuan2-13B beats ChatGPT on AGIEval.

Baichuan 2
The 2nd iteration the leading Chinese model is a major improvement.
Baichuan 2-13B beats ChatGPT on AGIEval.

Code: https://t.co/pwJIGDJ6nz
Paper: https://t.co/j7WRMJ5O5u


The effort is on a different scale.
The paper detailing the process for creating… pic.twitter.com/OpFfODhKyx

— Yam Peleg (@Yampeleg) September 13, 2023

This isn’t new. Whenever a new foundational model arrives, it often wants to show how it measures up to ChatGPT. However, the real question here was how Baichuan 2-13 B was able to do that.

Language matters

The rankings of LLMs on benchmarks often depend on the training dataset they use, and AGIEval is no exception. In AGIEval’s case, it primarily assesses foundational models within the framework of college entrance exams like SAT, LSAT, and various math competitions.

Surprisingly, the real reason for outperforming ChatGPT is that the Baichuan 2-13 B was trained on Chinese-English bilingual dataset comprising several million webpages from hundreds of reputable websites that represent various positive value domains, encompassing areas such as policy, law, vulnerable groups, general values, traditional virtues, and more.

Upon closer inspection of the AGIEval research paper, it becomes evident that, in addition to entrance exams like SAT and LSAT, it also encompasses Chinese entrance exams such as Gaokao. Furthermore, this benchmark extends its scope to include bilingual tasks in both Chinese and English.

On the other hand, open source models like LLaMA and Llama 2 have focused primarily on English. For instance, the main data source for LLaMA is Common Crawl, which comprises 67% of LLaMA’s pre-training data but is filtered to English content only.

As Baichuan is China based it has easy access to the chinese material to train its model. Recently, a report came out that said Chinese authorities have approved Baichuan Intelligent Technology and Zhipu AI’ requests to open its AI large language models to the public.

It’s apparent that Chinese authorities may not have intervened to prevent them from accessing data from the Chinese internet, distinct from the global internet used elsewhere.

Microsoft is Behind This

The AGIEval benchmark created by Microsoft says that evaluating the general abilities of foundational models to tackle human-level tasks is a vital aspect of their development and application in the pursuit of AGI.

Their paper casually dismisses traditional benchmarks, which rely on artificial datasets and says they may not accurately represent human-level capabilities. Does it mean that Baichuan 2-13B is closer to AGI than ChatGPT. If that’s indeed the case, it’s a significant achievement.

However if we introspect, AGIEval is no different from any other benchmark in a way that they all are based on a certain dataset on which they are evaluated.

Apart from AGIEval, if we check Baichuan 2 coding and math problem solving abilities, it is way behind ChatGPT. So how can we conclude that AGIEval is the criteria to judge AGI.

Recently, Baidu also claimed that Ernie 3.5, the latest version of its Ernie AI model, had surpassed “ChatGPT in comprehensive ability scores” and outperformed “GPT-4 in several Chinese capabilities.”. The Beijing-based company referred to a test conducted by the state newspaper China Science Daily, which included datasets like AGIEval and C-Eval.

Interestingly, Microsoft’s Orca earlier this year had also claimed that it performs better on AGIEval. In the Orca’s research paper it is specifically mentioned that “Evaluation benchmarks like AGIEval which relies on standardized tests such as GRE,SAT, LSAT etc offer more robust evaluation frameworks”. However, If we dig into Orca’s dataset, one finds out that it is also trained on Chinese dataset.

Orca scored higher than ChatGPT and was nearly identical to text-davinci-003 in the AGIEval benchmark. However, Orca still significantly lags behind GPT-4 in these metrics.

— Tiz (@tatendampofu4) June 16, 2023

The marketing of Orca revolved around the AGIEval benchmark. Similarly majority of the foundational models which are performing well on AGIeval have a Chinese dataset which gives them undue advantage. It isn’t fair to all the other models present out there.

In Conclusion

The performance of AI models on benchmarks like AGIEval is not solely indicative of their progress towards AGI. While models like Baichuan 2-13 B have showcased impressive scores, the underlying advantage often lies in their training data, particularly the accessibility of specific Chinese internet content.

AGIEval’s focus on real-world tasks is commendable, but it’s crucial to recognise a broader spectrum of abilities are equally vital in assessing AGI. Can we really say that if an LLM passes SAT, LSAT or any other exams, it is closer to AGI?

The post Will AGI Be Built in China? appeared first on Analytics India Magazine.