In today’s data-driven world, the ability to convey insights through compelling visuals is paramount. From business reports to academic research, data visualization plays a pivotal role in making complex information accessible and understandable. However, the journey from raw data to stunning charts and graphs can often be a tedious and time-consuming process, often requiring a deep understanding of not only your data and the business problem at hand but also the intricacies of dozens of data visualization tools and techniques.
Enter ChartGen AI, the groundbreaking new online application from Einblick, an AI-native data notebook based on 6 years of research at MIT and Brown University. ChartGen AI promises to change the way we create charts and graphs, making it easier, faster, and more accessible than ever before. In this blog post, we’ll explore what ChartGen AI can do and why AI-powered chart-making is the next revolution in data visualization.
From Data to Chart in Seconds
Over the past year, so many apps have sprung up claiming to leverage generative AI to speed up work inefficiencies. What sets ChartGen AI apart is how intuitive the UI is, and how quickly you can get results. It’s also completely free, with no account setup required.
Step 1: Upload Data
You can either upload your dataset or link to a Google Sheet. No need for complex data preparation or formatting.
Step 2: Describe Your Chart
Describe the chart you want in natural language. For example, “Create a pie chart of col_1” or “Plot a histogram of col_2.”
Step 3: Generate Chart
Click “Generate.” That’s it! In just three straightforward steps, you’ll have your chart ready for download as either a PNG or an SVG. If you want to tweak your chart’s appearance, simply add details to your prompt. For instance, “Use a pastel color palette” or “Make the axis labels bigger.”
Why ChartGen AI Matters: Saving Time and Boosting Productivity
Forget about dealing with finicky pivot tables. With ChartGen AI, you can directly upload your dataset or link to a Google Sheet, and get straight to chart creation. ChartGen AI offers customization at your fingertips. Want a specific color palette or more bins in your histogram? Just add a few words to your prompt, and it adapts accordingly.
Rapid Results without Data Preprocessing
ChartGen AI delivers charts in seconds. With ChartGen AI, you can directly upload your dataset or link to a Google Sheet. Say goodbye to data preprocessing, forget about dealing with finicky pivot tables, and get straight to chart creation.
Leverage Generative AI for Impact
Generative AI can be used in a variety of applications, but the space of AI charting is particularly appealing. After all, before you can create a chart, you have to know what it should look like–you have to be able to describe it. It is at this point in the process that ChartGen AI intervenes, and allows you to create visualizations faster than ever before.
Accessibility for All
ChartGen AI democratizes data visualization. It’s user-friendly, making it accessible to those without advanced technical skills, while benefiting those who have the technical skills, but lack the time resources.
Try ChartGen AI Today
If you’re looking to create stunning charts and graphs effortlessly, ChartGen AI should be your go-to tool. Whether you’re a data novice or an experienced analyst, ChartGen AI streamlines the chart making process.
Say farewell to the complexity of charting tools and embrace the simplicity of ChartGen AI. Try it today and experience the power of AI-driven chart creation. Unlock a new era of data visualization that’s easier, faster, and more accessible than ever before.
The post A New Free AI Tool Makes Creating Charts Easier Than Ever appeared first on Analytics India Magazine.
Cloud notebooks have become the norm these days for data scientists and analysts to run their code and generate analytical reports. Cloud notebooks provide a browser-based interface to write and execute code without the need to install anything locally. Additionally, high-end hardware is accessible for accelerating machine learning research and development.
Cloud notebook platforms provide more than just free computing and pre-built environments. They also offer third-party tool integrations, collaborations, and publication options. In this blog, we will explore the top seven cloud notebooks and their best features. Utilizing these features can enhance your current data science development stack.
1. Deepnote
Deepnote is on the top now. Why? Recently, they have introduced new features that will simplify your development experience. I love the platform, team, and the community. Moreover, I use it for every data science and machine learning project.
You can start a machine in less than a minute and benefit from a pre-built development environment. It also supports all kinds of programming languages, and you can create your own environment using Docker Hub.
I will highly recommend you create an account and experience it yourself. It has become easier even for non-technical professionals to write and debug the code using the Deepnote AI feature.
Image by Author 2. Kaggle
With Deepnote, Kaggle has also introduced new features this year. For example, they are adding new high-tier GPUs, scheduling runs, dedicated tabs for models, and quickly loading the dataset. The only thing they need to catch up on is live collaboration and commenting.
With Kaggle, you get high-tier CPUs, GPUs, and TPUs for free. Moreover, you get free storage, access to open-source datasets and code, Google Cloud integrations, and versioning.
It is my go-to platform when I am participating in a competition or experimenting with deep learning models.
Again, I highly recommend Kaggle due to its strong community and high-tier hardware for you AI projects.
Image by Author 3. Hex
Hex is now available to the public, and it is the popular option for your data science and analytics tasks. It provides a similar feature to Deepnote, but due to slow environment loading and code running. I have ranked it at third. It is also limited in many ways.
Hex is a modern Data Workspace that aims to make working with data easier and more collaborative. It allows users to connect to a variety of data sources, including databases, cloud storage, and APIs. Once the data is connected, users can analyze it using either SQL or Python directly within interactive notebooks.
Image by Author 4. Noteable
I came across Noteable when it was introduced as a ChatGPT plugin. Before that, I had no idea it existed. It is simple, fast, and comes with all kinds of features.
This platform provides data connection, loading, versioning, publishing of notebooks, live collaboration, and fast environment loading. The best part about this platform is its minimalist design. Additionally, you can connect it with ChatGPT to generate and run code with outputs. This feature makes it highly valuable in the notebook category.
Image by Author 5. Google Colab
Google Colab is the same old cloud notebook that we love and cherish. We use them to run our deep learning code, and sometimes it is an excellent and convenient tool. Things have changed over the years as they limit the free tier and focus on paid options.
Apart from easy access to free GPU and fast loading time, there is little to Google Colab. It is not a complete data science platform that you want to use everyday.
Image by Author 6. Naas
Naas is known for its data templates for all kinds of problems. This platform provides a low-code solution for creating powerful data products by combining automation, analytics, and AI.
It has limited computing and features. It does provide free credit every month for you to run and execute code. Apart from that, it is the JupyterLab on the cloud.
Image by Author 7. Datalore
JetBrains Datalore is similar to Noteable, but it is slow and lacks some key features. You are also limited by computing. I used to run my code on Datalore, but since its launch, there haven't been much improvement or changes to the platform. It seems like JetBrains have forgotten about it.
It has some features that you can get in Deepnote, but the UI is confusing for any beginner to get used to it. The only good part is that it provides free storage and computing.
Image by Author Conclusion
In conclusion, cloud notebooks have become essential tools for data scientists and analysts to efficiently carry out their work. The top options provide great value with free GPUs, easy setup, collaboration features, and integrations with other services. Deepnote stands out as the most full-featured option, with its fast environment loading, AI assistance, live collaboration, and publishing capabilities.
Kaggle is great for working with deep learning given its high-tier hardware access. Hex and Noteable offer modern interfaces and integrations like ChatGPT. While Google Colab and others have their niche uses, Deepnote seems to be leading the pack with its focus on the end-to-end data science workflow. Whichever platform you choose, cloud notebooks will undoubtedly enhance your data science projects and ability to derive insights. Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master's degree in Technology Management and a bachelor's degree in Telecommunication Engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.
More On This Topic
Top 5 Free Cloud Notebooks in 2022
New From Anaconda! Data Science Training and Cloud Hosted Notebooks
11 Best Practices of Cloud and Data Migration to AWS Cloud
The only Jupyter Notebooks extension you truly need
Analyze Python Code in Jupyter Notebooks
Snowflake and Saturn Cloud Partner To Bring 100x Faster Data Science to…
Private telecommunications companies and other internet service providers are advocating for app providers to contribute a “fair share” for building on their infrastructure. This stance challenges the fundamental principles of net neutrality.
In a response to these developments, the Minister of State for Electronics and Information Technology, Rajeev Chandrasekhar, took to Twitter today to reaffirm India’s commitment to net neutrality. He asserted that India was one of the pioneering nations to ensure net neutrality, firmly resisting attempts by telecom companies to act as gatekeepers of the internet. Chandrasekhar highlighted the significance of this decision in propelling India towards becoming a global leader in innovation, boasting a thriving startup ecosystem that adheres to international standards.
Over-the-top (OTT) services like Netflix and Amazon Prime, offering services like streaming, social networking, e-commerce, and gaming, serve as the lifeblood of India’s digital landscape. These services have not only provided users with high-quality content at minimal cost but have also catalysed the growth of data consumption and economic activities within India.
Proposing revenue-sharing mechanisms between OTT platforms and Internet Service Providers (ISPs) at this point could potentially reverse this positive trend.
The post Indian IT Minister Chandrasekhar Opposes Telcos Demands to Charge OTTs for ‘Fair Share’ appeared first on Analytics India Magazine.
Data science is a trendy buzz that every industry is aware of. As a data scientist, your main job is extracting meaningful insights from the data. But here is the downside — with data exploding at exponential rates, it is more challenging than ever. You will often get the feeling of finding the needle in a digital haystack. This is where the data science tools emerge as our saviors. They help you mine, clean, organize, and visualize the data to extract meaningful insights from it. Now, let's address the real problem. With the abundance of data science tools, how will you navigate to find the right ones? The answer to this question rests in this article. Through a careful blend of personal experience, invaluable community feedback, and the pulse of the data-driven world, I have curated a list that packs a punch. I have focused only on open-source data science tools because of their cost-effectiveness, agility, and transparency.
Without any further delay, let’s explore the top 10 open-source data science tools you need to have in your arsenal this year:
1. KNIME: Bridging Simplicity and Power
KNIME is a free and open-source tool that empowers both data science novices and experienced professionals by opening the door to effortless data analysis, visualization, and deployment. It's a canvas that transforms your data into actionable insights with minimal programming. It's a beacon of simplicity and power. You should consider using Knime for the following reasons:
GUI-based data preprocessing and pipelining empower users from various technical backgrounds to perform complex tasks without much hassle
Allows seamless integration into your current workflows and systems
The modular approach of KNIME enables the users to customize their workflows according to their need
2. Weka: Tradition Meets Modernity
Weka is a classic open-source tool that allows data scientists to preprocess data, build and test machine learning models, and visualize data using a GUI interface. Although it's quite old, it remains relevant in 2023 due to its adaptability to cater to model challenges. It provides support for various languages including R, Python, Spark, scikit-learn, etc. It is extremely handy and reliable. Here are some of the features of Weka that outshine:
It is not only suitable for data science practitioners but is also an excellent platform for teaching machine learning concepts thereby providing educational value.
Enables you to achieve sustainability effortlessly by cutting the data pipeline idle time resulting in reduced carbon emissions.
Delivers mind-bending performance by providing support for high I/O, low latency, small files, and mixed workloads with no tuning.
3. Apache Spark: Igniting Data Processing
Apache Spark is a well-known data science tool that offers real-time data analysis. It is the most widely used engine for scalable computing. I have mentioned it due to its lightning-fast data processing capabilities. You can easily connect to different data sources without being worried about where your data lives. Although it's impressive, it's not all sunshine and rainbows. Because of its speed, it needs a good amount of memory. Here is why you should choose Spark:
It is easy to use and offers a simple programming model that allows you to create applications using the languages that you are already familiar with.
You can get a unified processing engine for your workloads.
It’s a one-stop shop for batch processing, real-time updates, and machine learning.
4. RapidMiner: The Full Data Science Lifecycle
RapidMiner stands out due to its comprehensive nature. It's your true companion throughout your complete data science lifecycle. From data modeling and analysis to data deployment and monitoring, this tool covers it all. It offers a visual workflow design, eliminating the need for intricate coding. This tool can also be used to build custom data science workflows and algorithms from scratch. The extensive data preparation features in RapidMiner enable you to deliver the most refined version of data for modeling. Here are some of the key features:
It simplifies the data science process by providing a visual and intuitive interface.
RapidMiner's connectors make data integration effortless, regardless of size or format.
5. Neo4j Graph Data Science: Unveiling Hidden Connections
Neo4j Graph Data Science is a solution that analyzes the complex relationships between the data to discover hidden connections. It goes beyond rows and columns to identify how the data points are interacting with each other. It consists of pre-configured graph algorithms and automated procedures specifically designed for the Data Scientists to quickly demonstrate value from graph analysis. It is particularly useful for social network analysis, recommendation systems, and other scenarios where connections matter. Here are some of the additional benefits that it provides:
Improved predictions with a rich catalog of over 65 graph algorithms.
Allows seamless data ecosystem integration using ith 30+ connectors and extensions.
Its powerful tools allow fast-track deployment enabling you to quickly release workflows into the production environment.
6. ggplot2: Crafting Visual Stories
gglot2 is an amazing data visualization package in R. It turns your data into a visual masterpiece. It is built on the grammar of graphics offering a playground for customization. Even the default colors and aesthetics are much nicer. ggplot2 utilizes the layered approach to add details to your visuals. While it can turn your data into a beautiful story waiting to be told, it's important to acknowledge that dealing with complex figures can lead to cumbersome syntax. Here is why you should consider using it:
The ability to save plots as objects allows you to create different versions of the plot without repeating a lot of code.
Instead of juggling around the multiple platforms, ggplot2 provides a unified solution.
Plenty of helpful resources and extensive documentation to help you get started.
7. D3.js: Interactive Data Masterpiece
D3 is the short form of Data-Driven Documents. It is a powerful open-source javascript library that enables you to create stunning visuals by employing DOM manipulation techniques. It creates interactive visualizations that respond to the changes in data. However, it has a steep learning curve specifically for those who are new to JavaScript. Although its complexity can be a challenge the rewards it offers are invaluable. Some of them are listed below:
It offers customizability by providing a wealth of modules and APIs.
It is lightweight and doesn’t affect the performance of your web application.
It works well with the current web standards and can easily integrate with other libraries.
8. Metabase: Data Exploration Made Simple
Metabase is a drag-and-drop data exploration tool that is accessible to both technical and non-technical users. It simplifies the process of analyzing and visualizing the data. Its intuitive interface enables you to create interactive dashboards, reports, and visualizations. It is getting extremely popular among businesses. It provides several other benefits which are listed below:
Replaces the need for complex SQL queries with plain language queries.
Support for collaboration by enabling users to share their insights and findings with others.
Supports over 20 data sources, enabling users to connect to databases, spreadsheets, and APIs.
9. Great Expectations: Ensuring Data Quality
Great Expectations is a data quality tool that enables you to assert checks on your data and to catch any violations effectively. As the name suggests, you define some expectations or rules for your data and then it monitors your data against those expectations. It enables the data scientists to have more confidence in their data. It also provides data profiling tools to accelerate your data discovery. The key strengths of Great Expectations are as follows:
Generates detailed documentation for your data that is beneficial for both technical and non-technical users.
Seamless integration with different data pipelines and workflows.
Allows automated testing for detecting any issues or deviations earlier in the process
10. PostHog: Elevating Product Analytics
PostHog is an open-source primarily in the product analytics landscape enabling businesses to track user behavior to elevate product experience. It enables the data scientists and engineers to get the data much quicker removing the need for writing SQL queries. It’s a comprehensive product analysis suite with features like dashboards, trend analysis, funnels, session recording, and much more. Here are the key aspects of PostHog:
Provides an experimentation platform to data scientists through its A/B testing capabilities.
Allows seamless integration with data warehouses for both importing and exporting data.
Provides an in-depth understanding of user interaction with the product by capturing session replays, console logs, and network monitoring
Wrapping Up
One thing that I would like to mention is that as we are progressing more in the field of Data Science, these tools are not just mere choices now, they have become the catalyst guiding you toward informed decisions. So, please don’t hesitate to dive into these tools and experiment as much as you can. As I wrap up, I'm curious, Are there any tools you've come across or used that you'd like to add to this list? Feel free to share your thoughts and recommendations in the comments below. Kanwal Mehreen is an aspiring software developer with a keen interest in data science and applications of AI in medicine. Kanwal was selected as the Google Generation Scholar 2022 for the APAC region. Kanwal loves to share technical knowledge by writing articles on trending topics, and is passionate about improving the representation of women in tech industry.
More On This Topic
Overview of Albumentations: Open-source library for advanced image…
Closed Source VS Open Source Image Annotation
The Role of Open Source Tools in Accelerating Data Science Progress
How to Build Data Frameworks with Open Source Tools to Enhance Agility and…
Generate Synthetic Time-series Data with Open-source Tools
OLAP vs. OLTP: A Comparative Analysis of Data Processing Systems
Genpact has announced an expanded collaboration with Amazon Web Services (AWS) aimed at revolutionising financial crime risk operations through the integration of generative AI and LLMs. The partnership entails the integration of Genpact’s cloud-based financial crime suite, riskCanvas, with Amazon Bedrock, yielding significant efficiencies and benefits for clients, including Apex Fintech Solutions.
Leveraging its existing relationship with AWS, Genpact is merging its intellectual assets and industry expertise with AWS’s generative AI capabilities. The seamless integration of Amazon Bedrock FMs into Genpact’s riskCanvas financial crimes software suite aims to unlock substantial value, enhance speed, and improve accuracy in detecting, investigating, and preventing financial crime threats for businesses.
This integration enables experts to review outputs and adopt a guided decision-making process, offering comprehensive summaries and analyses of potential financial crime activities. This results in improved efficiency and precision.
Genpact has actively engaged with numerous riskCanvas clients to enhance the detection, investigation, and prevention of various financial crime threats. As a result, Genpact is delivering accelerated efficiencies and significant impact to clients in the finance and capital markets sectors, with a notable example being their collaboration with Apex Fintech Solutions.
Justin Morgan, Head of Financial Crimes Compliance at Apex Fintech Solutions, emphasised the importance of staying ahead of financial criminals: “With the addition of generative AI features to Genpact’s riskCanvas, our analysts will be able to produce Suspicious Activity Report (SAR) narratives and case summaries at the click of a button using inputs from millions of data points. We expect this will reduce time spent on case summarizations by 60%, allowing our analysts to spend more time identifying truly suspicious financial activity.”
By utilising approved client data from the secure riskCanvas ecosystem and Amazon Bedrock’s secure data handling capabilities, highly accurate outcomes are generated while maintaining robust data protection standards.
Atul Deo, General Manager, Amazon Bedrock at AWS, emphasised the importance of responsible AI implementation: “Amazon Bedrock is rooted in secure data handling, encrypting all data and allowing users to customise models privately. Integrated with Genpact’s riskCanvas, this powerful combination enables our mutual customers to enhance productivity in investigating, detecting, and preventing financial crime threats.”
BK Kalra, Global Business Leader, Financial Services, Consumer and Healthcare at Genpact, highlighted the growing need for generative AI in financial crime operations: “Genpact’s expanded relationship with AWS represents a pivotal step in redefining the operations landscape for enterprises. Together we can unlock untapped value, and fuel significant growth opportunities for our clients, solidifying our commitment to delivering valuable business impact.”
The post Genpact, AWS to Fight Financial Crime with Generative AI appeared first on Analytics India Magazine.
Meta AI, and everyone in its AI team, including Mark Zuckerberg and Yann LeCun, have been one of the biggest proponents of open source. It was indeed leading the open source LLM race all this while with LLaMA and then the pseudo open source, Llama 2. But since then, a lot of models have been coming up in the open source community that are outperforming Llama 2 on various benchmarks.
Now, it might be time for tech-giant to realise how to finally use its AI technology to earn some bucks. Why should it compete in the open source market? Zuckerberg, in his latest podcast with Lex Fridman in the metaverse, said that Meta might have to reconsider if it is going to open source the next iteration of Llama, which is Llama 3. “Right now, the priority is building that into a bunch of consumer products,” said Zuckerberg.
At the Meta Connect 2023, the team announced AI in a lot of consumer products such as AI chatbots and characters on WhatsApp and Instagram, based on celebrities, and a slew of other integrations. There is no doubt that open sourcing models and integrating them into its products has helped Meta move faster, but now there might be a chance of going the other way.
Not surprisingly, @ClementDelangue gets it. https://t.co/gOGWqC9KDZ
— Yann LeCun (@ylecun) September 29, 2023
Meta learns how to use Llama
Zuckerberg adds that Meta trained Llama 2 and released it as an open source, but it is not a consumer product, but just an AI infrastructure. Though Zuckerberg is all about open sourcing AI and loves what the community has been doing with Llama 2, when it comes to Llama 3, he said that the debate with Fridman was very helpful for open sourcing Llama 2, and the same would be needed for Llama 3.
“We would need a process to red team this, and make it safe. My hope is that we would be able to open source the next version when it is ready to, but we are not close to doing that this month. It’s a thing that we are still early in work now,” said Zuckerberg. This time, Llama 3 might even be better than GPT-4.
Thus, making Llama 3 safe sounds like a great bet, but it is somewhat in contrast to what Zuckerberg said in the last podcast with Fridman. “The stage we are in right now, the equities balance strongly in my view towards doing this more openly.” He believed that if companies think that they have gotten close to what they believe is “superintelligence”, it makes sense for them to discuss and think through a lot more.
“We want a lot more researchers working on it and for a lot of these reasons open source has a lot of advantages.” He explained how open source software is more secure because there are a lot of people criticising it, finding holes in it, and thus in the end making it safe. “I think this is how we will achieve more aligned models, by making it open source and letting people work on it,” said Zuckerberg.
There were still discussions if Llama 2 is in the works or not, but then it was actually open sourced by Meta. Yann LeCun also said that open sourcing is the best way to make progress in the LLM field, and thus promotes every company who goes this way.
More high-performance open source LLMs coming out to enrich the AI open ecosystem. This one from @MistralAI https://t.co/9Jg4RI8v3S
— Yann LeCun (@ylecun) September 27, 2023
But, there is always a doubt about how to make money out of open source AI. For example, Mistral AI, a startup founded by former DeepMind and Meta employees, is also a great proponent of open source. Just this week, the team released an open source model. But in 2024, the team aims to raise more fundings in 2024, to build a consumer product, apart from the open source ones.
Similarly, Meta’s competitors like Google or OpenAI have not released their models to the public and kept it closed. But earlier, when OpenAI released GPT-2, it was open source. The company decided to pivot to the closed source model approach after GPT-3. Probably because it realised that the revenue model with open source is not that fool proof.
Stepping up the game
With Zuckerberg’s latest comments, it seems like Meta might steer a little away from being the champion of open source and finally learn how to monetise it within its consumer products. Which is indeed fair. The company has been making strides in the metaverse and with its headsets, where it is implementing AI.
Hot take: LLaMA-3 model will be on par or better than GPT-4
— anton (@abacaj) September 29, 2023
Meta’s Llama 2 is arguably still “an order of magnitude behind what OpenAI and Google are doing,” said Zuckerberg. But still, many companies have been adopting the open source model and earning a few bucks for themselves. It would only make sense now for Meta to keep Llama 3 for itself, and use it for consumer products, at least for a little bit, just to catch up on the four years. Even if that means, ditching its true nature.
The post Llama 3 Might Not be Open Source appeared first on Analytics India Magazine.
Identifying the top nation-state actors can depend on who you ask, which underscores the need to gather threat intelligence from varied data sources.
In a climate where geopolitical issues now can drive industry discussions, organizations will be better served if they formulate their cybersecurity strategy on information that reflects threat activities on an international scale.
For some organizations, this requirement means gathering threat intel that is comprehensive and, in particular, diverse.
Most threat intelligence houses currently originate from the West or are Western-oriented, and this can result in bias or skewed representations of the threat landscape, noted Minhan Lim, head of research and development at Ensign Labs. The Singapore-based cybersecurity vendor was formed through a joint venture between local telco StarHub and state-owned investment firm, Temasek Holdings.
"We need to maintain neutrality, so we're careful about where we draw our data feeds," Lim said in an interview with ZDNET. "We have data feeds from all reputable [threat intel] data sources, which is important so we can understand what's happening on a global level."
Ensign also runs its own telemetry and SOCs (security operations centers), including in Malaysia and Hong Kong, collecting data from sensors deployed worldwide. Lim added that the vendor's clientele comprises multinational corporations (MNCs), including regional and China-based companies, that have offices in the U.S., Europe, and South Africa.
He noted that threat activities may not necessarily be motivated by geopolitical issues, as some threat activities originate from the region in which their targets are based. The vendor's Chinese customers, for example, have experienced cyberattacks that originate from Asia.
Rather that being politically driven, he added that most attacks tend to be financially motivated.
When nation-state threat actors are involved, the top three countries from where attacks originate are often China, Russia, and the U.S., said Candid Wuest, vice president of cyber-protection research for Acronis.
Also: Asian banks are a favorite target of cybercooks, and malicious bots their preferred tool
North Korea and Iran round out the top five locations from where nation-state attacks originate, and there has been little change in the countries that lead this pack, said Wuest, citing data points from Acronis' monitoring network. The data security vendor was founded in Singapore and is currently headquartered in Switzerland. Its founder Serguei Beloussov is Russian-born, but picked up Singapore citizenship in 2001. He left his role as CEO in 2021 and currently serves as Acronis' chief research officer.
Echoing Lim's views, Wuest told ZDNET that U.S. nation-state threat actors often are under-represented because most major news organizations are Western and generally will not make reference to this specific group of nation-state actors.
He noted that U.S. state-sponsored attacks also are very targeted, involving one to two victims rather than hundreds, and often go unnoticed. In comparison, the volume of nation-state attacks from China may seem higher because there are more organizations looking out for this group of threat actors, he said.
He added, though, that attacks from Chinese nation-state actors have also been higher in volume and frequency. In particular, there has been an increase in attacks targeting VPNs and firewalls, and known vulnerabilities in systems such as Microsoft Exchange Server.
The U.S. government last month released a report highlighting the top software vulnerabilities commonly exploited in 2022. These included several flaws previously highlighted in 2021 and used by China's state-sponsored cyber actors, according to the August 3 statement released by U.S. security agencies and their Five Eyes allies comprising Australia, New Zealand, Canada, and the U.K.
The Chinese government, on the other hand, has blamed U.S. intelligence agencies for a July 2023 cybersecurity attack on Wuhan Earthquake Monitoring Center, citing a "very complex" malware used in the incident. Beijing said the attack appeared to originate from government-backed hackers in the U.S. and had targeted network equipment that collected seismic intensity data. The datasets contained information concerning national security, such as details of military defense facilities, which are taken into account in determining seismic intensity.
Chinese officials suggest that accessing relevant data from seismic monitoring centers can enable hackers to estimate underground structures of a specific area and assess if it is a military base. This data will prove useful to foreign military intelligence agencies, such as the U.S. Department of Defense.
U.S. tech vendors, though, believe Chinese nation-state actors are armed with more sophisticated tools and tactics.
"Chinese cyber espionage has come a long way from the smash-and-grab tactics many of us are familiar with," said John Hultquist, Google Cloud Mandiant's chief analyst. "They have transformed their capability from one that was dominated by broad, loud campaigns that were far easier to detect. They were brash before, but now they are clearly focused on stealth."
"The result is an adversary much harder to track and detect. The reality is that we are facing a more sophisticated adversary than ever and we'll have to work much harder to keep up with them," Hultquist wrote in a commentary note on the recent Microsoft security breach, believed to be the work of cyberattackers based in China.
Also: This data platform will help banks share criminal intelligence
Codenamed Storm-0558, the adversary group gained access to Microsoft email systems for intelligence gathering, according to Microsoft. The tech giant said the breach impacted 25 organizations, including U.S. government agencies and accounts of individuals likely associated with these agencies.
Chinese cyber-espionage activities have increasingly leveraged initial access and post-compromise strategies aimed at minimizing detection, noted a July 2023 post by Google Cloud Mandiant. "Chinese threat groups' exploitation of zero days in security, networking, and virtualization software, and targeting of routers and other methods to relay and disguise attacker traffic both outside and inside victim networks," the U.S. vendor said. "We assess with high confidence that Chinese cyber espionage groups are using these techniques to avoid detection and complicate attribution."
In an interview with ZDNET, Hultquist reiterated the increased sophistication of attacks from Chinese nation-state actors, which have been notably more stealth and tougher to detect. He pointed to the complex attack instigated by Storm-0558 as indicative of skillsets that have improved dramatically.
He named China, Russia, North Korea, and Iran as the top nation-state actors most often encountered by Google Cloud Mandiant's customers. These threat actors also are targeting critical information infrastructures and, in some cases, working with governments or government-funded adversary groups and providing compromised information to intelligence agencies.
Also: The best security keys
Wuest, though, deemed the U.S. as the most sophisticated among nation-state actors, noting that the Equation Group — believed to be part of the Five Eyes alliance — used a specific encryption algorithm in a series of targeted attacks a decade ago that, until today, remains unbreakable.
He then ranked attacks by Russian nation-state actors second in terms of sophistication, followed by their counterparts in China.
Chinese threat actors often opt for volume in terms of their attacks, which are not usually sophisticated in nature, he said, suggesting that this simplicty might be because there is little reason to improve their tactics if these attacks continue to be effective. For instance, Chinese threat actors still use spear phishing, which is not difficult to execute and is largely dependent on humans as the weakest link, he said.
And while Wuest acknowledged that Chinese attacks are now more advanced, he noted that this was a natural progression as most skills would typically improve over two decades.
Describing Chinese nation-state actors as more sophisticated might be overstating the threat landscape and a rhetoric used by western governments to raise overall awareness among organizations about the need to have good cyberdefense, Wuest said.
Also: AI, trust, and data security are key issues for finance firms and their customers
He added that the majority of cyberattacks today still target old vulnerabilities that organizations have been left unpatched.
Lim also noted that threat actors have been targeting security appliances and attempting to evade detection. He said actors are now more targeted and aware of appliances used specifically by government agencies, and hence know which devices to target.
Attackers are also constantly updating their tactics, using the most current methods and targeting popular tools, such as video-conferencing applications and VPNs, he added.
The growing threat underscores the need for organizations to do the basics and adopt best practices, said Wuest, who noted that some companies are still failing to patch known vulnerabilities. With spear phishing and 90% of attacks coming through email, he also emphasised the importance of email security and other measures to combat such attacks.
Also: Can AI detectors save us from ChatGPT?
This defense will be especially critical as more hackers tap generative artificial intelligence (AI) to churn more convincing phishing email messages, scraping personal information from social media accounts to create fake personas, he said.
Lim concurred, noting that his team had begun detecting generative AI and machine-learning tools being used to craft emails, as well as to circumvent verification and access controls.
Wuest added that cyber criminals can leverage AI to gain scale and speed in launching attacks, and make it more difficult to verify and authenticate users.
He expressed concern that generative AI will disrupt many industries and significantly change how attacks are carried out. If generative AI is used to trigger false commands, for instance, the risk and potential threat will be significant, he said.
Hultquist urged industry leaders to transform their practices alongside threat adversaries or risk falling behind. "We need more data, better engineering, better tools, and better [defense] plans," he said. "Technology is always changing, as are the adversaries. We have to work hard to stay on top of that."
Any conversation about any LLMs is incomplete without how it fares on the various benchmarks like MMLU, HumanEval, AGIEval etc. Be it GPT-4 or Llama 2, the first criteria the creators try to showcase their LLMs is by putting up their benchmark scores on the research paper. Surprisingly, India, which is emerging as the new force in AI, does not have a benchmark of its own to evaluate LLMs.
In layman terms, we can say that LLMs are students and benchmarks are a kind of an examination for them on which they need to score good.
At present, discussions around generative AI predominantly revolve around China, the United States, and the Middle East, with notable players like TII (based in the UAE), Baidu (from China), and OpenAI (from the US) gaining significant attention.
There is still a need for a metric that India could utilise to evaluate LLMs, one that caters specifically to the country’s needs. Geography can also play a crucial role in understanding the purpose of an LLM, a factor that many people might not have paid attention to.
India is a diverse nation with a multitude of languages, cultures, and contexts. A benchmark specifically designed for India would take into account the linguistic diversity of the country. From addressing regional language nuances to recognizing cultural references, an Indian LLM benchmark could enable models to be more culturally sensitive.
A glimpse of LLM Leaderboard with list of LLMs fairing across various benchmarks (Source: Hugging Face)
Developers from around the globe create models and submit them to Hugging Face, which maintains an LLM Leaderboard. Here, users can check which model is performing best on different metrics like MMLU, ARC, HellaSwag, and more. Hugging Face even calculates an average based on these metrics. Undoubtedly, it is a very useful tool. However, the question remains: are these metrics enough to judge an LLM? Perhaps to some extent, yes.
Indian Datasets
Majority of the benchmarks which have evolved from the U.S. take their exams into consideration. For example, MMLU, covers 57 tasks including elementary mathematics, US history, computer science, and law. Similarly AGIEval, is derived from exams like SAT, LSAT and other tests like Chinese College Entrance Exam (Gaokao), law school admission tests, math competitions, lawyer qualification tests, and national civil service exams.
In India, there are several tough competitive exams like UPSC, NEET, JEE-Advanced, CAT etc. which essentially covers the essence of India. Creating a benchmark for LLMs that’s tailored to Indian competitive exams like the UPSC would enable the development of language models capable of understanding the unique requirements of these exams, which often involve complex questions, diverse languages, and a deep understanding of India’s history, culture, and governance, alongside critical thinking and logical reasoning.
Earlier this year, AIM tried to check GPT-3 for UPSC, surprisingly it failed. However, a few months later GPT-4 was able to cross the cutoff.
Work in Progress
It is unfortunate that India hasn’t been able to create an LLM from scratch until now. However, the situation appears to be changing as NVIDIA recently partnered with Indian giants like Reliance, TATA, and Infosys.
Furthermore, the Indian IT giant Tech Mahindra is currently working on an indigenous LLM known as Project Indus. This model will have the ability to speak in many Indic languages, notably Hindi, and is planned to cover 40 different Indic languages initially. More languages originating from the country will also be added subsequently.
The way India might use these LLMs can be completely different from how the rest of the world. To understand Indian models at performing various tasks we need to create a benchmark based on regional or vernacular dataset.
Moreover models like GPT-4, Llama 2, Falcon 180B and Mistral 7B are mostly trained on datasets which are in English. There is a dire need of datasets created in Indian languages. As of today, the Indian constitution recognizes 22 major languages of India.
Thanks to the Indian government, it launched Project Bhashini Initiative where it looks at developing a National Public Digital Platform for local languages, which can be leveraged to create products and services for citizens using generative AI.
Similar to Hugging Face’s LLM Leaderboard that tracks, ranks, and evaluates open LLMs and chatbots, Bhashini intends to create training and benchmarking datasets. It will lead to healthy competition and will kickstart an LLM race in India.
What next?
Majority of the benchmarks are created by the research wing of universities or tech companies. For example, MMLU was created by the University of California, Berkeley. It’s high time that Indian universities like IIT Delhi, IIT Madras and IISc Bengaluru start working towards creating a dataset which will further lead towards the creation of a benchmark. Also, universities have easy access to data related to competitive exams.
However, creating a benchmark is easier said than done. It requires a lot of computational infrastructures in terms of GPUs to run those models against the dataset in order to get their efficacy in different metrics. To kick start the process Indian universities can partner with Indian tech companies like Reliance, Tata and Infosys who will soon acquire tens of thousands of NVIDIA GH200 Grace Hopper Superchip.
On his recent visit to India, NVIDIA chief Jensen Huang said that he will partner with universities like IITs as well to create AI infrastructure. “We would like to work with every single university. The first thing we have to do is build AI infrastructure” he said.
Isn’t it a nice thought seeing OpenAI’s models claiming supremacy over an Indian benchmark.
The post The Need for LLM Benchmarks in India appeared first on Analytics India Magazine.
With a slew of announcements recently, OpenAI has been on a roll. A host of features which has enabled ChatGPT to see, hear and speak, were also made public. Now, OpenAI has brought back its Browse with Bing capability, removing the barrier of information, which was previously limited to November 2021.
While this was a welcome move by many, some believe that this could be the end of plugins, making its marketplace a little less reducent, and less attractive. Sam Altman took to X announcing ChatGPT going online, “We are back”, to which a user responded, “You are not just back but you are literally destroying everyone.”
Plugins become obsolete
The latest ‘Browse with Bing’ feature could indeed prove to be the plugin killer. For instance, why would anyone use a third-party plugin to browse the web with ChatGPT when you’ve got the Microsoft hallmark? It does sound safer and more reliable than a third-party plugin. For instance, a user tried the browsing option to extract a list of shoes under $100.
A number of CTOs and venture capitalists reiterated that ‘a whole bunch of plugins and wrappers became instantly obsolete.’
Back in May, OpenAI’s decision to embrace plugins for ChatGPT was met with enthusiasm. The ChatGPT plugin store filled up fast with over 100 pages featuring close to 900 plugins. Regardless, OpenAI’s Plugin Store is a hot mess. It lacks basic organisation and has no categorisation or ranking. In addition, it has been struggling with a host of security issues, compromising user data.
Moreover, amid this flurry of innovation, security researchers raised red flags. Johann Rehberger, a security researcher and red team director at Electronic Arts, has been diligently documenting the vulnerabilities plaguing ChatGPT’s plugins during his spare time. His findings are disconcerting. Rehberger has revealed that these plugins harbour the potential to compromise user data security, enabling malicious actors to pilfer chat histories, harvest personal information, and execute code on users’ devices.
The crux of the issue lies in plugins employing OAuth, a web standard designed for data-sharing across online accounts. Rehberger contends that these plugins fundamentally erode the trustworthiness of ChatGPT. Through the use of a plugin, a malicious website or document could orchestrate prompt injection attacks or introduce malicious payloads, effectively granting unauthorised access to sensitive information and systems.
OpenAI is not turning a blind eye to these concerns. While ChatGPT plugins have been in beta since their launch in March, the company has introduced warning mechanisms cautioning users to exercise prudence in trusting plugins
Nevertheless, the overarching questions regarding data security persist, which OpenAI seems to be getting rid of, or at least reducing the dependence on with this move.
Going online, good/bad for business?
In May, OpenAI had introduced and then disabled the ‘Browse with Bing’ feature after its beta release due to concerns about data leakage. The feature was introduced in beta for ChatGPT Plus subscribers, offering real-time data access as part of the $20 per month subscription. However, it inadvertently started displaying content, including full text from URLs, which raised privacy concerns for content owners.
In a blog post, OpenAI expressed its commitment to finding ways for creators, publishers, and content producers to benefit from their technology, but the identified bug in the browser feature contradicted this goal.
This privacy issue led to OpenAI facing challenges with various companies and countries in the European Union and others like Japan, warning against the use of ChatGPT due to data privacy concerns. Additionally, they had attracted flak from the likes of New York Times for copyright infringement, and there were rumours that the publication was likely to sue OpenAI.
A lot of people raised questions about the peculiar terminology ‘authoritative information’. As OpenAI has reintroduced the feature, ‘authoritative information; could be an attempt at patching these issues. However, browsing seems to be working for some, while others have complained that it doesn’t work for them—-which indicates that the company administering a slower rollout.
More lawsuits
Conclusively, the thing to ponder upon is what this would mean for OpenAI. It could bring back some customers it had conceded to other models with browsing capabilities, but it could fall back into the same pit of copyright lawsuits, not just from authors and artists but from large corporations which have blocked access to their website data.
A host of platforms, from Reddit to X have banned scraping their website for content, to build their own models(in Musk’s case), which could mean trouble for the plugin. However, OpenAI doesn’t need to scrape social media sites because it has inked partnerships with the likes of the Associated Press, a wire agency which gets its news and updates from multiple sources including social media platforms like X.
The post Did OpenAI Just Kill Plugins? appeared first on Analytics India Magazine.
The biggest developer conference is coming up and it’s not from Apple, Microsoft, Meta, or Google, it’s OpenAI Dev Day, the first ever conference from the company. The one-day event on November 6 will have a keynote address and breakout sessions led by the team of OpenAI technical staff, and a lot of announcements are awaited. But what exactly is left to announce?
OpenAI has been on a slew of announcements, most of them focused on developers. GPT-4V(ision), UI for fine-tuning, release of GPT 3.5 Turbo Instruct, and it even made ChatGPT Browse with Bing feature available again. Though Sam Altman has clearly said that there would be no announcements of GPT-5, there is very little left to announce.
on november 6, we’ll have some great stuff to show developers! (no gpt-5 or 4.5 or anything like that, calm down, but still i think people will be very happy…)https://t.co/QH1mpXzoqp
— Sam Altman (@sama) September 6, 2023
What to expect?
“Alright, who’s LLM do I need to RLHF to get an in-person invite to OpenAI dev day,” said Gred Kamaradt on X. It is hard to get an invite to the conference. If you are lucky enough to be approved, you would still need to be $450 to attend it. Regardless, expectations and predictions are on the way and out loud.
Before ruling out GPT-5, according to recent rumours OpenAI seems to be already in the possession of a powerful model named “Arrakis”. The word around the block is that OpenAI now has AGI. Which is still questionable, but Altman even went on to Reddit to make a joke about it.
Moreover, “We’re looking forward to showing our latest work to enable developers to build new things,” Altman said recently. There is a high speculation that there might be a developer competition, where OpenAI might select people to hire for their team. Altman said that this is an amazing time to join. The team is increasingly expanding its office and hiring for various roles such as engineer, data scientists, and developers.
sure 10x engineers are cool but damn those 10,000x engineer/researchers…
— Sam Altman (@sama) September 22, 2023
Apart from that, developers are asking for the cost reduction of GPT-4, including a session-based API, where developers do not have to send all messages, departing from the current stateless API.
What if OpenAI decides to upgrade all of its models to GPT-4 or even more, and make GPT-3.5 open source? That would be the biggest announcement, giving all the open source models a run for their money.
Venture into hardware
OpenAI is reportedly also now getting into the hardware business. WHOOP recently announced that it is partnering with OpenAI to launch WHOOP Coach, the most advanced generative AI feature to ever be released by a wearable. It is possible that the company might integrate more of its generative AI capabilities into various hardware designs, which might be an ideal step for the company.
According to a recent report, Jony Ive, the renowned designer of the iPhone, and OpenAI CEO Sam Altman have been discussing a new AI hardware. It is highly unlikely that the company would venture into the phone market as it would take years for the company to establish a manufacturing facility to compete with players like Apple in the market. On the flip side, it is easier for Apple to develop a GPT-4 like LLM and get into OpenAI’s market.
OpenAI might just want to stick with the software aspect of product development, and integrate its capabilities within other products, just like it did with WHOOP. Since OpenAI is all about keeping it simple and breaking things fast, it might turn to smaller edge use cases like a smart watch or even a smart ring to venture into the field.
A recent leak from Microsoft about an “AI bag” which users would be able to talk with and interact with AI models. This could also be powered by OpenAI’s GPT. But given all the tussles between OpenAI and Microsoft, the Altman led firm might want to start something of its own.
Furthermore, amidst all the announcements of XR headsets from Apple, Meta, and even Microsoft, it might be possible that OpenAI ventures into this field. It has already developed its GPT-4V(ision), which can be touted as a natural first step for the company to get more into computer vision products.
Making sense of all the recent acquisitions, OpenAI might even venture into the gaming market. The recent acquisition of Global Illumination, the 3D world designing company, is a hint that OpenAI is going all-in on a metaverse type simulation with AI bots. This can be a step towards getting into the gaming industry.
no one knows what happens next
— Sam Altman (@sama) August 10, 2023
All in all, though a lot of predictions for Dev Day have already been made, there is a lot that is expected. OpenAI might just announce GPT-5, who knows.
The post What is Left for OpenAI Dev Day appeared first on Analytics India Magazine.