Google’s new tools help users verify the authenticity of images online faster

Google AI about this image hero

AI text-to-image generators have the capability to produce some incredibly realistic images, like the viral image of Pope Francis wearing a white puffer jacket that nearly tricked the whole internet. Therefore, this technology poses a high risk of spreading misinformation — but Google's new tools are here to help.

At Google I/O, the company announced an "About this image" feature in Search that would help users discern whether an image is real by providing them with crucial image information.

Also: This new camera has Content Credentials built-in as a response to AI-generated images

On Wednesday, Google finally unveiled the "About this image" feature, which will provide users with the image's history, how other sites use and describe the image, and most importantly, the image's metadata at the tap of a button.

The metadata of an image is valuable because it contains information where you can find details that the image creator or publishers have added to the image themselves, assuring they get proper credit for their work.

Also, many AI image generators, such as Google's or Adobe's, include in the metadata that the image was generated by AI. Therefore, Google's easy access to the metadata information allows users to potentially access a clear marker informing them the photo was fake.

Also: Google's new AI-powered tool helps users learn English right in Search

To access the "About this image" tool, all users will have to do is click on the three dots that appear next to an image in Google Images results or click "more about this page" on the "About this result" tool in search results.

Google also introduced the "Fact Check Explorer" tool, which gives journalists a platform to easily learn more information about an image or topic.

Using the tool, all a user needs to do is drop an image or URL of what they are trying to fact-check. Then, the tool will see if the image has been featured anywhere in an existing fact-check and present that information.

According to Google, since the feature was released, over 70% of beta users reported that the new image features helped reduce their investigation time.

Also: Mid-career professionals, watch out. You're the most exposed to AI

Per testers' request, Google also unveiled the beta of a FactCheck Search API that users could integrate with their own in-house fact-checking solutions.

Google isn't the only company attempting to curve the misinformation from AI-generated images. The camera manufacturer Leica launched the world's first camera to have Content Credentials built in.

That means that at the point of capture, the photo will automatically include metadata that includes details, such as who captured the image as well as when and how it was captured.

Artificial Intelligence

Tech layoffs are back with a vengeance

Tech layoffs are back with a vengeance Haje Jan Kamps 9 hours

Welcome to Startups Weekly. Sign up here to get it in your inbox every Friday.

For my column this week, I told the story of how an ex-colleague was impersonated by an AI-powered spambot and almost tricked me. It a nutshell: AI is often used for good, but it is increasingly used for nefarious purposes as well. Of course, AI-powered spam is going to get really interesting, really fast: Some of the generative AIs are good enough to pass as humans. So, what happens when every spam message is customized to you, and different from every other spam message? Things are about to get really bad — before they hopefully get better.

Let me take you on a tour of the highlights and lowlights of startup world over the past week.

Layoffs are back

Wooden Jigsaw Puzzle with missing pieces; how to handle layoffs humanely

Image Credits: MirageC (opens in a new window) / Getty Images

Last month, Alex wrote that tech layoffs were pretty much a thing of the past. Shouldn’t have said that, buddy, you jinxed it.

Despite signs of economic recovery and predictions of avoiding a recession, tech companies continue to lay off employees. In October, Nokia announced it was laying off 14,000 employees following a quarter that saw profits drop by 69%, and other major tech companies like Qualcomm, Qualtrics, and LinkedIn also announced significant layoffs. Experts suggest that while the economy is improving, the recovery process is slow, leading many companies to prepare for a longer period of economic sluggishness. Moreover, a shift in investor mindset from growth to efficiency has led to cost-cutting measures, including layoffs. These trends, combined with tighter buying budgets and slower sales cycles, could continue to impact the tech sector into 2024. Ron has the full skinny on TC+ at “What’s behind the fresh round of tech layoffs?”

Product Hunt slashes staff: Product Hunt, a discovery site for startups, apps, and tech tools, has laid off approximately 60% of its team, including roles in design, product, and sales. The cuts were made for “strategic reasons,” Sarah reports.

The stack overfloweth: Stack Overflow, a developer community site owned by Prosus, has announced a 28% reduction in its workforce as part of its drive toward profitability. The company did not disclose the exact number of affected employees. It seems like AI may be the culprit, Ivan writes.

Don’t miss our comprehensive guide: The tech industry has faced a significant blow in 2023, with job losses exceeding 240,000, a 50% increase from the previous year. Major tech giants like Google, Amazon, Microsoft, Yahoo, Meta, and Zoom, along with numerous startups, have announced significant workforce reductions. We have our full guide here.

Transportation terror and triumphs

Tesla-Supercharger-EV

Image Credits: Tesla

The Rebelle Rally 2023, a 2,120-kilometer off-road and navigation competition for women, has become a testing ground for stock manufacturer vehicles, including electric vehicles. Out of the 65 teams that competed in the Rally’s eighth annual event, 10 were electrified vehicles, including four Rivian R1T pickups, marking a significant entry of EVs into this traditionally non-tech event. A Rivian team clinched first place in the 4×4 class, marking the first time an all-electric vehicle topped the podium.

Meanwhile, Tesla released its Q3 earnings report. And it wasn’t super pretty: The report showed a fall in gross margin to 17.9%, down from 25.1% last year. That caused a 44% profit drop (yikes). Tesla’s long-awaited Cybertruck is set to start initial deliveries, and Elon Musk warned that it will take 18 months for the pickup to become profitable.

It was stormy days for driverless taxis, too, as Cruise’s permit to operate as a robotaxi was joinked: The California Public Utilities Commission (CPUC) has suspended Cruise’s permits to operate and charge for its robotaxi service in San Francisco, following a similar move by the DMV. The DMV’s suspension came after Cruise allegedly withheld footage from an investigation into an incident wherein a pedestrian was hit and dragged by one of its autonomous vehicles. Cruise denied the claims. The suspension comes just three months after it granted the company the necessary permits to charge for rides. This led to more pushback against robotaxis in LA.

More from transportation startups:

All aboard the Tesla standard: Toyota and Lexus have announced plans to adopt Tesla’s chargers (NACS) for their electric vehicles starting in 2025. The only major automakers yet to adopt NACS are VW and Stellantis, but with the momentum toward Tesla’s standard, their conversion may be imminent, Harri reports.

Going places: Pebble has revealed a prototype of its flagship product, the Pebble Flow, an all-electric travel trailer designed for digital nomads. The 25-foot trailer, which can sleep up to four people, can be preordered for a refundable $500, with a starting cost of $109,000. An upgraded version, including a dual motor drivetrain, is available for $125,000, Kirsten reports.

Moar Tesla legal troubles: Tesla is under scrutiny from the U.S. Department of Justice again, related to the company’s advertised EV range, personnel decisions, and perks. This comes after an investigation suggested Tesla had been inflating its EV range estimates, Kirsten reports.

Who’s raising, and for what?

SAFE rounds, startups, venture capitalists

Image Credits: Getty Images

I Own My Data (IOMD), a startup founded by former PayPal executive Rohan Mahadevan, is aiming to revolutionize online shopping by eliminating the need for consumers to create new accounts with every purchase, Mary Ann reports. IOMD’s Node platform allows consumers to manage and store all their online interactions, purchases, and profiles on their own devices. Node has emerged from stealth with a $2.75 million seed funding. The startup points out it is not a payments company but an information company, storing users’ private information on their devices for instant transactions.

Navan (formerly known as TripActions), a fintech startup specializing in expense management, has partnered with Citi to provide a jointly branded travel and expense system for Citi Commercial cardholders, Mary Ann reports. The partnership is a huge deal, especially given Citi’s status as the third largest bank in the U.S, with over 25,000 global commercial card programs and 7 million cardholders, all of whom may soon be able to wave goodbye to expense reports. In other news, Darrell did his expenses in VR this week and actually enjoyed it. He is, truly, a strange human.

FFS, that’s not even a rounding error: Black founders in the U.S. raised a mere 0.13% of all capital allocated to startups in Q3, a significant drop from the $1 billion they raised in Q3 2022. The trend has been consistent since 2020 despite efforts, Dominic-Madori reports.

Back once again: Oh goodie, Tucker Carlson, the controversial former Fox News host, plans to launch a media startup called Last Country, following a $15 million investment, Rebecca reports.

That’s a big sack o’ cash, y’all: Global investment firm KKR has announced the final close of its third tech growth fund with approximately $3 billion in capital commitments. The group targets companies with strong long-term growth prospects, typically writing checks ranging from $50 million to $250 million, Connie reports.

Top reads on TechCrunch this week

Creating private AIs: ZenML, an open source AI framework, is helping companies build their own private AI models, reducing dependence on API providers like OpenAI and Anthropic. The Munich-based startup has raised $6.4 million since its inception.

Sod it, let’s build our own: For TC+, Ron took a deep dive to figure out why Monday.com, a company offering a suite of flexible business tools, has developed its own database solution, MondayDB, to meet unique customer needs.

Web Summit drama continues: Paddy Cosgrave, co-founder and CEO of Web Summit, has resigned amid controversy over his comments about Israel and Palestine. Despite his resignation, Cosgrave still owns 80% of the business. The conference organizers have confirmed that Web Summit 2023 in Lisbon and the February 2024 event in Qatar will proceed as planned.

AI Experts Propose ⅓ Investment Rule for Governments Globally

Prominent AI researchers, including Geoffrey Hinton, Yoshua Bengio, Andrew Yao, Daniel Kahneman, Dawn Song, and Yuval Noah Harari, have issued a call for concrete regulations to govern the development and deployment of AI.

This collective of experts, which includes three Turing Award winners, a Nobel laureate, and AI academics have called for immediate action, proposing that major organisations working on AI systems allocate at least one-third of their resources to ensure AI safety and ethical use, a commitment on par with their investment in capability development.

Geoffrey Hinton, widely regarded as one of the most influential AI scientists said: “There are companies planning to train models with 100x more computation than today’s state of the art, within 18 months. No one knows how powerful they will be. And there’s essentially no regulation on what they’ll be able to do with these models.”

In advance of the first international AI Safety Summit in London, these experts have released a concise paper outlining their consensus on how governments should approach the risks associated with AI. It represents the most concrete and comprehensive set of demands from leading AI academics to date as per the official statement.

The paper emphasizes that current AI models are potent and impactful, necessitating democratic oversight to prevent potential risks, including AI-generated misinformation, social injustice, power concentration, cyber warfare, and loss of control.

One of the 20 authors of the paper, computer scientist Yoshua Bengio stresses the urgency of these measures, given the rapid progression of AI technology.

“It’s time to get serious about advanced AI systems. These are not toys. Increasing their capabilities before we understand how to make them safe is utterly reckless. Companies will complain that it’s too hard to satisfy regulations— that “regulation stifles innovation.” That’s ridiculous. There are more regulations on sandwich shops than there are on AI companies,” stated Stuart Russell, professor of computer science at the University of California at Berkeley and a leading active voice in AI.

The post AI Experts Propose ⅓ Investment Rule for Governments Globally appeared first on Analytics India Magazine.

AI Safety Summit 2023 – Understanding Global AI Risks

AI Safety Summit 2023 – Understanding Global AI Risks October 27, 2023 by Ali Azhar

The AI Safety Summit 2023 is set to take place on 1st and 2nd November at the iconic Bletchley Park in the U.K. Some of the world's leading tech companies, AI experts, government officials, and civil society groups are taking part in the summit. The primary agenda at the summit is to highlight the risks of artificial intelligence with a focus on frontier AI and discuss how these risks can be mitigated through internationally coordinated actions.

Frontier models are highly advanced AI foundation models that hold enormous potential to power innovation, economic growth, and scientific progress. Unfortunately, frontier models also possess highly dangerous capabilities. Through misuse or accident, foundation models can cause significant harm and destruction on a global scale.

Understanding the risks posed by frontier AI is a key objective at the AI Safety Summit 2023. The goal is to have the world's leading minds on AI technology convene and discuss methods and strategies for international collaboration for AI safety. Another key objective at the summit is to showcase how the safety development of AI will help enable AI to be used for good globally. This will help promote innovation and growth of AI technology.

Rishi Sunak Warns of AI Dangers

Speaking ahead of the AI Safety Summit 2023, British Prime Minister Rishi Sunak said that governments must realize that they have the power to tackle risks posed by AI. Sunak emphasized the risks posed by AI include the capability to make weapons of mass destruction, spread misinformation and propaganda, and even escape human control. He also announced that Britain will set up an AI safety institute to gain a better understanding of different AI models, and how the risks posed by AI can be mitigated.

Sunak compared the rapid advances in AI technology to the industrial revolution, and warned that if government-level action is not taken, AI technology could make cyberattacks “faster, more effective, and large scale.” He also highlighted that misuse of AI could lead to “erosion of trust information”.

According to Sunak, deepfakes and hyper-realistic bots can be used to manipulate financial markets, spread fake news, and even undermine the criminal justice system. Sunak believes that AI could disrupt the labor market by displacing human workers and suggested a “robot tax” on businesses profiting from the replacement of workers by AI.

The summit is set to host key figures from around the globe including Google DeepMind CEO Demis Hassabis, U.S. Vice President, Kamala Harris, and other leaders from G7 economies. China has been invited to the summit, but there has been no confirmation if their representative will attend the summit.

Oxford AI Experts Disagree with Sunak

(Blue Planet Studio/Shutterstock)

Brent Mittelstadt, Associate Professor and Director of Research at the Oxford Internet Institute, University of Oxford disagrees with Sunak’s view that the UK should not be in a rush to regulate AI due to a lack of proper understanding of the technology. Mittelstadt believes that an incredible range of research has been completed on AI technology is enough to start regulating the industry.

According to Mittelstadt, not taking a proactive approach to regulation will be a mistake, and it will allow the private sector and technology development to dictate when the regulation starts and how they are policied in the future. The government should be in charge of deciding when and how to regulate the industry.

Carissa Veliz, Associate Professor at the Faculty of Philosophy and the Institute for Ethics in AI, University of Oxford is also concerned about Sunak’s remarks on AI technology. Veliz says that tech executives are not the right people to advise governments on how to regulate AI and Sunak’s views sounded similar to those voiced by big tech firms. Several other AI experts at the University of Oxford share similar concerns and are urging the U.K. government to not delay regulating the AI industry.

Related Items

Why Trusting AI is All a Matter of the Right Data at the Right Time

Unraveling the Fiction and Reality of AI’s Evolution: An Interview with Paul Muzio

Navigating the AI Skills Revolution in the Age of GenAI: LinkedIn Report

Related

Kochi-based Nimo Planet Unveils “Apple Vision Pro-like” Spatial Computer That Fits in Your Pocket

Kochi-based Nimo Planet, a company specialising in spatial computers, recently introduced Nimo Spacial OS and Nimo Core. The pre-production version of Nimo is now available to select early users through their website.

Nimo OS and Nimo Core, along with Nimo Glass, to create a portable spatial computing system for users worldwide. These components are designed to provide a personalised, multi-screen workspace experience for the modern workforce.

The generative AI inside the OS is NimoGPT, which the company says is the first step towards creating deep GenAI for spatial computers.

Rohildev, the founder of the company said, “Nimo 1 is highly optimised for Keyboard, Mouse and Trackpad interactions. We felt that for productivity these interfaces are better than hand tracking. But we are working on something that will replace keyboard and mouse in the future.”

Their technology combines a pocket-sized computer and a virtual private display to eliminate the limitations of using various devices like smartphones, tablets, and laptops. This setup ensures better usability and portability, addressing the challenges of mobile computing for the average office worker.

The idea was to optimise resources efficiently, reducing CPU and memory usage while rendering up to six high-fidelity 3D screens. This approach enhances performance, battery life, and minimises heat generation, especially during multitasking.

Nimo Planet’s was founded in 2018 by Rohildev Nattukallingal with the aim of creating a pocket-sized computer that enhances productivity with impressive multi-screen workspaces. This unique computer and its accompanying OS represent their response to this challenge, addressing the problem in a singular way.

Nimo 1 glasses offer similar features to Apple Vision Pro at half the price. It lacks built-in speakers or Apple’s M2 chips but uses Qualcomm’s Snapdragon XR1 processor, suitable for spatial computing. While not as powerful as Apple Vision Pro, Nimo offers a similar user interface with wireless mouse functionality for application management.

The company stands out by allowing users to switch between six virtual screens by turning their head. Widgets like a calendar can be placed at the top, and settings like Wi-Fi are easily accessible below, resembling Apple’s promises. Nimo is sleeker than Apple Vision Pro, although its arms remain bulky, supporting touch inputs for Bluetooth keyboard and mouse convenience. The company has raised $150 thousand dollars from 7 investors.

The post Kochi-based Nimo Planet Unveils “Apple Vision Pro-like” Spatial Computer That Fits in Your Pocket appeared first on Analytics India Magazine.

KDnuggets News, October 27: 5 Free Books to Master Data Science • 7 Steps to Mastering LLMs

Featured Posts

  • 5 Free Books to Master Data Science
  • 7 Steps to Mastering Large Language Models (LLMs)

From Our Partners

  • Secure Your Seat for Canada’s #1 Data AI Conference from Corp Agency
  • Accelerate Your Machine Learning Journey with Uplimit’s Metaflow Mastery Course from Uplimit
  • In-demand SAS Certifications for Your Success from SAS
  • Data Science Methods Drive Business Success from Northwestern University
  • Semantic Layer: The Backbone of AI-powered Data Experiences from Cube

This Week's Posts

  • Comparing Natural Language Processing Techniques: RNNs, Transformers, BERT
  • RAG vs Finetuning: Which Is the Best Tool to Boost Your LLM Application?
  • Why SQL is THE Language to Learn for Data Science
  • Best Practices for Building ETLs for ML
  • Are Kaggle Competitions Useful for Real World Problems?
  • Exploring Data Mesh: A Paradigm Shift in Data Architecture
  • Fasten Your Seatbelt: Falcon 180B is Here!
  • Rust Burn Library for Deep Learning
  • How To Fine-Tune ChatGPT 3.5 Turbo
  • Mastering the Art of Data Cleaning in Python
  • The Generative AI Bubble Will Burst Soon
  • Unlocking Reliable Generations through Chain-of-Verification: A Leap in Prompt Engineering
  • DALL·E 3 is Here with ChatGPT Integration
  • ChatGPT vs. BARD
  • Some Kick Ass Prompt Engineering Techniques to Boost our LLM Models
  • 40% of Labour Force Will be Affected by AI in 3 Years
  • 7 Best Cloud Database Platforms
  • More Tips for Successfully Navigating Beginner Data Science Job Interviews
  • Gradient Descent: The Mountain Trekker's Guide to Optimization with Mathematics
  • Top Companies in India to Consider for Employment
  • Introduction to Databases with SQL: Free Harvard Course
  • A Brief History of the Neural Networks
  • 3 Ways to Make Money with ChatGPT and AI
  • 10 Basic Statistical Concepts in Plain English
  • Greening AI: 7 Strategies to Make Applications More Sustainable
  • The Growth Behind LLM-based Autonomous Agents
  • Graph of Thoughts: A New Paradigm for Elaborate Problem-Solving in Large Language Models
  • 7 Platforms for Getting High Paying Data Science Jobs
  • How Predictive Analytics is Revolutionizing Decision-Making in Tech
  • Beyond Skynet: Crafting the Next Frontier in AI Evolution
  • Python f-Strings Magic: 5 Game-Changing Tricks Every Coder Needs to Know!

More On This Topic

  • 7 Steps to Mastering Large Language Models (LLMs)
  • KDnuggets News, June 22: Primary Supervised Learning Algorithms Used in…
  • 5 Free Books to Master Data Science
  • KDnuggets News, October 5: Top Free Git GUI Clients for Beginners • A Day…
  • 5 Free Books to Master Machine Learning
  • 5 Free Books to Help You Master Python

Google expands bug bounty program to include rewards for AI attack scenarios

Abstract 3D render of wavy thin wires and particles forming a hand

In cybersecurity, threats change quickly. Add rapidly evolving generative AI tech to the mix, and security concerns evolve by the minute. Google is one of the biggest players in artificial intelligence technology, and it recognizes the need to adapt to this threat.

Google is expanding its existing Vulnerability Rewards Program (VRP) to include vulnerabilities specific to generative AI and considering the unique challenges that generative AI poses, like biases, model manipulations, data misinterpretations, and other adversarial attacks.

Also: Cybersecurity 101: Everything on how to protect your privacy and stay safe online

The VRP is a bug bounty program that rewards external security researchers for testing and reporting software vulnerabilities in Google's products and services. Now, this will include generative AI products. Some of Google's most popular generative AI products include Bard, Lens, and other AI integrations in Search, Gmail, Docs, and more.

As generative AI becomes more integrated into different Google tools and programs, the potential risks increase and Google already has internal Trust and Safety teams working to foresee these risks. With the expansion of the bug bounty program to include generative AI, Google is trying to encourage research in AI safety to ensure responsible AI becomes the norm.

Also: Beyond passwords: 4 key security steps you're probably forgetting

Google also offered more information on its reward criteria for reporting bugs in AI products so users can easily determine what is in scope and what isn't.

External security researchers are tasked with finding these vulnerabilities in exchange for financial gain, which in turn gives Google, the company behind the bug bounty program, the opportunity to fix these threats before bad actors exploit them. This ensures a more secure product for users.

Aside from encompassing generative AI into its VRP, Google introduced the Secure AI Framework to support creating responsible and safe AI applications. It also announced it's collaborating with the Open Source Security Foundation to ensure the integrity of AI supply chains.

Also: WormGPT: What to know about ChatGPT's malicious cousin

Users who want to join Google's bug bounty program can submit a bug or security vulnerability directly to the company. In 2022, Google issued over $12 million in rewards to security researchers as part of its bug bounty program.

Security

Uni3D: Exploring Unified 3D Representation at Scale

Scaling up representations of text and visuals has been a major focus of research in recent years. Developments and research conducted in the recent past have led to numerous revolutions in language learning and vision. However, despite the popularity of scaling text and visual representations, the scaling of representations for 3D scenes and objects has not been sufficiently discussed.

Today, we will discuss Uni3D, a 3D foundation model that aims to explore unified 3D representations. The Uni3D framework employs a 2D-initialized ViT framework, pretrained end-to-end, to align image-text features with their corresponding 3D point cloud features.

The Uni3D framework uses pretext tasks and a simple architecture to leverage the abundance of pretrained 2D models and image-text-aligned models as initializations and targets, respectively. This approach unleashes the full potential of 2D models and strategies to scale them to the 3D world.

In this article, we will delve deeper into 3D computer vision and the Uni3D framework, exploring the essential concepts and the architecture of the model. So, let’s begin.

Uni3D and 3D Representation Learning : An Introduction

In the past few years, computer vision has emerged as one of the most heavily invested domains in the AI industry. Following significant advancements in 2D computer vision frameworks, developers have shifted their focus to 3D computer vision. This field, particularly 3D representation learning, merges aspects of computer graphics, machine learning, computer vision, and mathematics to automate the processing and understanding of 3D geometry. The rapid development of 3D sensors like LiDAR, along with their widespread applications in the AR/VR industry, has resulted in 3D representation learning gaining increased attention. Its potential applications continue to grow daily.

Although existing frameworks have shown remarkable progress in 3D model architecture, task-oriented modeling, and learning objectives, most explore 3D architecture on a relatively small scale with limited data, parameters, and task scenarios. The challenge of learning scalable 3D representations, which can then be applied to real-time applications in diverse environments, remains largely unexplored.

Moving along, in the past few years, scaling large language models that are pre-trained has helped in revolutionizing the natural language processing domain, and recent works have indicated a translation in the progress to 2D from language using data and model scaling which makes way for developers to try & reattempt this success to learn a 3D representation that can be scaled & be transferred to applications in real world.

Uni3D is a scalable and unified pretraining 3D framework developed with the aim to learn large-scale 3D representations that tests its limits at the scale of over a billion parameters, over 10 million images paired with over 70 million texts, and over a million 3D shapes. The figure below compares the zero-shot accuracy against parameters in the Uni3D framework. The Uni3D framework successfully scales 3D representations from 6 million to over a billion.

The Uni3D framework consists of a 2D ViT or Vision Transformer as the 3D encoder that is then pre-trained end-to-end to align the image-text aligned features with the 3D point cloud features. The Uni3D framework makes use of pretext tasks and simple architecture to leverage the abundance of pretrained 2D models and image text aligned models as initialization and targets respectively, thus unleashing the full potential of 2D models, and strategies to scale them to the 3D world. The flexibility & scalability of the Uni3D framework is measured in terms of

  1. Scaling the model from 6M to over a billion parameters.
  2. 2D initialization to text supervised from visual self-supervised learning.
  3. Text-image target model scaling from 150 million to over a billion parameters.

Under the flexible and unified framework offered by Uni3D, developers observe a coherent boost in the performance when it comes to scaling each component. The large-scale 3D representation learning also benefits immensely from the sharable 2D and scale-up strategies.

As it can be seen in the figure below, the Uni3D framework displays a boost in the performance when compared to prior art in few-shot and zero-shot settings. It is worth noting that the Uni3D framework returns a zero-shot classification accuracy score of over 88% on ModelNet which is at par with the performance of several state of the art supervision methods.

Furthermore, the Uni3D framework also delivers top notch accuracy & performance when performing other representative 3D tasks like part segmentation, and open world understanding. The Uni3D framework aims to bridge the gap between 2D vision and 3D vision by scaling 3D foundational models with a unified yet simple pre-training approach to learn more robust 3D representations across a wide array of tasks, that might ultimately help in the convergence of 2D and 3D vision across a wide array of modalities.

Uni3D : Related Work

The Uni3D framework draws inspiration, and learns from the developments made by previous 3D representation learning, and Foundational models especially under different modalities.

3D Representation Learning

The 3D representation learning method uses cloud points for 3D understanding of the object, and this field has been explored by developers a lot in the recent past, and it has been observed that these cloud points can be pre-trained under self-supervision using specific 3D pretext tasks including mask point modeling, self-reconstruction, and contrastive learning.

It is worth noting that these methods work with limited data, and they often do not investigate multimodal representations to 3D from 2D or NLP. However, the recent success of the CLIP framework that returns high efficiency in learning visual concepts from raw text using the contrastive learning method, and further seeks to learn 3D representations by aligning image, text, and cloud point features using the same contrastive learning method.

Foundation Models

Developers have exhaustively been working on designing foundation models to scale up and unify multimodal representations. For example, in the NLP domain, developers have been working on frameworks that can scale up pre-trained language models, and it is slowly revolutionizing the NLP industry. Furthermore, advancements can be observed in the 2D vision domain as well because developers are working on frameworks that use data & model scaling techniques to help in the progress of language to 2D models, although such frameworks are difficult to replicate for 3D models because of the limited availability of 3D data, and the challenges encountered when unifying & scaling up the 3D frameworks.

By learning from the above two work domains, developers have created the Uni3D framework, the first 3D foundation model with over a billion parameters that makes use of a unified ViT or Vision Transformer architecture that allows developers to scale the Uni3D model using unified 3D or NLP strategies for scaling up the models. Developers hope that this method will allow the Uni3D framework to bridge the gap that currently separates 2D and 3D vision along with facilitating multimodal convergence.

Uni3D : Method and Architecture

The above image demonstrates the generic overview of the Uni3D framework, a scalable and unified pre-training 3D framework for large-scale 3D representation learning. Developers make use of over 70 million texts, and 10 million images paired with over a million 3D shapes to scale the Uni3D framework to over a billion parameters. The Uni3D framework uses a 2D ViT or Vision Transformer as a 3D encoder that is then trained end-to-end to align the text-image data with the 3D cloud point features, allowing the Uni3D framework to deliver the desired efficiency & accuracy across a wide array of benchmarks. Let us now have a detailed look at the working of the Uni3D framework.

Scaling the Uni3D Framework

Prior studies on cloud point representation learning have traditionally focused heavily on designing particular model architectures that deliver better performance across a wide range of applications, and work on a limited amount of data thanks to small-scale datasets. However, recent studies have tried exploring the possibility of using scalable pre-training in 3D but there were no major outcomes thanks to the availability of limited 3D data. To solve the scalability problem of 3D frameworks, the Uni3D framework leverages the power of a vanilla transformer structure that almost mirrors a Vision Transformer, and can solve the scaling problems by using unified 2D or NLP scaling-up strategies to scale the model size.

Reliance Jio Unveils JioSpaceFiber: India’s First Satellite Broadband

Reliance Jio unveiled India’s first satellite-based giga-fibre service, JioSpaceFiber, at the India Mobile Congress held in New Delhi on Friday. This groundbreaking service is designed to provide high-speed broadband connectivity to previously hard-to-reach regions throughout India.

Reliance Jio’s chairman, Akash Ambani, stated: “Jio has enabled millions of homes and businesses in India to experience broadband internet for the first time. With JioSpaceFiber, we expand our reach to cover the millions yet to be connected.”

The new service is poised to empower individuals and businesses across India, granting them access to essential online government services, education, healthcare, and entertainment with gigabit internet speeds. Ambani demonstrated the new satellite broadband at the Congress in the presence of Modi.

As one of India’s leading broadband service providers, Jio already serves over 450 million consumers with both fixed-line and wireless options. With JioSpaceFiber, the company aims to foster digital inclusivity by further extending its suite of broadband services, which already includes JioFiber and JioAirFiber.

Jio’s offerings ensure that individuals and businesses gain unparalleled access to reliable, low-latency, high-speed internet and entertainment services, regardless of their geographic location. The satellite network underpinning JioSpaceFiber will expand mobile backhaul capacity, thereby enhancing the reach and scalability of Jio True5G, even in the most remote areas of India.

Jio has partnered with SES, granting it access to the world’s latest medium earth orbit (MEO) satellite technology. SES’s O3b and new O3b mPOWER satellites enable Jio to provide scalable and affordable broadband connectivity across India. This unique satellite technology allows Jio to offer truly unique gigabit, fibre-like services from space.

To demonstrate its capabilities and extensive reach, JioSpaceFiber has already connected four of the most remote locations in India: Gir in Gujarat, Korba in Chhattisgarh, Nabarangpur in Odisha, and ONGC-Jorhat in Assam. This achievement highlights JioSpaceFiber’s potential to bring high-speed broadband services to underserved areas in the country.

John-Paul Hemingway, chief strategy officer at SES, noted, “Together with Jio, we are honoured to support the government of India’s Digital India initiative with a unique solution that aims at delivering multiple gigabits per second of throughput to any location in India.”

SES’s fibre-like services from space are already deployed in parts of India and are poised to drive digital transformation even in the country’s most rural areas.

The post Reliance Jio Unveils JioSpaceFiber: India’s First Satellite Broadband appeared first on Analytics India Magazine.

7 Steps to Mastering Data Wrangling with Pandas and Python

7 Steps to Mastering Data Wrangling with Pandas and Python
Image generated with DALLE 3

Are you an aspiring data analyst? If so, learning data wrangling with pandas, a powerful data analysis library, is an essential skill to add to your toolbox.

Almost all data science courses and bootcamps cover pandas in their curriculum. Though pandas is easy to learn, its idiomatic usage and getting the hang of common functions and method calls requires practice.

This guide breaks down learning pandas—into 7 easy steps—starting with what you probably are familiar with and gradually exploring the powerful functionalities of pandas. From prerequisites—through various data wrangling tasks—to building a dashboard, here’s a comprehensive learning path.

Step 1: Python and SQL Fundamentals

If you’re looking to break into data analytics or data science, you first need to pick up some basic programming skills. We recommend starting with Python or R, but we’ll focus on Python in this guide.

Learn Python and Web Scraping

To refresh your Python skills you can use one of the following resources:

  • Free Python for Everybody video lecture — freeCodeCamp
  • Python course — Kaggle

Python is easy to learn and start building. You can focus on the following topics:

  • Python basics: Familiarize yourself with Python syntax, data types, control structures, built-in data structures, and basic object-oriented programming (OOP) concepts.
  • Web scraping fundamentals: Learn the basics of web scraping, including HTML structure, HTTP requests, and parsing HTML content. Familiarize yourself with libraries like BeautifulSoup and requests for web scraping tasks.
  • Connecting to databases: Learn how to connect Python to a database system using libraries like SQLAlchemy or psycopg2. Understand how to execute SQL queries from Python and retrieve data from databases.

While not mandatory, using Jupyter Notebooks for Python and web scraping exercises can provide an interactive environment for learning and experimenting.

Learn SQL

SQL is an essential tool for data analysis; But how will learning SQL help you learn pandas?

Well, once you know the logic behind writing SQL queries, it's very easy to transpose those concepts to perform analogous operations on a pandas dataframe.

Learn the basics of SQL (Structured Query Language), including how to create, modify, and query relational databases. Understand SQL commands such as SELECT, INSERT, UPDATE, DELETE, and JOIN.

To learn and refresh your SQL skills you can use the following resources:

  • SQL tutorial — Khan Academy
  • Intro to SQL — Kaggle Learn
  • Advanced SQL — Kaggle Learn

By mastering the skills outlined in this step, you will have a solid foundation in Python programming, SQL querying, and web scraping. These skills serve as the building blocks for more advanced data science and analytics techniques.

Step 2: Loading Data From Various Sources

First, set up your working environment. Install pandas (and its required dependencies like NumPy). Follow best practices like using virtual environments to manage project-level installations.

As mentioned, pandas is a powerful library for data analysis in Python. Before you start working with pandas, however, you should familiarize yourself with the basic data structures: pandas DataFrame and series.

To analyze data, you should first load it from its source into a pandas dataframe. Learning to ingest data from various sources such as CSV files, excel spreadsheets, relational databases, and more is important. Here’s an overview:

  • Reading data from CSV files: Learn how to use the pd.read_csv() function to read data from Comma-Separated Values (CSV) files and load it into a DataFrame. Understand the parameters you can use to customize the import process, such as specifying the file path, delimiter, encoding, and more.
  • Importing data from Excel files: Explore the pd.read_excel() function, which allows you to import data from Microsoft Excel files (.xlsx) and store it in a DataFrame. Understand how to handle multiple sheets and customize the import process.
  • Loading data from JSON files: Learn to use the pd.read_json() function to import data from JSON (JavaScript Object Notation) files and create a DataFrame. Understand how to handle different JSON formats and nested data.
  • Reading data from Parquet files: Understand the pd.read_parquet() function, which enables you to import data from Parquet files, a columnar storage file format. Learn how Parquet files offer advantages for big data processing and analytics.
  • Importing data from relational database tables: Learn about the pd.read_sql() function, which allows you to query data from relational databases and load it into a DataFrame. Understand how to establish a connection to a database, execute SQL queries, and fetch data directly into pandas.

We’ve now learned how to load the dataset into a pandas dataframe. What’s next?

Step 3: Selecting Rows and Columns, Filtering DataFrames

Next, you should learn how to select specific rows and columns from a pandas DataFrame, as well as how to filter the data based on specific criteria. Learning these techniques is essential for data manipulation and extracting relevant information from your datasets.

Indexing and Slicing DataFrames

Understand how to select specific rows and columns based on labels or integer positions. You should learn to slice and index into DataFrames using methods like .loc[], .iloc[], and boolean indexing.

  • .loc[]: This method is used for label-based indexing, allowing you to select rows and columns by their labels.
  • .iloc[]: This method is used for integer-based indexing, enabling you to select rows and columns by their integer positions.
  • Boolean indexing: This technique involves using boolean expressions to filter data based on specific conditions.

Selecting columns by name is a common operation. So learn how to access and retrieve specific columns using their column names. Practice using single column selection and selecting multiple columns at once.

Filtering DataFrames

You should be familiar with the following when filtering dataframes:

  • Filtering with conditions: Understand how to filter data based on specific conditions using boolean expressions. Learn to use comparison operators (>, <, ==, etc.) to create filters that extract rows that meet certain criteria.
  • Combining filters: Learn how to combine multiple filters using logical operators like '&' (and), '|' (or), and '~' (not). This will allow you to create more complex filtering conditions.
  • Using isin(): Learn to use the isin() method to filter data based on whether values are present in a specified list. This is useful for extracting rows where a certain column's values match any of the provided items.

By working on the concepts outlined in this step, you’ll gain the ability to efficiently select and filter data from pandas dataframes, enabling you to extract the most relevant information.

A Quick Note on Resources

For steps 3 to 6, you can learn and practice using the following resources:

  • 10 minutes to pandas — pandas user guide
  • Pandas and Python for Data Analysis by Example — freeCodeCamp
  • Intro to pandas — Kaggle Learn

Step 4: Exploring and Cleaning the Dataset

So far, you know how to load data into pandas dataframes, select columns, and filter dataframes. In this step, you will learn how to explore and clean your dataset using pandas.

Exploring the data helps you understand its structure, identify potential issues, and gain insights before further analysis. Cleaning the data involves handling missing values, dealing with duplicates, and ensuring data consistency:

  • Data inspection: Learn how to use methods like head(), tail(), info(), describe(), and the shape attribute to get an overview of your dataset. These provide information about the first/last rows, data types, summary statistics, and the dimensions of the dataframe.
  • Handling missing data: Understand the importance of dealing with missing values in your dataset. Learn how to identify missing data using methods like isna() and isnull(), and handle it using dropna(), fillna(), or imputation methods.
  • Dealing with duplicates: Learn how to detect and remove duplicate rows using methods like duplicated() and drop_duplicates(). Duplicates can distort analysis results and should be addressed to ensure data accuracy.
  • Cleaning string columns: Learn to use the .str accessor and string methods to perform string cleaning tasks like removing whitespaces, extracting and replacing substrings, splitting and joining strings, and more.
  • Data type conversion: Understand how to convert data types using methods like astype(). Converting data to the appropriate types ensures that your data is represented accurately and optimizes memory usage.

In addition, you can explore your dataset using simple visualizations and perform data quality checks.

Data Exploration and Data Quality Checks

Use visualizations and statistical analysis to gain insights into your data. Learn how to create basic plots with pandas and other libraries like Matplotlib or Seaborn to visualize distributions, relationships, and patterns in your data.

Perform data quality checks to ensure data integrity. This may involve verifying that values fall within expected ranges, identifying outliers, or checking for consistency across related columns.

You now know how to explore and clean your dataset, leading to more accurate and reliable analysis results. Proper data exploration and cleaning are super important or any data science project, as they lay the foundation for successful data analysis and modeling.

Step 5: Transformations, GroupBy, and Aggregations

By now, you are comfortable working with pandas DataFrames and can perform basic operations like selecting rows and columns, filtering, and handling missing data.

You’ll often want to summarize data based on different criteria. To do so, you should learn how to perform data transformations, use the GroupBy functionality, and apply various aggregation methods on your dataset. This can further be broken down as follows:

  • Data transformations: Learn how to modify your data using techniques such as adding or renaming columns, dropping unnecessary columns, and converting data between different formats or units.
  • Apply functions: Understand how to use the apply() method to apply custom functions to your dataframe, allowing you to transform data in a more flexible and customized way.
  • Reshaping data: Explore additional dataframe methods like melt() and stack(), which allow you to reshape data and make it suitable for specific analysis needs.
  • GroupBy functionality: The groupby() method lets you group your data based on specific column values. This allows you to perform aggregations and analyze data on a per-group basis.
  • Aggregate functions: Learn about common aggregation functions like sum, mean, count, min, and max. These functions are used with groupby() to summarize data and calculate descriptive statistics for each group.

The techniques outlined in this step will help you transform, group, and aggregate your data effectively.

Step 6: Joins and Pivot Tables

Next, you can level up by learning how to perform data joins and create pivot tables using pandas. Joins allow you to combine information from multiple dataframes based on common columns, while pivot tables help you summarize and analyze data in a tabular format. Here’s what you should know:

  • Merging DataFrames: Understand different types of joins, such as inner join, outer join, left join, and right join. Learn how to use the merge() function to combine dataframes based on shared columns.
  • Concatenation: Learn how to concatenate dataframes vertically or horizontally using the concat() function. This is useful when combining dataframes with similar structures.
  • Index manipulation: Understand how to set, reset, and rename indexes in dataframes. Proper index manipulation is essential for performing joins and creating pivot tables effectively.
  • Creating pivot tables: The pivot_table() method allows you to transform your data into a summarized and cross-tabulated format. Learn how to specify the desired aggregation functions and group your data based on specific column values.

Optionally, you can explore how to create multi-level pivot tables, where you can analyze data using multiple columns as index levels. With enough practice, you’ll know how to combine data from multiple dataframes using joins and create informative pivot tables.

Step 7: Build a Data Dashboard

Now that you’ve mastered the basics of data wrangling with pandas, it's time to put your skills to test by building a data dashboard.

Building interactive dashboards will help you hone both your data analysis and visualization skills. For this step, you need to be familiar with data visualization in Python. Data Visualization — Kaggle Learn is a comprehensive introduction.

When you’re looking for opportunities in data, you need to have a portfolio of projects—and you need to go beyond data analysis in Jupyter notebooks. Yes, you can learn and use Tableau. But you can build on the Python foundation and start building dashboards using the Python library Streamlit.

Streamlit helps you build interactive dashboards—without having to worry about writing hundreds of lines of HTML and CSS.

If you’re looking for inspiration or a resource to learn Streamlit, you can check out this free course: Build 12 Data Science Apps with Python and Streamlit for projects across stock prices, sports, and bioinformatics data. Pick a real-world dataset, analyze it, and build a data dashboard to showcase the results of your analysis.

Next Steps

With a solid foundation in Python, SQL, and pandas you can start applying and interviewing for data analyst roles.

We’ve already included building a data dashboard to bring it all together: from data collection to dashboard and insights. So be sure to build a portfolio of projects. When doing so, go beyond the generic and include projects that you really enjoy working on. If you are into reading or music (which most of us are), try to analyze your Goodreads and Spotify data, build out a dashboard, and improve it. Keep grinding!

Bala Priya C is a developer and technical writer from India. She likes working at the intersection of math, programming, data science, and content creation. Her areas of interest and expertise include DevOps, data science, and natural language processing. She enjoys reading, writing, coding, and coffee! Currently, she's working on learning and sharing her knowledge with the developer community by authoring tutorials, how-to guides, opinion pieces, and more.

More On This Topic

  • KDnuggets™ News 22:n05, Feb 2: 7 Steps to Mastering Machine Learning…
  • 7 Steps to Mastering Python for Data Science
  • Revamping Data Visualization: Mastering Time-Based Resampling in Pandas
  • 7 Steps to Mastering Machine Learning with Python in 2022
  • 7 Steps to Mastering Data Cleaning and Preprocessing Techniques
  • Mastering the Data Universe: Key Steps to a Thriving Data Science Career