Internet Of Things (IOT):  Application In Hazardous Locations

Panorama aerial view in the cityscape skyline with smart service

Introduction to Internet of Things (IOT):

Internet of Things (IoT) represents the fourth-generation technology that facilitates the connection and transformation of products into smart, intelligent and communicative entities. IoT has already established its footprint in various business verticals such as medical, heath care, automobile, and industrial applications. IoT empowers the collection, analysis, and transmission of information across various networks, encompassing both server and edge devices. This information can then undergo further processing and distribution to multiple inter-connected devices through cloud connectivity.

IoT Application in Oil & Gas Industry:

IoT is used in the Oil and Gas Industry for two basic reasons: First – low power design, a fundamental requirement for intrinsically safe products, Second – two-way wireless communication. These two advantages are a boon for the products used in Oil and Gas industries. The only challenge is for the product design to meet the hazardous location certification.

An intrinsic safe certification is mandatory for any device placed in hazardous locations. The certification code depends on the type of protection, zone, and the region where the product shall be installed.

In the North American and Canadian markets, the area classification is done in three classes:

Class I: Location where flammable gases and vapors are present.

Class II: Location where combustible dust is present.

Class III: Location where flying is present.

The hazardous area is further divided into two divisions, based upon the probability that a dangerous fuel to air mixture will occur or not.

Dvision-1: Location is where there is a high probability (by underwriting standards) that an explosive concentration of gas or vapor is present during normal operation of the plant.

Division-2: Location is where there is a very low probability that the flammable material is present in the explosive concentration during normal operation of the plant; so, an explosive concentration is expected only in case of a failure of the plant containment system.

The GROUP is also one of the meaningful nomenclatures of the hazardous area terms.

The four gas groups were created so that electrical equipment intended to be used in hazardous (classified) locations could be rated for families of gases and vapors and tested with a designated worst-case gas/air mixture to cover the entire group.

The temperature class definitions are used to designate the maximum operating temperatures on the surface of the equipment, which should not exceed the ignition temperature of the surrounding atmosphere.

Areas classified per NEC Article 505 are divided into three zones based on the probability of an ignitable concentration being present, rather than into two divisions as per NEC article 501. Areas that would be classified division 1 are further divided into zone 0 and zone 1. A zone 0 area is more likely to contain an ignitable atmosphere than zone 1 area. Division 2 and zone 2 areas are essentially equivalent.

Zone-0: Presence of ignitable concentration of combustible gases and vapors continuously, or for long periods of time.

Zone-1: Intermittent hazard may be present.

Zone-2: Hazard will be present under abnormal conditions.

IoT-based products can be designed for various applications, a few of them are listed below:

  1. Temperature Sensor
  2. Pressure Monitoring
  3. Gas Monitoring
  4. Flow Monitoring

A typical block diagram of the IoT application is shown below:

Figure 1: IOT Block Diagram

An IoT product might consist of a battery as a power source or can be powered externally from either 9V ~ 36V DC supply available in the process control applications or 110/230Vac input.

The microcontroller can be selected based on the applications, power consumption, and the peripheral requirements. The microcontroller converts the analog signal to digital and based on the configuration can send the data on wired/wireless to the remote station. Analog signal conditioning stands as a pivotal component of the product, bridging the connection between the sensor and facilitating the conversion of analog signals for compatibility with the microcontroller. The Bluetooth interface suggested in the example is due to its wide acceptance and low power consumption. The wireless interface depends on the end-application of the product.

Electronics Design Consideration

The electronics design of an IoT product for a hazardous location is very complex and needs a careful selection of the architecture and base components as compared to the IoT developed for commercial applications. In case the IoT is for a hazardous location, the product must be intrinsically safe and should not cause an explosion under fault conditions. The product architecture should be designed considering various mechanical, and electronics requirements as defined in the IEC 60079 standards, certification requirements and the functional specifications.

Power Source: This is one of the main elements in an IoT-based product. Battery selection should meet the overall power budget of the product, followed by the battery lifetime. In case of intrinsic safety, special consideration is required for where the battery in charged. IEC 60079-11 clause 7.4 provide details for the type of battery and its construction details. Separation distance from the battery and electrical interface should be done as per Table-5 of IEC 60079-11. If the battery is used in the compartment, sufficient ventilation must be provided to ensure that no dangerous gas accumulation occurs during discharge or inactivity periods. In scenarios where IoT operates on DC power sources such as 9~36Vdc (nominal 24Vdc), the selection of power supply barrier protection becomes a critical consideration, particularly when catering to intrinsic safety norms. This necessitates a thorough analysis of the product’s prerequisites and the mandatory certifications. Adding to the complexity is the existence of IoT devices functioning on 230Vac, demands intrinsic safe calculations and certifications aligned with Um = 250V.

Microcontroller: Its central processing unit of the IoT product. The architecture of the microcontroller, power, and clock frequency processing must be carefully selected for a particular application. The Analog to Digital Conversion (ADC) part of the microcontroller should be selected based on the required accuracy, update rate, and resolution. Microcontroller should have enough sleep modes so that the power is optimally utilized for IoT applications and should have sufficient memory/peripheral interface to meet the product specifications.

Analog Signal Conditioning: The front-end block should meet the intrinsic safe requirements as per the IEC 60079 standards and should also protect the product from EMI-EMC testing. Barrier circuit should provide enough isolation for meeting the spark-gap ignition requirements and impedance requirement of the transducer. Also, along with the safety requirements, the designer should ensure that extracted sensor signal is not degraded from the excessive noise present in outside environment. All the sensors used for collecting data from the process parameters to the signal conditioning block must be certified for the particular zone.

Wireless Communications: There are various wireless options available for sending data from the IoT product to the sensor such as (6LOWPAN, ZigBEE, ZWave, Bluetooth, Wi-Fi, Wireless HART). Selection of a particular wireless interface requires knowledge of end application, RF-power, antenna, and protocol. Selection of the interface for a particular IoT application should be done keeping these basic things in mind:

  1. The amount of data to be shared to the server.
  2. RF power.
  3. Power consumed for each bit of data transferred.
  4. Update rate of the data and distance of communication.
  5. Security of data.

In case of intrinsic safe applications, it’s important to note that the use of certified modules does not directly confer suitability for deployment in hazardous locations. The product must undergo fresh testing within an intrinsic safe lab to assess both quantifiable and non-quantifiable ffaults, along with spark testing. or the countable and non-countable faults and spark testing. The RF power transmitted from the devices should be limited as per Table-1x of IEC 60079-0.

Conclusion

When building IoT solutions for hazardous locations, special conditions relating to creepage and clearance, encapsulation, and separation distance must be carefully considered. Also, when battery and RF signals are used, it’s expected the designer should be aware of the applicable standards and limitation of these standards for such products.

With more than 25 years of experience in designing mission-critical and consumer-grade embedded hardware designs, eInfochips is well poised to make products which are smaller, faster, reliable, efficient, intelligent and economical. We have worked on developing complex embedded control systems for avionics and industrial solutions. At the same time, we have also developed portable and power efficient systems for wearables, medical devices, home automation and surveillance solutions.

eInfochips, as an Arrow company, has a strong ecosystem of manufacturing partners who can help right from electronic prototype design, manufacturing, production, and certification. eInfochips works closely with the contract manufacturers to make sure that the designs are optimized for testing (DFT) and manufacturing (DFM) to reduce design alterations on production transfer. To know more about this contact us.

einfochips can help product-based companies to develop intrinsically safe products and get the product certified by lab for various certifications like ATEX/IECEx/CSA.

References

  1. “IEC 60079–0” in Explosive Atmospheres – Part 0: General Requirements, Geneva. Switzerland.
  2. “IEC 60079–11 Part 11” in Equipment Protection by Intrinsic Safety “i”, Geneva, Switzerland.
  3. “UL 2225” in Standard for Safety; Cables and Cable Fittings for Use In Hazardous (Classified) Locations, Northbrook. IL:UL.
  4. “CSA C22.1–18 Rule 18–092” in Canadian Electrical Code Part I, Toronto, Canada:CSA Group.
  5. “NFPA 70” in National Electrical Code, Quincy, MA: National Fire Protection Association.
  6. “CAN/CSA C22.2 No.60079–0” in Explosive Atmospheres – Part 0: General Requirements, Toronto, Canada:CSA Group.

About Authors:

Kartik Gandhi, currently serving in the capacity of Senior Director of Engineering, possesses a distinguished career spanning over two decades, marked by a profound expertise in fields including Business Analysis, Presales, and Embedded Systems. Throughout his professional journey, Mr. kartik has demonstrated his proficiency across diverse platforms, notably Qualcomm and NXP, and has contributed his talents to several esteemed product-based organizations.

Dr. Suraj Pardeshi has more than 20 years of experience in Research & Development, Product Design & Development, and testing. He has worked on various IoT-enabled platforms for Industrial applications. He has more than 15 publications in various National and International journals. He holds two Indian patents, Gold Medalist and Ph.D (Electrical) from M.S University, Vadodara.

China’s Ernie AI is just as good as — or better than — ChatGPT, says its creators

Colorful chatbot

There's a new name in the top spot of the ever-changing AI arms race. Well, at least if you believe the creator behind it.

Speaking at an event in Beijing, Baidu founder Robin Li, whose company is often called the Chinese equivalent of Google, said that the latest version of his company's AI chatbot matches ChatGPT in overall capability.

Also: The best AI chatbots: ChatGPT and alternatives

In real time, Li asked Ernie Bot 4.0 several questions, presented math problems, and asked it to write a martial arts novel. He also asked Ernie to create a few posters and videos on the spot. If there was any doubt about his challenge to OpenAI, Li made sure his beliefs were clear when he said "Ernie is not inferior in any respect to GPT-4."

And there may just be evidence to prove what he's saying is true.

Just a few months ago, Li went even further about the previous version of the chatbot, Ernie 3.5, saying that it actually surpassed ChatGPT in several areas, including those in the Chinese language. To back up that claim, Baidu quoted a report from a Chinese national newspaper that ran a test using two benchmarks, AGIEval and C-Eval, to measure AI performance. One of those tests used a standard admission test someone might take to get into college. Ernie 3.5 scored better than GPT-4.

Also: What is GPT-4? Here's everything you need to know

Why does it matter if Ernie is better than ChatGPT? Like what is happening in the US, AI is quickly becoming a part of life in China, being inserted into products like online search, file-sharing, work collaboration, and maps.

It's clear that AI is going to have potentially massive applications, and as AI makes its way into the workplace, it's important for companies to be on the leading edge. After all, more success for individual companies means a country's whole economy benefits.

It's clear that AI isn't going anywhere, and that everyone is going to have to adapt to its increasing usage. What remains unclear is who's going to do it best.

Artificial Intelligence

DSC Weekly 17 October 2023

Announcements

  • AI is crucial for today’s businesses, providing analysis that empowers and facilitates better decision-making. It’s even been shown that organizations that leverage insights provided by AI experience higher ROI, sales growth, more efficient operations, faster time to market, and increased customer engagement and satisfaction. However, AI teams have a wide variety of enabling technologies and platforms to choose from, ranging from generative AI to machine learning and everything in between. Attend the Leveraging AI in the Enterprise summit for advice to help you make better buying decisions, get the most of your AI investments, and help AI teams increase their value.
  • Cloud computing offers benefits like scalability, cost efficiency and storage capacity, but can introduce many security threats due to its large attack surface and overall complexity. Register for the Building a Secure Cloud Environment summit to hear leading experts discuss security strategies for multi-cloud architecture, best practices to tackle the challenges of securing the multi cloud, and leading tools and platforms to ward off common threats and keep cloud environments safe.

Top Stories

  • 12 Generative AI Trends to Watch Out for
    October 12, 2023
    by Roger Brown
    The advent of generative AI is empowering everyone alike – organizations, small businesses, individuals, students, and medical professionals, to name a few. The last couple of years have been revolutionary for artificial intelligence innovation and transformation. How will 2024 shape up for AI, AI tools, and related professionals?
  • AI’s Kryptonite: Data Quality
    October 14, 2023
    by Bill Schmarzo
    The ability of Generative AI (GenAI) tools to deliver accurate and reliable outputs entirely depends on the accuracy and reliability of the data used to train the Large Language Models (LLMs) that power the GenAI tool. Unfortunately, the Law of GIGO – Garbage In, Garbage Out – threatens the widespread adoption of GenAI.
  • Internet Of Things (IOT): Application In Hazardous Locations
    October 17, 2023
    by Kartik Gandhi, Dr. Suraj Pardeshi
    Internet of Things (IoT) represents the fourth-generation technology that facilitates the connection and transformation of products into smart, intelligent and communicative entities. IoT has already established its footprint in various business verticals such as medical, heath care, automobile, and industrial applications.
Education_DSC_160x600-2

In-Depth

  • Uncharted digital landscapes and the quest for timeless identity
    October 17, 2023
    by Michael Peres
    In a recent podcast episode, Lex Freedman and Mark Zuckerberg convened in the Metaverse, where the digital realm intertwines with reality. Their astonishingly realistic interaction, while highlighting technological advancements, also prompted deeper contemplations. As the line between digital recreations and reality becomes increasingly blurred, it beckons questions about the definitions of identity and consciousness.
  • The digital evolution in aviation: how big data and analytics are transforming the industry
    October 17, 2023
    by Jeremy Bowen
    Long before passengers sit back, relax, and enjoy their flight, data has played a critical role in getting them to their seats. It has been a cornerstone of the aviation industry since the early days of air travel. Indeed, from the early 20th century, data was collected through manual processes such as pilots logging information.
  • Explainable Artificial Intelligence (XAI) for AI & ML Engineers
    October 16, 2023
    by Shanthababu Pandian
    Hello AI&ML Engineers, as you all know, Artificial Intelligence (AI) and Machine Learning Engineering are the fastest growing fields, and almost all industries are adopting them to enhance and expedite their business decisions and needs; for the same, they are working on various aspects and preparing the data for the AIML platform with the help of SMEs.
  • Significance of AI in the development of software products
    October 13, 2023
    by Ryan Williamson
    Artificial Intelligence (AI) is emerging as a formidable force, revolutionizing how we conceive, create, and deliver software solutions. As technology advances at an unprecedented pace, the role of AI in this domain has become increasingly significant. It’s no longer just a buzzword; it’s a fundamental tool that promises to reshape the entire software development process.
  • Future of AI and data science – How to secure a bright career
    October 13, 2023
    by Aileen Scott
    Companies, more often, pay attention to automation and innovation over proficiency and productivity. However, firms can maintain a balance between both due to the extensive usage of AI and data science programs. Here are the stats that show the impact of AI and data science in diverse sectors.
  • DSC Weekly 10 October 2023
    October 10, 2023
    by Scott Thompson
    Read more of the top articles from the Data Science Central community.

Uncharted digital landscapes and the quest for timeless identity

Mark Zuckerberg and Lex Fridman on Lex Fridmans recent podcast

In a recent podcast episode, Lex Freedman and Mark Zuckerberg convened in the Metaverse, where the digital realm intertwines with reality. Their astonishingly realistic interaction, while highlighting technological advancements, also prompted deeper contemplations. As the line between digital recreations and reality becomes increasingly blurred, it beckons questions about the definitions of identity and consciousness and the very essence of being. Navigating through history reveals varying interpretations of self, while philosophical debates on free will present further intricacies. Meanwhile, the looming horizon of digital immortality introduces a plethora of ethical challenges.

Tracing the Self: Evolving interpretations across time

Throughout history, cultures and civilizations have grappled with the fundamental question: What defines us as individuals? Is it our thoughts, emotions, memories, or perhaps something more tangible?

Ancient cultures, in their quest to understand the self, often attributed the essence of humanity to the heart. For instance, in ancient Egypt, the heart was revered as the focal point of intelligence, emotions, and even the soul’s destiny in the afterlife. This reverence was so profound that during mummification, the heart was preserved, while the brain, deemed irrelevant, was discarded.

However, as understanding evolved, so did perceptions of identity. Modern science shifted the seat of consciousness from the heart to the brain, a complex organ with billions of neurons intricately networked. Today, the brain stands as a testament to our thoughts, memories, and personality, often described as the bastion of our identity.

But this assertion raises further quandaries. When parts of the brain, such as the frontal lobe, suffer damage, does one’s identity undergo a transformation? Many who have witnessed loved ones endure brain injuries would argue that the essence of the individual certainly remains.

Thus, in attempting to locate and define the self, both historical and contemporary interpretations pose more questions than they answer, paving the way for even more profound contemplations on the nature of free will and determinism.

The age-old journey: Unraveling the essence of self

Throughout history, cultures and civilizations have grappled with the fundamental question: What defines us as individuals? Is it our thoughts, emotions, memories, or perhaps something more tangible?

Ancient cultures, in their quest to understand the self, often attributed the essence of humanity to the heart. For instance, in ancient Egypt, the heart was revered as the focal point of intelligence, emotions, and even the soul’s destiny in the afterlife. This reverence was so profound that during mummification, the heart was preserved, while the brain, deemed irrelevant, was discarded.

However, as understanding evolved, so did perceptions of identity. Modern science shifted the seat of consciousness from the heart to the brain, a complex organ with billions of neurons intricately networked. Today, the brain stands as a testament to our thoughts, memories, and personality, often described as the bastion of our identity.

However, this understanding invites deeper questions. If, for instance, damage occurs to the frontal lobe of the brain, does it alter our identity? While specific abilities may change due to such injuries, many argue that the foundational essence of an individual remains intact.

Thus, in attempting to locate and define the self, both historical and contemporary interpretations pose more questions than they answer, paving the way for even more profound contemplations on the nature of free will and determinism.

Facing Mortality

Death is an inevitable facet of the human experience, shaping our perspectives, aspirations, and fears. Evolutionarily crafted for survival, our desire to leave legacies has driven us from writing memoirs to monumental achievements. Now, digital recreation offers a form of virtual immortality once only dreamt of, letting us persist in digital realms even as the natural world moves on without us.

The Dance of Determinism: Free Will in the Crosshairs

Historically, the debate between free will and determinism has swayed philosophical thought. Ancient Stoics believed everything unfolded according to a divine plan, while Epicureans saw randomness as a potential influencer of human actions, granting a sense of agency.

The modern era, brimming with advances in neuroscience, presents more complex understandings. Sam Harris, a neuroscientist and philosopher, offers a compelling viewpoint in his book “Free Will.” He argues, “Free will is an illusion. Our wills are simply not of our own making. Thoughts and intentions emerge from background causes of which we are unaware and over which we exert no conscious control.” Harris contends that our choices, rather than emanating from some intrinsic self-driven intent, are determined by prior events, chiefly our genetics and the environment we’re subjected to. He challenges the romanticized notion of free will by stating that our actions, decisions, and even our desires are an outcome of myriad factors beyond our conscious control.

Uncharted digital landscapes and the quest for timeless identity

Sam Harris, author of Free Will

Harris’s perspective raises a significant question for our digital age: If our behaviors are largely predetermined, then what truly separates our organic selves from a digital avatar constructed to mimic our decisions and actions? This contemplation not only unsettles our sense of self but also thrusts us into the ethical quagmire of digital consciousness reconstruction.

Navigating the Digital Afterlife: Immortality’s Ethical Labyrinth

The progression of technology now allows for the creation of detailed digital replicas of individuals. Such advancements bring forth complex ethical considerations:

Consent and Posthumous Rights

Is it ethically sound to recreate an individual’s digital presence without their explicit agreement? Who should control these virtual identities once the original is no longer present?

Reality vs. Representation

Despite their detail and sophistication, digital avatars remain crafted representations. The challenge lies in ensuring they authentically reflect the individual they represent.

Emotions at Play

While some might find solace in interacting with digital iterations of loved ones, for others, it could be a distressing experience. The emotional landscape becomes even more intricate.

Questioning the End

When digital extensions become possible, it redefines our perceptions of life’s transient nature. This raises questions about the intrinsic value of fleeting moments.

Identity in the Age of Data

With an increasing focus on data and analytics, behaviors and decisions can be viewed as mere outcomes of genetic and environmental interplays. This perspective compels a reevaluation of the concept of ‘self’ and its placement in a digital realm.

In this evolving landscape, society must grapple with these challenges, balancing the allure of digital immortality with the nuances and ethics it introduces.

V. Life’s Impermanence: The Heart of Its Beauty

Across cultures and histories, one of the most profound ways we have understood the value of moments is through their ephemerality. The same way sunsets captivate with their fleeting beauty, the relationships we cherish gain depth because of their inherent temporality.

This sentiment rings true especially with loved ones. The time spent with our parents, for instance, becomes all the more precious because we’re acutely aware that life will one day take them away from us. It’s this knowledge of life’s impermanence that infuses such moments with deeper meaning, urging us to hold on, to savor, and to cherish.

Throughout our evolutionary history, humans have developed mechanisms to deal with the concept of mortality. We’ve either avoided confronting it directly or have found more rudimentary ways to cope, such as creating lasting legacies or believing in afterlives. Yet, as technology presents us with the tantalizing possibility of a digital continuation, we’re compelled to ask: How will this reshape our age-old strategies of coping? Will these extensions strip life’s fleeting moments of their inherent value? If the horizon of mortality shifts, how will it redefine the essence of our shared experiences?

In the grand scheme of human existence, a central element of our consciousness has been grappling with mortality. By potentially offering an alternative, technology not only challenges our current perceptions but also beckons us to reassess the evolutionary constructs we’ve long held.

Other potential titles:

  • The Uncanny Digital Frontier: New Realities of Identity
  • Between Pixels and Perceptions: The Future of Our Digital Selves
  • From Heartbeats to Algorithms: The Evolution of Identity
  • Into the Digital Abyss: The New Landscape of Self and Soul.


Michael Peres
(Mikey Peres) is a software engineer, writer, and author. Peres founded multiple startups in the tech industry and writes extensively on topics relating to technology, leadership, entrepreneurship, and scientific research.

– https://linkedin.com/in/mikeyperes

– https://facebook.com/mikeyperes

– https://twitter.com/mikeyperes

Your AI experiments will fail if you don’t focus on this special ingredient

datas-gettyimages-1300497464

Data is often called the new oil because of its value to modern business. In the case of ExxonMobil, senior IT executive Andrew Curry says data isn't just the new oil, but the fuel for a whole range of AI-led initiatives across the company.

As manager of the oil and gas giant's Central Data Office, Curry is responsible for executing enterprise-wide data principles — and the rapid growth of interest in emerging technologies, such as AI and machine learning (ML), during the past 12 months hasn't escaped his attention.

Also: If AI is the future of your business, should the CIO be the one in control?

"Everybody's talking about AI and ML, but you can't be successful in that area without having high-quality data," he says, speaking with ZDNET in an interview at the recent Snowflake Summit 2023 in Las Vegas. "The more opportunities that we're seeing in AI, the more that's putting pressure on us to ask, 'Do we have the quality data to be successful in this area?'"

Curry has refined his answer to this question while working for ExxonMobil since 1999. After filling a plethora of roles, and helping the company establish its Central Data Office, he's learned that one element is crucial to success in emerging technology: data strategy.

Businesses won't be able to make the most of AI and ML unless they first ensure they have a strong, secure data strategy, which is something Curry has established at ExxonMobil, both during his earlier roles and since moving to his current position as the firm's data chief in May 2023.

"We understand the quality of our data," he says. "We know which data is ready for advanced AI and ML capabilities. We also know, and this is equally important, which data isn't quite there as well in terms of quality."

Also: AI gold rush makes basic data security hygiene critical

As manager of the Central Data Office, Curry is establishing a long-term data strategy that defines the technologies, processes, skillsets, and rules to help the organization manage its information assets successfully in an age of AI.

As part of its data strategy, ExxonMobil has a technology ecosystem that uses Snowflake's cloud-based platform. Curry says Snowflake has given the business a consolidated data foundation for the first time.

More generally, his team helps ensure any data initiatives are right for the business. A big part of that effort focuses on the areas where ExxonMobil can differentiate itself through data — and the policies that will allow people to exploit emerging technology in a safe and secure manner.

"You need to know where you can leverage this technology and know where you're not ready yet," he says. "We're really positioning ourselves in the right areas."

Also: Generative AI is everything, everywhere, all at once

To the end, Curry says his team is evaluating large language models (LLMs) and generative AI capabilities as part of its long-term data strategy.

Right now, there's not much to talk about in terms of public developments — once again, the focus is on ensuring the firm's data quality is high before plowing into projects.

"I'll say there's nothing in production yet for us," he says. "But we continue to position that technology and to understand where the business opportunities are, and we're making sure the data is ready for that work."

Like so many other executives at large enterprises, Curry is taking a cautious approach to the use of LLMs and generative AIs as part of his company's data strategy.

Also: Open source is actually the cradle of artificial intelligence. Here's why

"Generative AI and LLMs is an area where we've put a lockdown on it," he says. "At ExxonMobil, you can't just go out and use this technology."

Other research suggests a cautious approach is far from unusual. In fact, many firms aren't even ready to explore emerging technology.

While 98% of executives polled in a recent survey by Workday believe there are potentially big benefits from deploying AI and ML, almost half (49%) report their organization is unprepared to exploit these rewards due to a lack of tools, skills, and knowledge.

For Curry, these are significant hurdles that must be overcome before a blue-chip company like ExxonMobil starts experimenting with generative AI and LLMs in a production environment.

Also: AI aims to predict and fix developer coding errors before disaster strikes

"I think there's a lot of issues about securing your data and thinking about what becomes public information when you leverage some of these public tools, for example," he says.

"So, we're taking a cautious approach for broad use within the company, but at same time, there are active plans along the way in a more controlled environment."

Samantha Searle, director analyst at Gartner, agrees that most blue-chip businesses are taking a careful approach to generative AI.

Like Curry, she says getting your data strategy right is an essential first step for organizations before they start pushing data into LLMs.

"Absolutely, because people aren't going to like it if you are recommending the wrong things to them," she said to ZDNET in a video interview. "That's going to have the adverse effect in terms of customer retention."

Also: How trusted generative AI can improve the connected customer experience

Searle also points to the importance of data quality. Right now, digital leaders should be concentrating on how their information assets are collected, stored, and exploited.

"It's very important to make sure that what the AI models predict is accurate. We still need humans to verify the results are right — we know with generative AI that these tools can hallucinate," she says. "So, you still need verification steps. And making sure you've got optimum data quality is a key step to mitigating these risks."

Back at ExxonMobil, Curry says it's important to recognize that, while his business is cautious about generative AI, it has a long history of utilizing other areas of ML and AI.

He expects the company to use emerging technologies in several key areas going forward.

First, finance and trading: ExxonMobil has a significant cash flow, and AI and ML could help refine the firm's processes.

Also: Generative AI tops Gartner's top 25 emerging technologies for 2023

"Are we exposed in certain areas? Do we have too much cash in certain areas? Should we be investing in other areas? We think we can make improvements on those margins with our cash if we can leverage our ML capabilities."

Second, supply chain: "Being able to understand when things are being impacted is very important, and being able to react in a timely manner means a lot to the business, and so that is a key area that we're investing in."

Finally, Curry points to the potential use of emerging technology at the subsurface level to help interpret trends in seismic surveys.

"That's an area we're working to progress," he says. "It's probably one area that's going to have a longer lead time for us as we work to understand it."

Some Kick Ass Prompt Engineering Techniques to Boost our LLM Models

Some Kick Ass Prompt Engineering Techniques to Boost our LLM Models
Image created with DALL-E3

Artificial Intelligence has been a complete revolution in the tech world.

Its ability to mimic human intelligence and perform tasks that were once considered solely human domains still amazes most of us.

However, no matter how good these late AI leap forwards have been, there’s always room for improvement.

And this is precisely where prompt engineering kicks in!

Enter this field that can significantly enhance the productivity of AI models.

Let’s discover it all together!

The Essence of Prompt Engineering

Prompt engineering is a fast-growing domain within AI that focuses on improving the efficiency and effectiveness of language models. It’s all about crafting perfect prompts to guide AI models to produce our desired outputs.

Think of it as learning how to give better instructions to someone to ensure they understand and execute a task correctly.

Why Prompt Engineering Matters

  • Enhanced Productivity: By using high-quality prompts, AI models can generate more accurate and relevant responses. This means less time spent on corrections and more time leveraging AI’s capabilities.
  • Cost Efficiency: Training AI models is resource-intensive. Prompt engineering can reduce the need for retraining by optimizing model performance through better prompts.
  • Versatility: A well-crafted prompt can make AI models more versatile, allowing them to tackle a broader range of tasks and challenges.

Before diving into the most advanced techniques, let’s recall two of the most useful (and basic) prompt engineering techniques.

A Glimpse into Basic Prompt Engineering Methods

Sequential Thinking with “Let’s think step by step”

Today it is well-known that LLM models’ accuracy is significantly improved when adding the word sequence “Let’s think step by step”.

Why… you might ask?

Well, this is because we are forcing the model to break down any task into multiple steps, thus making sure the model has enough time to process each of them.

For instance, I could challenge GPT3.5 with the following prompt:

If John has 5 pears, then eats 2, buys 5 more, then gives 3 to his friend, how many pears does he have?

The model will give me an answer right away. However, if I add the final “Let’s think step by step”, I am forcing the model to generate a thinking process with multiple steps.

Few-Shot Prompting

While the Zero-shot prompting refers to asking the model to perform a task without providing any context or previous knowledge, the few-shot prompting technique implies that we present the LLM with a few examples of our desired output along with some specific question.

For example, if we want to come up with a model that defines any term using a poetic tone, it might be quite hard to explain. Right?

However, we could use the following few-shot prompts to steer the model in the direction we want.

Your task is to answer in a consistent style aligned with the following style.

<user>: Teach me about resilience.

<system>: Resilience is like a tree that bends with the wind but never breaks.

It is the ability to bounce back from adversity and keep moving forward.

<user>: Your input here.

If you have not tried it out yet, you can go challenge GPT.

However, as I am pretty sure most of you already know these basic techniques, I will try to challenge you with some advanced techniques.

Advanced Prompt Engineering techniques

1. Chain of Thought (CoT) Prompting

Introduced by Google in 2022, this method involves instructing the model to undergo several reasoning stages before delivering the ultimate response.

Sounds familiar right? If so, you are totally right.

It is like merging both Sequential Thinking and Few-Shot Prompting.

How?

Essentially, CoT prompting directs the LLM to process information sequentially. This means we exemplify how to solve a first problem with multiple steps reasoning and then send to the model our real task, expecting it to emulate a comparable chain of thought when responding to the actual query we want it to solve.

Some Kick Ass Prompt Engineering Techniques to Boost our LLM Models
Image taken from the paper.

So, following the very first example, we could provide a prompt with the first problem and a sequential explanation of how to solve it. Then, we can send our input and see how magic happens!

<user> If John has 5 pears, then eats 2, buys 5 more, then gives 3 to his friend, how many pears does he have? Let’s think step by step.

<agent> Let’s go step by step. John starts with 5 pears. He eats 2 pears, so he has 5–2 = 3 pears left. Then, he buys 5 more pears. So, he now has 3 + 5 = 8 pears. Finally, he gives 3 pears to his friend. Therefore, he has 8–3 = 5 pears left. So, after all these steps, John has 5 pears.

<user> Your input here.

However, most of you must be thinking… Do I have to think of a sequential way to solve any problem any time I want to ask something on ChatGPT?

Well… you are not the first one! And this leads us to…

2. Automatic Chain-of-Thought (Auto-CoT)

In 2022, Zhang and colleagues introduced a method to avoid this manual process. There are two main reasons to avoid any manual task:

  • It can be boring.
  • It can yield bad results — for instance, when our mental process is wrong.

They suggested using LLMs combined with the "Let's think step by step" prompt to sequentially produce reasoning chains for each demonstration.

This means asking ChatGPT how to solve any problem sequentially and then using this very same example to train it how to solve any other problem.

3. Self-Consistency

Self-consistency is another interesting prompting technique that aims to improve the chain of thought prompting for more complex reasoning problems.

So… what’s the main difference?

The main idea under Self-Consistency is being aware that we can train the model with a wrong example. Just imagine I solve the previous problem with a wrong mental process:

<User> If John has 5 pears, then eats 2, buys 5 more, then gives 3 to his friend, how many pears does he have? Let’s think step by step.

<Agent> Start with 5 pears. John eats 2 pears. Then, he gives 3 pears to his friend. These actions can be combined: 2 (eaten) + 3 (given) = 5 pears in total affected. Now, subtract the total pears affected from the initial 5 pears: 5 (initial) — 5 (affected) = 0 pears left.
Then, any other task I send to the model will be wrong.

This is why Self-Consistency involves sampling from various reasoning paths, each of them containing a chain of thought, and then letting the LLM choose the best and most consistent path to solve the problem.

Some Kick Ass Prompt Engineering Techniques to Boost our LLM Models
Image taken from the paper

In this case, and following the very first example again, we can show the model different ways of solving the problem.

<User> If John has 5 pears, then eats 2, buys 5 more, then gives 3 to his friend, how many pears does he have?

<Agent> Start with 5 pears. John eats 2 pears, leaving him with 5–2 = 3 pears. He buys 5 more pears, which brings the total to 3 + 5 = 8 pears. Finally, he gives 3 pears to his friend, so he has 8–3 = 5 pears left.

<User> If John has 5 pears, then eats 2, buys 5 more, then gives 3 to his friend, how many pears does he have?

<Agent> Start with 5 pears. He then buys 5 more pears. John eats 2 pears now. These actions can be combined: 2 (eaten) + 5 (bought) = 7 pears in total. Subtract the pear that Jon has eaten from the total amount of pears 7 (total amount) — 2 (eaten) = 5 pears left.

<User> Your input here.

And here comes the last technique.

4. General Knowledge Prompting

A common practice of prompt engineering is augmenting a query with additional knowledge before sending the final API call to GPT-3 or GPT-4.

According to Jiacheng Liu and Co, we can always add some knowledge to any request so the LLM knows better about the question.

Some Kick Ass Prompt Engineering Techniques to Boost our LLM Models
Image taken from the paper.

So for instance, when asking ChatGPT if part of golf is trying to get a higher point total than others, it will validate us. But, the main goal of golf is quite the opposite. This is why we can add some previous knowledge telling it “The player with the lower score wins”.

Some Kick Ass Prompt Engineering Techniques to Boost our LLM Models

So.. what’s the funny part if we are telling the model exactly the answer?

In this case, this technique is used to improve the way LLM interacts with us.

So rather than pulling supplementary context from an outside database, the paper's authors recommend having the LLM produce its own knowledge. This self-generated knowledge is then integrated into the prompt to bolster commonsense reasoning and give better outputs.

So this is how LLMs can be improved without increasing its training dataset!

Concluding Thoughts

Prompt engineering has emerged as a pivotal technique in enhancing the capabilities of LLM. By iterating and improving prompts, we can communicate in a more direct manner to AI models and thus obtain more accurate and contextually relevant outputs, saving both time and resources.

For tech enthusiasts, data scientists, and content creators alike, understanding and mastering prompt engineering can be a valuable asset in harnessing the full potential of AI.

By combining carefully designed input prompts with these more advanced techniques, having the skill set of prompt engineering will undoubtedly give you an edge in the coming years.

Josep Ferrer is an analytics engineer from Barcelona. He graduated in physics engineering and is currently working in the Data Science field applied to human mobility. He is a part-time content creator focused on data science and technology. You can contact him on LinkedIn, Twitter or Medium.

More On This Topic

  • Kick Ass Midjourney Prompts with Poe
  • RAG vs Finetuning: Which Is the Best Tool to Boost Your LLM Application?
  • Web LLM: Bring LLM Chatbots to the Browser
  • What did COVID do to all our models?
  • The Art of Prompt Engineering: Decoding ChatGPT
  • Mastering Generative AI and Prompt Engineering: A Free eBook

8 ways Google just upgraded accessibility in its apps and Pixel devices

Guided Frames

Even though the latest devices on the market are meant to make people's everyday lives easier, not everyone can take advantage of the features due to accessibility limitations. Today, Google unveils eight new features to help people with disabilities make the most of their devices.

Also: The best Android phones you can buy

The first feature helps connect users with disabled-owned businesses by displaying a new identity attribute for the disability community that will display next to the business listing in Google Maps and Search.

Ultimately, this feature will help potential customers connect with and support the disability community while also strengthening community ties between people.

Other Google identity attributes include Asian-owned, Black-owned, Latino-owned, LGBTQ+-owned, veteran-owned, and women-owned.

Also coming to Maps are more accessible walking routes. Users can now opt to get walking routes that are wheelchair-accessible and stair-free.

In addition, Google is expanding where wheelchair-accessible information is displayed. The information will be included in business and place pages on Maps for Android Auto, and Google's open-source Android Automative operating system.

As a result, the wheelchair icon will display next to search results, making it easier to find destinations with appropriate accommodations.

Also: 5 quick tips to strengthen your Android phone security today

Next, Google is upgrading its Search with Live View feature to include screen-reading capabilities.

Now, when a user is using Search in Live View with screen reading on, they will get auditory feedback on the places around them, including helpful information, such as the place's name and how far away it is.

Google's Assistant Routines is also getting an accessibility upgrade with more personalization options. Users will now be able to select their own Routines shortcut style, customize it with their own images, and select different sizing.

"Research has shown that this personalization can be particularly helpful for people with cognitive differences and disabilities and hope it will bring the helpfulness of Assistant Routines to even more people," said Google in the release.

Also: How to add reading mode to your Android devices

Earlier this year, Google introduced a feature on Google Chrome that detects typos and displays suggestions based on what Google thinks you meant. The feature is now expanding to Chrome on Android and iOS to help users who are prone to typos get to the content they need faster.

Google Pixel phones are getting their own wave of updates, too. The first update is a new Magnifier app that uses Pixel's advanced camera technology to act as a physical magnifying glass that can zoom in.

Google says the app was designed in collaboration with partners at the Royal National Institute of Blind People and the National Federation of the Blind to help the low-vision community see details up close.

The Magnifier app is available on the Google Play Store for Pixel 5 and up.

In early October, with the release of the Google Pixel 8, the company unveiled the newest version of Guided Frame, which makes it easier for people who are blind or have low vision to take selfies.

The newest update allows Guided Frame to be used for more than just your front camera, but also rear-facing camera photos on Pixel 6 phones and up.

Innovation

AudioSep : Separate Anything You Describe

LASS or Language-queried Audio Source Separation is the new paradigm for CASA or Computational Auditory Scene Analysis that aims to separate a target sound from a given mixture of audio using a natural language query that provides the natural yet scalable interface for digital audio tasks & applications. Although the LASS frameworks have advanced significantly in the past few years in terms of achieving desired performance on specific audio sources like musical instruments, they are unable to separate the target audio in the open domain.

AudioSep, is a foundational model that aims to resolve the current limitations of LASS frameworks by enabling target audio separation using natural language queries. The developers of the AudioSep framework have trained the model extensively on a wide variety of large-scale multimodal datasets, and have evaluated the performance of the framework on a wide array of audio tasks including musical instrument separation, audio event separation, and enhancing the speech amongst many others. The initial performance of AudioSep satisfies the benchmarks as it demonstrates impressive zero-shot learning capabilities and delivers strong audio separation performance.

In this article, we will be taking a deeper dive into the working of the AudioSep framework as we will evaluate the architecture of the model, the datasets used for training & evaluation, and the essential concepts involved in the working of the AudioSep model. So let’s begin with a basic introduction to the CASA framework.

CASA, USS, QSS, LASS Frameworks : The Foundation for AudioSep

The CASA or the Computational Auditory Scene Analysis framework is a framework used by developers to design machine listening systems that have the ability to perceive complex sound environments in a way similar to the way humans perceive sound using their auditory systems. Sound separation, with a special focus on target sound separation, is a fundamental area of research within the CASA framework, and it aims to solve the “cocktail party problem” or separating real-world audio recordings from individual audio source recordings or files. The importance of sound separation can be attributed mainly to its widespread applications including music source separation, audio source separation, speech enhancement, target sound identification, and a lot more.

Most of the work on sound separation done in the past revolves mainly around the separation of one or more audio sources like music separation or speech separation. A new model going by the name of USS or Universal Sound Separation aims to separate arbitrary sounds in real world audio recordings. However, it is a challenging & restrictive task to separate every sound source from an audio mixture primarily because of the wide array of different sound sources existing in the world which is the major reason why the USS method is not feasible for real-world applications working in real-time.

A feasible alternative to the USS method is the QSS or the Query-based Sound Separation method that aims to separate an individual or target sound source from the audio mixture based on a particular set of queries. Thanks to this, the QSS framework allows developers & users to extract the desired sources of audio from the mixture based on their requirements that makes the QSS method a more practical solution for digital real-world applications like multimedia content editing or audio editing.

Furthermore, developers have recently proposed an extension of the QSS framework, the LASS framework or the Language-queried Audio Source Separation framework that aims to separate arbitrary sources of sound from an audio mixture by making use of the natural language descriptions of the target audio source. As the LASS framework allows users to extract the target audio sources using a set of natural language instructions, it might become a powerful tool with widespread applications in digital audio applications. When compared against traditional audio-queried or vision-queried methods, using natural language instructions for audio separation offers a greater degree of advantage as it adds flexibility, and makes the acquisition of query information much more easier & convenient. Furthermore, when compared with label query-based audio separation frameworks that make use of a predefined set of instructions or queries, the LASS framework does not limit the number of input queries, and has the flexibility to be generalized to open domain seamlessly.

Originally, the LASS framework relies on supervised learning in which the model is trained on a set of labeled audio-text paired data. However, the main issue with this approach is the limited availability of annotated & labeled audio-text data. In order to reduce the reliability of the LASS framework on annotated audio-text labeled data, the models are trained using the multimodal supervision learning approach. The primary aim behind using a multimodal supervision approach is to use multimodal contrastive pre-training models like the CLIP or Contrastive Language Image Pre Training model as the query encoder for the framework. Since the CLIP framework has the ability to align text embeddings with other modalities like audio or vision, it allows developers to train the LASS models using data-rich modalities, and allows the interference with the textual data in a zero-shot setting. The current LASS frameworks however make use of small-scale datasets for training, and applications of the LASS framework across hundreds of potential domains are yet to be explored.

To resolve the current limitations faced by the LASS frameworks, developers have introduced AudioSep, a foundational model that aims to separate sound from an audio mixture using natural language descriptions. The current focus for AudioSep is to develop a pre-trained sound separation model that leverages existing large-scale multimodal datasets to enable the generalization of LASS models in open-domain applications. To summarize, the AudioSep model is : “A foundational model for universal sound separation in open domain using natural language queries or descriptions trained on large-scale audio & multimodal datasets”.

AudioSep : Key Components & Architecture

The architecture of the AudioSep framework comprises two key components: a text encoder, and a separation model.

The Text Encoder

The AudioSep framework uses a text encoder of the CLIP or Contrastive Language Image Pre Training model or the CLAP or Contrastive Language Audio Pre Training model to extract text embeddings within a natural language query. The input text query consists of a sequence of “N” tokens that is then processed by the text encoder to extract the text embeddings for the given input language query. The text encoder makes use of a stack of transformer blocks to encode the input text tokens, and the output representations are aggregated after they are passed through the transformer layers that results in the development of a D-dimensional vector representation with fixed length where D corresponds to the dimensions of CLAP or the CLIP models while the text encoder is frozen during the training period.

The CLIP model is pre-trained on a large-scale dataset of image-text paired data using contrastive learning which is the primary reason why its text encoder learns mapping textual descriptions on the semantic space that is also shared by the visual representations. The advantage the AudioSep gains by using CLIP’s text encoder is that it can now scale up or train the LASS model from unlabeled audio-visual data using the visual embeddings as an alternative, thus enabling the training of LASS models without the requirement of annotated or labeled audio-text data.

The CLAP model works similar to the CLIP model and makes use of contrastive learning objective as it uses a text & an audio encoder to connect audio & language, thus bringing text & audio descriptions on an audio-text latent space joined together.

Separation Model

The AudioSep framework makes use of a frequency-domain ResUNet model that is fed a mixture of audio clips as the separation backbone for the framework. The framework works by first applying an STFT or a Short-Time Fourier Transform on the waveform to extract a complex spectrogram, the magnitude spectrogram, and the Phase of X. The model then follows the same setting and constructs an encoder-decoder network to process the magnitude spectrogram.

The ResUNet encoder-decoder network consists of 6 residual blocks, 6 decoder blocks, and 4 bottleneck blocks. The spectrogram in each encoder block uses 4 residual conventional blocks to downsample itself into a bottleneck feature whereas the decoder blocks make use of 4 residual deconvolutional blocks to obtain the separation components by upsampling the features. Following this, each of the encoder blocks & its corresponding decoder blocks establish a skip connection that operates at the same upsampling or downsampling rate. The residual block of the framework consists of 2 Leaky-ReLU activation layers, 2 batch normalization layers, and 2 CNN layers, and furthermore, the framework also introduces an additional residual shortcut that connects the input & output of every individual residual block. The ResUNet model takes the complex spectrogram X as the input, and produces the magnitude mask M as the output with the phase residual being conditioned on text embeddings that controls the magnitude of scaling, and rotation of the angle of the spectrogram. The separated complex spectrogram can then be extracted by multiplying the predicted magnitude mask & phase residual with STFT (Short-Time Fourier Transform) of the mixture.

In its framework, AudioSep uses a FiLm or Feature-wise Linearly modulated layer to bridge the separation model & the text encoder after the deployment of the convolutional blocks in the ResUNet.

Training and Loss

During the training of the AudioSep model, developers use the loudness augmentation method, and train the AudioSep framework end-to-end by making use of an L1 loss function between the ground truth & predicted waveforms.

Datasets and Benchmarks

As mentioned in previous sections, AudioSep is a foundational model that aims to resolve the current dependency of LASS models on annotated audio-text paired datasets. The AudioSep model is trained on a wide array of datasets to equip it with multimodal learning capabilities, and here is a detailed description of the dataset & benchmarks used by developers to train the AudioSep framework.

AudioSet

AudioSet is a weakly-labeled large-scale audio dataset comprising over 2 million 10-second audio snippets extracted directly from YouTube. Each audio snippet in the AudioSet dataset is categorized by the absence or presence of sound classes without the specific timing details of the sound events. The AudioSet dataset has over 500 distinct audio classes including natural sounds, human sounds, vehicle sounds, and a lot more.

VGGSound

The VGGSound dataset is a large-scale visual-audio dataset that just like AudioSet has been sourced directly from YouTube, and it contains over 2,00,000 video clips, each of them having a length of 10 seconds. The VGGSound dataset is categorized into over 300 sound classes including human sounds, natural sounds, bird sounds, and more. The use of the VGGSound dataset ensures that the object responsible for producing the target sound is also describable in the corresponding visual clip.

AudioCaps

AudioCaps is the largest audio captioning dataset available publicly, and it comprises over 50,000 10-second audio clips that are extracted from the AudioSet dataset. The data in the AudioCaps is divided into three categories: training data, testing data, and validation data, and the audio clips are humanly-annotated with natural language descriptions using the Amazon Mechanical Turk platform. It’s worth noting that each audio clip in the training dataset has a single caption, whereas the data in the testing & validation sets each have 5 ground-truth captions.

ClothoV2

The ClothoV2 is an audio captioning dataset that consists of clips sourced from the FreeSound platform, and just like AudioCaps, each audio clip is humanly-annotated with natural language descriptions using the Amazon Mechanical Turk platform.

WavCaps

Just like AudioSet, WavCaps is a weakly-labeled large-scale audio dataset comprising over 400,000 audio clips with captions, and a total runtime approximating to 7568 hours of training data. The audio clips in the WavCaps dataset are sourced from a wide array of audio sources including BBC Sound Effects, AudioSet, FreeSound, SoundBible, and more.

Training Details

During the training phase, the AudioSep model randomly samples two audio segments sourced from two different audio clips from the training dataset, and then mixes them together to create a training mixture where the length of each audio segment is about 5 seconds. The model then extracts the complex spectrogram from the waveform signal using a Hann window of size 1024 with a 320 hop size.

The model then makes use of the text encoder of the CLIP/CLAP models to extract the textual embeddings with text supervision being the default configuration for AudioSep. For the separation model, the AudioSep framework uses a ResUNet layer consisting of 30 layers, 6 encoder blocks, and 6 decoder blocks resembling the architecture followed in the universal sound separation framework. Furthermore, each encoder block has two convolutional layers with a 3×3 kernel size with the number of output feature maps of encoder blocks being 32, 64, 128, 256, 512, and 1024 respectively. The decoder blocks share symmetry with the encoder blocks, and the developers apply the Adam optimizer to train the AudioSep model with a batch size of 96.

Evaluation Results

On Seen Datasets

The following figure compares the performance of AudioSep framework on seen datasets during the training phase including the training datasets. The below figure represents the benchmark evaluation results of the AudioSep framework when compared against baseline systems including Speech Enhancement models, LASS, and CLIP. The AudioSep model with CLIP text encoder is represented as AudioSep-CLIP, whereas the AudioSep model with CLAP text encoder is represented as AudioSep-CLAP.

As it can be seen in the figure, the AudioSep framework performs well when using audio captions or text labels as input queries, and the results indicate the superior performance of the AudioSep framework when compared against previous benchmark LASS and audio-queried sound separation models.

On Unseen Datasets

To assess the performance of AudioSep in a zero-shot setting, developers continued to evaluate the performance on unseen datasets, and the AudioSep framework delivers impressive separation performance in a zero-shot setting, and the results are displayed in the figure below.

Furthermore, the image below shows the results of evaluating the AudioSep model against Voicebank-Demand speech enhancement.

The evaluation of the AudioSep framework indicates a strong & desired performance on unseen datasets in a zero-shot setting, and thus makes way for performing sound operation tasks on new data distributions.

Visualization of Separation Results

The below figure shows the results obtained when the developers used the AudioSep-CLAP framework to perform visualizations of spectrograms for ground-truth target audio sources, and audio mixtures and separated audio sources using text queries of diverse audios or sounds. The results allowed developers to observe that the spectrogram’s separated source pattern is close to the source of the ground truth that further supports the objective results obtained during the experiments.

Comparison of Text Queries

The developers evaluate the performance of AudioSep-CLAP and AudioSep-CLIP on AudioCaps Mini, and the developers make use of the AudioSet event labels , the AudioCaps captions, and re-annotated natural language descriptions to examine the effects of different queries, and the following figure shows an example of the AudioCaps Mini in action.

Conclusion

AudioSep is a foundational model that is developed with the aim of being an open-domain universal sound separation framework that uses natural language descriptions for audio separation. As observed during the evaluation, the AudioSep framework is capable of performing zero-shot & unsupervised learning seamlessly by making use of audio captions or text labels as queries. The results & evaluation performance of AudioSep indicate a strong performance that outperforms current state of the art sound separation frameworks like LASS, and it might be capable enough to resolve the current limitations of popular sound separation frameworks.

AMD Extends PyTorch Support on RDNA-3 For Local ML Training

AMD Extends PyTorch Support on RDNA-3 Graphics For Local ML Training

In a move aimed at enhancing accessibility to AI for developers and researchers, AMD had unveiled ROCm 5.7 last month, the latest iteration of its open software ecosystem for accelerated computing. Now, the company has now extended its support for PyTorch machine learning (ML) to AMD RDNA 3-based graphics via ROCm 5.7.

This means that individuals engaged in working with ML models and algorithms in PyTorch can leverage AMD ROCm 5.7 on Ubuntu Linux. This enables them to harness the parallel computing capabilities offered by AMD Radeon RX 7900 XTX and Radeon PRO W7900 GPUs, which come equipped with 192 dedicated AI accelerators.

These accelerators promise up to 2X higher AI performance per compute unit in comparison to the preceding generation. The integration of ML on desktops provides users with a local, private, and cost-effective avenue for supporting ML training and inference, thereby reducing the reliance on cloud-based solutions.

Read: AMD’s Attempt to Break NVIDIA’s CUDA

AMD is also trying to break NVIDIA’s CUDA monopoly in the AI parallel computing segment. AMD has made a huge leap with its recent bet on Nod.ai, an open source AI software firm. Nod.ai has been known for developing a portfolio of tools and systems for boosting AI applications on AMD hardware.

The software bet has been going on at AMD for some time now. In August, the company also announced the acquisition of Mipsology, a French AI startup, which has also been a long-standing AMD partner and developing AI software for the chipmaker, similar to Nod.ai.

The post AMD Extends PyTorch Support on RDNA-3 For Local ML Training appeared first on Analytics India Magazine.

AI is Disrupting Culture in Organisations

Silicon Valley has never been famous for its moral practices. In the latest series of (unfortunate) events OpenAI, the ChatGPT maker has been second guessing its core values as it recently tweaked them secretly. The AI company’s revised career page mentions its values as “AGI focus,” “intense and scrappy,” “scale,” “make something people love,” and “team spirit.”

Less than a month ago’s archived page shows the capped-profit company’s original core values — including attributes like “audacious,” “thoughtful,” “unpretentious,” “impact-driven,” “collaborative,” and “growth-oriented.”

In 2019, the same year GPT-2 was released in the wild, OpenAI announced turning from a non-profit to capped profit model to “allow it to rapidly increase investments in compute and talent.“ The company’s co-founder and tech industry’s self proclaimed troubleshooter, Elon Musk also fumed when OpenAI became a ‘$30 billion market cap-company’ after his $100 million donation.

The budding startup among the who’s who of tech is not the first one to re-do its root principles and move away from its original purpose — the betterment of humanity instead of minting money. In the case of OpenAI, the company has publicly never shied away from voting for an AGI or “the equivalent of a median human that you could hire as a co-worker.” Hence, the revised language marks a bigger shift in OpenAI’s focus.

Google’s Mischief Managed

A year before OpenAI was to be founded Google’s Larry Page admitted that the company outgrew its original mission statement to “organise the world’s information and make it universally accessible and useful” from the launch of the company, but he did not know how to redefine it.

Back in 2014, Page set out to do some really big things with the resources at his disposal. He wanted to push the boundaries of technology and the market in various areas, like AI, robotics, health, disease, and biotechnology. This effort was led by Sergey Brin and his Google X labs, which were known for taking on daring “moonshot” projects.

A few years later, the duo moved away from their responsibilities at the company. Google also figured out its love-to-make-money identity and has crossed various lines to do so. The recent episode is when Google updated its privacy policy amid the ongoing ‘bigger the better’ language model war. The updated policy allows the company to collect and analyse information its users share online to “improve its services and develop new AI-powered products.”

More recently, behind the users’ back, Google rewrote its ‘Helpful Content Update’, from “written by people” to “content created for people” to rank sites on its search engine.The linguistic pivot shows that the company does recognise the significance of AI tools.

The Other End

As some companies have witnessed severe backlash for changing its Terms of Service (TOS), new questions are being raised about customer privacy, choice and trust as some users have no option to opt out of the updates.

On the contrary, software king Microsoft, which funnelled eyebrows raising $10 billion dollars in OpenAI, has also revamped its terms and conditions to handle AI. The latest five changes limit the user’s from using Microsoft service’s data to create, train, or improve (directly or indirectly) any other AI service.

The Satya-Nadella run corporation has previously said that it doesn’t save conversations or use that data to train its AI models for its Bing Enterprise Chat mode. Its policies are less clear for its Microsoft 365 Copilot; although it doesn’t appear to use customer data or prompts for training, it does store information.

Joining the veteran’s club, OpenAI lost no time and consolidated core values which now feels like a bunch of fluff. However, it shows OpenAI’s hand when it comes to its single-minded focus looking forward. “We are committed to building safe, beneficial AGI that will have a massive positive impact on humanity’s future,” the OpenAI job postings page now explains. “Anything that doesn’t help with that is out of scope.”

The post AI is Disrupting Culture in Organisations appeared first on Analytics India Magazine.