Why is GitHub Bullish About AI in Cybersecurity?

GitHub foresees the pivotal role of AI in the software development lifecycle, including security. Over the past year, the company incorporated over 70 features into GitHub Advanced Security. But at GitHub Universe 2023, held in November, it announced that they are adding generative AI to the mix.

The company now believes security vulnerabilities can be identified at the very stage when the code is being written. By leveraging a Large Language Model (LLM), GitHub now not only identifies potential vulnerabilities but also provides developers with secure code suggestions from the start.

“With auto fix, we’re now going to suggest the fix in the pull request for them. So the developers will not just see the alert, but also a suggested fix powered by AI right there,” Jacob Depriest, VP and Deputy Chief Security Officer at GitHub, told AIM.

These are not ordinary fixes. They are concise, actionable suggestions for swiftly comprehending and addressing vulnerabilities. Developers can now resolve issues faster and prevent new vulnerabilities from creeping into your codebases.

Depriest said GitHub has already seen great success with this new feature, “This implies that when developers receive an alert while working, they address the issue approximately 50% of the time before it reaches production and this is huge.”

Protecting secrets with AI

GitHub is not just using LLMs solely for uncovering potential code vulnerabilities, but the company also uses these powerful models to detect leaked passwords with reduced false positives.

Nearly 80 percent of the breaches originate through credential leakage or secrets being leaked, according to Depriest. “Secret scanning has been integral to GitHub’s advanced security, forming a key component of our security programme.

“Now with AI, we’re also going to detect generic secrets and low confidence patterns in code as well, which is going to really improve that capability and capture more and protect secrets before they’re even hitting production.”

Moreover, Depriest believes security starts with the developer and to be more precise, with the developer’s account. Given the high number of credential leakages, it opted to lean heavily into enabling multi-factor authentication for all contributors on github.com.

“That wasn’t an easy thing to do. That wasn’t a quick thing to do. It took a lot of planning and a lot of investment to make that work. But we really believe that’s the right thing to do,” he said.

And now, the introduction of the new secret scanning feature allows GitHub to detect generic or unstructured secrets in the code.

Protecting against AI Vulnerabilities

Despite GitHub’s optimistic stance on leveraging AI in cybersecurity, the era of generative AI has presented various instances where it emerges as a substantial cybersecurity threat.

For instance, prompt injection attacks remain a significant challenge to address for cybersecurity teams. Over time, we have seen LLMs are vulnerable to prompt injection attacks.

Given GitHub’s close alliance with Microsoft, it might be leveraging OpenAI’s GPT models, most notably GPT-4, the most advanced LLM to date. However, GPT-4 has also been found to be vulnerable to prompt injection attacks.

Depriest believes responsible integration and security measures within the tooling are crucial in safeguarding against such manipulations. This approach is fundamental to ensuring protection in scenarios involving prompt injection and similar vulnerabilities.

Moreover, according to Depriest, safeguarding both the infrastructure and the overall network workspace remains a key aspect of cybersecurity, even in the AI age.

“We approach this responsibility with the same diligence as we do for the rest of our infrastructure, including protecting github.com. This entails threat detection, security operations, and ensuring code security, and this extends to AI models. We uphold uniform controls and compliance principles across all facets of our core responsibilities.”

Does AI fundamentally change cybersecurity?

While GitHub is banking on generative AI capabilities in a big way, including for cybersecurity, Depriest does not believe it will fundamentally change the cybersecurity landscape.

“The reality is every new technology is dual-use. This pattern has been consistent across various technologies throughout the past two decades. I still don’t think fundamentally it will change how we approach the security of what we need to do, what our job is, and how we will keep the platform safe.”

He holds the belief that the advantages of generating secure code from the outset and consistently maintaining its security far exceed the potential risks associated with certain threats in the landscape.

“We strongly believe that our efforts at GitHub, incorporating AI into the developer workflow, will yield a substantial and highly valuable impact, far outweighing any potential risks in the threat landscape.

While scaling remains an industry challenge, we envision this as the future for enhancing security in the realm of developers.”

The post Why is GitHub Bullish About AI in Cybersecurity? appeared first on Analytics India Magazine.

Why Taking a Year Off to Build a Startup is a Bad Idea

A lot of Indian universities are giving a year off to students for building startups. BITS Pilani, IIT Madras, IIT Hyderabad, DIT University, IIT Bombay, IIT Kharagpur, and a few more are giving a year off, which some call “a temporary withdrawal programme” allowing students to take a break from their academic studies and pursue their entrepreneurial calling.

“Bad idea,” said a user on LinkedIn talking about the universities’ move to allow a year off. He highlighted how typically, only 2 to 6 percent of startups achieve profitability within the first year.

What about tech entrepreneurs or students whose startups fail or prove non-viable in the market?

AIM Recruits believes that stability concerns may arise with candidates who have entrepreneurial experience, as they may not always seamlessly integrate into the traditional corporate structure. “Entrepreneurs often tend to revert to launching their ventures shortly after joining an organisation,” said the spokesperson.

At the same time, the failure of the startup also signals a potential deficit in comprehending market dynamics or industry trends. Most investors would not want to fund startups started by students, as they would believe they would end up losing all the money.

No more jobs?

Experts say that in the product-based or startup sectors, skills outweigh degrees. “Entrepreneurial experience holds particular significance, indicating a strong commitment and work ethic.”

The main reason for a year off from universities was allowing the students to gain real-world experience and apply theoretical knowledge in practical settings. This experience will not only enhance their entrepreneurial skills but also make them more adaptable and resilient professionals.

Candidates with startup backgrounds often showcase leadership or dedicated individual contributor traits, providing a dual advantage for organisational success. Additionally, their technical expertise, gained from working on their projects, enhances proficiency in essential skills.

But the same is not the case with big companies such as Microsoft, Google, or Indian IT giants. “They don’t hire such people at all. They are very particular about pedigree,” said the AIM Recruits expert. “Their main target pool are folks from IITs, NITs, BITs, DTU, MIT, without any backlogs or gaps.”

Adding to this, Indian IT companies have put a freeze on hiring new employees. Many of them have a huge bench size, and the companies believe that they require a lot more training of the existing employees before they can hire new ones. According to a recent report, private engineering institutes have started to report a 50-70% drop in placements in 2023.

Not a toss of a coin

One of the most significant drawbacks of taking a year off is the potential impact on future career opportunities. Employers often value a consistent academic track record, and a gap year might be perceived as a red flag, potentially affecting the chances of securing desirable job placements after graduation.

Even though universities are the best places for hiring and promoting AI and entrepreneurship, it seems like a risky move to move out of college for a year unless you are absolutely sure that the startup would work. For the universities, it becomes very hard to determine if the AI venture would be worth it or not.

It is also better for the bigger companies to hire more freshers and promote their entrepreneurial callings through incubators and startup programmes, instead of letting them out in the open market. Though some universities provide resources during the gap, it is still not enough to make a profitable business.

Funding winter?

The government of India announced in August its scheme of funding 1,200 startups under the Digital India Initiative, giving more boost to the entrepreneurs in universities. Despite this, the startup funding in India dropped significantly in 2023.

According to Tracxn data, in 2023, the startup funding declined by 73% compared to 2022. Drop from $25 billion to $7 billion, which is the lowest figure in the last five years. Most importantly, seed-stage funding dropped 60% from $1.7 billion in 2022 to $678 million in 2023.

All in all, it might not entirely be a bad idea to take a year off, but it would also be wiser to get a job (which is also difficult) and then possibly join their startup programme. Since the market is down, the percentages show that the chance of failure is significantly higher. Think twice.

The post Why Taking a Year Off to Build a Startup is a Bad Idea appeared first on Analytics India Magazine.

Indian AI and Robotics Startup Claims to Achieve Level 5 Autonomy 

Indian autonomous driving company Swaayatt Robots has announced that it achieved the world’s first Level 5 autonomous driving capability demo. In the demonstration, its autonomous vehicle learned to negotiate complex traffic dynamics in the Toll-Plaza and successfully crossed highly unstructured toll-gates.

The company said that this achievement, second only to previous demos in 2017 and 2023, signifies a significant leap in Level-5 capabilities. The team enabled autonomous driving to negotiate tight and dynamic adversarial environments using multi-reinforcement learning agents. Notably, the demo on October 22, 2023, involved bidirectional traffic negotiation on single-lane roads.

Swayaatt Robots posted a video which displays the vehicle entering the toll-gate region through a highway, navigating bidirectional traffic dynamics in an open area without strict driving rules. The complexity is heightened by randomly parked trucks, a common occurrence in Toll-Plaza areas, and the vehicle’s ability to decide which toll-gate-passage to commit to while avoiding overtaking trucks.

To illustrate the complexity, a tractor tire was placed behind a large 18-wheeler truck, testing the vehicle’s capability to detect obstacles at night. Impressively, the vehicle successfully negotiated this scenario while another truck overtook from the right.

The vehicle demonstrated advanced decision-making and motion planning algorithms, crucial for navigating such intricate scenarios. The team emphasized the scalability of this framework with unsupervised deep learning and teased an upcoming showcase in February, promising an end-to-end negotiation of daytime traffic.

Notably, the vehicle slowed and paused approaching a speed-breaker, aligning itself on-the-fly with the selected toll-gate, adhering to driving rules and speed limits. The system even adapted to unexpected obstacles, such as a broken traffic-police barricade, showcasing a robust and adaptive autonomous driving system.

Interestingly, Tesla cars currently fall under Level 2 of the six levels of vehicle automation defined by the Society of Automotive Engineers (SAE).

The post Indian AI and Robotics Startup Claims to Achieve Level 5 Autonomy appeared first on Analytics India Magazine.

PyPy Shifts to GitHub

On December 29, 2023, the team behind PyPy announced the migration of their canonical repository and issue tracker from Heptapod to GitHub, a move that changes their project management and community engagement strategies.

This migration from Mercurial to GitHub is motivated by several factors. The team faced difficulties in tracking issues and contributions on Heptapod. They recognised the need for greater visibility and easier access, something GitHub is renowned for in the open-source community according to their blog post.

Moreover, GitHub’s dominance in hosting open-source projects presents an opportunity for PyPy to integrate into a more vibrant and active developer ecosystem. The unified platform of GitHub also simplifies the tracking of issues and code, streamlining the development process.

The process of migration involved several technical steps, including the transfer of code, issues, and merge requests. Tools like git-remote-hg were used for code conversion, and node-gitlab-2-github was employed for migrating issues and merge requests. However, the transition was not without its challenges. The PyPy team encountered issues such as branch naming discrepancies and complexities in converting Mercurial branches to GitHub’s format.

For developers, this move to GitHub brings a host of implications. It promises easier access and opportunities for contributing to PyPy, owing to GitHub’s widespread use and familiarity within the developer community.

Additionally, the migration is expected to enhance the visibility of issues and the overall progress of development. While this may require some adaptation in terms of workflow to align with GitHub’s environment, the PyPy team is actively working to refine the migration process and is inviting community feedback and involvement to make the transition as smooth as possible.

In the broader context of open-source software development, PyPy’s migration reflects a trend of consolidation on GitHub.

Discussions on Hacker News indicate mixed feelings within the community regarding GitHub’s dominance. While some view it as a positive development for open-source projects, others express concerns about the centralisation of power and the challenges faced by smaller platforms in attracting and retaining large projects. PyPy’s move could potentially influence other open-source projects, further centralising software development on GitHub and shaping the future landscape of open-source collaborations.

The post PyPy Shifts to GitHub appeared first on Analytics India Magazine.

Top Generative AI Search Alternatives for Google 

Google has been at the forefront of search engines from an early stage. 2023 has been a great year for Google’s advancements in generative AI. Several groundbreaking models were introduced in 2023. From Gemini to Imagen 2, Google witnessed a drastic drift in AI. Although being one of the most used and widely known search engines, it has competitors and alternatives that provide the same use case and provide more safety and privacy than Google.

Privacy apprehensions regarding Google’s extensive data collection practices have prompted users to seek alternatives that prioritize data protection. By diversifying digital experiences, users often opt for alternatives to mitigate concerns related to Google’s monopoly in various service sectors.

The desire for greater customisation and control over digital environments, innovative features, ethical considerations, and platform independence further fuels the demand for alternatives. Cost considerations, technological advancements, global accessibility, and security concerns also contribute to users exploring and adopting alternative services.

Here is a list of generative AI search alternatives for Google.

Bing Chat

Bing Chat, Microsoft’s AI-powered search companion, represents a significant leap forward in the evolution of search engines. With its array of features, such as contextual summaries, insights and analysis, consideration of multiple perspectives, and seamless integration with Bing Search, Bing Chat offers users a comprehensive and efficient search experience.

By analysing search queries, assessing relevance, extracting key information, and generating clear summaries alongside additional insights, it not only saves time but also enhances comprehension and aids in making more informed decisions. The benefits range from time-saving and personalised experiences to improved research efficiency. Looking ahead, Bing GPT holds promise for further advancements, including conversational search capabilities, personalised learning assistance, and support for researchers in uncovering hidden patterns and generating new ideas.

DuckDuckGo

DuckDuckGo stands out as a privacy champion in search engines, utilising generative algorithms to deliver search results while upholding user privacy. Its distinctive feature lies in its commitment to not tracking user data, offering a compelling alternative for individuals prioritising online privacy.

Unlike many other search engines, DuckDuckGo goes the extra mile with built-in features such as tracker blocking and private browsing, reinforcing its dedication to shielding users from intrusive elements on the internet. By incorporating generative AI into its operations, DuckDuckGo ensures a user-friendly search experience and establishes itself as a trustworthy option for those seeking a more private and secure online journey.

Download it on Play Store: DuckDuckGo

Brave Browser

Brave Browser emerges as a standout choice in the digital landscape, embodying speed, privacy, and an ad-free experience. A notable feature of Brave lies in its incorporation of generative AI, actively blocking intrusive ads and trackers to bolster user privacy and expedite browsing by eliminating superfluous content. What sets Brave apart is its innovative take on online advertising. Rather than subjecting users to invasive ads, the browser introduces the Basic Attention Token (BAT), a unique rewards system.

BAT incentivizes users to engage with privacy-respecting ads voluntarily. This pioneering approach empowers users to control their ad experience and introduces a novel way for advertisers to connect with an audience genuinely interested in their content. With cutting-edge technology and user-centric features, Brave Browser redefines the browsing experience for those seeking speed, privacy, and a fair approach to online advertising.

Download it on Play Store: Brave Browser

You.com

You.com emerges as a promising contender in the search engine landscape with a distinctive focus on reinstating objectivity and trust in online information. Leveraging AI, it pioneers fact-checking by identifying misinformation, revealing sources, and actively seeking diverse perspectives to foster a balanced understanding of complex issues. Beyond fact-checking, You.com excels in personalisation, employing explainable AI to transparently communicate its assessments and encouraging community-driven efforts to improve accuracy.

Challenges such as adapting to evolving misinformation tactics, maintaining neutrality, and scaling up are acknowledged, emphasising the platform’s commitment to continuous improvement. Despite these hurdles, You.com’s commitment to making online information more reliable and trustworthy positions it as a transformative force in reshaping how users navigate the vast digital landscape, emphasising accuracy and transparency over sheer engagement metrics.

Perplexity

Perplexity emerges as a transformative force in search engines, redefining the user experience through generative AI. With its commitment to comprehensive synthesis, Perplexity delivers detailed summaries and direct answers, transcending traditional snippets to provide a profound understanding of search results. The emphasis on exploring different perspectives is a standout feature, diversifying sources and combating biased results, fostering a more balanced comprehension of topics.

Including Copilot Assistant and a Pro Version further enhances the user experience, offering refined search queries, sentiment analysis, and advanced functionalities. While Perplexity acknowledges its AI-generated nature, users must exercise caution regarding accuracy and information overload, especially considering its early developmental stage.

The post Top Generative AI Search Alternatives for Google appeared first on Analytics India Magazine.

Generative AI business model disruption: The NYT lawsuit posturing 

Generative AI business model disruption: The NYT lawsuit posturing 

2024 will be all about changing business models due to the massive disruption of generative AI.

There will be new winners and many losers. The incumbents especially have a lot to lose – but permissionless innovation has always been the hallmark of American innovation.

We see the usual vanguard action from the incumbents who find themselves disrupted – and we must not forget that regulation favours the incumbent, when viewed from this perspective of innovation, most of the elements of the NYT lawsuit against OpenAI and Microsoft are predictable.

Here are some thoughts and implications for the business model:

  1. Billions in lost revenue? I doubt it! Have you ever seen anyone read a newspaper recently? Billions over a year for a left-leaning and non-business newspaper is overkill. Probably thousands if that. So, this is all posturing (and has precedence as we see below)
  2. If Internet scraping were illegal – so would search engines. So, while many people talk of scraping as a concern – it may be a red herring.
  3. These lawsuits are part of a wider trend. Getty Images vs. Stability AI; Sarah Anderson, Kelly McKernan, and Karla Ortiz vs. Stability AI; Sarah Silverman vs. Open AI and FaceBook; Novelists Mona Awad and Paul Tremblay vs. Open AI; 5000 Authors’ Petition; Federal Trade Commission Investigation:
  4. The most interesting aspect is the memorization i.e. the idea that somehow Gen AI has memorized whole article verbatim. But even here, questions remain. Do we see the same phenomenon in other LLMs? For other publications? So far, there is no news of others experiencing this. Even if it were, its hard to prove when the document was uploaded and by whom. Like the famous Samsung case where patent documents were uploaded to GPT – it would be hard to show who exactly uploaded a document and when

Hence, this could be all posturing. Again there is a precedence with Viacom and YouTube.

Viacom’s early objections to YouTube centered around copyright infringement and the lack of control over their content with a $1 Billion Lawsuit: In 2007 alleging massive copyright infringement. After years of legal battles, the lawsuit was settled with no money changing hands. The terms weren’t disclosed, but the settlement likely involved some agreement on how Viacom’s content would be managed on YouTube in the future.

Industry-Wide Changes:

The Viacom vs. YouTube case was a landmark in the evolution of digital content. It influenced how other media companies approached online platforms, leading to more collaborations and innovative content distribution strategies.

In summary, Viacom’s early objections to YouTube were based on legitimate concerns over copyright infringement and control over their content. However, these objections were eventually overcome through legal rulings, technological solutions like the Content ID system, and strategic partnerships that allowed both parties to benefit from the digital distribution of content.

This evolution marked a significant shift in the media landscape, reflecting broader changes in how content is created, distributed, and monetized in the digital age.

This may sound less exciting – but we are likely to see something similar here.

Image source: what else OpenAI – an elephant reading the NYT and trying to memorise it – elephants have long memories 🙂

Views expressed in this article are personal and are not associated with any organisation I am associated with

Mastering IoT Data Management for Business Success

thinkstockphotos-610749178-100734913-large

In today’s tech-driven landscape, the proliferation of Internet of Things (IoT) devices has revolutionized how businesses collect and utilize data. The interconnectivity of these devices has created an unprecedented influx of data, requiring efficient management strategies to harness its full potential.

Understanding IoT data

IoT devices span a vast array, from sensors in machinery to smart appliances and wearable gadgets. These devices continuously generate data streams encompassing diverse types, such as sensor readings, location information, user interactions, and more. Harnessing this data presents unparalleled opportunities, but it also poses significant challenges.

The challenge of volume, velocity, and variety

The crux of IoT data management lies in handling the 3Vs: Volume, Velocity, and Variety. The sheer volume of data inundating systems demands robust storage and processing capabilities. The velocity at which data arrives requires real-time or near-real-time processing to derive timely insights. Moreover, the variety of data formats and sources complicates integration and analysis.

Strategies for efficient IoT data management

  1. Scalable infrastructure using cloud and edge computing:

IoT data often overwhelms traditional infrastructure. Cloud platforms like AWS IoT, Microsoft Azure IoT Hub, or Google Cloud IoT provide scalable storage and computing resources. For instance, AWS IoT Core allows seamless integration, storing, and processing of IoT data at scale, enabling businesses to adapt to fluctuating data volumes.

Additionally, edge computing brings computation closer to the data source, reducing latency and bandwidth usage. Consider a factory employing edge devices to analyze machinery data locally, sending only critical insights to the cloud for further analysis. This minimizes latency for real-time decision-making.

  1. Data filtering and prioritization:

Filtering data at the source ensures that only relevant information is transmitted. This is crucial in scenarios where bandwidth is limited or costly. For instance, IoT sensors in agriculture collect various data points. Filtering out non-critical data allows farmers to focus on essential factors like soil moisture levels or temperature variations, optimizing crop yields.

Prioritizing data based on urgency or importance ensures that critical insights are addressed promptly. In healthcare, wearable devices monitoring vital signs prioritize sending emergency alerts to healthcare providers, providing immediate attention when necessary.

  1. Robust security measures:

Securing IoT data is paramount due to its sensitivity. Employing encryption methods like AES (Advanced Encryption Standard) or TLS (Transport Layer Security) safeguards data during transmission and storage. For example, encrypted communication protocols in smart homes secure user data transmitted between smart devices and central hubs, preventing unauthorized access.

Access controls and authentication mechanisms prevent unauthorized access to IoT devices. Multifactor authentication ensures that only authorized personnel can access and manage IoT devices or data, enhancing overall security.

  1. Streamlined data integration and interoperability:

Integration platforms and standardized protocols facilitate seamless communication and data exchange among diverse IoT devices. For instance, MQTT (Message Queuing Telemetry Transport) protocol enables efficient data transmission between IoT devices and servers in industrial IoT applications, ensuring interoperability.

Implementing APIs and middleware solutions allows different IoT devices to communicate and share data, creating a unified ecosystem. This is evident in intelligent cities where various IoT devices, such as traffic sensors and environmental monitors, seamlessly transfer data to optimize city operations.

  1. Predictive analytics and AI-driven insights:

Machine learning algorithms and predictive analytics glean actionable insights from IoT data. These tools enable predictive maintenance in industries like transportation, where IoT sensors on vehicles predict potential component failures, allowing proactive maintenance, minimizing downtime, and optimizing operations.

AI-driven anomaly detection in IoT data helps detect irregular patterns, flagging potential security breaches or operational issues. For instance, in the energy sector, AI algorithms analyze IoT sensor data from power grids to identify anomalies indicative of potential failures or cyber-attacks, allowing preemptive measures to be taken.

The future of IoT data management

As IoT adoption proliferates across industries, the future of IoT data management is poised for evolution. Advancements in edge computing, data automation, AI-driven analytics, and 5G technology will further enhance data processing capabilities, enabling more sophisticated real-time applications.

  1. Advancements in edge computing:

The evolution of edge computing will revolutionize IoT data management. Edge devices with enhanced processing capabilities will perform complex computations locally, reducing latency and dependency on centralized cloud systems. For instance, edge AI chips in autonomous vehicles will process vast amounts of sensor data in real time, enabling faster decision-making without relying extensively on cloud servers.

  1. 5G technology integration:

The integration of 5G technology will significantly impact IoT data management. Its ultra-fast, low-latency networks will enable seamless transmission of large volumes of IoT data. For instance, remote surgeries utilizing IoT devices and 5G connectivity in healthcare will rely on instant, high-bandwidth data transmission to ensure precision and real-time responsiveness.

  1. AI-driven analytics and edge AI:

The convergence of AI-driven analytics and edge AI will empower IoT devices to perform more sophisticated data processing and analysis at the device level. For example, smart cameras with edge AI capabilities will analyze video data locally, recognizing objects, anomalies, or potential security threats without continuously relying on cloud-based AI systems, minimizing latency and bandwidth requirements.

  1. Blockchain for IoT security and data integrity:

Integrating blockchain technology will enhance security and data integrity in IoT ecosystems. Immutable ledgers and decentralized architectures will ensure transparent and tamper-proof data transactions. For instance, in supply chain management, blockchain-enabled IoT devices will track and validate each step of a product’s journey, ensuring authenticity and minimizing fraud.

  1. Federated learning for privacy-preserving insights:

Federated learning, a privacy-preserving technique, will gain prominence in IoT data management. This approach enables collaborative model training across multiple IoT devices without sharing raw data. For instance, in smart homes, IoT devices will collaboratively train AI models to improve energy efficiency without compromising individual user data privacy.

  1. Augmented Reality (AR) and IoT integration:

Integration of IoT with AR technologies will create immersive experiences and enhance IoT data visualization. For instance, in manufacturing, IoT sensors integrated with AR glasses worn by technicians will overlay real-time equipment performance data onto machinery, enabling quick diagnostics and maintenance.

  1. Quantum computing for complex IoT analytics:

The emergence of quantum computing will revolutionize complex IoT data analytics. Quantum computing’s immense processing power will handle intricate data analysis tasks, such as optimizing large-scale IoT networks or performing complex simulations for predictive maintenance in industrial settings.

  1. Evolving regulatory frameworks:

As IoT adoption grows, regulatory frameworks around data privacy, security, and interoperability will continue to evolve. Standards and regulations will ensure the ethical use, interoperability, and security of IoT devices and the vast amounts of data they generate.

Conclusion

Effectively managing IoT data is a cornerstone for businesses aiming to thrive in the data-driven era. By embracing scalable infrastructure, robust security measures, streamlined integration, and advanced analytics, organizations can leverage insights derived from IoT data to drive innovation, enhance operational efficiency, and stay ahead in an increasingly competitive landscape.

Operational, real-time edge analytics for developers

Interview podcast with Rahul Pradhan, VP of Product and Strategy at Couchbase

Operational, real-time edge analytics for developers
Image by Gerd Altmann from Pixabay

Operational and analytics systems are coming together with the help of new database management innovations. A recent step from Couchbase’s point of view has been to bring a real-time analytics capability to the operational applications that developers use Couchbase to create.

With real-time data and analytics together in one operational platform, developers can ensure the data freshness and associated context necessary to address the LLM hallucination problem.

Couchbase features an edge component so that users at remote sites can work either offline or online. Mobile applications often encounter last-mile constraints, which is how the offline-online synchronization capability becomes essential.

At the edge, inference can run as close as possible to where the data is generated. Users can reduce resource requirements with quantization (constraining the input by limiting it to integers only, for example), ensuring good enough quality for a number of use cases. Otherwise, Couchbase is designed to provide the same user experience as with the cloud.

With fresh data, app developers can focus on building these machine learning-related functions into their applications.

Couchbase is a multimodal NoSQL database that began with a native JSON document mode in 2011. It’s built with a high-performance storage engine and distributed caching to serve demanding environments.

Couchbase blends operational transactional, analytics, machine learning, generative and predictive capabilities in a single platform, delivering in near real-time with submillisecond latency. It includes integrated key-value and SQL interfaces and search and analytics functionality.

Interviewee Rahul Pradhan’s current role is VP of Product and Strategy. He first became interested in distributed computing as a software engineer at Nortel Networks back in the 2000s. He has a considerable infrastructure engineering and product management background, building network and security software at Nortel Networks, then at the storage layer for EMC (later Dell EMC) in file and block storage.

Pradhan joined Couchbase six years ago. His interest on the product side has been in high-performance, real-time, highly personalized cloud use cases.

Hope you enjoy the interview and catching up on what Couchbase has been up to as much as I have.

Podcast interview with Rahul Pradhan, VP of Product and Strategy at Couchbase

Social Impact of Generative AI: Benefits and Threats

Featured image for generative AI

Today, Generative AI is wielding transformative power across various aspects of society. Its influence extends from information technology and healthcare to retail and the arts, permeating into our daily lives.

As per eMarketer, Generative AI shows early adoption with a projected 100 million or more users in the USA alone within its first four years. Therefore, it is vital to evaluate the social impact of this technology.

While it promises increased efficiency, productivity, and economic benefits, there are also concerns regarding the ethical use of AI-powered generative systems.

This article examines how Generative AI redefines norms, challenges ethical and societal boundaries, and evaluates the need for a regulatory framework to manage the social impact.

How Generative AI is Affecting Us

Generative AI has significantly impacted our lives, transforming how we operate and interact with the digital world.

Let's explore some of its positive and negative social impacts.

The Good

In just a few years since its introduction, Generative AI has transformed business operations and opened up new avenues for creativity, promising efficiency gains and improved market dynamics.

Let’s discuss its positive social impact:

1. Fast Business Procedures

Over the next few years, Generative AI can cut SG&A (Selling, General, and Administrative) costs by 40%.

Generative AI accelerates business process management by automating complex tasks, promoting innovation, and reducing manual workload. For example, in data analysis, models like Google's BigQuery ML accelerate the process of extracting insights from large datasets.

As a result, businesses enjoy better market analysis and faster time-to-market.

2. Making Creative Content More Accessible

More than 50% of marketers credit Generative AI for improved performance in engagement, conversions, and faster creative cycles.

In addition, Generative AI tools have automated content creation, making elements like images, audio, video, etc., just a simple click away. For example, tools like Canva and Midjourney leverage Generative AI to assist users in effortlessly creating visually appealing graphics and powerful images.

Also, tools like ChatGPT help brainstorm content ideas based on user prompts about the target audience. This enhances user experience and broadens the reach of creative content, connecting artists and entrepreneurs directly with a global audience.

3. Knowledge at Your Fingertips

Knewton’s study reveals students utilizing AI-powered adaptive learning programs demonstrated a remarkable 62% improvement in test scores.

Generative AI brings knowledge to our immediate access with large language models (LLM) like ChatGPT or Bard.ai. They answer questions, generate content, and translate languages, making information retrieval efficient and personalized. Moreover, it empowers education, offering tailored tutoring and personalized learning experiences to enrich the educational journey with continuous self-learning.

For example, Khanmigo, an AI-powered tool by Khan Academy, acts as a writing coach for learning to code and offers prompts to guide students in studying, debating, and collaborating.

The Bad

Despite the positive impacts, there are also challenges with the widespread use of Generative AI.

Let's explore its negative social impact:

1. Lack of Quality Control

People can perceive the output of Generative AI models as objective truth, overlooking the potential for inaccuracies, such as hallucinations. This can erode trust in information sources and contribute to the spread of misinformation, impacting societal perceptions and decision-making.

Inaccurate AI outputs raise concerns about the authenticity and accuracy of AI-generated content. While existing regulatory frameworks primarily focus on data privacy and security, it's difficult to train models to handle every possible scenario.

This complexity makes regulating each model's output challenging, especially where user prompts may inadvertently generate harmful content.

2. Biased AI

Generative AI is as good as the data it's trained on. Bias can creep in at any stage, from data collection to model deployment, inaccurately representing the diversity of the overall population.

For instance, examining over 5,000 images from Stable Diffusion reveals that it amplifies racial and gender inequalities. In this analysis, Stable Diffusion, a text-to-image model, depicted white males as CEOs and women in subservient roles. Disturbingly, it also stereotyped dark-skinned men with crime and dark-skinned women with menial jobs.

Addressing these challenges requires acknowledging data bias and implementing robust regulatory frameworks throughout the AI lifecycle to ensure fairness and accountability in AI generative systems.

3. Proliferating Fakeness

Deepfakes and misinformation created with Generative AI models can influence the masses and manipulate public opinion. Moreover, Deepfakes can incite armed conflicts, presenting a distinctive menace to both foreign and domestic national security.

The unchecked dissemination of fake content across the internet negatively impacts millions and fuels political, religious, and social discord. For example, in 2019, an alleged deepfake played a role in an attempted coup d'état in Gabon.

This prompts urgent questions about the ethical implications of AI-generated information.

4. No Framework for Defining Ownership

Currently, there is no comprehensive framework for defining ownership of AI-generated content. The question of who owns the data generated and processed by AI systems remains unresolved.

For example, in a legal case initiated in late 2022, known as Andersen v. Stability AI et al., three artists joined forces to bring a class-action lawsuit against various Generative AI platforms.

The lawsuit alleged that these AI systems utilized the artists' original works without obtaining the necessary licenses. The artists argue that these platforms employed their unique styles to train the AI, enabling users to generate works that may lack sufficient transformation from their existing protected creations.

Additionally, Generative AI enables widespread content generation, and the value generated by human professionals in creative industries becomes questionable. It also challenges the definition and protection of intellectual property rights.

Regulating the Social Impact of Generative AI

Generative AI lacks a comprehensive regulatory framework, raising concerns about its potential for both constructive and detrimental impacts on society.

Influential stakeholders are advocating for establishing robust regulatory frameworks.

For instance, the European Union proposed the first-ever AI regulatory framework to instill trust, which is expected to be adopted in 2024. With a future-proof approach, this framework has rules tied to AI applications that can adapt to technological change.

It also proposes establishing obligations for users and providers, suggesting pre-market conformity assessments, and proposing post-market enforcement under a defined governance structure.

Additionally, the Ada Lovelace Institute, an advocate of AI regulation, reported on the importance of well-designed regulation to prevent power concentration, ensure access, provide redress mechanisms, and maximize benefits.

Implementing regulatory frameworks would represent a substantial stride in addressing the associated risks of Generative AI. With profound influence on society, this technology needs oversight, thoughtful regulation, and an ongoing dialogue among stakeholders.

To stay informed about the latest advances in AI, its social impact, and regulatory frameworks, visit Unite.ai.

GenAI: Synthesizing DNA Sequences with LLM Techniques

When people talk about Large Language Models, the most common topics are text summarization, text generation, and answering prompts with GPT. Yet, this is just the tip of the iceberg. What if the language has an unusual alphabet? In this article, I discuss creating meaningful, synthetic sentences — very long ones with millions of letters — in a peculiar language. Its alphabet has 4 letters: A, C, G, T. Each one represents a protein: adenine (A), cytosine (C), guanine (G), and thymine (T). This is the DNA language, and the long sentences are DNA sequences.

Patterns Found in DNA Sequences

Just like standard English, letter and word combinations are not random. Some happen frequently, and some not at all. Then, rare combinations may indicate a particular genetic condition, similar to typos in the English language. But the big difference is the absence of separators (commas, spaces, question marks), creating long, uninterrupted sequences of letters. In this context, a word is any combination of 2, 3 or more consecutive letters. Because of this, all words overlap.

In some sense, DNA synthetization is somewhat easier than generating English text. Thus, this could be a first project for professionals starting with LLMs. Yet, there are long-range autocorrelations and non-probabilistic rules. For instance, how and where do I insert the correct word that identifies the gender of a human being? I will not answer these questions here. Instead, I focus on rather short-range patterns. In the end, it is not that much different from standard LLMs focusing mostly on adjacent tokens and autoregressive predictors. In both cases, word embeddings are a critical component.

Figure 1 shows three DNA subsequences: a real one, a synthetic one, and a random sequences of letters A, C, G, T. Can you identify patterns in the two top ones?

gen
Figure 1: comparing real, synthetic, and random DNA

The Training Set

The training set consists of a number of human DNA sequences, and it is publicly available. The original project consisted in classifying different subsequences showcasing different genetic traits, based on the patterns and words found, including their statistical distributions. I grouped the data, producing a sequence with a few millions letters. In addition to the 4 letters A, C, G, T, there were a number of occurrences of the letter N, presumably representing missing data.

The training set and Python code are on my GitHub repository. You can access them by downloading the technical document available as paper #34, here. An interesting ethical question is about privacy: the individual subsequences may be enough to identify the corresponding persons, and thus the diseases that they might have. Thankfully, the synthetic DNA should make this reverse-engineering — matching DNA against existing databases — impossible.

Note that the technology discussed here also works in other contexts. For instance, to synthesize drugs or molecules.

Synthesizing DNA Sequences

The algorithm has two steps. First looking at pairs of adjacent words and compute occurrences and conditional probabilities. Then sequentially generate new words based on the last generated word and on the probabilities in question. In the end, you end up with a large table of word embedding, each embedding having hundreds or thousands of components (related words), just like in standard LLM.

More specifically, it works as follows. Look at a DNA word S1 consisting of n1 consecutive letters, to identify potential candidates for the next word S2 consisting of n2 symbols. Then, assign a probability to each word S2 conditionally on S1, and use these transition probabilities to sample S2 given S1, then move to the right by n2 letters, do it again, and so on. Eventually you build a synthetic sequence of arbitrary length. There is some analogy to Markov chains and autoregressive processes.

An interesting application consists of generating synthetic sequences — thus artificial human beings — then produce pictures of the artificial people in question based on their artificial DNA. To do this, you may infer physical appearances based on the DNA: eye and hair color, nose shape, and so on.

Evaluating the Quality of Synthetic DNA

To evaluate the results, I compared the frequencies of a large number of words, both in a synthetic and real sequence. I used the Hellinger distance, returning a value between 0 (best) and 1 (worst). Figure 2 is a QQ plot, with each dot representing a word. The X-axis represents the frequency of the word in question in the synthetic data, while the Y-axis represents its frequency in real DNA.

genome6
Figure 2: QQ plot: random (orange) and synthetic (blue) vs real DNA (red diagonal)

Clearly, if you look at Figure 2, DNA sequences are anything but random. The synthetic DNA replicates the correct distribution, while the random DNA is completely off. However, the evaluation metric depends on the length of the words used. In this case, all 4000 selected words had 6 letters. A QQ plot with ten thousand 8-letter words is featured in paper #34, here. The fit is not as great a in Figure 2. In general, these QQ plots do not capture long-range interactions.

Author

Towards Better GenAI: 5 Major Issues, and How to Fix Them

Vincent Granville is a pioneering GenAI scientist and machine learning expert, co-founder of Data Science Central (acquired by a publicly traded company in 2020), Chief AI Scientist at MLTechniques.com and GenAItechLab.com, former VC-funded executive, author and patent owner — one related to LLM. Vincent’s past corporate experience includes Visa, Wells Fargo, eBay, NBC, Microsoft, and CNET.

Vincent is also a former post-doc at Cambridge University, and the National Institute of Statistical Sciences (NISS). He published in Journal of Number Theory, Journal of the Royal Statistical Society (Series B), and IEEE Transactions on Pattern Analysis and Machine Intelligence. He is the author of multiple books, including “Synthetic Data and Generative AI” (Elsevier, 2024). Vincent lives in Washington state, and enjoys doing research on stochastic processes, dynamical systems, experimental math and probabilistic number theory. He recently launched a GenAI certification program, offering state-of-the-art, enterprise grade projects to participants.