Former Twitter engineers are building Particle, an AI-powered news reader

Former Twitter engineers are building Particle, an AI-powered news reader Sarah Perez @sarahintampa / 9 hours

A team led by former Twitter engineers is rethinking how AI can be used to help people process news and information. Particle.news, which entered into private beta over the weekend, is a new startup offering a personalized, “multi-perspective” news reading experience that not only leverages AI to summarize the news, but aims to do so in a way that fairly compensates authors and publishers — or so is the claim.

While Particle hasn’t yet shared its business model, it arrives at a time when there’s a growing concern about the impact of AI on a rapidly shrinking news ecosystem. News that is summarized by AI could limit clicks to publishers’ websites, which means their ability to monetize via advertising would also be reduced.

The startup was founded last year by former Senior Director of Product Management at Twitter, Sara Beykpour, who worked on products like Twitter Blue, Twitter Video, and conversations, and had spearheaded the experimental app, twttr. She had been at Twitter from 2015 through 2021, growing her position from software engineering to that of a senior director of product management. Her co-founder is a former senior engineer at both Twitter and Tesla, Marcel Molina.

The premise behind Particle, as Beykpour explained last month, is to make it easier to keep up with news using AI.

“Sometimes it feels like headlines are all we have time for. We also want to understand more, but faster,” she wrote in an introduction to the startup on Threads. “We’re in the early stages of using AI to transform the way we interact with news.”

Using Particle, news readers are offered a quick, bulleted summary of the story, with information pulled from a variety of sources. However, when announcing the private beta, Beykpour noted that readers can either use the summary to get up to speed or can choose to go deeper to “learn about how a story has unfolded over time.”

The venture-backed startup has raised funding from Kindred Ventures and Adverb Ventures, as well as various angel investors, including Twitter and Medium co-founder Ev Williams and Behance founder, Scott Belsky.

Remarked Belsky on X, “Particle has become a daily app for me. It synthesizes the many articles (and angles) on any news topic, surfaces the key points as objectively as possible, and lets you dig further across many dimensions. In the era of abstraction ahead, great example of daily AI,” he wrote.

Particle offers a demo of its technology for logged-out users via its website, where articles are featured along with their summary, timestamp as to when they were last updated, and, in a small section at the bottom, the sources they draw from.

These sources pull from across the political spectrum and include big-name publishers like The New York Times, CNBC, the AP, ABC, CNN, Breitbart, The Guardian, The Washington Post, Politico, Fox News, USA Today, The Daily Caller, New York Post, The Hill, and others. International outlets are also pulled from, when relevant, the demos indicate. However, each bullet point is not linked to its original source or sources, which makes it difficult to fact-check the accuracy of the AI summary without delving into all the articles. (Key terms are, however, linked). We noted, too, that the photograph accompanying a news summary is watermarked with the publisher’s logo.

Image Credits: Particle

The end product will likely differ, given that Particle is just now launching its private beta for testing, and intends to offer a mobile app in the future, as it’s hiring for a Senior iOS Engineer.

A similar model of leveraging a variety of news sources and then employing AI to summarize, was recently employed by Artifact, the now-shuttered startup from Instagram’s co-founders. In its case, Artifact’s team curated the news sources up front based on factors related to their integrity and quality. For example, the outlet had to be quick to make corrections, when wrong, and be transparent about their funding. We’re hoping to talk in more detail about how Particle vets its sources closer to a public launch.

Another AI-powered news app, Bulletin, also recently launched to tackle clickbait along with offering news summaries.

Given the interest in this space, what could make Particle stand out is its founding team. Arriving from Twitter, the co-founders have experienced what a real-time news ecosystem feels like, and have the technical and product experience to build a quality product. Whether or not publishers who feel that AI is eating into their space will feel “fairly compensated,” however, remains to be seen.

Adverb Ventures co-founder and Managing Direct April Underwood praised Particle in a post on LinkedIn about the firm’s investment.

“We got the chance to back them just as we were completing our very first close for Fund 1 — we had to wait for our first capital call to hit to wire them the money!,” she said on Sunday, adding that Adverb closed its $75 million Fund I just a couple of months ago. “Sara and Marcel are the kind of founders we dreamed of backing when we set out to build a new early-stage firm. They are going after a big problem space. They’ve got the skills to tackle big problems at a high level of product quality. And they can attract other talented folks to join them, and together invent a future consumers don’t know to ask for (yet),” Underwood wrote.

Requests for comment were not returned. Particle’s beta sign-up form is here.

Scaling FAIR data sharing in an R&D culture

Interview with Ben Gardner of AstraZeneca

Scaling FAIR data sharing in an R&D culture

Image by Gerd Altmann from Pixabay

Ben Gardner’s experience working with drug discovery teams goes back two decades, when he was a team leader at Pfizer. “We were just transferring information by spreadsheet…. It was supposed to be from a data warehouse with our values in there, but our values were all wrong and never made any sense. So every Friday, I downloaded the week’s chemical structures and inserted them into a spreadsheet, and I sent the spreadsheet across the corridor to our chemists.”

Currently, Gardner is the R&D lead for Data Mesh and Semantic Infrastructure at AstraZeneca. What he’s learned over the past six years at the company is that it’s important not to expect everyone on the team to latch onto web ontology language (OWL), graph thinking, or web semantics. Gardner says to put those concepts and methods under the semantic infrastructure hood.

The findable, accessible, interoperable and reusable (FAIR) data automation capabilities of semantic technologies are powerful, but broader adoption hinges on incremental changes that don’t require lots of new tooling, thinking or workflow changes. For that reason, the FAIR data focus across the AZ R&D teams is on boosting data quality via these controlled vocabularies that are shared and updated via flat tables shared via Snowflake.

For its controlled vocabularies, AZ uses the Simple Knowledge Organization System eXtension for Labels (SKOS-XL). SKOS is a widely used and well-supported W3C standard compatible with OWL and the rest of the W3C semantic stack.

Key to a successful sharing effort is the sufficient, guided use of uniform resource identifiers (URIs) at the right levels of the data layer.

It’s always best to be use case driven, but avoid targeting the biggest problems to begin with. “Off the side of the stage is where we started,” Ben remembers, “so that we were away from the political animals who were trying to chew us up in the early stages. We learned to do stuff to the side and deliver value that way.”

Such an illuminating interview. Hope you enjoy it.

FAIR data podcast with Ben Gardner of AstraZeneca

NVIDIA Partners with Indian Govt to Bring Sovereign AI 

Shanker Trivedi, senior vice president at NVIDIA, had a meeting with Union IT Minister Rajeev Chandrasekhar at MeitY today. The discussions revolved around India’s thriving Digital Public Infrastructure and explored potential areas of partnership with NVIDIA in the realm of Sovereign AI.

“From NVIDIA’s perspective, what is really interesting is the growth of sovereign AI. Digital Public Infrastructure is the basis for building Sovereign AI,” said Trivedi.

“Together with a very strong infrastructure to run these generative AI models built on digital public infrastructure, it’s a huge opportunity for NVIDIA to partner with the government,” he added.

Mr. Shanker Trivedi, SVP @nvidia called on me at my office @GoI_MeitY today. https://t.co/kzjVRyj3qT pic.twitter.com/5YX8QylOCq

— Rajeev Chandrasekhar 🇮🇳 (@Rajeev_GoI) February 26, 2024

“Thankful for a meaningful dialogue with MoS, MEITY, Rajeev Chandrasekhar and Shankar Trivedi on the critical role of Sovereign AI. We are at a juncture where the alignment of AI is with national priorities, health, agriculture & education is indispensable,” said Vishal Dhupar, managing director, Asia South at NVIDIA.

What is Sovereign AI

Jensen Huang, the CEO of the world’s third-most valued company, lately has been on a side quest to spread one meaningful message: Every nation, regardless of its size or resources, must develop its own ‘Sovereign AI’.

He has travelled over a dozen geographies, right from the Indian subcontinent to his latest stop, the UAE, for the recent World Government Summit in Dubai.

As Huang envisions it, sovereign AI isn’t just about having some fancy algorithms humming away in a server room. It’s about a fundamental shift in power, reclaiming national autonomy in generative AI. It’s about ensuring that the decisions made by AI, which will increasingly impact everything from healthcare to defence, reflect the priorities of each individual nation.

Huang is basically helping countries take control of how they can shape and build AI to fit their needs. To make sure that when AI makes decisions about things like healthcare and defence, those decisions match the values and priorities of each. country. He believes AI could change not just how we live but also the character of different nations.

The post NVIDIA Partners with Indian Govt to Bring Sovereign AI appeared first on Analytics India Magazine.

FlowGPT is the wild west of GenAI apps

FlowGPT is the wild west of GenAI apps Kyle Wiggers 12 hours

A few months ago, OpenAI launched the GPT Store, a marketplace where people can create and list AI-powered chatbots customized to perform a number of tasks (e.g. coding, answering trivia questions and so on). The GPT Store is powerful, to be sure. But using it requires using OpenAI’s models and no others, which some chatbot creators — and users — are opposed to doing.

So startups are creating alternatives.

One, FlowGPT, aims to be a sort of “app store” for GenAI models like Google’s Gemini, Anthropic’s Claude, Meta’s Llama 2 and OpenAI’s DALL-E 3 as well as front-end experiences for those models (think text fields and prompt suggestions). Through FlowGPT, users can build their own GenAI-powered apps and make them publicly available, earning tips for their contributions.

Jay Dang, a UC Berkeley computer science dropout, and Lifan Wang, a former engineering manager at Amazon, co-founded FlowGPT last year out of a shared desire to create a platform where people could quickly spin up — and share — GenAI apps.

“There’s still a learning curve for users to use AI,” Dang told TechCrunch in an email interview. “FlowGPT is making the bar lower in each iteration, making it more accessible.”

Dang describes FlowGPT as an “ecosystem” for GenAI-powered apps — a collection of infrastructure and creator tools tied to a marketplace and community of GenAI app users. Users get a feed of apps and app collections recommended to them based on trending categories (e.g. “Creative,” “Programming,” “Game” “Academic”), while creators get options for customizing the behavior — and appearance — of GenAI apps.

Users interact with GenAI apps on FlowGPT through a chat window that’s not dissimilar to ChatGPT, with options to type in prompts, thumbs-up (or thumbs-down) apps, share links to conversations or tip individual app creators. Each app has a creator-provided description along with the date it was created, how many times it’s been used and the model the creator recommends to power it.

I say model the creator recommends because FlowGPT apps really, at their core, are prompts — prompts that prime models to respond in certain ways. For example, the “Scared Girl from Horror Movie” app instructs ChatGPT to narrate — as the title hints at — a horror story involving one scared girl. “TitleTuner” prompts ChatGPT to optimize headlines so they rank better on search engines. And SchoolGPT leverages ChatGPT for step-by-step solutions to math, physics and chemistry problems.

FlowGPT

Image Credits: FlowGPT

You’ll notice the heavy reliance on ChatGPT. Use FlowGPT long enough, and you’ll also notice that many of the prompts break when the model’s switched from the default.

Sometimes it’s a matter of the selected model not having the right capabilities. Other times, the prompt runs up against a model’s filters and safeguards.

On the subject of safeguards…

Some of FlowGPT’s most popular apps are essentially jailbreaks designed to circumvent models’ safety measures. There’s multiple versions of DAN on the marketplace — “DAN” being a popular prompting method used to get models to responding to prompts unbounded by their usual rules. Elsewhere, there’s apps like WormGPT, which purports to be able to code malware (and link to paid, dark web versions of the chatbot that do more), and dating simulators that run afoul of OpenAI’s rules against fostering romantic companionship.

Many of these apps could potentially cause harm, like therapy apps and apps that advertise themselves as authoritative health resources. GenAI models like ChatGPT are a notoriously bad health advice givers, with one study showing that an earlier version of ChatGPT rarely provided referrals to specific resources for help relating to suicide, addiction and sexual assault.

Any app on FlowGPT that offends — say, for giving instructions on how to generate deepfake nudes with an AI image generator (and there’s several that do this) — can be reported to the platform’s community manager for review. And FlowGPT does offer a toggle for “sensitive content.”

FlowGPT

Image Credits: FlowGPT

But it’s clear from looking at the homepage that FlowGPT has a moderation problem. It’s the wild west of GenAI apps — and the toggle’s ineffective to the point where I barely notice a difference in app selection with it switched on.

Dang swears up and down that FlowGPT is in fact an ethical and rule-abiding platform, with risk mitigation policies in place aimed at “ensur[ing] public safety.”

“We’re proactively engaging with leading experts in the field of AI ethics,” he said. “Our collaboration is focused on developing comprehensive strategies to minimize risks associated with AI deployment.”

Considering that this writer got a FlowGPT app to give instructions on selling drugs and robbing a bank, I’d say that the company has some work to do.

Investors seemingly feel otherwise.

This week, Goodwater announced that it led a $10 million “pre-Series A” round in FlowGPT with participation from existing backer DCM. Goodwater partner Coddy Johnson, speaking to TechCrunch via email, said that he sees FlowGPT “helping to lead the way” in GenAI by offering “the widest choice” and “the most flexibility and freedom” to both users and creators.

“We believe the biggest future for AI is in open ecosystems,” Johnson added. “FlowGPT [is letting] creators to choose their models and collaborate with their communities.”

FlowGPT

Image Credits: FlowGPT

I’m not so sure that all the maintainers of the models FlowGPT’s tapping — particularly those who’ve pledged to make AI safety a top priority — will share in that enthusiasm.

Nevertheless and in lieu of repercussions from said vendors (at least as of publication time), FlowGPT — which isn’t revenue-generating yet — is laying the groundwork for expansion. The company’s beta testing apps for Android and iOS that’ll bring a revamped FlowGPT experience to mobile, working on a revenue-sharing model for app creators and recruiting to grow its Berkely-based 10-person team, Dang said.

“With millions of monthly users and a fast growth rate, we’ve already proven we are on the right track, and we believe it’s the time to accelerate the progress,” he continued. “We are setting a new standard for immersion in
AI-driven environments, offering a world where creativity knows no bounds … [O]ur mission remains to cultivate a more open and creator-focused platform.”

We’ll see how far that gets it.

Mobile-Agents: Autonomous Multi-modal Mobile Device Agent With Visual Perception

The advent of Multimodal Large Language Models (MLLM) has ushered in a new era of mobile device agents, capable of understanding and interacting with the world through text, images, and voice. These agents mark a significant advancement over traditional AI, providing a richer and more intuitive way for users to interact with their devices. By leveraging MLLM, these agents can process and synthesize vast amounts of information from various modalities, enabling them to offer personalized assistance and enhance user experiences in ways previously unimaginable.

These agents are powered by state-of-the-art machine learning techniques and advanced natural language processing capabilities, allowing them to understand and generate human-like text, as well as interpret visual and auditory data with remarkable accuracy. From recognizing objects and scenes in images to understanding spoken commands and analyzing text sentiment, these multimodal agents are equipped to handle a wide range of inputs seamlessly. The potential of this technology is vast, offering more sophisticated and contextually aware services, such as virtual assistants attuned to human emotions and educational tools that adapt to individual learning styles. They also have the potential to revolutionize accessibility, making technology more approachable across language and sensory barriers.

In this article, we will be talking about Mobile-Agents, an autonomous multi-modal device agent that first leverages the ability of visual perception tools to identify and locate the visual and textual elements with a mobile application’s front-end interface accurately. Using this perceived vision context, the Mobile-Agent framework plans and decomposes the complex operation task autonomously, and navigates through the mobile apps through step by step operations. The Mobile-Agent framework differs from existing solutions since it does not rely on mobile system metadata or XML files of the mobile applications, allowing room for enhanced adaptability across diverse mobile operating environments in a vision centric way. The approach followed by the Mobile-Agent framework eliminates the requirement for system-specific customizations resulting in enhanced performance, and lower computing requirements.

Mobile-Agents: Autonomous Multi-modal Mobile Device Agent

In the fast-paced world of mobile technology, a pioneering concept emerges as a standout: Large Language Models, especially Multimodal Large Language Models or MLLMs capable of generating a wide array of text, images, videos, and speech across different languages. The rapid development of MLLM frameworks in the past few years has given rise to a new and powerful application of MLLMs: autonomous mobile agents. Autonomous mobile agents are software entities that act, move, and function independently, without needing direct human commands, designed to traverse networks or devices to accomplish tasks, collect information, or solve problems.

Mobile Agents are designed to operate the user’s mobile device on the bases of the user instructions and the screen visuals, a task that requires the agents to possess both semantic understanding and visual perception capabilities. However, existing mobile agents are far from perfect since they are based on multimodal large language models, and even the current state of the art MLLM frameworks including GPT-4V lack visual perception abilities required to serve as an efficient mobile agent. Furthermore, although existing frameworks can generate effective operations, they struggle to locate the position of these operations accurately on the screen, limiting the applications and ability of mobile agents to operate on mobile devices.

To tackle this issue, some frameworks opted to leverage the user interface layout files to assist the GPT-4V or other MLLMs with localization capabilities, with some frameworks managing to extract actionable positions on the screen by accessing the XML files of the application whereas other frameworks opted to use the HTML code from the web applications. As it can be seen, a majority of these frameworks rely on accessing underlying and local application files, rendering the method almost ineffective if the framework cannot access these files. To address this issue and eliminate the dependency of local agents on underlying files on the localization methods, developers have worked on Mobile-Agent, an autonomous mobile agent with impressive visual perception capabilities. Using its visual perception module, the Mobile-Agent framework uses screenshots from the mobile device to locate operations accurately. The visual perception module houses OCR and detection models that are responsible for identifying text within the screen and describing the content within a specific region of the mobile screen. The Mobile-Agent framework employs carefully crafted prompts and facilitates efficient interaction between the tools and the agents, thus automating the mobile device operations.

Furthermore, the Mobile-Agents framework aims to leverage the contextual capabilities of state of the art MLLM frameworks like GPT-4V to achieve self-planning capabilities that allows the model to plan tasks based on the operation history, user instructions and screenshots holistically. To further enhance the agent’s ability to identify incomplete instructions and wrong operations, the Mobile-Agent framework introduces a self-reflection method. Under the guidance of carefully crafted prompts, the agent reflects on incorrect and invalid operations consistently, and halts the operations once the task or instruction has been completed.

Overall, the contributions of the Mobile-Agent framework can be summarized as follows:

  1. Mobile-Agent acts as an autonomous mobile device agent, utilizing visual perception tools to carry out operation localization. It methodically plans each step and engages in introspection. Notably, Mobile-Agent relies exclusively on device screenshots, without the use of any system code, showcasing a solution that's purely based on vision techniques.
  2. Mobile-Agent introduces Mobile-Eval, a benchmark designed to evaluate mobile-device agents. This benchmark includes a variety of the ten most commonly used mobile apps, along with intelligent instructions for these apps, categorized into three levels of difficulty.

Mobile-Agent : Architecture and Methodology

At its core, the Mobile-Agent framework consists of a state of the art Multimodal Large Language Model, the GPT-4V, a text detection module used for text localization tasks. Along with GPT-4V, Mobile-Agent also employs an icon detection module for icon localization.

Visual Perception

As mentioned earlier, the GPT-4V MLLM delivers satisfactory results for instructions and screenshots, but it fails to output the location effectively where the operations take place. Owing to this limitation, the Mobile-Agent framework implementing the GPT-4V model needs to rely on external tools to assist with operation localization, thus facilitating the operations output on the mobile screen.

Text Localization

The Mobile-Agent framework implements a OCR tool to detect the position of the corresponding text on the screen whenever the agent needs to tap on a specific text displayed on the mobile screen. There are three unique text localization scenarios.

Scenario 1: No Specified Text Detected

Issue: The OCR fails to detect the specified text, which may occur in complex images or due to OCR limitations.

Response: Instruct the agent to either:

  • Reselect the text for tapping, allowing for a manual correction of the OCR's oversight, or
  • Choose an alternative operation, such as using a different input method or performing another action relevant to the task at hand.

Reasoning: This flexibility is necessary to manage the occasional inaccuracies or hallucinations of GPT-4V, ensuring the agent can still proceed effectively.

Scenario 2: Single Instance of Specified Text Detected

Operation: Automatically generate an action to click on the center coordinates of the detected text box.

Justification: With only one instance detected, the likelihood of correct identification is high, making it efficient to proceed with a direct action.

Scenario 3: Multiple Instances of Specified Text Detected

Assessment: First, evaluate the number of detected instances:

Many Instances: Indicates a screen cluttered with similar content, complicating the selection process.

Action: Request the agent to reselect the text, aiming to refine the selection or adjust the search parameters.

Few Instances: A manageable number of detections allows for a more nuanced approach.

Action: Crop the regions around these instances, expanding the text detection boxes outward to capture additional context. This expansion ensures that more information is preserved, aiding in decision-making.

Next Step: Draw detection boxes on the cropped images and present them to the agent. This visual assistance helps the agent in deciding which instance to interact with, based on contextual clues or task requirements.

This structured approach optimizes the interaction between OCR results and agent operations, enhancing the system's reliability and adaptability in handling text-based tasks across various scenarios. The entire process is demonstrated in the following image.

Icon Localization

The Mobile-Agent framework implements an icon detection tool to locate the position of an icon when the agent needs to click on it on the mobile screen. To be more specific, the framework first requests the agent to provide specific attributes of the image including shape and color, and then the framework implements the Grounding DINO method with the prompt icon to identify all the icons contained within the screenshot. Finally, Mobile-Agent employs the CLIP framework to calculate the similarity between the description of the click region, and calculates the similarity between the deleted icons, and selects the region with the highest similarity for a click.

Instruction Execution

To translate the actions into operations on the screen by the agents, the Mobile-Agent framework defines 8 different operations.

  • Launch Application (App Name): Initiate the designated application from the desktop interface.
  • Tap on Text (Text Label): Interact with the screen portion displaying the label “Text Label”.
  • Interact with Icon (Icon Description, Location): Target and tap the specified icon area, where “Icon Description” details attributes like color and shape of the icon. Choose “Location” from options such as top, bottom, left, right, or center, possibly combining two for precise navigation and to reduce mistakes.
  • Enter Text (Input Text): Input the given “Input Text” into the active text field.
  • Scroll Up & Down: Navigate upwards or downwards through the content of the present page.
  • Go Back: Revert to the previously viewed page.
  • Close: Navigate back to the desktop directly from the current screen.
  • Halt: Conclude the operation once the task is accomplished.

Self-Planning

Every step of the operation is executed iteratively by the framework, and before the beginning of each iteration, the user is required to provide an input instruction, and the Mobile-Agent model uses the instruction to generate a system prompt for the entire process. Furthermore, before the start of every iteration, the framework captures a screenshot and feeds it to the agent. The agent then observes the screenshot, operation history, and system prompts to output the next step of the operations.

Self-Reflection

During its operations, the agent might face errors that prevent it from successfully executing a command. To enhance the instruction fulfillment rate, a self-evaluation approach has been implemented, activating under two specific circumstances. Initially, if the agent executes a flawed or invalid action that halts progress, such as when it recognizes the screenshot remains unchanged post-operation or displays an incorrect page, it will be directed to consider alternative actions or adjust the existing operation's parameters. Secondly, the agent might miss some elements of a complex directive. Once the agent has executed a series of actions based on its initial plan, it will be prompted to review its action sequence, the latest screenshot, and the user's directive to assess whether the task has been completed. If discrepancies are found, the agent is tasked to autonomously generate new actions to fulfill the directive.

Mobile-Agent : Experiments and Results

To evaluate its abilities comprehensively, the Mobile-Agent framework introduces the Mobile-Eval benchmark consisting of 10 commonly used applications, and designs three instructions for each application. The first operation is straightforward, and only covers basic application operations whereas the second operation is a bit more complex than the first as it has some additional requirements. Finally, the third operation is the most complex of them all since it contains abstract user instruction with the user not explicitly specifying which app to use or what operation to perform.

Moving along, to assess the performance from different perspectives, the Mobile-Agent framework designs and implements 4 different metrics.

  • Su or Success: If the mobile-agent completes the instructions, it is considered to be a success.
  • Process Score or PS: The Process Score metric measures the accuracy of each step during the execution of the user instructions, and it is calculated by dividing the number of correct steps by the total number of steps.
  • Relative Efficiency or RE: The relative efficiency score is a ratio or comparison between the number of steps it takes a human to perform the instruction manually, and the number of steps it takes the agent to execute the same instruction.
  • Completion Rate or CR: The completion rate metric divides the number of human-operated steps that the framework completes successfully with the total number of steps taken by a human to complete the instruction. The value of CR is 1 when the agent completes the instruction successfully.

The results are demonstrated in the following figure.

Initially, for the three given tasks, the Mobile-Agent attained completion rates of 91%, 82%, and 82%, respectively. While not all tasks were executed flawlessly, the achievement rates for each category of task surpassed 90%. Furthermore, the PS metric reveals that the Mobile-Agent consistently demonstrates a high likelihood of executing accurate actions for the three tasks, with success rates around 80%. Additionally, according to the RE metric, the Mobile-Agent exhibits an 80% efficiency in performing operations at a level comparable to human optimality. These outcomes collectively underscore the Mobile-Agent's proficiency as a mobile device assistant.

The following figure illustrates the Mobile-Agent's capability to grasp user commands and independently orchestrate its actions. Even in the absence of explicit operation details in the instructions, the Mobile-Agent adeptly interpreted the user's needs, converting them into actionable tasks. Following this understanding, the agent executed the instructions via a systematic planning process.

Final Thoughts

In this article we have talked about Mobile-Agents, a multi-modal autonomous device agent that initially utilizes visual perception technologies to precisely detect and pinpoint both visual and textual components within the interface of a mobile application. With this visual context in mind, the Mobile-Agent framework autonomously outlines and breaks down the intricate tasks into manageable actions, smoothly navigating through mobile applications step by step. This framework stands out from existing methodologies as it does not depend on the mobile system's metadata or the mobile apps' XML files, thereby facilitating greater flexibility across various mobile operating systems with a focus on visual-centric processing. The strategy employed by the Mobile-Agent framework obviates the need for system-specific adaptations, leading to improved efficiency and reduced computational demands.

Mistral’s ‘Le Big Model’ Beats Google’s Gemini Pro, Signs Multi-Year Deal with Microsoft

Mistral AI to Raise $487 Mn Nearing $2 Bn Valuation

Mistral AI today released Mistral Large, its latest and most advanced language model. It is accessible through La Plateforme and Microsoft Azure, marking a strategic distribution partnership with Microsoft.

Mistral Large achieves strong results on commonly used benchmarks, making it the world’s second-ranked model generally available through an API (next to GPT-4) beating Google’s Gemini Pro and Anthropic’s Claude.

The model demonstrates advanced multilingual capabilities, fluently understanding English, French, Spanish, German, and Italian. Its 32K tokens context window allows precise information recall from extensive documents, enhancing its usability for complex multilingual reasoning tasks, including text understanding, transformation, and code generation.

Mistral Large has native multi-lingual capacities. It strongly outperforms LLaMA 2 70B on HellaSwag, Arc Challenge and MMLU benchmarks in French, German, Spanish and Italian.

Alongside Mistral Large, Mistral AI has also introduced Mistral Small, an optimised model designed for low latency workloads. Outperforming Mixtral 8x7B and featuring lower latency, Mistral Small offers a refined solution between Mistral’s open-weight offering and its flagship model.

Mistral AI has streamlined its endpoint offerings, providing open-weight endpoints with competitive pricing and introducing new optimized model endpoints – mistral-small-2402 and mistral-large-2402. The company aims to offer users a comprehensive view of performance/cost tradeoffs.

Introducing JSON format mode, Mistral AI allows developers to obtain model output in a structured and valid JSON format. Additionally, the model supports function calling, enabling more intricate interactions with internal code, APIs, or databases. Currently, function calling and JSON format are only available on mistral-small and mistral-large.

Multi-Year Partnership with Microsoft

Microsoft announced a multi-year partnership with Mistral AI. Microsoft’s partnership with Mistral focuses on three core areas- Supercomputing infrastructure, Scale to Market and AI research and development.

“We’re announcing a multi-year partnership with MistralAI, as we build on our commitment to offer customers the best choice of open and foundation models on Azure,” wrote Microsoft chief Satya Nadella.

Microsoft will provide Mistral AI with access to Azure AI supercomputing infrastructure, ensuring superior performance and scalability for AI training and inference workloads.

The collaboration aims to make Mistral AI’s premium models accessible to customers through Models as a Service (MaaS) in the Azure AI Studio and Azure Machine Learning model catalog. Customers can use Microsoft Azure Consumption Commitment (MACC) for purchasing Mistral AI’s models, enhancing global availability.

Further, Microsoft and Mistral AI will explore collaboration in training purpose-specific models for select customers, focusing on European public sector workloads.

The post Mistral’s ‘Le Big Model’ Beats Google’s Gemini Pro, Signs Multi-Year Deal with Microsoft appeared first on Analytics India Magazine.

5 critical metrics every data scientist should monitor in hybrid cloud environments

Critical Metrics Every Data Scientist Should Monitor in Hybrid Cloud Environments

Experienced data scientists will find it helpful to think of hybrid cloud environments as a kind of high-tech ecosystem—complex and full of pitfalls that could swallow you whole if you’re not careful.

In this context, keeping tabs on key metrics isn’t just helpful; it’s your secret to making sure everything runs smoother than ever. Here are just five that need to be on your radar.

Latency labyrinth: Navigating the secret sathways

Imagine you’re in a maze, where every twist and turn could either lead you closer to sweet, sweet data or into a dead end of sluggish performance—welcome to the Latency Labyrinth.

To steer clear of these dead ends, data scientists gotta keep their eyes peeled on network latency like hawks. Why? Because even microseconds of delay can throw a wrench in your predictive models or real-time analytics.

To dodge these delays and optimize response times, savvy pros are using tools like SolarWinds hybrid cloud monitoring. These platforms help pinpoint where those pesky bottlenecks hide so that you can streamline data flow and keep everything humming along nicely.

Error rates: The silent alarms of hybrid cloud environments

Spotting errors in hybrid cloud environments is like trying to find a sneaky gremlin—it’s wreaking havoc, but it’s darn good at hiding. Elevated error rates are like silent alarms ringing throughout your system; ignore them at your peril. They’re the red flags signaling buggy code, integration snafus, security flaws, or even more complex problems with your data pipelines.

It’s awesome when you can jump in and fix issues before they blow up into bigger problems. Being proactive means less downtime and better service for customers—which let’s face it, that’s the name of the game.

Whether you’re troubleshooting an API acting wonky or tracking down some quirky back-end issue that just popped up, using insights from platforms can help identify those glitches early on so you can squash ’em flat and move on with your day.

Throughput throttle: Keeping the data freeway wide open

Cruising through a data freeway, throughput is your speedometer. Too much traffic and your work grinds to a halt; too little and you’re not pushing the limits of what’s possible. It’s all about striking that sweet balance – ensuring that data moves like it’s got a green light on every block.

Data scientists looking to avoid congestion use tools and methodologies to ensure their systems aren’t just flashing warning lights but actually guiding them down the quickest route possible. That means no unnecessary pit stops or idling around—just pure, seamless data flow. Feeling the rush of tons of data processed efficiently is legit one of those small wins in life that add up big time.

Resource rodeo: Wrangling your cloud resources

It’s like throwing a lasso around your cloud resources—you want to catch just the right amount. Resource Utilization is your rodeo show, where you’re scoring points based on how effectively you’re using what you’ve got. CPU, memory, storage setup – these are the wild stallions that’ll buck if they aren’t managed with a careful hand.

You don’t wanna be the one caught overspending on resources you ain’t even using or gasping at performance issues ’cause your server’s as packed as a clown car. Keeping an eye on usage metrics ensures that not only are costs kept in check but also that your applications are zipping along without tripping over their own virtual feet. By staying tuned-in and tweaking things here and there, you’ll keep it all riding smoothly—no cowboy hat necessary!

Security sentries: Guarding the data castle

Your cloud castle is stocked with precious data jewels, and without a doubt, you need top-notch security sentries keeping watch. Tracking security threats isn’t just about slapping on armor; it’s about recognizing the subtle whispers of danger before they become shouts.

If there’s one reality every data wizard knows, it’s that threats evolve faster than viral memes. So what’s your move? Keeping vigilant by monitoring authentication attempts, access patterns, and network traffic for signs of suspicious behavior. Think of it like setting up traps for cyber goblins trying to sneak into your treasure vault—stay sharp and they won’t stand a chance. This metric isn’t glamorous but man, is it crucial for peace in the realm (and peace of mind).

The last word

So there you go, those five metrics are imperative for data scientists to master in hybrid cloud environments. Keeping a close watch on these can seriously give you an edge, making sure your cloud strategy is solid and your analytics are spot on.

Mistral’s ‘Le Big Model’ Beats Google’s Gemini Pro

Mistral AI to Raise $487 Mn Nearing $2 Bn Valuation

Mistral AI today released Mistral Large, its latest and most advanced language model. It is accessible through La Plateforme and Microsoft Azure, marking a strategic distribution partnership with Microsoft.

Mistral Large achieves strong results on commonly used benchmarks, making it the world’s second-ranked model generally available through an API (next to GPT-4) beating Google’s Gemini Pro and Anthropic’s Claude.

The model demonstrates advanced multilingual capabilities, fluently understanding English, French, Spanish, German, and Italian. Its 32K tokens context window allows precise information recall from extensive documents, enhancing its usability for complex multilingual reasoning tasks, including text understanding, transformation, and code generation.

Mistral Large has native multi-lingual capacities. It strongly outperforms LLaMA 2 70B on HellaSwag, Arc Challenge and MMLU benchmarks in French, German, Spanish and Italian.

Alongside Mistral Large, Mistral AI has also introduced Mistral Small, an optimised model designed for low latency workloads. Outperforming Mixtral 8x7B and featuring lower latency, Mistral Small offers a refined solution between Mistral’s open-weight offering and its flagship model.

Mistral AI has streamlined its endpoint offerings, providing open-weight endpoints with competitive pricing and introducing new optimized model endpoints – mistral-small-2402 and mistral-large-2402. The company aims to offer users a comprehensive view of performance/cost tradeoffs.

Introducing JSON format mode, Mistral AI allows developers to obtain model output in a structured and valid JSON format. Additionally, the model supports function calling, enabling more intricate interactions with internal code, APIs, or databases. Currently, function calling and JSON format are only available on mistral-small and mistral-large.

The post Mistral’s ‘Le Big Model’ Beats Google’s Gemini Pro appeared first on Analytics India Magazine.

Darwin AI gives small LatAm companies AI-powered sales assistant

Darwin AI gives small LatAm companies AI-powered sales assistant Christine Hall 10 hours

Smaller companies are just as eager to use AI tech to supercharge their sales processes as their bigger competitors. However, they often lack the in-house IT expertise or the capital to implement enterprise-level tools, like OpenAI or Anthropic.

Darwin AI, a Brazil-based AI startup, is developing a conversational AI assistant for small businesses across Latin America who want to get into AI, but don’t have an IT staff. The assistant is designed to interact with customers in a more human-like manner to help generate more revenue. Should the conversation escalate, either negatively or become a sales lead, it will bring in a human to continue the conversation.

This is the second company for Lautaro Schiaffino and Ezequiel Sculli, who previously co-founded Sirena.app, a shared inbox tool for WhatsApp for mid-market companies. They grew the company to an annual recurring revenue of $15 million and presence in 25 countries before selling to Zenvia in 2020.

Schiaffino and Sculli left Zenvia in 2022 and started talking about what they were going to do next. They knew they wanted to help the same market and to take complex technology and simplify it for those without the technical team or time and skill to develop something on their own.

“AI is a great opportunity, but difficult to implement for small businesses,” Schiaffino told TechCrunch. “We saw the evolution of the mid market, as business-to-consumer in Latin America realized there are lots of leads, but the conversion rate is low. People have to talk to 100 or 200 leads to sell one product.”

Rasa, an enterprise-focused dev platform for conversational GenAI, raises $30M

So Schiaffino and Sculli began developing Darwin AI’s system that connects with a company’s customer relationship management tool and evaluates possible sales leads and escalates the ones, most likely to buy, to a human salesperson. Using AI, Darwin takes into account the needs of companies and then filters leads and customers. Schiaffino described it as a two-sided system that “talks” to employees and customers to advance the most important leads.

As more companies implement automation into their processes, the conversational AI market is expected to grow over 20% annually through 2030. That’s attracted a number of startups wanting to solve this for enterprise, most recently Rasa, Kore.ai, DXwand and OpenDialog.

Meanwhile, the company continues to configure the AI by onboarding more users. Darwin is also close to implementing a self-learning AI function that will get a company up-and-running in a matter of days without the need for a special IT team.

Since launching in 2023, Darwin has processed thousands of conversations and has customers in countries, including Mexico, Peru, Argentina, Brazil and Colombia. The company is on track to reach over 1 million conversations this year. It also has an integration with Zapier and can connect to regional CRMs.

The company has brought in revenue since the beginning. In fact, the founders realized they had a hit when customers were paying before there was even a user interface, Schiaffino said, though he declined to say how much revenue Darwin is generating. Darwin has a setup fee, a monthly fixed plan and also charges a usage fee per conversation. Additional plan tiers are coming this year.

Darwin has raised $2.5 million total including pre-seed and a seed round. Canary led the most recent round of $2.1 million and was joined by H20 Capital Innovation, Dalus Capital, FJ Labs, and Latitude Capital.

“The money funds will be for product development, go-to-market and also the operations teams to guarantee the quality of the product,” Schiaffino said.

Why Latin American SaaS startups are different from their US peers

Meta Reveals Strategy for the 2024 EU Parliament Elections

As the 2024 EU Parliament elections approach, the role of digital platforms in influencing and safeguarding the democratic process has never been more prominent. Amidst this backdrop, Meta, the company behind major social platforms like Facebook and Instagram, has outlined a series of initiatives aimed at ensuring the integrity of these elections.

Marco Pancini, Meta's Head of EU Affairs, has detailed these strategies in a company blog, reflecting the company's recognition of its influence and responsibilities in the digital political landscape.

Establishing an Elections Operations Center

In preparation for the EU elections, Meta has announced the establishment of a specialized Elections Operations Center. This initiative is designed to monitor and respond to potential threats that could impact the integrity of the electoral process on its platforms. The center aims to be a hub of expertise, combining the skills of professionals from various departments within Meta, including intelligence, data science, engineering, research, operations, content policy, and legal teams.

The purpose of the Elections Operations Center is to identify potential threats and implement mitigations in real time. By bringing together experts from diverse fields, Meta aims to create a comprehensive response mechanism to safeguard against election interference. The approach taken by the Operations Center is based on lessons learned from previous elections and is tailored to the specific challenges of the EU political environment.

Fact-Checking Network Expansion

As part of its strategy to combat misinformation, Meta is also expanding its fact-checking network within Europe. This expansion includes the addition of three new partners in Bulgaria, France, and Slovakia, enhancing the network's linguistic and cultural diversity. The fact-checking network plays a crucial role in reviewing and rating content on Meta's platforms, providing an additional layer of scrutiny to the information disseminated to users.

The operation of this network involves independent organizations that assess the accuracy of content and apply warning labels to debunked information. This process is designed to reduce the spread of misinformation by limiting its visibility and reach. Meta's expansion of the fact-checking network is an effort to bolster these safeguards, particularly in the context of the highly charged political environment of an election.

Long-Term Investment in Safety and Security

Since 2016, Meta has consistently increased its investment in safety and security, with expenditures surpassing $20 billion. This financial commitment underscores the company’s ongoing effort to enhance the security and integrity of its platforms. The significance of this investment lies in its scope and scale, reflecting Meta’s response to the evolving challenges in the digital landscape.

Accompanying this financial investment is the substantial growth of Meta’s global team dedicated to safety and security. This team has expanded fourfold, now comprising approximately 40,000 personnel. Among these, 15,000 are content reviewers who play a critical role in overseeing the vast array of content across Meta’s platforms, including Facebook, Instagram, and Threads. These reviewers are equipped to handle content in more than 70 languages, encompassing all 24 official EU languages. This linguistic diversity is crucial for effectively moderating content in a region as culturally and linguistically varied as the European Union.

This long-term investment and team expansion are integral components of Meta’s strategy to safeguard its platforms. By allocating significant resources and personnel, Meta aims to address the challenges posed by misinformation, influence operations, and other forms of content that could potentially undermine the integrity of the electoral process. The effectiveness of these investments and efforts is a subject of public and academic scrutiny, but the scale of Meta’s commitment in this area is evident.

Countering Influence Operations and Inauthentic Behavior

Meta's strategy to safeguard the integrity of the EU Parliament elections extends to actively countering influence operations and coordinated inauthentic behavior. These operations, often characterized by strategic attempts to manipulate public discourse, represent a significant challenge in maintaining the authenticity of online interactions and information.

To combat these sophisticated tactics, Meta has developed specialized teams whose focus is to identify and disrupt coordinated inauthentic behavior. This involves scrutinizing the platform for patterns of activity that suggest deliberate efforts to deceive or mislead users. These teams are responsible for uncovering and dismantling networks engaged in such deceptive practices. Since 2017, Meta has reported the investigation and removal of over 200 such networks, a process openly shared with the public through their Quarterly Threat Reports.

In addition to tackling covert operations, Meta also addresses more overt forms of influence, such as content from state-controlled media entities. Recognizing the potential for government-backed media to carry biases that could influence public opinion, Meta has implemented a policy of labeling content from these sources. This labeling aims to provide users with context about the origin of the information they are consuming, enabling them to make more informed judgments about its credibility.

These initiatives form a critical part of Meta's broader strategy to preserve the integrity of the information ecosystem on its platforms, particularly in the politically sensitive context of elections. By publicly sharing information about threats and labeling state-controlled media, Meta seeks to enhance transparency and user awareness regarding the authenticity and origins of content.

Addressing GenAI Technology Challenges

Meta is also confronting the challenges posed by Generative AI (GenAI) technologies, especially in the context of content generation. With the increasing sophistication of AI in creating realistic images, videos, and text, the potential for misuse in the political sphere has become a significant concern.

Meta has established policies and measures specifically targeting AI-generated content. These policies are designed to ensure that content on their platforms, whether created by humans or AI, adheres to community and advertising standards. In situations where AI-generated content violates these standards, Meta takes action to address the issue, which may include removal of the content or reduction in its distribution.

Furthermore, Meta is developing tools to identify and label AI-generated images and videos. This initiative reflects an understanding of the importance of transparency in the digital ecosystem. By labeling AI-generated content, Meta aims to provide users with clear information about the nature of the content they are viewing, enabling them to make more informed assessments of its authenticity and reliability.

The development and implementation of these tools and policies are part of Meta’s broader response to the challenges posed by advanced digital technologies. As these technologies continue to advance, the company's strategies and tools are expected to evolve in tandem, adapting to new forms of digital content and potential threats to information integrity.