CoRover.ai is the Silent Winner of Indian LLM Race

CoRover.ai is the Silent Winner of Indian LLM Race

Ankush Sabharwal, the co-founder of CoRover.ai, has had a busy few months building BharatGPT. Most recently, his company launched an educational tablet called Milkyway, which would be powered by CoRover’s BharatGPT virtual assistant, including video and chatbots for students.

Notwithstanding a packed schedule, Sabharwal spoke to AIM about how CoRover.ai’s BharatGPT was built and what exactly it offers that the government is on board to implement in its services. Starting its AI journey in 2016, CoRover has been building virtual assistants for various partners and government agencies such as IRCTC, MaxLife, Chennai Police, and LIC, to name a few.

“Ours is the only platform that supports informational, transactional, advisory, and multilingual support of 14 Indian languages, including audio, video, and text format,” emphasised Sabharwal. He claims that over 130 crore users have used its virtual assistant not directly, but through their customers.

‘Why build from scratch?’

Right from the start, Sabharwal said that they had heaps of multilingual data from their customers and were able to build simple assistants for them all these years. In the beginning, CoRover.ai used Microsoft’s Gordon model which was very small but useful enough to build these assistants.

“Now with generative AI, our data became very useful to build LLM-based chatbots that could solve larger problems,” said Sabharwal. He said that with the success of ChatGPT, the company started fine-tuning Instruct GPT-based Pythia model, which is a 6.9 billion parameter open source model by Allen AI Institute. This was after they experimented with several open source models.

Sabharwal said that CoRover.ai’s BharatGPT is used only to power other virtual assistants, and they do not charge extra for building their own models. “You can come with your data and we can build an assistant for your business with very simple steps,” said Sabharwal. He said that this includes assistants in 14 languages including voice.

“The luxury we had while building the model was that in almost all of our contracts, we said that we would also have the right to collect the data from our customers,” he said, talking about how difficult it is to collect Indic language data. “Most of this is anonymous data, so that is not a problem.”

“We are building something out of open source, fine-tuning it on our data, and making it proprietary,” said Sabharwal. He added that the company has been looking to acquire more GPUs as they have already finished building their first model on open source and aim to build a foundational model soon. “We have the data which is unique and would be able to create a foundational model as soon as we have the resources,” said Sabharwal.

The best way forward

Sabharwal, giving an example of Java script, said that there is not exactly a requirement for building models from scratch. “We have all these foundational models that are good enough to build virtual assistants for specific use cases. That is what we are doing by building models for specific use cases, instead of a broad generalised model,” said Sabharwal, highlighting how Corover.ai believes in solving focused and accurate problems in different sectors and domains.

“When Java was launched, there was no urge to create another Java,” explained Sabharwal, saying that we have enough technology to implement LLMs in useful ways and solve problems in useful ways. “India is the biggest producer of data and if we focus on creating more products with AI, I think we would have solved the problems for India,” he said.

Corover.ai is backed by Google, and leverages its cloud for building LLMs. “Like the government of India’s mission of Make AI in India and Make AI Work for India, our BharatGPT ensures that the data is from India, and remains in India,” highlighted Sabharwal, saying that the company is also currently renting GPUs from Google for building its models.

“We have over 400 leads from companies in India, Korea, and other parts of the world, and our motto is to provide human centric conversation virtual assistants to the world, which we have been doing since 2016,” said Sabharwal.

“We are agnostic with how we want to implement our technology. The Milkyway tablet is the first step towards integrating it into hardware devices,” concluded Sabharwal.

The post CoRover.ai is the Silent Winner of Indian LLM Race appeared first on Analytics India Magazine.

OpenAI’s Sora Generates Photorealistic Videos

OpenAI released on Feb. 15 an impressive new text-to-video model called Sora that can create photorealistic or cartoony moving images from natural language text prompts. Sora isn’t available to the public yet; instead, OpenAI released Sora to red teamers — security researchers who mimic techniques used by threat actors — to assess possible harms or risks. OpenAI also offered Sora to selected designers and audio and visual artists to get feedback on how Sora can best be optimized for creative work.

OpenAI’s emphasis on safety around Sora is standard for generative AI nowadays, but it also shows the importance of precautions when it comes to AI that could be used to create convincing fake images, which could, for instance, damage an organization’s reputation.

What is Sora?

Sora is a generative AI diffusion model. Sora can generate multiple characters, complex backgrounds and realistic-looking movements in videos up to a minute long. It can create multiple shots within one video, keeping the characters and visual style consistent, allowing Sora to be an effective storytelling tool.

In the future, Sora could be used to generate videos to accompany content, to promote content or products on social media, or to illustrate points in presentations for businesses. While it shouldn’t replace the creative minds of professional video makers, Sora could be used to make some content more quickly and easily. While there’s no information on pricing yet, it’s possible OpenAI will eventually have an option to incorporate Sora into its ChatGPT Enterprise subscription.

“Media and entertainment will be the vertical industry that may be early adopters of models like these,” Gartner Analyst and Distinguished VP Arun Chandrasekaran Chandrasekaran told TechRepublic in an email. “Business functions such as marketing and design within technology companies and enterprises could also be early adopters.”

How do I access Sora?

Unless you have already received access from OpenAI as part of its red teaming or creative work beta testing, it’s not possible to access Sora now. OpenAI released Sora to selected visual artists, designers and filmmakers to learn how to optimize Sora for creative uses specifically. In addition, OpenAI has given access to red team researchers specializing in misinformation, hateful content and bias. Gartner Analyst and Distinguished VP Arun Chandrasekaran said OpenAI’s initial release of Sora is “a good approach and consistent with OpenAI’s practices on safe release of models.”

“Of course, this alone won’t be sufficient, and they need to put in practices to weed out bad actors getting access to these models or nefarious uses of it,” Chandrasekaran said.

How does Sora work?

Sora is a diffusion model, meaning it gradually refines a nonsense image into a comprehensible one based on the prompt, and uses a transformer architecture. The research OpenAI performed to create its DALL-E and GPT models — particularly the recapturing technique from DALL-E — were stepping stones to Sora’s creation.

SEE: AI engineers are in demand in the U.K. (TechRepublic)

Sora videos don’t always look completely realistic

Sora still has trouble telling left from right or following complex descriptions of events that happen over time such as prompts about a specific movement of the camera. Videos created with Sora are likely to be spotted through errors in cause-and-effect, OpenAI said, such as a person taking a bite out of a cookie but not leaving a bite mark.

For instance, interactions between characters may show blurring (especially around limbs) or uncertainty in terms of numbers (e.g., how many wolves are in the below video at any given time?).

What are OpenAI’s safety precautions around Sora?

With the right prompts and tweaking, the videos Sora makes can easily be mistaken for live-action videos. OpenAI is aware of possible defamation or misinformation problems arising from this technology. OpenAI plans to apply the same content filters to Sora as the company does to DALL-E 3 that prevent “extreme violence, sexual content, hateful imagery, celebrity likeness, or the IP of others,” according to OpenAI.

If Sora is released to the public, OpenAI plans to watermark content created with Sora with C2PA metadata; the metadata can be viewed by selecting the image and choosing the File Info or Properties menu options. People who create AI-generated images can still remove the metadata on purpose, or may do so accidentally. OpenAI does not currently have anything in place to prevent users of its image generator, DALL-E 3, from removing metadata.

“It is already [difficult] and increasingly will become impossible to detect AI-generated content by human beings,” Chandrasekaran said. “VCs are making investments in startups building deepfake detection tools, and they (deepfake detection tools) can be part of an enterprise’s armor. However, in the future, there is a need for public-private partnerships to identify, often at the point of creation, machine-generated content.”

What are competitors to Sora?

Sora’s photorealistic videos are quite distinct, but there are similar services. Runway provides ready-for-enterprise text-to-video AI generation. Fliki can create limited videos with voice synching for social media narration. Generative AI can now reliably add content to or edit videos taken in the conventional way as well.

On Feb. 8, Apple researchers revealed a paper about Keyframer, its proposed large language model that can create stylized, animated images.

TechRepublic has reached out to OpenAI for more information about Sora.

Microsoft expands Copilot data protection so more users can chat with ease

copilot

Copilot has proven to be a worthy ChatGPT competitor, boasting the same features as OpenAI's chatbot with extra perks, including GPT-4, internet access, and footnotes. As a result, working professionals and students can greatly benefit from using the chatbot, and this update makes it possible for them to do it safely.

On Tuesday, Microsoft announced its "next wave of commercial data protection availability," allowing more users with eligible work or student Entra ID accounts to access commercial data protection in Copilot just by signing in at no additional cost.

Also: How to use Copilot (formerly called Bing Chat)

Microsoft's commercial data protection addresses major confidentiality issues with using generative AI chatbots, ensuring that prompts and responses are not saved, that the data is not used to train the underlying models, and Microsoft has no eyes-on access to user or chat data, according to the release.

Users will know that the data protection is on because there will be a "Protected" badge next to the user's profile icon, and there is the text that reads "Your personal and company data are protected" above the text box.

Users will also be able to enjoy the additional protection on mobile in the Copilot app, on all major web browsers, in the sidebar using Copilot in Microsoft Edge, and via their computer's taskbar in Windows.

Starting in late February, users with the following licenses will be eligible for commercial data protection in Copilot at no additional cost: Microsoft 365 F1, Microsoft Office 365 E1/ E1 Plus/E3/E5/F3, Microsoft 365 Business Basic, Microsoft 365 Apps for enterprise, and Microsoft 365 Apps for business, according to the release.

Also: How ChatGPT (and other AI chatbots) can help you write an essay

"This next wave of availability significantly expands the number of users eligible for commercial data protection in Copilot, paving the way for organizations and schools to be even more productive and creative," said Microsoft.

Despite the already broad expansion, Microsoft shares its goal is to provide commercial data protection in Copilot to every Entra ID user.

Artificial Intelligence

Poor people need not apply for this dating app

Poor people need not apply for this dating app Haje Jan Kamps 10 hours

Welcome to Startups Weekly — your weekly recap of everything you can’t miss from the world of startups. Sign up here to get it in your inbox every Friday.

I’m having one of those weeks where I’m just constantly, very slowly shaking my head at people. As I sat down to read all the stories on TechCrunch and write the Startups Weekly newsletter, well, things didn’t get better.

Just when you thought the dating scene couldn’t get any more exclusive, along comes Score, the app that says, “Love is in the air . . . but only if you’ve got the credit score to breathe it.” That’s right, folks, in a world where swiping right could mean finding your soul mate or the next person to ghost you, Score ensures that at least they won’t be ghosting you due to your bad credit. Launched by a financial platform (of course — this smells like a marketing stunt), this app is for those who’ve managed to navigate the treacherous waters of adulting with a half-decent credit. Because nothing says true love like a solid financial report, right? But wait, there’s a twist! The app is not just exclusive — it’s temporary. For those who don’t make the cut? Well, they’re sent off to financial literacy boot camp, because nothing heals a bruised ego like being told you’re not financially savvy enough for love.

America, ladies and gentlemen.

Anyway. Elsewhere in the land of unicorns . . .

Most interesting startup stories this week

Tesla Autopilot NTSB FSD software

Image Credits: Tesla

In the latest episode of “How Not to Win Friends and Influence Government Agencies,” the Dawn Project, a safety advocacy group that’s been on Tesla’s case for a while, decided to spice up their Super Bowl ad with an ad that was essentially a call to arms against Tesla’s Full Self-Driving software. It was meant to be a mic drop. Instead, it turned into a facepalm moment when the National Transportation Safety Board (NTSB) was like, “Um, excuse me, we didn’t sign up for this.” The NTSB is known for many a thing — appearances in Super Bowl ads ain’t it, and the org was quick to issue a “take our seal off your homework” order to the Dawn Project. They pointed out that the Dawn Project did not have permission to use the seal, and its inclusion falsely implied the NTSB’s endorsement of the campaign. Dramaaaaaaaa.

Oh, but there was plenty more drama where that came from:

Some smoke, some mirrors: Boston Dynamics’s secret sauce is a blend of advanced robotics and marketing genius, served with a side of “don’t try this at home” warnings. But beware, not all that glitters in robot videos is gold: Many robot demo videos are bending the truth to varying degrees.

Everything is fine, AI promise: In the latest episode of “AI’s Musical Chairs,” Andrej Karpathy, the AI maestro who was one of the founding members of OpenAI, has once again exited the company. No, it’s not a dramatic soap opera twist or a covert AI uprising; Karpathy insists it’s all smooth sailing, devoid of drama or clandestine plots.

Shut your piehole, AI: The Federal Communications Commission (FCC) has officially declared AI-voiced robocalls as the latest public enemy, branding them illegal. If you were looking forward to a personal, albeit fake, call from a presidential candidate or two, you might want to adjust your expectations. The FCC’s message is clear: AI voice clones, you’re officially on the naughty list.

Most interesting fundraises this week

hands of two people tearing money with flames in the background

Image Credits: Derek Berwin / Getty Images

In a twist that’s got the venture capital world buzzing, Foundry Group, the Boulder-based VC firm known for backing hits like Fitbit and Zynga, is hanging up its investment top hat. After 18 years and nearly $3.5 billion under management, Foundry has decided its latest $500 million fund, Foundry 2022, will be its swan song. Foundry still plans to lead Series A and B financings with the remaining third of its last fund, but the decision to not raise more funds raises eyebrows and questions about the future for its portfolio companies.

This move follows a similar unexpected announcement from Boston-based OpenView at the tail end of last year. Two closures don’t mark a trend, of course, but I’ll bet you billions of dollars to millions of donuts that the TechCrunch team will keep a veeeeeery close eye on this one.

Big raise for banking small companies: Finom, a European challenger bank tailored for SMEs and freelancers, has successfully secured $54 million in a Series B funding round. This funding round underscores the growing demand for specialized financial services for SMEs.

Lettuce raise some more money: Indoor farming, once the darling of the startup world with a $3 billion investment influx from 2012 to 2022, is facing a harsh reality check. Companies like AppHarvest and Fifth Season have hit bankruptcy, while others like Iron Ox have been forced into layoffs and valuation cuts. Despite these challenges, Hippo Harvest emerges as a beacon of hope, securing $21 million in Series B funding.

Well done — have a cookie: SocialCrowd, a performance management startup, has successfully raised $1.6 million in a pre-seed funding round led by Bread & Butter Ventures. Founded in 2020, SocialCrowd offers a SaaS platform akin to a “Fitbit for work,” enabling companies to set and reward employee goals.

This week’s big trend: Hardware

Image Credits: Cory Green/Yahoo

Okay, fine, so perhaps I’m a little bit biased — in the past week I’ve changed gears a little, and I’m going to start writing about hardware a bit more again (here’s what I cover and how to pitch me). The hardware desk is hella punching above its weight, especially this past week — there’s a lot of things happening in the business of atoms.

The industrial robotics sector, after enjoying a surge in orders during the pandemic, experienced a significant downturn in 2023, with orders dropping by nearly one-third, according to the Association for Advancing Automation (A3). This 30% decrease underscores a cooling period for what was once a booming industry, although the decline was not entirely unexpected given the record sales in the preceding years.

More hardware startup nuggets:

Technically, all phones are foldable: And now, Apple is rumored to want to make ones that work after you fold them. As opposed to the last time that was happening. We’ve been asking for folding iPhones for a while, come to think of it.

Dry powder for the big guns: Despite the controversial nature of firearms, Biofire has managed to attract institutional VC backing, raising a $7 million round from notable investors. This funding achievement highlights a shift in the venture capital landscape, where deep tech and defense tech startups are increasingly gaining attention.

Open this app with your face: Brian has been doing an extraordinary job covering all things Apple Vision Pro. This week, he breaks down his favorite apps (so far).

Other unmissable TechCrunch stories . . .

Every week, there’s always a few stories I want to share with you that somehow don’t fit into the categories above. It’d be a shame if you missed ’em, so here’s a random grab bag of goodies for ya:

Dirty money, those cleaning fees: Airbnb’s recent earnings report reveals a significant shift toward more transparent pricing, with nearly 300,000 listings eliminating or reducing cleaning fees. This move, affecting almost 40% of active listings, addresses long-standing customer grievances regarding unexpected costs at checkout.

Notion, but secret-er: Notion recently expanded its suite with a privacy-centric acquisition, purchasing Skiff, a platform known for its end-to-end encrypted file storage, documents, calendar events, and email services.

Mozilla hits the brakes: Mozilla, the organization renowned for its Firefox browser, is undergoing significant strategic shifts. The company plans to reduce its investment in several products, resulting in layoffs impacting around 60 employees.

Put down the LSD, AI: Oh, the wonders of modern technology, where Google’s Gemini chatbot, once known as Bard, and Microsoft’s Copilot are now apparently time travelers. Ahead of the 2024 Super Bowl, the bots had stats and results, before the game had even started. Whoops.

Burning rubber. And more: A Waymo robotaxi found itself the target of a fiery attack in San Francisco. The incident saw a crowd turn their boredom or perhaps technophobia into an act of vandalism that ended with the autonomous vehicle in flames. To its credit, it didn’t try to defend itself, so I guess there’s that.

Add ChatGPT to your WordPress website for $60 with this deal

gpt-stack

This $60 WordPress plugin adds ChatGPT to your website.

With artificial intelligence taking the world by storm and ChatGPT paving its path, it's time to take full advantage of AI capabilities wherever possible. One option is integrating ChatGPT into your WordPress website.

This deal gives you a ChatGPT WordPress plugin with lifetime access for just $60.

Install the ChatGPT plugin to your WordPress website and explore its capabilities on both admin and visitor ends. Whether you run a blog, e-commerce business, or portfolio showcase, ChatGPT could improve user experience, search functionality, and engagement.

On the backend, you could generate content like product listings, blog posts, or SEO descriptions, interpret website analytics, generate reports, and manage user-generated content like comments or reviews.

When used on the front end of your site, you could provide a live customer service chatbot with natural language responses, interactive FAQs to help visitors find information, and product recommendations based on customer preferences. You also have the option to make any features exclusively available to logged-in users.

The WordPress plugin connects to your OpenAI account. That means having this plugin doesn't get you the paid version of ChatGPT, but you can certainly use it with the free version.

Don't miss this low price for the ChatGPT WordPress plugin.

ZDNET Recommends

Tech giants sign voluntary pledge to fight election-related deepfakes

Tech giants sign voluntary pledge to fight election-related deepfakes Kyle Wiggers 8 hours

Tech companies are pledging to fight election-related deepfakes as policymakers amp up pressure.

Today at the Munich Security Conference, vendors including Microsoft, Meta, Google, Amazon, Adobe and IBM signed an accord signaling their intention to adopt a common framework for responding to AI-generated deepfakes intended to mislead voters. Thirteen other companies, including AI startups OpenAI, Anthropic, Inflection AI, ElevenLabs and Stability AI and social media platforms X (formerly Twitter), TikTok and Snap, joined in signing the accord, along with chipmaker Arm and security firms McAfee and TrendMicro.

The undersigned said they’ll use methods to detect and label misleading political deepfakes when they’re created and distributed on their platforms, sharing best practices with one another and providing “swift and proportionate responses” when deepfakes start to spread. The companies added that they’ll pay special attention to context in responding to deepfakes, aiming to “[safeguard] educational, documentary, artistic, satirical and political expression” while maintaining transparency with users about their policies on deceptive election content.

The accord is effectively toothless and, some critics may say, amounts to little more than virtue signaling — its measures are voluntary. But the ballyhooing shows a wariness among the tech sector of regulatory crosshairs as they pertain to elections, in a year when 49% of the world’s population will head to the polls in national elections.

“There’s no way the tech sector can protect elections by itself from this new type of electoral abuse,” Brad Smith, vice chair and president of Microsoft, said in a press release. “As we look to the future, it seems to those of us who work at Microsoft that we’ll also need new forms of multistakeholder action … It’s abundantly clear that the protection of elections [will require] that we all work together.”

No federal law in the U.S. bans deepfakes, election-related or otherwise. But 10 states around the country have enacted statutes criminalizing them, with Minnesota’s being the first to target deepfakes used in political campaigning.

Elsewhere, federal agencies have taken what enforcement action they can to combat the spread of deepfakes.

This week, the FTC announced that it’s seeking to modify an existing rule that bans the impersonation of businesses or government agencies to cover all consumers, including politicians. And the FCC moved to make AI-voiced robocalls illegal by reinterpreting a rule that prohibits artificial and pre-recorded voice message spam.

In the European Union, the bloc’s AI Act would require all AI-generated content to be clearly labeled as such. The EU’s also using its Digital Services Act to force the tech industry to curb deepfakes in various forms.

Deepfakes continue to proliferate, meanwhile. According to data from Clarity, a deepfake detection firm, the number of depfakes that have been created increased 900% year over year.

Last month, AI robocalls mimicking U.S. President Joe Biden’s voice tried to discourage people from voting in New Hampshire’s primary election. And in November, just days before Slovakia’s elections, AI-generated audio recordings impersonated a liberal candidate discussing plans to raise beer prices and rig the election.

In a recent poll from YouGov, 85% of Americans said they were very concerned or somewhat concerned about the spread of misleading video and audio deepfakes. A separate survey from The Associated Press-NORC Center for Public Affairs Research found that nearly 60% of adults think AI tools will increase the spread of false and misleading information during the 2024 U.S. election cycle.

Meet Sora, OpenAI’s impressive new video generation tool

Screenshot 2024-02-16 112620

OpenAI made waves in 2021 when they announced DALL-E, a text-to-image generative AI tool that gave select beta participants the ability to generate images in real time. The results were crude, visually distinct as AI-generated, and certainly needed more time. But despite the quality of the images, there were hopes that the model could be refined. For many, this first generation of DALL-E was like a toddler first making human figures. No one expected perfection, but to be able to see so clearly the silhouette of the intended subject completely generated by a computer was inspiring.

Just yesterday, OpenAI unveiled the new model they dubbed “Sora,” which is capable of generating video clips from text input. Currently, only a small group of testers has access to Sora while they determine where the safety limitations should be. From the examples OpenAI has shared, some of these videos already can pass as real footage. In particular, the footage where the subject is a location, animal, or object. Take a look at the example below:

The prompt given to generate this 20-second video is “A litter of golden retriever puppies playing in the snow. Their heads pop out of the snow, covered in.” If you’ve ever used a generative AI to create an image before, you’ll know that shorter prompts tend to give strange results and that verbose prompts with specific imagery to generate will tend to be closer to the image you have in your head. But as impressive as this video is, there are still a handful of tells for this first iteration of the tool. The physics of the snow still has an uncanny feeling to it as it looks to move on its own in some instances.

However, I am not viewing these videos in the way they are usually intended. I am viewing these with the intent of finding flaws in their presentation because I opened them knowing fully that these are AI-generated videos. I would assume that once the tool is fully released and these clips are used sparingly as stock video, most people would have trouble determining if it’s AI-generated. Even now with ChatGPT being released just over a year ago, people are having a harder time determining if text is AI-generated and the detection tools that are available fall short of being reliable.

While earlier generations of AI-generated content were more obvious to the casual viewer who happened to come across them without the context of them being AI-generated, I don’t have the same expectations going forward. With it being an election year in the US, and with the rise of AI-generated political misinformation, there are serious concerns about the ethical use of AI generation for OpenAI to consider before releasing this tool to the public. There is already precedent for using AI to swing elections, but will AI regulation be capable of reigning it in, or will any legislation be too little, too late?

OpenAI’s Sora release can be found here, and the technical report can be found here.

Sora: OpenAI’s Text-to-Video Generation Model Takes the Internet by Storm 

OpenAI has released text to video generation model Sora. It can generate videos up to a minute long while maintaining visual quality and adherence to the user’s prompt.

Introducing Sora, our text-to-video model.
Sora can create videos of up to 60 seconds featuring highly detailed scenes, complex camera motion, and multiple characters with vibrant emotions. https://t.co/7j2JN27M3W
Prompt: “Beautiful, snowy… pic.twitter.com/ruTEWn87vf

— OpenAI (@OpenAI) February 15, 2024

OpenAI’s Sora is designed to understand and simulate complex scenes, featuring multiple characters, specific motions, and intricate details of the subject and background. The model not only interprets user prompts accurately but also ensures the persistence of characters and visual style throughout the generated video.

One of Sora’s standout features is its ability to take existing still images and breathe life into them, animating the content with precision and attention to detail. Additionally, it can extend or fill in missing frames in an existing video, showcasing its versatility in manipulating visual data.

Sora builds on past research in DALL·E and GPT models. It uses the recaptioning technique from DALL·E 3, which involves generating highly descriptive captions for the visual training data.

While Sora’s capabilities are impressive, OpenAI acknowledges certain weaknesses, such as challenges in accurately simulating the physics of complex scenes and occasional confusion regarding spatial details in prompts.

OpenAI is taking proactive safety measures, engaging with red teamers to assess potential harms and risks. The company is also developing tools to detect misleading content generated by Sora and plans to include metadata for better transparency.

For now, Sora will be available to red teamers and select creative professionals. The company aims to gather feedback from diverse users to refine and enhance Sora, ensuring its responsible integration into various applications.

The team behind Sora is led by Tim Brooks, a research scientist at OpenAI, Bill Peebles, also a research scientist at OpenAI, and Aditya Ramesh, the creator of DALL·E and the head of videogen.

The unveiling of Sora follows Google’s recent release of Lumiere, a text-to-video diffusion model designed to synthesise videos, creating realistic, diverse, and coherent motion. Unlike existing models, Lumiere generates entire videos in a single, consistent pass, thanks to its cutting-edge Space-Time U-Net architecture.

Google today also released Gemini 1.5. This new model outperforms ChatGPT and Claud with 1 million token context window — the largest ever seen in natural processing models. In contrast, GPT-4 Turbo has 128K context windows and Claude 2.1 has 200K context windows.

Gemini 1.5 can process vast amounts of information in one go, including 1 hour of video, 11 hours of audio, codebases with over 30,000 lines of code, or over 700,000 words.

The post Sora: OpenAI’s Text-to-Video Generation Model Takes the Internet by Storm appeared first on Analytics India Magazine.

Master The Art Of Command Line With This GitHub Repository

Master The Art Of Command Line With This GitHub Repository
Image by Author

As a professional who works with data, I understand the importance of being efficient and accurate in the workplace. That's why I believe mastering the command line is an essential skill for streamlining data analysis tasks and improving productivity. It's equally important for regular users who want to optimize their operating system usage and automate various tasks.

In this blog, we will review a popular (144k ?) one-page guide available on GitHub. The guide is designed to equip you with essential command-line skills that can enhance your workflow.

What is the Command Line?

The Command Line (CLI), also known as the terminal or console, is a text-based interface that allows users to interact with a computer's operating system through the use of typed commands. It offers an alternative to graphical user interfaces (GUIs) and provides a more direct and precise way to access and manipulate files, directories, and system resources.

Master The Art Of Command Line With This GitHub Repository
Screenshot by Author

Users can enter commands in a terminal that allows users to perform tasks with precision and automation, such as scripting, software development, data processing, and system administration. The terminal enables users to execute multiple complex operations with just one command.

Why is this Guide Important?

Mastering the art of the command line is a journey that can significantly enhance your productivity and understanding of your computer system. Whether you're a beginner or an experienced user, the command line offers a powerful way to navigate, customize, and automate tasks on your computer.

It is particularly beneficial for data scientists. Through the command line, data professionals can streamline data cleaning, execute data pipelines, automate data-related tasks, and use various command line tools for testing and model development.

Master The Art Of Command Line With This GitHub Repository
Screenshot from jlevy/the-art-of-command-line

This guide aims to provide essential command-line knowledge in one page, with a focus on Linux but also including tools for macOS and Windows users. It covers basic commands, processing files and data, system debugging, and commands that are only available on Mac and Windows. The guide is available in multiple languages, thanks to the contributions of various authors and translators.

Languages: Čeština ∙ Deutsch ∙ Ελληνικά ∙ English ∙ Español ∙ Français ∙ Indonesia ∙ Italiano ∙ 日本語 ∙ 한국어 ∙ polski ∙ Português ∙ Română ∙ Русский ∙ Slovenščina ∙ Українська ∙ 简体中文 ∙ 繁體中文

Content of the Guide

The scope of this guide is broad yet concise, aiming to cover everything important, provide specific examples, and avoid unnecessary details. It's designed for interactive Bash use, but many tips apply to other shells and Bash scripting as well.

Basics

It's essential to learn basic Bash commands and understand their documentation `man <command>` and a master at least one text-based editor (e.g. Vim, Emacs, nano) for efficient terminal-based editing. Additionally, it's important to learn about file and output manipulation, including redirection (>, <, |), and file globbing.

Everyday Use

For efficient command completion and history, use Tab and Ctrl-R, respectively. To navigate and manage files, understand directory navigation using ls, cd , ln, chmod, and chown.

Processing Files and Data

Learn to use text processing tools: grep, awk, sed, cut, sort, uniq, and wc. For file searching, learn to use find and locate to locate files and directories.

System Debugging

Get familiar with system monitoring and debugging tools such as top, ps, netstat, dmesg, and iotop. Use strace, ltrace, and system logs for performance analysis and issue diagnosis.

One-liners

One-liners are powerful command sequences that perform complex tasks quickly. Examples include sorting and counting occurrences in text files, batch renaming, and system monitoring.

Batch renaming script for changing .txt to .md for all files in a directory:

for file in *.txt; do mv "$file" "${file%.txt}.md"; done

Obscure but Useful

Specialized commands like expr, cal, yes, env, and printenv offer useful functionalities for specific scenarios.

macOS Only

Mac users have access to unique tools like Homebrew for package management, pbcopy and pbpaste for clipboard interaction, and specific file and system utilities (mdfind, mdls).

Windows Only

Windows users can turn to Cygwin, Windows Subsystem for Linux (WSL), or MinGW for Unix-like command-line environments. Tools like wmic, ipconfig, and PowerShell scripts extend command-line capabilities on Windows.

Playful Commands

By using tools like curl, egrep, tr, and cowsay, you can fetch, process, and display information creatively, showcasing the power and flexibility at your fingertips.

Conclusion

This guide is a useful cheat sheet for learning about new CLI tools and their applications in various scenarios. It is actively maintained, and you can even contribute to the project by creating a pull request. The Master The Art Of Command Line guide is by the community and for the community, so if you find any mistakes or learn something new that is missing, please update the main README.md file.

I hope you learn about new tools and utilities from this guide and apply them to your projects. In my experience, I have used more command-line tools than actual Python code for data projects, especially if you are a data engineer or MLOps engineer.

Further Read

  • 5 More Command Line Tools for Data Science
  • How to Clean Text Data at the Command Line
  • Data Science at the Command Line: The Free eBook (Must Read)

Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master's degree in Technology Management and a bachelor's degree in Telecommunication Engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.

More On This Topic

  • Data Science at the Command Line: The Free eBook
  • 5 More Command Line Tools for Data Science
  • ChatGPT CLI: Transform Your Command-Line Interface Into ChatGPT
  • 10 GitHub Repositories to Master Machine Learning
  • A Lightning Fast Look at Single Line Exploratory Data Analysis
  • Data storytelling — the art of telling stories through data

Meta’s V-JEPA Video Model Learns by Watching

Along with Open AI’s Sora, Meta released a new AI model called Video Joint Embedding Predictive Architecture (V-JEPA) yesterday. V-JEPA improves machines’ understanding of the world by analysing interactions between objects in videos. The model continues Yann LeCun, Meta’s VP & Chief AI Scientist’s vision, for creating machine intelligence that learns similarly to humans.

The fifth iteration of I-JEPA which was released mid last year has seen developments from comparing abstract representations of images rather than the pixels themselves and extending it to videos. It advances the predictive approach by learning from learning from images to videos, which introduces the complexity of temporal (time-based) dynamics in addition to spatial information.

V-JEPA predicts missing parts of videos without needing to recreate every detail. It learns from unlabeled videos, which means it doesn’t require data that’s been categorised by humans to start learning.

This method makes V-JEPA more efficient, requiring fewer resources to train. The model is particularly good at learning from a small amount of information, making it faster and less resource-intensive compared to older models.

The model’s development involved masking large sections of videos. This approach forces V-JEPA to make guesses based on limited context, helping it understand complex scenarios without needing detailed data. V-JEPA focuses on the general idea of what’s happening in a video rather than specific details, like the movement of individual leaves on a tree.

V-JEPA has shown promising results in tests, where it outperformed other video analysis models using a fraction of the data typically required. This efficiency is seen as a step forward in AI, making it possible to use the model for various tasks without extensive retraining.

Looking ahead, Meta plans to expand V-JEPA’s capabilities, including adding sound analysis and improving its ability to understand longer videos.

This work supports Meta’s broader goal of advancing machine intelligence to perform complex tasks more like humans. V-JEPA is available under a Creative Commons NonCommercial licence, allowing researchers worldwide to explore and build upon this technology.

The post Meta’s V-JEPA Video Model Learns by Watching appeared first on Analytics India Magazine.