IN-SPACe Advances Private Space Industry with 5 Tech Transfers

This Indian Startup is Going to Space

In a significant move towards bolstering the Indian space industry, this past week, IN-SPACe carried out a technology transfer on Wednesday. The occasion marked a noteworthy juncture in the nation’s space endeavours, as five vital technologies developed at the Space Applications Centre under ISRO were passed on to the private sector.

The first technology handed over was the X-band mini-Synthetic Aperture Radar (SAR), commonly mounted on Beechcraft Super King B-200 aircraft flying at altitudes of 8 kilometres. This miniaturised Airborne SAR operates at X-Band and has multifaceted applications encompassing agriculture, disaster management, oceanography, land-use analysis, geology, and cartography. It has found new stewards in Data Patterns, headquartered in Chennai, and Astra Microwave, situated in Hyderabad.

Another game-changing technology that changed hands was the ground penetration radar (GPR). This innovation holds immense potential in security applications such as landmine detection and locating buried objects at depths of 6-10 meters below the surface. It has been entrusted to a firm based in Thiruvananthapuram.

The third technology under the spotlight was the six-port monopole antenna feed technology, which was conferred upon a Hyderabad-based company. Additionally, the Optical Imaging System, a technology born from the SAC-Isro stable, was bequeathed to Optimise Solution Private Limited in Ahmedabad.

A mobile satellite service (MSS) terminal designed for two-way communication and tracking, especially for fishermen, was handed over to Azista Aerospace Pvt Ltd, an Ahmedabad-based company renowned for its contributions to satellite-enabled ground systems, including hydrometeorology solutions.

IN-SPACe, functioning as a single-window, autonomous agency within the Department of Space (DOS), serves as a pivotal catalyst in facilitating private sector participation in India’s ambitious space ventures.

Nilesh Desai, Director of SAC-ISRO, expressed enthusiasm about the technology transfers, stating, “The five crucial technologies were transferred to private players who have been working in the space sector. There is a strong focus on application-driven programs and utilizing space technologies for public purposes.”

This momentous move not only strengthens India’s position in the global space race but also underscores the nation’s commitment to harnessing space technology for the betterment of society. As private enterprises embrace these technologies, India’s space sector is poised for significant growth and innovation.

The post IN-SPACe Advances Private Space Industry with 5 Tech Transfers appeared first on Analytics India Magazine.

Why Meta Should Rush the Launch of LlaMa 3

Meta, the uncrowned king of open-source is going through tough times. The recent releases of Llama and Llama 2, praised for being open-source language models, led to the departure of some scientists and engineers who had worked on Llama.

The reason behind their departure was an internal battle for computing resources with another Meta research team developing a rival model.

While the tech giant is grappling with internal issues, it is facing stiff competition from others willing to contribute to open-source.

Open source LLMs have welcomed a new king and it comes from the Middle East. TII’s recent release of the rendition of its Falcon model is leading the charts.

With a 180-billion-parameter size and trained on a massive 3.5-trillion-token dataset, Falcon 180b has forced the community to consider it. In terms of performance, Falcon 180B has secured its position, dominating the leaderboard for open-access models. While definitive rankings are challenging to establish at this early stage, Falcon 180B’s performance is already drawing comparisons to PaLM-2, a testament to its prowess.

This is high time Meta needs to rush the release of LlaMa 3 if it wants to be at pace with the competition and doesn’t want to be left behind.

And Then There Were Two

Meta, however, isn’t competing with anyone else but OpenAI, which is talking about multimodal capabilities and looking to integrate the iteration of its image generation model DALLE 3.

In such an environment, the discussion surrounding LLaMa 3 is filled with diverse expectations and predictions. Many anticipate LLaMa 3 using high-quality training data, like Phi 1.5 to enhance its performance. There’s also excitement about the potential for more tokens and further exploration of scaling laws. Additionally, there’s the conversation around Mixture-of-Architecture, a statistical approach to architecture that can improve over the drawbacks of parametric architecture, which couldoutperform individual experts or submodels.

LLaMa 3 is also expected to bring multimodal capabilities to open-source. Meta could tap into its own ecosystem of multimodal models built on LLaMa like mPLUG-Owl, llava, minigpt4 and blip2 based on LLaMa.

Open source Banks on LLaMa

Meta has taken on as a pivotal player for smaller initiatives that depend on open-source LLMs. The open-source LLM leaderboard is filled with models fine-tuned on LlaMa, six from the top at least are LLaMa-based—Uni-TianYan, FashionGPT, sheep-duck, Orca, to GenZ Model—by Indian developers.

While Falcon presents a very good and powerful alternative, there are doubts over its licensing.

In essence, the implication of these clauses means that the Licensor reserves the right to modify the Acceptable Use Policy without explicitly notifying users, and users are expected to adapt their usage to conform to the latest version of the policy. Failure to do so could potentially lead to a breach of the licence terms.

Several forums from Reddit to hacker news agreed on the significance of Meta’s role in open-source LLM development and reacted with disappointment to the delay in LLaMa 3. According to a WSJ article, the conglomerate has not even started training it yet and will kickstart the project in early to mid-2024.

The delay also meant that the open-source community would lag behind. Commenters reiterated that Meta’s actions have a substantial impact on the availability of such models. If Meta chooses not to release open-source LLaMa 3, it’s unlikely that any other experienced and well-funded team would be willing to give away a model that has cost millions to develop, bar Falcon.

The post Why Meta Should Rush the Launch of LlaMa 3 appeared first on Analytics India Magazine.

DALL.E 3 and Midjourney Battle It Out

The cat’s out of the bag! After secretly working on an image generation tool for months, OpenAI finally announced DALL.E 3. Not only that, DALL.E 3 will be integrated with ChatGPT Plus and ChatGPT Enterprise in October – finally fulfilling the multimodal tag for GPT-4. With text and image generation now available on the famous chatbot, does the fate of other image-generation tools such as Midjourney seem murky?

Should Midjourney Be Worried?

When the image generation tool was being tested within Discord users a few months ago, the output generated by it was considered far superior to Midjourney. A user even mentioned that they have ‘zero interest in using Midjourney after using it.’ To what extent that holds true has been tested through a side-by-side comparison. Creative director and community developer in AI and art, Nick St. Pierre made a comparison with images generated from both DALL.E-3 and Midjourney by giving them the same set of prompts.

Prompt: Close-up photograph of a hermit crab nestled in wet sand, with seafoam nearby and the details of its shell and texture of the sand accentuated.

DallE. 3 (top) and Midjourney (bottom)

Prompt: A 2D animation of a folk music band composed of anthropomorphic autumn leaves, each playing traditional bluegrass instruments, amidst a rustic forest setting dappled with the soft light of a harvest moon.

Dall.E 3 (left) and Midjourney (right)

It is noticeable that in the images generated by DALL.E 3, minute intricacies related to input prompts are followed. They are more closely aligned to the specific details mentioned in the prompt instructions.

Simpler Language Prompts

The USP of DALL.E 3 is the simplicity in its usage of text prompts. With ChatGPT integration, users can easily input via simple, conversational-kind of prompts, via simple sentences or detailed paragraphs, that will output relevant images. Users can continue to have conversations with the chatbot to further tweak the generated output.

A user has called out the superiority of DALL.E 3 in terms of image quality, prompt coherence, and an accessible UI. However, Midjourney is working on its latest version (V6) which is said to have a better understanding of natural language understanding. They are even looking to bring it on the web and mobile platforms.

Tool Versatility

Midjourney, a pure generative AI platform for creating images has all features that are built for image editing. The output can be tweaked to an extent of providing better colour, contrast or composition. With a list of features such as Zoom, Pan, Remix, etc. Midjourney V5.2 (the last released version) is synonymous with any image generation/editing tool and can be compared to the likes of Adobe too. However, DALL.E 3 does not have any of these features.

Multimodality In Pieces

With DALL.E 3 integration on ChatGPT, multimodality has been addressed, however, images can be produced only as an output. In Midjourney, you can upload images as reference with text prompts to get a desired image as per user needs. This is not available on DALL.E 3, and input is still in the form of text only.

Midjourney can be accessed only through Discord, something that is being highly criticised, especially after DALL.E-3’s release. The onboarding process and usage of the application on Discord seems complicated to many, which dissuades people from using it.

Safety First

During the initial testing phase with Discord users, OpenAI’s tool had no control on its safety feature. Gore, inappropriate images along with trademarked brand logos were generated. This has however been updated before the release of DALL.E 3. OpenAI announced in its latest blog about how safety has been prioritised in the image-generation feature by collaborating with red teamers and removing harmful biases related to visual representation. The model declines public figure requests, and even requests for images in the style of a living artist.

This is an area where Midjourney has tricky boundaries. In the past, a number of images of public figures such as Pope Francis, Donald Trump, Elon Musk and many more have been created using Midjourney. It led to wide criticism and copyright infringement lawsuits too. However, this issue is not yet addressed in spite of releasing five versions of the application. Furthermore, OpenAI is researching and building an internal tool, a provenance classifier, to help identify if an image is generated by DALL.E 3.

Starting as an avocado chair to becoming an avocado patient, OpenAI’s fascination with the fruit to show how far DALL E has come is quite impressive. Integrating it on a chatbot that has over 100M users is the best way for higher reach too. However, Midjourney with its impressive image quality with every version release, and with another version in the pipeline that will probably have a web and mobile presence, will certainly be a game-changer.

How it started 🥑 How it’s going #Dalle3 pic.twitter.com/w2r7Zpvra4

— Red (@adventurared) September 20, 2023

The post DALL.E 3 and Midjourney Battle It Out appeared first on Analytics India Magazine.

ServiceNow Releases Now Assist, an AI Suite for Enterprises  

ServiceNow, the digital workflow leader, has launched Now Assist, a suite of AI-powered solutions that enhance the productivity and experience of its customers across IT, customer service, HR, and development.

Now Assist is powered by ServiceNow’s own generative AI engine and the company has released a domain specific, ‘Now LLM’, built for productivity and data privacy at an enterprise level.

“Now Assist is a game-changer for our customers who want to optimise their workflows and deliver better outcomes for their employees and customers,” said CJ Desai, chief product officer at ServiceNow. “By using AI to automate and augment tasks, Now Assist empowers our customers to focus on what matters most and drive digital transformation at scale.”

The AI suite consists of 4 solutions tailored for different domains:

Now Assist for ITSM helps IT leaders improve the IT experience for their agents and employees by providing instant answers and summaries of incidents, requests, and live chat interactions.

Now Assist for CSM streamlines the customer service process from beginning to end by generating summaries for cases and chats, reducing manual work and allowing agents to resolve customer issues faster.

Now Assist for HRSD helps HR leaders drive productivity and operational efficiency by reducing redundant, manual tasks for HR teams, and getting employees the answers they need quickly with little disruption to their day.

For creators, it helps development teams create and scale apps more quickly on the Now Platform by using AI to generate high-quality code suggestions from natural language text.

In the same industry, companies such as Oracle have also introduced advanced AI assistants tailored for enterprises, aiming to enhance workflow efficiency and overall productivity. A report from Goldman Sachs predicts that generative AI will significantly boost human productivity, contributing nearly $7 trillion to the global GDP over the next decade. Consequently, an increasing number of startups are actively pursuing the development of their own Large Language Models (LLMs) to seize a share of this burgeoning opportunity.

The post ServiceNow Releases Now Assist, an AI Suite for Enterprises appeared first on Analytics India Magazine.

Open Source will Break All Shackles and Win

Open Source will Break All Shackles and Win

The more we get into AI, the more we get into the debate about if it should be open source or closed source. The open source community says that we should take the power away from the tech giants and not let them control it. On the other hand, closed source advocates like OpenAI or Google talk extensively about the dangers around open source. Well, who is right?

Regulators are hell-bent on putting stops around open source AI models. But, the models keep rising up and keep beating the closed source ones on various fronts.

Regulators: Open source AI must be stopped
Open Source AI: pic.twitter.com/BCoj4Tjfdh

— Yam Peleg (@Yampeleg) September 20, 2023

Open source is the harbinger of AI revolution

Early this week, during the US Senate of Intelligence hearing, Yann LeCun, the Meta AI chief, wearing the red bowtie of the National Academy of Engineering, presented his case for open source. “AI is going to become a common platform, and because it’s a common platform, it needs to be open source if you want it to be a platform on top of which a whole ecosystem can be built,” said LeCun.

He gives the example of the infrastructure of the internet. “It didn’t start out as open source, but as commercial. But then the open source platforms won because they are more secure, easy to customise, and safer.” He also points out how open source is necessary for the larger world. Unlike Silicon Valley, where information flows very very quickly because of a close-knit ecosystem, other countries in Europe still have to work in silos.

AI systems are fast becoming a basic infrastructure.
Historically, basic infrastructure always ends up being open source (think of the software infra of the internet, including Linux, Apache, JavaScript and browser engines, etc)
It's the only way to make it reliable, secure, and… https://t.co/p5jBkKQgyq

— Yann LeCun (@ylecun) September 21, 2023

This is definitely true for a lot of companies that are focusing on building native AI products. Regardless of the ease of use and simplicity, the closed source AI products cannot guarantee privacy and security. Take the case of Stable Diffusion and Llama 2, even though Midjourney, DALL-E or GPT offer great usage in the easiest ways, it is still not as controllable as the former ones, which are indeed open source.

Is there a hassle to build on top of open source products? Yes, for sure. But is it going to be beneficial for the companies? Without a doubt!

Open source: For or Against

Every business needs their core AI products. Arguably, if a company is building one of these products that is relying on AI, the first step should be to opt for open source and build it from scratch.

On the other hand, using GPT wrappers means outsourcing of core business, often based on confidential data, which isn’t a desirable long-term strategy for companies.

One common argument against open source AI is that it can’t compete with the vast resources of industry labs. Building foundational AI models is undeniably expensive, and it’s believed that only well-capitalised teams can produce groundbreaking results.

Another argument against open source AI is its supposed inability to reason. It’s often claimed that open source models perform poorly on benchmarks and lack emergent capabilities required for complex tasks. However, this argument overlooks the fact that reasoning doesn’t matter for the majority of AI use cases.

Most users and developers don’t have a pressing need for advanced reasoning capabilities. Open source models are highly proficient at handling the most valuable and common tasks, such as summarisation and ELI5, and with fine-tuning and adequate labelled data, they can cover the vast majority of use cases effectively.

Moreover, even closed source models are bound to get leaked too. All the alignment that the creators argue for is going to go down the drain.

Closed source LLMs will eventually leak, too. The key is to align LLMs well enough that it doesn’t matter if they leak, or even if they’re open sourced. We may even need open source LLMs that we can trust to fortify and defend against renegade AI software

— Alan Cowen (@AlanCowen) August 10, 2023

Blaze Your open source Glory

Last month, the frontrunners of the generative AI revolution, OpenAI raised the alarm over open source AI dangers. “An important test for humanity will be whether we can collectively decide not to open source LLMs that can reliably survive and spread on their own. Once spreading, LLMs will get up to all kinds of crime, it’ll be hard to catch all copies, and we’ll fight over who’s responsible,” said Jan Leike, from the alignment team of OpenAI.

The AI question neatly separates top-down from bottom-up progressives. Top-down ones want government control. Bottom-up ones want open source. May the latter win.

— Pedro Domingos (@pmddomingos) September 20, 2023

It’s true that there are ethical implications about generative AI, but that applies to everyone—both closed source and open source. The best way to go forward, just like any other technology is to let it go its way without putting shackles around it.

The post Open Source will Break All Shackles and Win appeared first on Analytics India Magazine.

Executives need better tech skills. Here are six ways to educate upward

exec-vr-gettyimages-1367549726

Every company may be now a technology company, but it may only be a clever catchphrase. A recent survey out of Accenture finds only 21% of companies have advanced their strategy development to "integrate technology in a meaningful way."

The problem is business executives may be able to use laptops and mobile phones, but that's where their grasp of technology trails off.

Also: One in four tech professionals are ready to leave their jobs

"The stumbling block may be the way you and your leadership team think — or limit your thinking — about technology," the Accenture authors relate. "You expect your executive team to demonstrate a comprehensive understanding of your products and customers, your business model, and your balance sheet. There should be a similar expectation that those same executives have an equally fluent understanding of technology and its potential for your business. Yet that's not the case at many companies."

In some organizations, tech-savvy goes all the way to the top, other industry experts observe. At JPMorgan Chase, for instance, the bank's CEO, Jamie Dimon, demonstrated an intimate understanding of technology in its annual report. "He talks about the journey to the cloud, about how difficult it is, and how many applications they have to refactor," observes Gur Steif, president of digital business automation at BMC Software.

"What struck me isn't the fact that he has 4,000 applications, or that it's easy or difficult, or any of that," Steif continues. "The thing that struck me is that five years ago, you wouldn't have the CEO of a bank talking about refactoring applications. Now we have the CEO of a bank talking about how many applications they have, and how they have to refactor them in order to take advantage of the cloud."

Also: 5 ways to attract top tech talent, according to these business leaders

If you are fluent in technology, your role may soon expand beyond coding and development. Technology professionals need to become educators, providing executives with a crash course — along with regular updates, of course — in technology essentials. Here are some ways to educate upward.

1. Present data on business success stories

The Accenture authors identified "tech-forward" enterprises — with business leaders who have a solid grasp of technology — and found they tend to outperform their peers. "Before and during the pandemic, these companies were 2.3X more likely to outperform peers in terms of revenue growth and return on invested capital."

2. Get receptive sponsors on board

If you are not in the C-suite yet, you need a champion with access to his or her colleagues.

3. Develop estimates of the impact of technologies

Technologists in tech-forward companies "have explicitly taken it on themselves to develop outlooks on the business potential of technology," the Accenture team notes.

4. Look to low-code, and preferably no-code, solutions

It's imperative to make it as easy as possible for executives to leverage technology through self-service environments. It isn't their job to fuss with underlying technologies.

5. Executives don't have to know Python, but they need to be aware of what it can do

"By doing so, they can speak with confidence about how nascent technologies, such as generative AI, relate to the company's current strategy and strategic alternatives," the Accenture team states.

6. Offer challenging workshops

Don't just conduct training sessions — present tech pitch sessions. For example, technologists at tech-forward companies "challenge executives' orthodoxies during recurring 'reinvention sessions,' to imagine a different future for the company." These technologists "are prepared to be challenged and debate potential risks and mitigation — much like a start-up pitch to a venture capital committee."

Also: Everyone wants responsible AI, but few people are doing anything about it

There is plenty of concern about reskilling workers for the tech-driven, AI-infused economy ahead. But it's just as important to provide reskilling to executive leadership as well.

Featured

OpenAI Wins Again 

Google has been talking about Gemini for a while now, and people are growing impatient. It appears that Google is all talk and no action. Meanwhile, OpenAI grabbed the opportunity and recently announced that it plans to integrate Dall-E 3 with ChatGPT Plus and ChatGPT Enterprise.

This is surely a game changer move by OpenAI as it would make GPT-4 the first functional multimodal out in the market which creates text and image both, similar to what Gemini aspires.

I think DALL·E 3 is not just a stance against MidJourney. It's actually a sneak peak of the upcoming, epic battle of massively multimodal LLMs, against DeepMind Gemini.
Quote: "DALL·E 3 is built natively on ChatGPT". This is the key phrase.
DALL·E 3's extraordinary language… pic.twitter.com/08rpGTgD5h

— Jim Fan (@DrJimFan) September 20, 2023

To make up for the absence of Gemini, Google recently added extensions to Bard along with the ability to upload images with Lens and get Search images in responses. It was Google’s attempt to make Bard multimodal. However, only time will tell if it will be able to withstand upcoming competition from Dall-E integrated ChatGPT Plus, scheduled to come in October.

That being said, OpenAI has the potential to impact not only Google Bard and Gemini but also put pressure on other text-to-image generation models out there like Midjourney and Stable Diffusion as Dall-E 3 has shown promise by creating images of higher quality.

Users Advantage

Integrating Dall-E 3 with ChatGPT Plus gives OpenAI an edge as compared to other image generation tools as it has the largest user base as compared to all other models out there in any segment.

At the moment, ChatGPT is one of the world’s popular websites, attracting a staggering 1.4 billion visits globally in August. Meanwhile, during the same month, Bard received 183.5 million visits. On the other hand, Midjourney has over 15 million active users and saw 21 million visits in August. Stable Diffusion has more than 10 million daily active users across all channels, according to Stability AI chief Emad Mostaque.

My personal opinion is that if Dall E will be integreted into Chatgpt
It will be much more used in the wider aspect of the world.
Simply because Chatgpt is huge.
However i love midjourney and its so fun and good to use.
So i guess we will see

— Junior (@juniorforesight) September 21, 2023

Also, viewing it from the perspective of users, Dall-E 3 on ChatGPT gives them freedom to generate text as well as image on a single platform. If the users are getting the benefit of text generation and image generation on the same platform, they would any day prefer ChatGPT.

If we look at the numbers, it is very evident that ChatGPT boasts of a huge user base who won’t shy away from using newer versions of ChatGPT plus at a price of $20 or a little more. MidJourney, however, has a huge price difference and sells monthly plans ranging from $10 to $120,

It can be said that OpenAI is paving the way for a unified multimodal model which is capable of handling a wide range of tasks. Additionally, there have been user complaints regarding the user interface of Midjourney, which is presently hosted on Discord.

Midjourney needs to get out of Discord ASAP

— Nick St. Pierre (@nickfloats) September 20, 2023

Multimodal Market is Scattered

If we examine the currently available multimodal models, we find that they are quite scattered. There isn’t a single model that can perform all tasks. Alongside closed-source models, there are also various open-source models claiming to be multimodal. It is however still not clear which model deserves to claim that it is the real multimodal.

For example, Hugging Face recently introduced a multimodal model named IDEFICS. It has the ability to process both text and image inputs and generate descriptions for the images. Similarly, Bard also possesses the capability to accept image inputs.

Also, Meta recently launched SeamlessM4T, a foundational speech/text translation and transcription model with all-in-one system that performs multiple tasks such as speech-to-speech, speech-to-text, text-to-text translation, and speech recognition. OpenAI and Google have also developed their own speech-to-text models, namely Whisper and AudioPaLM-2, respectively.

If OpenAI adds text-to-speech and speech-to-text features as well to ChatGPT Plus, it could race ahead of other models, making it challenging for others to catch up. Meanwhile, OpenAI doesn’t seem to have any plans to stop here. According to recent reports, it is also planning to integrate GPT-Vision into GPT-4, indicating that it is here to stay.

The post OpenAI Wins Again appeared first on Analytics India Magazine.

OpenAI beefs up its image generating AI tool with DALL-E 3

DALL-E 3 creates a ship

OpenAI has unveiled the next generation of its image creation tool. Known as DALL-E 3, the new version is designed to better understand your text descriptions to create more precise and faithful images. On its new DALL-E 3 webpage, OpenAI didn't reveal much about the tool but did provide hints as to how it aims to surpass its predecessor DALL-E 2.

DALL-E 3 is designed to better grasp the nuances and details in your descriptions, thereby creating more accurate images, OpenAI said. Current AI-powered image generators sometimes ignore words in your descriptions, resulting in images that miss the mark. Based on the images displayed on the DALL-E 3 page, the new version seems capable of creating more accurate, detailed, and imaginative images.

Also: The best AI image generators of 2023

With the buzz around AI, image generators have become popular among individuals and businesses. Such tools as DALL-E 2, Microsoft's Bing Image Creator, Midjourney, Stable Diffusion, DreamStudio, and Craiyon all work more or less the same. Using a prompt, you describe the image you want generated. Choose a style and other attributes. In response, the tool creates one or more images that hopefully match your request.

But like many of today's AI bots, these image generators can be challenging to use. Typically, you have to phrase your prompt in just the right way. And even then, they don't always interpret your requests correctly. Acknowledging that modern text-to-image systems force you to learn prompt engineering, OpenAI said that DALL-E 3 would be a leap forward in generating images that better adhere to your descriptions.

Built on ChatGPT, DALL-E 3 will be accessible through the ChatGPT platform. The benefit here is that you'll be able to use ChatGPT to brainstorm your image ideas and prompts. You can then pose a request to create an image using a simple sentence or a more detailed paragraph.

Also: My two favorite ChatGPT Plus plugins and the remarkable things I can do with them

In the examples offered on the DALL-E 3 webpage, OpenAI showed how the new version would work.

One image was generated based on the description: "Tiny potato kings wearing majestic crowns, sitting on thrones, overseeing their vast potato kingdom filled with potato subjects and potato castles."

A second was created from the description: "An illustration of an avocado sitting in a therapist's chair, saying 'I just feel so empty inside,' with a pit-sized hole in its center. The therapist, a spoon, scribbles notes."

And two images were generated based on a description that read: "An expressive oil painting of a basketball player dunking, depicted as an explosion of a nebula." One image used DALL-E 2, while the other used DALL-E 3 .

OpenAI also stressed that it has limited DALL-E 3's ability to create violent, adult, or hateful content — as it has with previous versions. Safety improvements have been made in areas such as the creation of public figures and certain harmful biases. For example, the tool will decline prompts that ask for a public figure by name.

Also: Who owns code, images, and narratives generated by AI?

AI-generated images can also pose a problem when used to depict a real person or event, misleading people into thinking that the image is real. To combat that issue, OpenAI said that it's testing a new internal tool that can tell whether or not an image was created by DALL-E 3.

Currently in closed testing, DALL-E 3 is scheduled to roll out to ChatGPT Plus and Enterprise customers in early October.

Artificial Intelligence

Microsoft Introduces Multimodal Kosmos-2.5

Microsoft Introduces Multimodal Kosmos-2.5

Microsoft is breaking new ground in the realm of multimodal AI with the introduction of Kosmos-2.5, a literate model designed for the intricate task of machine reading of text-intensive images. Building on the success of its predecessor, Kosmos-1, and Kosmos-2, Microsoft’s Kosmos-2.5 boasts an impressive array of features and capabilities that are set to transform the landscape of image-text understanding.

Click here to read the paper

Kosmos-2.5 has been meticulously pre-trained on vast datasets containing text-intensive images. This extensive training equips Kosmos-2.5 with exceptional proficiency in two closely intertwined transcription tasks:

Spatially-Aware Text Blocks: Kosmos-2.5 can expertly generate text blocks within images while accurately assigning each block its precise spatial coordinates. This breakthrough capability enhances the model’s understanding of text in images, enabling it to provide structured and coherent textual descriptions of image content.

Structured Markdown Text Output: In addition to spatial awareness, Kosmos-2.5 excels in producing structured text output in markdown format. This ensures that not only is the text extracted from images, but it is also presented in a structured and stylized manner.

Summary – key points, training objectives, their impact on the Kosmos-2.5 overall performance, and results (especially interesting comparison with the Nougat model 👀) https://t.co/qi5R18hEvK

— Igor Tica (@ITica007) September 21, 2023

The remarkable capabilities of Kosmos-2.5 are achieved through a shared Transformer architecture, task-specific prompts, and adaptable text representations. This multimodal literate model is a versatile tool that can be harnessed for a wide range of real-world applications involving text-rich images.

The model has undergone extensive testing, demonstrating its proficiency in end-to-end document-level text recognition and image-to-markdown text generation. Furthermore, Kosmos-2.5 can be effortlessly adapted to various text-intensive image understanding tasks using different prompts through supervised fine-tuning.

The introduction of Kosmos-2.5 signifies a significant step towards the future scaling of multimodal large language models. This groundbreaking work by Microsoft is poised to have a transformative impact on the field of AI and image-text understanding.

Kosmos-1 showed that Language is not all that you need. It showcased the potential of integrating language, action, multimodal perception, and world modeling for the advancement of artificial general intelligence (AGI). Kosmos-2.5 is the next step.

The post Microsoft Introduces Multimodal Kosmos-2.5 appeared first on Analytics India Magazine.

Cisco to buy Splunk in $28B bid to secure enterprises in AI era

Cisco logo

Cisco has announced plans to acquire data analytics vendor, Splunk, as it looks to offer enterprises deeper visibility and threat detection capabilities amid the growing adoption of artificial intelligence (AI).

Estimated to be worth $28 billion in equity value, at $157 per share in cash, the acquisition will form one of the world's largest software companies, Cisco said in a statement Thursday. The networking equipment vendor added that the deal also will boost its recurring revenue.

Also: Executives need better tech skills. Here are six ways to educate upward

The union will fuel "the next generation of AI-enabled security and observability," pushing companies from threat detection and response to threat prediction and prevention, said Cisco's CEO and chair Chuck Robbins said in a post.

The IT landscape will continue to evolve as organizations digitize their business and the adoption of AI accelerates, Robbins said, noting that these new technologies bring along new opportunities as well as greater complexity.

"Data is one of the most powerful resources in business, with every organization relying on it to help stay securely connected, run their business, and make mission-critical decisions," he said. "However, customers need a better way to manage, protect, and unlock data's true value, while staying resilient and secure in a world that is constantly changing."

Also: The ethics of generative AI: How we can harness this powerful technology

To help companies do this, he pitched the need for Cisco Security Cloud, which offers visibility into various forms of security data, including network data, identifies, email, web traffic, endpoint devices, and processes.

The addition of Splunk's data analytics platforms will further bolster Cisco's security offerings, providing coverage across applications, devices, and clouds, the vendor said.

This will be crucial in today's hyperconnected world, Cisco said, where data is everywhere and organizations are relying on it to run their business. "Factoring in the acceleration and adoption of generative AI, expanding threat surfaces, and multiple cloud environments, it creates a level of complexity that is unlike anything organizations have faced," Cisco said.

Also: What technology analysts are saying about the future of generative AI

"On top of the data and security challenges, generative AI is rapidly transforming industries and creating new opportunities," Robbins noted. "Together, Cisco and Splunk see a broad range of data across applications, security, and the network."

The two vendors will look to help enterprise customers tap these AI opportunities and gain visibility to their data, he said.

The acquisition is expected to be finalized by the end of the third quarter of next year, subject to regulatory approval, following which, Splunk President and CEO Gary Steele will be part of Cisco's executive leadership team reporting to Robbins.

Artificial Intelligence