Apple’s Scary Fast Event: New M3 Chip Family, MacBook Pro Line & iMac

Apple revealed its next-generation M3 silicon, the latest MacBook Pro family and a new iMac at a Halloween-themed Scary Fast event on October 30.

Jump to:

  • Apple’s M3 chips boast graphics and AI performance increases
  • New MacBook Pro line with M3 chips can be preordered now
  • New iMac running on Apple’s M3 can be preordered now

Apple’s M3 chips boast graphics and AI performance increases

Apple’s next-generation silicon — the M3, M3 Pro and M3 Max (Figure A) — are what Apple CEO Tim Cook called a “new family of breakthrough chips.”

Figure A

Apple Oct 2023 Event M3 Chips Family
The M3 chip, Apple’s next-generation silicon, was unveiled during the company’s Scary Fast event on Oct. 30, 2023. Image: TechnologyAdvice

The M3 line is the first family of chips for a personal computer built using 3 nanometer technology for extraordinarily small transistors, said Johny Srouji, senior vice president of hardware technologies at Apple. In the GPU, the M3 family has dynamic caching, which adjusts GPU utilization depending on how memory is being used and increases performance for demanding apps and games. It includes hardware accelerated mesh shading for geometry in animation and hardware accelerated ray tracing for gaming.

SEE: Curious about iCloud? We note its pros and cons so you can decide if iCloud is right for your work. (TechRepublic)

The Apple M3 adds increased capabilities, a 12-core CPU and 18-core GPU, and is up to 40% faster than the M1 Pro. It supports up to 24 GB of memory. M3 Max has a 16-core CPU, a 40-core GPU and up to 128 GB of memory for the largest AI workloads.

New MacBook Pro line with M3 chips can be preordered now

The M3 chip family is at the core of the new line of Apple MacBook Pro laptops. There are three tiers of options.

  • The 14″ inch Apple MacBook Pro with the M3 chip is up to 60% faster than the 13″ MacBook Pro with M1, Apple said. This level of performance is good for 3D modeling or augmented reality, Apple said.
  • The 14″ and 16″ Apple MacBook Pro with M3 Pro have greater performance and more unified memory than the MacBook Pros using M3 chips and are suitable for coders, medical researchers and other professionals. These new Apple laptops can be attached to two displays.
  • The Apple MacBook Pro with M3 Max is good for the largest use cases like 3D animation or generative AI and includes up to 128 GB of unified memory. It can attach to four external displays.

The MacBook Pro, which can be preordered now and goes on sale November 7, will start at $1,599 for the 14″ model and $2,499 for the 16″ model.

New iMac running on Apple’s M3 can be preordered now

Lastly, Apple announced a new 24″ iMac desktop with the M3 chip (Figure B).

Figure B

The new lineup of iMac desktops will include the M3 chip. Image: Apple
The new lineup of iMac desktops will include the M3 chip. Image: Apple

The new iMac starts at $1,299. It can be preordered today and will be available next week on November 7.

Subscribe to the Apple Weekly Newsletter

Whether you want iPhone and Mac tips or the latest enterprise-specific Apple news, we've got you covered.

Delivered Tuesdays Sign up today

L&T Technology Services Partners With AWS to Unleash Generative AI Smart Vehicles 

L&T Uses Artificial Intelligence To Support 20 Cities In Combating Against COVID-19

L&T Technology Services (LTTS), a leading global digital engineering and research and development (R&D) services company, announced a strategic collaboration with Amazon Web Services (AWS) to expedite the shift towards Software Defined Vehicles (SDVs) using generative AI.The move aims to enhance the efficiency and safety of automotive operations through innovative software-driven solutions.

SDVs rely on advanced software to manage their functions, prioritise safety and performance. By leveraging AWS, LTTS will facilitate development of next-generation SDVs. Through tailored safety and security solutions covering digital cockpit, connected services, and autonomous driving, LTTS aims to accelerate the time to market for these innovative automotive products by 25%.

“To fuel this transformation, we’re committed to training 1,000 engineers on generative AI with AWS by March 2024, ensuring that the future of mobility is shaped by the brightest minds and the most cutting-edge technology.” said Alind Saxena, president, sales, and whole time director at LTTS.

Utilising Amazon CodeWhisperer, an AI-powered code recommendation tool, LTTS engineers would swiftly develop intelligent applications, such as driver warnings and automated braking systems. These applications are designed to enhance a vehicle’s performance on the road, enabling real-time interaction for stakeholders via large language models (LLMs) built on AWS.

“We’re thrilled to help innovators like LTTS harness the full potential of cloud and generative AI to drive innovation in the industry with a digitally-skilled workforce. Together, we’re not just building cars – we’re building experiences, driving efficiency, and creating a smarter tomorrow.” said Vaishali Kasture, Director (Commercial Sales) of AWS India and South Asia.

Additionally, LTTS harnesses Amazon Bedrock, a fully managed service for building and scaling generative AI applications, to implement cloud-based vehicle test automation. This approach enables the reuse of proven, high-quality, safety-compliant code, reducing the time needed to develop new software applications significantly.

Moreover, LTTS will adopt AWS IoT FleetWise, a service that facilitates the collection, transformation, and real-time transfer of vehicle data to the cloud. This implementation would enhance vehicle quality, safety, and autonomy.
LTTS is not the first to explore generative AI to enhance the car experience. Earlier this year, Mercedes-Benz announced that it is set to add OpenAI’s ChatGPT within its in-house multimedia system, MBUX Voice Assistant, as part of an exclusive beta program available only in the US. Moreover, Tesla also leverages generative AI to train its fleet of its FSD cars.

The post L&T Technology Services Partners With AWS to Unleash Generative AI Smart Vehicles appeared first on Analytics India Magazine.

Fear Not, Startups: Every OpenAI Update Can Be Your Next Stepping Stone

This has become like a plague. Whenever OpenAI releases new features to ChatGPT, it is often linked to affecting startups that were built around similar functionalities. OpenAI recently introduced two new features to ChatGPT, namely ‘Upload many types of documents’ and ‘Use Tools without Switching.’

This new ‘multimodal’ update, according to many, is expected to kill over thousands of startups, if not dozens. Some of the popular names include ChatPDF, AskYourPDF, and PDF.ai, and many more.

With the latest update, ChatGPT now not only reads PDFs, but it also supports a variety of document types within the same conversation, including PDFs, images, CSVs, and more. Previously, users were limited to uploading images in the default mode, but now they can seamlessly upload documents and immediately begin asking questions, expanding the platform’s versatility and usability.

Moreover, users are no longer required to specify the ChatGPT mode they want to use. Browsing, Advanced Data Analysis (formerly known as Code Interpreter), and DALL-E 3 will all be available within the same conversation. GPT will determine the appropriate time to activate each mode, such as invoking DALL-E when the user intends to create an image.

PDF.ai’s founder Damon Chen was one of the first ones to react, and humorously shared on X:

“Last night, I had a conversation with my wife on ChatGPT news. I asked, ‘What if PDF.ai doesn’t succeed in the end?’ To which, she casually responded, ‘It’s just one project. Start a new one.”

Not all is lost for GenAI Startups

Despite the development, Chen is still optimistic about PDF.ai. “I don’t think ChatGPT will ever implement small PDF-related features that customers desperately ask for.” However, he believes that small players will go away or not even want to get started and big players with VC money will die once they burn all their money soon. But about his own startup he is pretty much assured.

“PDF.ai is bootstrapped and profitable with a very healthy margin. We don’t have a mission to become another unicorn; several million dollars in ARR is good enough for me, and that’s my goal in the next 1-2 years. I’m 1000% confident we can make it!” he added.

ChatGPT has been in the market for almost a year, and OpenAI will continue to add new features gradually. OpenAI’s ultimate goal is to achieve AGI, and small features like the ability to read PDFs are just a small part of that larger objective.

Sahar Mor, product lead at stripe said “UI and ease of use are still valid propositions so vertical startups targeting specific segments can still prevail. Those are the horizontal AI startups that are at risk.”

Mark Zahm, Founder, Glass Acres LLC said “GPT wrappers will exist and thrive as long as GPT exists. Parent platforms can’t fill every niche and use-case. See every extension and plugin..ever.”

Rowan Cheung shared on X, “Why is nobody talking about the people getting filthy rich on GPT wrappers right now?”

According to him, some startups are outperforming billion-dollar companies in web traffic. The list was almost evenly split between wrappers, fine-tunes, and proprietary models. This means that some GPT wrappers are receiving more monthly visits than companies with billion-dollar valuations.

Source: a16z

GPT wrapper-based startups often offer a cost-effective and efficient approach for businesses looking to incorporate AI capabilities without the complexities of building a model from scratch.

OpenAI is not alone

Besides OpenAI, numerous startups are emerging, and generative AI businesses are playing a significant role in the rise of unicorns. According to a recent report from venture capital firm Accel, 60% of these new unicorns belong to the Generative AI sector.

Investment in GenAI startups based in Europe and Israel amounted to nearly $1 billion in the past year. This figure, however, is significantly lower than the funding received by U.S. Generative AI startups, which exceeded $14 billion during the same period.

However, it’s important to note that this data is influenced by a substantial $10 billion funding received by OpenAI alone, as highlighted in the report.

Moreover, as per the recent reports there are several photo AI apps and even other AI chatbots are actually making more money than ChatGPT. In September, both ‘Chat & Ask AI‘ and ‘ChatOn — AI Chat Bot Assistant‘ generated substantial revenue, raking in almost $3.38 million and $2.11 million, respectively.

Additionally, ‘AI Chatbot — Nova’ and ‘AI Chatbot: AI Chat Smith’ were not far behind, earning approximately $1.44 million and $1.72 million during the same period. Moreover, Character.ai, a chatbot startup backed by a16z, is making waves with 2.39 million downloads recorded as of September.

OpenAI might be ten steps ahead of its competitors with its multimodal capabilities, but that doesn’t mean other startups should simply give up. Instead they can take inspiration from it to make a better product.

The post Fear Not, Startups: Every OpenAI Update Can Be Your Next Stepping Stone appeared first on Analytics India Magazine.

Open Source RedPajama-Data-v2 with 30 Trillion Tokens is Here

RedPajama

RedPajama has unveiled the latest version of its dataset, RedPajama-Data-v2, which is a colossal repository of web data aimed at advancing language model training. This dataset encompasses a staggering 30 trillion tokens, meticulously filtered and deduplicated from a raw pool of over 100 trillion tokens, sourced from 84 CommonCrawl data dumps in five languages, including English, French, Spanish, German, and Italian.

Click here to check out the GitHub repository.

RedPajama-Data-v2 comes with a remarkable addition of 40+ pre-computed data quality annotations that offer invaluable tools for further data filtering and weighting.

The dataset covers 5 languages, with 40+ pre-computed data quality annotations that can be used for further filtering and weighting. Here is one example of how to filter RedPajama-Data-v2 in a similar way as Gopher: pic.twitter.com/VqKObX9Iqr

— Together AI (@togethercompute) October 30, 2023

Over the past six months, the impact of RedPajama’s previous release, RedPajama-1T, has been profound in the language model community. This 5TB dataset of high-quality English tokens has been downloaded by more than 190,000 individuals, who have harnessed its potential in creative ways.

RedPajama-1T served as a stepping stone towards the goal of creating open datasets for language model training, but RedPajama-Data-v2 takes this ambition to new heights with its mammoth 30 trillion token web dataset.

RedPajama-Data-v2 stands out as the largest public dataset specifically crafted for LLM training, significantly contributing to the field. Most notably, it introduces 40+ pre-computed quality annotations, empowering the community to enhance the dataset’s utility. This release encompasses over 100 billion text documents derived from 84 CommonCrawl data dumps, constituting a total of 100+ trillion raw tokens.

Together.AI says that the dataset offers a solid foundation for advancing state-of-the-art open LLMs such as Llama, Mistral, Falcon, MPT, and the RedPajama models.

RedPajama-Data-v2 primarily focuses on CommonCrawl data, while data sources such as Wikipedia are available in RedPajama-Data-v1. To further enrich the dataset, users are encouraged to integrate Stack (by BigScience) for code-related content and s2orc (by AI2) for scientific articles. RedPajama-Data-v2 is meticulously crafted from publicly available web data, comprising the core elements of plain text source data, 40+ quality annotations, and deduplication clusters.

The process of creating the source data begins with each CommonCrawl snapshot passing through the CCNet pipeline, chosen for its light processing approach, preserving raw data integrity. This results in the generation of 100 billion individual text documents, maintaining alignment with the overarching principle of data preservation.

The post Open Source RedPajama-Data-v2 with 30 Trillion Tokens is Here appeared first on Analytics India Magazine.

Apple Announces M3, M3 Pro, M3 Max Chips

Apple

At the much awaited “Scary Fast” launch event, Apple has introduced the M3, M3 Pro, and M3 Max chips, representing a monumental leap in computing power for Mac systems. Built using 3-nanometer by TSMC process technology, these chips pack more transistors into a smaller space, revolutionizing speed and efficiency.

The M3 family of chips showcases a next-generation GPU, marking a historic advancement in graphics architecture for Apple silicon. This GPU introduces Dynamic Caching technology, alongside hardware-accelerated ray tracing and mesh shading, a first for Mac. Rendering speeds surge to 2.5 times faster than the preceding M1 chips, setting a new standard.

The CPU performance cores and efficiency cores exhibit impressive gains, boasting 30% and 50% faster speeds, respectively, compared to their M1 counterparts. Moreover, the Neural Engine exhibits a 60% boost in performance. Additionally, a new media engine enhances video experiences with AV1 decode support.

Johny Srouji, Apple‘s SVP of Hardware Technologies, remarked, “M3, M3 Pro, and M3 Max are the most advanced chips ever built for a personal computer.

Key Features of the M3 Chip Family

Next-Gen GPU: The M3 chips feature a next-generation GPU, employing Dynamic Caching, hardware-accelerated ray tracing, and mesh shading, pushing graphics performance to unprecedented levels.

Enhanced CPU: The M3 family boasts architectural improvements in performance and efficiency cores, delivering up to 30% and 50% faster speeds, respectively.

Unified Memory Architecture: With support for up to 128GB of memory, M3 chips offer high bandwidth, low latency, and unmatched power efficiency, unlocking new possibilities for users.

Custom Engines for AI and Video: The enhanced Neural Engine accelerates machine learning models by up to 60%, while the advanced media engine provides hardware acceleration for popular video codecs.

M3, M3 Pro, and M3 Max: Catering to Diverse Needs

M3: Featuring 25 billion transistors, a 10-core GPU, and an 8-core CPU, M3 showcases phenomenal performance with up to 24GB of unified memory.

M3 Pro: With 37 billion transistors, an 18-core GPU, and a 12-core CPU, M3 Pro caters to users requiring additional graphics-intensive capabilities and supports up to 36GB of unified memory.

M3 Max: Boasting an impressive 92 billion transistors, M3 Max hosts a 40-core GPU, a 16-core CPU, and supports up to 128GB of unified memory, making it ideal for the most demanding pro workloads.

Just like how it was showcased at the Apple event, The M3 series aligns with Apple’s commitment to environmental sustainability. Its power-efficient design contributes to the new MacBook Pro’s impressive battery life, reaching up to 22 hours. Apple’s overarching goal is to achieve net-zero climate impact across the entire business by 2030, ensuring every chip in every Mac is carbon neutral from design to manufacturing.

The post Apple Announces M3, M3 Pro, M3 Max Chips appeared first on Analytics India Magazine.

Overseeing generative AI: New software leadership roles emerge

double-gettyimages-621856560

A majority of software leaders are already — or soon will be — incorporating generative AI into their day-to-day work activities. By 2025, more than half of all software-engineering leadership role descriptions will explicitly require oversight of generative AI, according to a Gartner analysis.

This shift in responsibilities brings an urgency to the need to extend the scope of software leadership well beyond the bounds of application development and maintenance. Team management, talent management, business development, and enforcing ethics will be part of generative AI oversight, according to Gartner analyst Haritha Khandabattu.

While generative AI will not replace developers, "it has the ability to automate certain aspects of software engineering," she adds. And while it "cannot replicate the creativity, critical thinking and problem-solving abilities that humans possess," AI serves as a force multiplier that can enhance efficiency.

Also: Everyone wants responsible AI, but few people are doing anything about it

Other experts also recognize the importance of software engineering leadership positions. "The role of managers in the burgeoning societal transformation involving AI cannot be overstated," states Nicholas Berente of the University of Notre Dame and Bin Gu of Boston University, writing in MIS Quarterly.

"It is the managers that make all key decisions about AI. They oversee the development and implementation of AI-based systems, managers use them in their decision making, leverage them to target customers, and monitor and adjust the decisions, processes, and routines that appropriate AI. Managers allocate resources, oversee AI projects, and govern the organizations that are shaping the future."

Challenges for managers include mapping AI against business strategies, promoting human-AI interfaces, as well as paying attention to "data, privacy, security, ethics, labor, human rights, and national security," Berente and his co-authors point out.

Business alignment will be another key leadership capability. Industry leaders suggest AI in its leading forms — generative and operational — is not only a productivity tool for developers, but that this emerging technology also presents business opportunities that software leaders need to understand and push forward. "AI projects aren't just technology projects," says John Roese, global chief technology officer at Dell Technologies.

"The good ones are aligned to business outcomes. AI projects almost inevitably interrupt organizational structures and those aren't technical decisions. Every investment and shift to automation causes legacy jobs to disappear and creates new jobs charged with making that automation operate."

Also: AI will change the role of developers forever, but leaders say that's good news

The demand for new leadership skills means IT professionals should expect an expansion of the teams in which software leaders participate or lead. "AI breakthroughs have given rise to a new level of technical expertise such as AI specialists and machine learning engineers who develop and deploy AI algorithms and neural networks," says Bryan Madden, global head of AI marketing at AMD.

"AI and its deployment are evolving at a rapid pace. AI projects need a rounded approach to make sure, not only are practical and technological factors considered, but that governance, policy, and ethics are also following suit."

It's also important to remember that the leadership of AI is likely to be a team game. While most AI efforts are generally led by the CEO, CIO, or head of engineering, "employees from various departments should collaborate together, building internal use cases to accelerate product capabilities for customers," says Naveen Zutshi, CIO of Databricks.

"Teams from the business side of the organization can work with engineers, those under the CIO, and IT to build internal large language models that improve business processes in all departments."

Also: AI will change software development in massive ways, says MongoDB CTO

This demand for collaboration means the success of AI "will depend on open partnerships and collaboration across technology, business, and society," says AMD's Madden.

"As AI becomes more ubiquitous across industries such as healthcare, finance, and education, there will be a need for domain experts to provide context and insights for AI application developers. Those insights will help the technology community hone their application of AI in the best way for the best return for their customer base. There will be roles emerging that bring policy experts into the realm of application development."

In addition to line-of-business expertise, the rise of AI will mean there is also a growing focus on prompt engineering and in-context learning capabilities. Databricks' Zutshi says, "This is a newer ability for developers to optimize prompts for large language models and build new capabilities for customers, further expanding the reach and capability of AI tools."

Also: Software developers work best in teams. Here's how AI is helping

Yet another area where software leaders will need to take the lead is AI ethics. Software engineering leaders "must work with, or form, an AI ethics committee to create policy guidelines that help teams responsibly use generative AI tools for design and development," Gartner's Khandabattu reports in her analysis. Software leaders will need to identify and help "to mitigate the ethical risks of any generative AI products that are developed in-house or purchased from third-party vendors."

Finally, recruiting, developing, and managing talent will also get a boost from generative AI, Khandabattu adds. Generative AI applications can speed up hiring tasks, such as performing a job analysis and transcribing interview summaries. For example, she says software leaders "can enter a prompt requesting keywords or key phrases related to skills or experience for platform engineering." Generative AI will also support skills management and development. Khandabattu says: "This will help software engineering leaders rethink roles by identifying skills that can be combined to create new positions and eliminate redundancies."

Artificial Intelligence

ChatGPT for career growth? Practica introduces AI-based career coaching and mentorship

ChatGPT for career growth? Practica introduces AI-based career coaching and mentorship Sarah Perez @sarahintampa / 8 hours

Can an AI be your mentor? That’s what a startup called Practica believes. The company, which evolved out of a marketplace for one-on-one executive coaching, has now launched an AI system built on top of an existing knowledge base it developed over the years spent with a large group of human coaches. The resulting AI chatbot experience functions as a personalized workplace mentor and coach that can aid professionals in bettering their skills across a dozen different topics, including management, strategy, sales, personal development, growth, customer success, marketing, data, design, finance, and more.

Originally co-founded in January 2020 by Dave Whittemore, former Thinkful Head of Product (which exited to Chegg), and former Dropbox engineering manager, Andy Scheff, Practica initially tackled the problem around continuous upskilling throughout your career with a traditional executive coaching marketplace.

“We both became executive coaches,” explains Whittemore. “Andy and I both plunged ourselves into it…I coached product managers and Andy coached engineers,” he says.

Later that year, the site grew by adding a marketplace for other executive coaches, and today it still offers 250 human coaches across domain-specific expertise. 90% of that business is B2B — that is, selling Practica’s services to employers. That business is also now profitable on an operating basis.

However, the founders realized that pricing prevented people from being able to access its services.

Image Credits: Practica

“The average per-hour price was $200 an hour and the per-hour price varied based on the seniority level of the person being coached and the coach…If you got coached for the full year, the average person who does a full year of coaching spends about $3,000 a year,” Whittemore said. “That price is the barrier, and that’s what led us to AI coaching,” he adds.

The idea was to take what they had learned through personal one-on-one coaching and blend that together with AI technology. Over the years, they had learned what makes coaching relationships successful and that helped them understand what kind of constructs to build around and on top of an LLM. Their knowledge base consists of a large number of publicly available learning materials, ranging from blog posts to conference talks to videos to podcasts to books and more, which have been hand-picked and curated into hundreds of different skills across a variety of topics.

Image Credits: Practica

Of course, many websites today are now blocking AI web crawlers as they don’t want to be aggregated into AI chatbot experiences. So we asked if there was a concern that some of these sites where the materials were found would do the same. But unlike ChatGPT or some other AI chatbots, Practica says it’s focused on pushing its users out to the sources where the content comes from, not just providing answers. This results in increased traffic for the publisher or content provider it indexes and references.

That said, the company has not yet made any formal licensing arrangements with its educational source providers — some of which could be as simple as a helpful blog post from an engineer, or someone writing about how they overcame challenges as a manager.

In addition, in terms of the AI models under the hood, the goal is to be vendor-agnostic and only focus on the application layer on top of that — which means Practica isn’t directly competing with companies like OpenAI or Google, but rather capable of working with them.

To coach its professional users, the company uses a technique called Retrieval Augmented Generation (RAG) to match the best learning resources for the situation a given learner is in, the team tells TechCrunch. The AI coach explains what’s in the sources it retrieved and why they’re helpful — just as a human coach would. It cites the sources and encourages the user to go and read them as “homework.”

Practica’s use of third-party content, then, is more of a curated search engine, rather than machine learning model training. But the AI is doing the coaching.

That part of the AI’s methodology mixes together a number of coaching tools, including instruction, questioning for context, finding present challenges in the learner’s job to use as learning materials, mapping learning progress to the learner’s career goals, and celebrating wins along the way. It finds the appropriate material, organizes the insights from those materials into a list that the user can interact with, and adds notes you can reference later.

Plus, unlike today’s generalized AI chatbots, Practica’s AI remembers the learner’s history so it can build on their skills as they continue to use the service.

“You can have generalized system instructions, but it’s not really memorizing what it needs to know about you from one session to the next,” says Whittemore, comparing Practica to general-purpose AI chatbots. “We try to be very intentional about that…we remember how you’ve been developing over time so that we can continue to coach really well — the same way a human coach would.”

The system has been in private testing since July of this year and is now opening up to individual learners for between $10 to $20 per month per user. (A “teams” version from employers is also in limited testing.)

At this price point, Practica hopes to make executive coaching more accessible.

“We’re optimistic that it’s actually going to expand executive coaching,” says Whittmore. “Our hope is that we get to get more people touching this at the AI coaching level, and learning about how effective executive coaching is and therefore upgrading to the premium product,” he says, referring to the company’s human coaching services.

Practica has raised $1.5 million in outside funding from across two rounds in 2021 and 2022, before it shifted into AI-based coaching. Both rounds were led by Script Capital and included over 40 individual angel investors, many of whom were leaders in the domains where the company focuses its coaching.

EasyPhoto: Your Personal AI Photo Generator

EasyPhoto : Your Personal AI Portrait Generator

Stable Diffusion Web User Interface, or SD-WebUI, is a comprehensive project for Stable Diffusion models that utilizes the Gradio library to provide a browser interface. Today, we're going to talk about EasyPhoto, an innovative WebUI plugin enabling end users to generate AI portraits and images. The EasyPhoto WebUI plugin creates AI portraits using various templates, supporting different photo styles and multiple modifications. Additionally, to enhance EasyPhoto’s capabilities further, users can generate images using the SDXL model for more satisfactory, accurate, and diverse results. Let's begin.

An Introduction to EasyPhoto and Stable Diffusion

The Stable Diffusion framework is a popular and robust diffusion-based generation framework used by developers to generate realistic images based on input text descriptions. Thanks to its capabilities, the Stable Diffusion framework boasts a wide range of applications, including image outpainting, image inpainting, and image-to-image translation. The Stable Diffusion Web UI, or SD-WebUI, stands out as one of the most popular and well-known applications of this framework. It features a browser interface built on the Gradio library, providing an interactive and user-friendly interface for Stable Diffusion models. To further enhance control and usability in image generation, SD-WebUI integrates numerous Stable Diffusion applications.

Owing to the convenience offered by the SD-WebUI framework, the developers of the EasyPhoto framework decided to create it as a web plugin rather than a full-fledged application. In contrast to existing methods that often suffer from identity loss or introduce unrealistic features into images, the EasyPhoto framework leverages the image-to-image capabilities of the Stable Diffusion models to produce accurate and realistic images. Users can easily install the EasyPhoto framework as an extension within the WebUI, enhancing user-friendliness and accessibility to a broader range of users. The EasyPhoto framework allows users to generate identity-guided, high-quality, and realistic AI portraits that closely resemble the input identity.

First, the EasyPhoto framework asks users to create their digital doppelganger by uploading a few images to train a face LoRA or Low-Rank Adaptation model online. The LoRA framework quickly fine-tunes the diffusion models by making use of low-rank adaptation technology. This process allows the based model to understand the ID information of specific users. The trained models are then merged & integrated into the baseline Stable Diffusion model for interference. Furthermore, during the interference process, the model uses stable diffusion models in an attempt to repaint the facial regions in the interference template, and the similarity between the input and the output images are verified using the various ControlNet units.

The EasyPhoto framework also deploys a two-stage diffusion process to tackle potential issues like boundary artifacts & identity loss, thus ensuring that the images generated minimizes visual inconsistencies while maintaining the user’s identity. Furthermore, the interference pipeline in the EasyPhoto framework is not only limited to generating portraits, but it can also be used to generate anything that is related to the user’s ID. This implies that once you train the LoRA model for a particular ID, you can generate a wide array of AI pictures, and thus it can have widespread applications including virtual try-ons.

Tu summarize, the EasyPhoto framework

  1. Proposes a novel approach to train the LoRA model by incorporating multiple LoRA models to maintain the facial fidelity of the images generated.
  2. Makes use of various reinforcement learning methods to optimize the LoRA models for facial identity rewards that further helps in enhancing the similarity of identities between the training images, and the results generated.
  3. Proposes a dual-stage inpaint-based diffusion process that aims to generate AI photos with high aesthetics, and resemblance.

EasyPhoto : Architecture & Training

The following figure demonstrates the training process of the EasyPhoto AI framework.

As it can be seen, the framework first asks the users to input the training images, and then performs face detection to detect the face locations. Once the framework detects the face, it crops the input image using a predefined specific ratio that focuses solely on the facial region. The framework then deploys a skin beautification & a saliency detection model to obtain a clean & clear face training image. These two models play a crucial role in enhancing the visual quality of the face, and also ensure that the background information has been removed, and the training image predominantly contains the face. Finally, the framework uses these processed images and input prompts to train the LoRA model, and thus equipping it with the ability to comprehend user-specific facial characteristics more effectively & accurately.

Furthermore, during the training phase, the framework includes a critical validation step, in which the framework computes the face ID gap between the user input image, and the verification image that was generated by the trained LoRA model. The validation step is a fundamental process that plays a key role in achieving the fusion of the LoRA models, ultimately ensuring that the trained LoRA framework transforms into a doppelganger, or an accurate digital representation of the user. Additionally, the verification image that has the optimal face_id score will be selected as the face_id image, and this face_id image will then be used to enhance the identity similarity of the interference generation.

Moving along, based on the ensemble process, the framework trains the LoRA models with likelihood estimation being the primary objective, whereas preserving facial identity similarity is the downstream objective. To tackle this issue, the EasyPhoto framework makes use of reinforcement learning techniques to optimize the downstream objective directly. As a result, the facial features that the LoRA models learn display improvement that leads to an enhanced similarity between the template generated results, and also demonstrates the generalization across templates.

Interference Process

The following figure demonstrates the interference process for an individual User ID in the EasyPhoto framework, and is divided into three parts

  • Face Preprocess for obtaining the ControlNet reference, and the preprocessed input image.
  • First Diffusion that helps in generating coarse results that resemble the user input.
  • Second Diffusion that fixes the boundary artifacts, thus making the images more accurate, and appear more realistic.

For the input, the framework takes a face_id image(generated during training validation using the optimal face_id score), and an interference template. The output is a highly detailed, accurate, and realistic portrait of the user, and closely resembles the identity & unique appearance of the user on the basis of the infer template. Let’s have a detailed look at these processes.

Face PreProcess

A way to generate an AI portrait based on an interference template without conscious reasoning is to use the SD model to inpaint the facial region in the interference template. Additionally, adding the ControlNet framework to the process not only enhances the preservation of user identity, but also enhances the similarity between the images generated. However, using ControlNet directly for regional inpainting can introduce potential issues that may include

  • Inconsistency between the Input and the Generated Image : It is evident that the key points in the template image are not compatible with the key points in the face_id image which is why using ControlNet with the face_id image as reference can lead to some inconsistencies in the output.
  • Defects in the Inpaint Region : Masking a region, and then inpainting it with a new face might lead to noticeable defects, especially along the inpaint boundary that will not only impact the authenticity of the image generated, but will also negatively affect the realism of the image.
  • Identity Loss by Control Net : As the training process does not utilize the ControlNet framework, using ControlNet during the interference phase might affect the ability of the trained LoRA models to preserve the input user id identity.

To tackle the issues mentioned above, the EasyPhoto framework proposes three procedures.

  • Align and Paste : By using a face-pasting algorithm, the EasyPhoto framework aims to tackle the issue of mismatch between facial landmarks between the face id and the template. First, the model calculates the facial landmarks of the face_id and the template image, following which the model determines the affine transformation matrix that will be used to align the facial landmarks of the template image with the face_id image. The resulting image retains the same landmarks of the face_id image, and also aligns with the template image.
  • Face Fuse : Face Fuse is a novel approach that is used to correct the boundary artifacts that are a result of mask inpainting, and it involves the rectification of artifacts using the ControlNet framework. The method allows the EasyPhoto framework to ensure the preservation of harmonious edges, and thus ultimately guiding the process of image generation. The face fusion algorithm further fuses the roop(ground truth user images) image & the template, that allows the resulting fused image to exhibit better stabilization of the edge boundaries, which then leads to an enhanced output during the first diffusion stage.
  • ControlNet guided Validation : Since the LoRA models were not trained using the ControlNet framework, using it during the inference process might affect the ability of the LoRA model to preserve the identities. In order to enhance the generalization capabilities of EasyPhoto, the framework considers the influence of the ControlNet framework, and incorporates LoRA models from different stages.

First Diffusion

The first diffusion stage uses the template image to generate an image with a unique id that resembles the input user id. The input image is a fusion of the user input image, and the template image, whereas the calibrated face mask is the input mask. To further increase the control over image generation, the EasyPhoto framework integrates three ControlNet units where the first ControlNet unit focuses on the control of the fused images, the second ControlNet unit controls the colors of the fused image, and the final ControlNet unit is the openpose (real-time multi-person human pose control) of the replaced image that not only contains the facial structure of the template image, but also the facial identity of the user.

Second Diffusion

In the second diffusion stage, the artifacts near the boundary of the face are refined and fine tuned along with providing users with the flexibility to mask a specific region in the image in an attempt to enhance the effectiveness of generation within that dedicated area. In this stage, the framework fuses the output image obtained from the first diffusion stage with the roop image or the result of the user’s image, thus generating the input image for the second diffusion stage. Overall, the second diffusion stage plays a crucial role in enhancing the overall quality, and the details of the generated image.

Multi User IDs

One of EasyPhoto’s highlights is its support for generating multiple user IDs, and the figure below demonstrates the pipeline of the interference process for multi user IDs in the EasyPhoto framework.

To provide support for multi-user ID generation, the EasyPhoto framework first performs face detection on the interference template. These interference templates are then split into numerous masks, where each mask contains only one face, and the rest of the image is masked in white, thus breaking the multi-user ID generation into a simple task of generating individual user IDs. Once the framework generates the user ID images, these images are merged into the inference template, thus facilitating a seamless integration of the template images with the generated images, that ultimately results in a high-quality image.

Experiments and Results

Now that we have an understanding of the EasyPhoto framework, it is time for us to explore the performance of the EasyPhoto framework.

The above image is generated by the EasyPhoto plugin, and it uses a Style based SD model for the image generation. As it can be observed, the generated images look realistic, and are quite accurate.

The image added above is generated by the EasyPhoto framework using a Comic Style based SD model. As it can be seen, the comic photos, and the realistic photos look quite realistic, and closely resemble the input image on the basis of the user prompts or requirements.

The image added below has been generated by the EasyPhoto framework by making the use of a Multi-Person template. As it can be clearly seen, the images generated are clear, accurate, and resemble the original image.

With the help of EasyPhoto, users can now generate a wide array of AI portraits, or generate multiple user IDs using preserved templates, or use the SD model to generate inference templates. The images added above demonstrate the capability of the EasyPhoto framework in producing diverse, and high-quality AI pictures.

Conclusion

In this article, we have talked about EasyPhoto, a novel WebUI plugin that allows end users to generate AI portraits & images. The EasyPhoto WebUI plugin generates AI portraits using arbitrary templates, and the current implications of the EasyPhoto WebUI supports different photo styles, and multiple modifications. Additionally, to further enhance EasyPhoto’s capabilities, users have the flexibility to generate images using the SDXL model to generate more satisfactory, accurate, and diverse images. The EasyPhoto framework utilizes a stable diffusion base model coupled with a pretrained LoRA model that produces high quality image outputs.

Interested in image generators? We also provide a list of the Best AI Headshot Generators and the Best AI Image Generators that are easy to use and require no technical expertise.

White House Executive Order on AI Provides Guidelines for AI Privacy and Safety

The White House press conference podium.
Image: Maksym Yemelyanov/Adobe Stock

Today, U.S. President Joe Biden released an executive order on the use and regulation of artificial intelligence. The executive order features wide-ranging guidance on maintaining safety, civil rights and privacy within government agencies while promoting AI innovation and competition throughout the U.S.

Although the executive order doesn’t specify generative artificial intelligence, it was likely issued in reaction to the proliferation of generative AI, which has become a hot topic since the public release of OpenAI’s ChatGPT in November 2022.

Jump to:

  • What does the executive order on safe, secure and trustworthy AI cover?
  • Is this AI executive order a law, and how will its guidelines be used?
  • Global discussions of AI safety continue

What does the executive order on safe, secure and trustworthy AI cover?

The executive order’s guidelines about AI are broken up into the following sections:

Safety and security

Any company developing ” … any foundation model that poses a serious risk to national security, national economic security, or national public health and safety … ” must keep the U.S. government informed of their training and red team safety tests, the executive order states. In red team tests, security researchers attempt to break into an organization to test the organization’s defenses. New standards will be created for companies using AI to develop biological materials.

Privacy

The development and use of privacy-preserving techniques will be prioritized in terms of federal support. Privacy guidance for federal agencies will be strengthened with AI risks in mind.

Equity and civil rights

Landlords, federal benefits programs and federal contractors will receive guidelines to keep AI algorithms from exacerbating discrimination. Best practices will be developed for the use of AI in the criminal justice system.

Consumers, patients and students

AI use will be assessed in healthcare and education.

Supporting workers

Principles and best practices will be developed to reduce harm from AI in terms of job displacement, labor equity, collective bargaining and other potential labor impacts.

Promoting innovation and competition

The federal government will encourage AI innovation in the U.S., including streamlining visa criteria, interviews and reviews for immigrants highly skilled in AI.

Advancing American leadership abroad

The federal government will work with other countries on advancing AI technology, standards and safety.

Responsible and effective government use of AI

The executive order promotes helping federal agencies access AI and hire AI specialists. The government will issue guidance for agencies’ use of AI.

Is this AI executive order a law, and how will its guidelines be used?

An executive order isn’t a law and may be modified. The executive order on AI security doesn’t include revoking the right of any existing AI company to operate, an anonymous senior official from the Biden administration told The Verge.

The executive order directs the way specific government agencies should be involved in AI regulation going forward. The National Institute of Standards and Technology will lead the way on establishing standards for red team testing for high-risk AI foundation models. The Department of Homeland Security will be responsible for applying those standards in critical infrastructure sectors and will create an AI Safety and Security Board. AI threats to critical infrastructure and other major risks will be the purview of the Department of Energy and the Department of Homeland Security.

SEE: It’s important to balance the benefits of AI with the downsides of the “dehumanization” of work, Gartner says. (TechRepublic)

The federal AI Cyber Challenge will be used as groundwork for an advanced cybersecurity program to discover and mitigate vulnerabilities in critical software.

The National Security Council and White House Chief of Staff will work on a National Security Memorandum to direct future guidelines for the federal government related to AI, particularly in the military and intelligence agencies. The National Science Foundation will work with a Research Coordination Network to advance work on privacy-related research and technologies.

The Department of Justice and federal civil rights officers will coordinate on combating algorithm-based discrimination.

“Recommendations are not regulations, and without mandates, it’s hard to see a path towards accountability when it comes to regulating AI,” Forrester Senior Analyst Alla Valente told TechRepublic in an email. “Let’s recall that when Colonial Pipeline experienced a ransomware attack that triggered a domino effect of negative consequences, pipeline operators had cybersecurity guidelines that were voluntary, not mandatory.”

She compared the executive order to the EU AI Act, which offers a more “risk-based” approach.

“For this executive order to have teeth, requirements must be clear, and actions must be mandated when it comes to ensuring safe and compliant AI practices,” Valente said. “Otherwise, the order will be simply more suggestions that will be ignored by those standing to benefit from them most.”

“We believe reasonable regulatory oversight is inevitable for AI, just as we’ve seen implemented for broadcasting, aviation, pharmaceuticals — all the key transformative tech of the past 150 years,” wrote Graham Glass, CEO of AI education company CYPHER Learning, in an email to TechRepublic. “Compliance with eventual ‘rules of [the] road’ for AI will improve with international coordination.”

Global discussions of AI safety continue

U.K. Prime Minister Rishi Sunak stated on Oct. 26 that he would set up a governmental body to assess risks from AI. The research network would include buy-in from multiple countries, including China. The U.K. will hold an AI Safety Summit on November 1 and November 2, where international governments will discuss the safety and risks of generative AI. The EU is still working on finalizing its AI Act.

Subscribe to the Developer Insider Newsletter

From the hottest programming languages to commentary on the Linux OS, get the developer and open source news and tips you need to know.

Delivered Tuesdays and Thursdays Sign up today

ICMR Data Leak Exposes 81.5M Indians’ Personal Information

In what could potentially be the largest data breach in India’s history, sensitive details of 81.5 million Indians have surfaced on the dark web as per reports. One of the most concerning aspects of this breach is that the epicenter of the leakage has not been pinpointed. The ICMR has been under cyber-attacks since February, with over 6,000 attempted breaches recorded last year.

This alarming development has prompted India’s investigative agency, the Central Bureau of Investigation (CBI), to prepare for a thorough probe into the incident, pending an official complaint from the Indian Council of Medical Research (ICMR).

The breach was brought to public attention when a ‘threat actor’ using the pseudonym ‘pwn0001’ advertised the stolen database on a breached forum in the dark web. The compromised information includes Aadhaar and passport details, along with names, phone numbers, and addresses. According to the ‘threat actor,’ this extensive dataset was obtained from the Covid-19 testing records collected by ICMR.

Central agencies and the council were aware of the continuous threats and had urged the ICMR to strengthen its cybersecurity measures to prevent any data leaks.

The seriousness of this incident prompted the involvement of the Computer Emergency Response Team of India (CERT-In), which notified the ICMR about the breach. The verification of sample data for sale matched with the actual data from ICMR, triggering an immediate response from relevant government agencies.

As the breach is suspected to involve foreign actors, the case has gained significant attention at the highest levels of government. Multiple agencies and ministries have been mobilized to address the crisis and investigate the breach thoroughly. Remedial measures are already in place, and Standard Operating Procedures have been deployed to mitigate further damage.

The Covid-19 test data in question is dispersed among several government entities, including the National Informatics Centre (NIC), ICMR, and the Ministry of Health, making it difficult to trace the source of the breach.

The American cyber security and intelligence agency Resecurity was the first to identify the data leak. ‘pwn0001’ posted information about the breach on Breach Forums on October 9, offering access to 815 million “Indian Citizen Aadhaar & Passport” records. To provide perspective, this volume of compromised data exceeds the entire population of India, which stands at just over 1.486 billion people.

Analysts found that one of the leaked samples contained 100,000 records of personally identifiable information related to Indian residents. Some of these records were cross-verified through a government portal’s “Verify Aadhaar” feature, confirming the authenticity of Aadhaar credentials.

The post ICMR Data Leak Exposes 81.5M Indians’ Personal Information appeared first on Analytics India Magazine.