More fun with DALL-E 3 in ChatGPT: Can it design a T-shirt?

robot-shirtscreenshot-2023-10-23-133606.png

The other day, I introduced you to the DALL-E 3 text-to-image system, working inside of ChatGPT Plus. I had a lot of fun playing with the add-on, so I decided to see if I could talk it into making a T-shirt design for me.

Also: DALL-E 3 in ChatGPT Plus is helpful but also gave me images of laptops from 1900

I had some limited luck. It's fun, but it can also be oh-so-annoying.

You might as well come along for the ride. Prepare to be amazed, enthralled, and completely blown away. Get ready to be gobsmacked by the stubbornness, obtuseness, and uncooperativeness. It'll be fun. It'll be frustrating. And yes, I do this for a living.

My idea was to make a ZDNET-themed T-shirt design with a robot. I like the retro robot designs, so that was where I was going to start.

Here's my initial request. I didn't just want a retro robot. I also wanted it to feel a bit more modern. So I asked for it to be rendered in a cyberpunk-style environment. I was actually quite happy with the results.

I particularly liked the third image, so I asked for it on a T-shirt. What I got was definitely not what I wanted. It's almost like it's a caricature of the previous image.

I clarified I wanted it on a black T-shirt, and that I wanted it to keep the original style. I got back something fairly acceptable. Of course, the robot changed again.

But I wanted the image to blend into the shirt, not have a neon border. I tried again, and got yet another robot image.

This wasn't quite as nice as the original robot I chose from the four generated, but I would have been happy with this on my T-shirt. Let me point out that I could have just saved this image (assuming I could get the robot without the shirt), taken it into Photoshop, and generated the rest of my design in something like five minutes.

But no. I had to try doing it with the help of an AI (or two, if you consider DALL-E and ChatGPT as separate AIs conspiring with each other).

I tried nailing down this robot with this background by assigning it a name. It was ultimately a futile effort, but worth the try.

What I wanted to do was keep that exact robot, but change the colors of the graphic to reflect the green used by ZDNET. So I fed it ZDNET's hex color value and told it to redraw the robot. To prevent confusion, I told it to use the robot design I defined as robot1.

As you can see, it's a completely different robot. The previous one was much more of a close-up, while this latest one shows much more background. Why, DALL-E? Why do you ignore me?

So, I tried again. I told DALL-E there was an error, and instructed that the robot be torso and head. DALL-E responded with a thoroughly terrifying robot, not the one that I previously defined as the one we wanted to use.

I tried again, this time specifying I wanted a more 1950's-style robot like my original instruction. This design, I liked.

But once again, I wanted to soften the edges of the design. And once again, DALL-E decided to give me a completely new robot.

By this point, I had pretty much lost patience trying to get DALL-E to use the robot I wanted. Instead, I moved ahead and asked it to include ZDNET's slogan on the shirt.

And once again, I got a less desirable robot design, plus an unintelligible saying, a headache, and chills at the back of my neck.

Aggregating the instruction

I decided to try a different approach. By this point, I figured out the parameters of the prompt. I wanted:

  • A 1950s pop culture robot
  • Torso and head rendered in a photorealistic style
  • Cyberpunk environment
  • Use the #D0FF4B hex color code as the predominant color
  • Blend the image into a black T-shirt
  • Add ZDNET's slogan, "Tomorrow belongs to those who embrace it today" to the shirt

Here's the prompt

I got back four designs. None of them were on a T-shirt, but they were interesting.

Finally, here are each of the four images at full size, starting with image 1:

Here's image 2:

Here's image 3:

And, finally, here's image 4:

My conclusion is that DALL-E in ChatGPT, much like Midjourney, can create compelling images. Midjourney has much better image modification tools, as does the standalone DALL-E.

So, can you create a T-shirt design using DALL-E? Without question. But you have to be flexible. You probably won't get what you want, but if you happen to like one that DALL-E generates, and don't need to make any modifications or corrections, you're golden.

Which do you like best? Let us know in the comments below.

You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter on Substack, and follow me on Twitter at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, and on YouTube at YouTube.com/DavidGewirtzTV.

CIOs assess generative AI’s risk and reward for software engineers

Coding in waves

There's tremendous hype about the potential impact of generative artificial intelligence (AI) tools in software development and engineering.

Some experts believe these tools cloud boost productivity by reducing the repetitive tasks that slow IT professionals down.

Also: Can AI code? In baby steps only

Other experts believe the rapid rise of generative AI could mean the end of software development and engineering as we know it.

So, what's the truth?

Jarrod Phipps, CIO at auto specialist Holman, says a sense of perspective is crucial.

Yes, generative AI tools, such as OpenAI's ChatGPT and GitHub Copilot, have the potential to transform the work activities of developers and engineers.

However, that transformation isn't going to happen overnight. What's more, these AI tools won't work in isolation but will instead generate benefits as an adjunct to human IT professionals.

Also: AI is great at coding, but there are some massive caveats

"I call it an exoskeleton," says Phipps, who talks with ZDNET about the potential impact of generative AI. "It makes you stronger, faster, more agile. The way AI could wrap around all the pieces of our business is an exoskeleton that makes people better at what they do. Generative AI is not necessarily a direct threat, it's a compliment. And we want to wrap an exoskeleton around our developers to make them more efficient at writing code."

While some generative products can already write code, Phipps is not focused on the ability of these tools to provide an all-encompassing approach to software development.

"I'm interested in how these tools can help guide the development process, so the developer is still in full control and has some level of creative responsibility," he says.

Phipps says the idea of a personal assistant for software developers is a "no-brainer" for most enterprises.

Also: Is AI in software engineering reaching an 'Oppenheimer moment'? Here's what you need to know

At the other extreme, he says the thought of letting AI go off and write code by itself is simply a no-go: "I'm not necessarily sure when generative AI is going to write all our code. In fact, I don't see a time when that would happen."

Mukul Agrawal, director of technology at Vistaprint, has a similar view: "Never think about AI replacing people. Some of the tasks might get replaced, but not people."

Agrawal explained to ZDNET how he — like every other IT professional right now — is trying to figure out what AI means for developers and engineers.

"My two cents is that AI will have its own space, and some of the mundane tasks will go away," he says. "And then our teams will have the opportunity to focus on higher value work."

Also: AI is transforming organizations everywhere. How these 6 companies are leading the way

Agrawal says big tech-focused organizations like Vistaprint will eventually benefit from AI-enabled software development and engineering — but not yet, and the explanation comes down to key reasons: costs and risks.

In terms of costs, he says businesses will need to see a return on investment: "You have to really think about the long-term value of any investment in this space."

When it comes to risks, Agrawal says Vistaprint must be careful about data privacy.

"Given our business has so much secret sauce, we worry about it, because anything that goes to ChatGPT is being fed into a public system," he says. "You cannot use those tools for your secret sauce."

Avivah Litan, distinguished VP analyst at Gartner, also recognizes that while generative AI could lead to code-generation productivity increases, there are also significant challenges that need to be overcome before the tools can be used in an enterprise context.

Also: Six skills you need to become an AI prompt engineer

"You have three main risks," she says. "Number one, your code is full of bugs, number two, your code is full of vulnerabilities and security errors, and number three, you're infringing on someone's licensed code."

Litan told ZDNET in an interview that now is the time for senior managers to start talking with their personnel about how generative AI might be exploited safely in the longer term.

"Companies need to spend time educating their personnel, including their developers, about the opportunities and the risks," she says.

While most CIOs are choosing to keep generative AI tools away from production environments, it might not be long before IT professionals start using generative AI for disparate elements of the software development and engineering process.

"The main message I have is to get your staff up to date and put the resources into training, and then take advantage of it," she says. "It's incredible what you can do with code generation now. I could build an entire application without knowing any JavaScript or how to code. But you must be educated on all the pluses and the minuses — and that doesn't happen overnight."

Also: Two divergent skills that matter in an AI world: Math and business development

That's a sentiment that resonates with Omer Grossman, global CIO at CyberArk. In an interview with ZDNET, he suggests now is the time to start exploring generative AI.

"Leaders should make decisions," he says. "And I'm emphasizing that point because if you don't make any decisions because you are risk-averse, you risk missing out."

For business leaders who are thinking about how to use generative AI in areas such as software development and engineering, Grossman suggests a range of steps. "The first thing is to make sure you build responsible guardrails that promote innovation while keeping it secure," he says.

At CyberArk, Grossman has put in place a framework and guidelines that are adjusted as new challenges and opportunities in AI emerge.

"I decided we will promote innovation no matter what, but we'll do it responsibly," he says. One of the key supporting elements for this approach is a cross-organization tiger team, which meets on a bi-weekly basis to discuss new developments and potential implications.

"You need to make sure this team is not only full of tech guys, but also legal, because there are some fresh risks you need to mandate," Grossman says. "Having a bi-weekly meeting ensures you don't have a backlog that's big and that you're responsive to the requests for AI as they evolve."

Also: If you use AI-generated code, what's your liability exposure?

Grossman says generative AI vendors will continue to push out new services and features — and business leaders must develop a strategy that gives professionals in key areas, such as software development and engineering the opportunity to explore the tools safely.

"Every time OpenAI or Microsoft comes up with their next product, we get many requests — everybody wants to experiment," he says. "You must be responsible for the education of employees. As an executive, you must be more agile in the way you think and less waterfall-like. And generative AI is a great example of how that approach can pay dividends."

Artificial Intelligence

MiniGPT-5: Interleaved Vision-And-Language Generation via Generative Vokens

Over the past few years, Large Language Models (LLMs) have garnered attention from AI developers worldwide due to breakthroughs in Natural Language Processing (NLP). These models have set new benchmarks in text generation and comprehension. However, despite the progress in text generation, producing images that coherently match textual narratives is still challenging. To address this, developers have introduced an innovative vision and language generation approach based on “generative vokens,” bridging the gap for harmonized text-image outputs.

The foundation behind MiniGPT-5 is a two-staged training strategy that focuses heavily on description-free multimodal data generation where the training data does not require any comprehensive image descriptions. Furthermore, to boost the model’s integrity, the model incorporates a classifier-free guidance system that enhances the effectiveness of a voken for image generation. In the initial phase, the MiniGPT-5 framework has demonstrated powerful performance and a substantial improvement over the baseline Divter model that is trained on the MMDialog dataset, and has constantly demonstrated its ability to deliver comparable & even superior multimodal outputs in the human evaluations performed on the VIST dataset that further highlights its performance & efficiency across various benchmarks.

MiniGPT5 : An Introduction

With the recent developments of the LLM frameworks, and applications based on these LLM frameworks, multimedia feature integration is a field that has witnessed a rise in its popularity as it also proves to be a vital advancement that powers a wide array of applications from state-of-the-art content creation tools to cutting-edge multimodal dialogue agent. With continuous research and development, language and vision models are at the point where work is going on to facilitate them to generate both text & visual data seamlessly. The ability of LLM to generate multimodal data seamlessly will help in enhancing interactions across different domains including e-commerce, media, and virtual reality.

Ultimately, the aim is to allow models to synthesize, recognize, and respond in a consistent & logical way using both textual & visual modalities, thus playing a crucial role in harmonizing the flow of information, and creating logical & consistent narratives. The need to achieve a blend of textual & visual modalities is fueled primarily by the need of more fluid, integrated & interactive multimodal interactions in LLMs, and ultimately achieving the alternating language and vision generation. However, achieving integrated & interactive multimodal interactions in LLMs is a complicated task riddled with numerous challenges including

  1. Although current LLM are extremely efficient & capable when it comes to text generation, and processing text-image pairs, they do not deliver satisfactory performance when it comes to generating images.
  2. The development of these vision and language models relies heavily on topic-focused data that makes it challenging for models to align the generated text with its corresponding images.
  3. Finally, there is a need to come up with more effective strategies as with an increase in their capabilities, the memory requirements of LLMs also increase especially when performing downstream tasks.

The MiniGPT-5 framework, an interleaved language & vision generating algorithm technique that introduces the concept of “generative vokens” in an attempt to address the challenges mentioned above. The MiniGPT-5 framework proposes a new approach for multimodal data generation by amalgamating Large Language Models with Stable Diffusion techniques by using special visual tokens. The proposed two-stage training method used by the MiniGPT-5 framework highlights the importance of a foundational stage free of descriptions, and preparing the model to deliver efficient performance even in scenarios with limited data.

But what separates the MiniGPT-5 model from current existing frameworks is that the generic stages of the MiniGPT-5 framework do not consist of domain specific annotations. Furthermore, to ensure that the generated text, and their corresponding images are in harmony with one another, the MiniGPT-5 framework deploys a dual-loss strategy that further enhances MiniGPT-5’s approach of using classifier-free guidance and generative vokens. The MiniGPT-5 framework optimizes training efficiency, and addresses the memory constraints thanks to their parameter-efficient strategy for fine tuning the model.

To provide you with a quick summary, the MiniGPT-5 framework

  1. Proposes a method that uses multimodal encoders that represent a novel & generic method that has historically proved to be more effective than traditional LLMs, and uses generative tokens combined with Stable Diffusion techniques to generate interleaved language & visual outputs.
  2. Proposes a dual-stage training strategy for generation of description-free multimodal output, and the inclusion of classifier-free guidance during training to further refine the quality of data generated.

The MiniGPT-5 model is inspired heavily from the previous research & work done in the fields of

  • Text to Image Generation : To facilitate the transformation of textual descriptions into their respective visual representations, and text to image models.
  • MLLMs or Multimodal Large Language Models : Using pre-trained LLM models to explore their applications & effectiveness in generating multimodal data.
  • Multimodal Generation with Large Language Models : To augment the capabilities of a LLM to seamlessly integrate language & visual data generation.

MiniGPT-5 : Method, Architecture, and Framework

To facilitate large language models with multimodal data generation capabilities, the MiniGPT-5 model introduces a framework that aims to integrate text to image generation models and pretrained multimodal large language models. The MiniGPT-5 framework further introduces the “generative vokens”, special visual tokens that allows developers to address the discrepancies that appear across different domains by being able to train directly on raw images. To further enhance the quality of the multimodal data generated by the LLMs, the MiniGPT-5 framework introduces a classifier-free strategy coupled with an advanced two-stage training method. Let’s have a detailed look at the MiniGPT-5 framework.

MultiModal Input Stage

Developments of LLMs in the recent past have brought LLMs multimodal comprehension abilities to light, enabling processing images as a sequential input. The MiniGPT-5 framework makes use of specially designed generative vokens for outputting visual features in an attempt to expand LLMs multimodal comprehension abilities to multimodal data generation. Furthermore, the MiniGPT-5 framework makes use of parameter efficient and cutting edge fine tuning techniques for multimodal output learning with the LLM framework.

Multimodal Encoding

The pretrained visual encoder in the MiniGPT-5 framework transforms each input image into a feature, and each text token is embedded as a vector, and the input prompt features are generated when these embeddings are concatenated with one another.

Adding Vokens in Large Language Models

Traditionally, Large Language Model vocabulary consists only of textual tokens which is why the developers working on the MiniGPT-5 framework had to bridge the gap between the generative & the traditional LLMs. The MiniGPT-5 framework introduces a set of special tokens as generative tokens into the vocabulary of the LLM. The framework then harnesses the hidden output state of the LLM for these special vokens for subsequent image generation, and the insertion of interleaved images is represented by the position of the vokens.

PEFT or Parameter Efficient Fine Tuning

PEFT or Parameter Efficient Fine Tuning is a crucial concept used to train LLMs, and yet, the applications of PEFT in multimodal settings is still unexplored to a fairly large extent. The MiniGPT-5 framework uses the Parameter Efficient Fine Tuning over the encoder of the MiniGPT-4 framework in order to train the model to understand prompts or instructions better, and even enhancing the overall performance of the model in a zero-shot or novel environments.

Multimodal Output Generation

To align the generative model with the generative tokens accurately, the MiniGPT-5 framework formulates a compact mapping module for matching the dimensions, and incorporating supervisory losses including latent diffusion model loss, and text space loss. The latent diffusion supervisory loss aligns the appropriate visual features with the tokens directly whereas the text space loss helps the model learn the correct positions of the tokens. Because the generative vokens in the MiniGPT-5 framework are guided directly by the images, the MiniGPT-5 framework does not require images to have a comprehensive description, resulting in a description-free learning.

Text Space Generation

The MiniGPT-5 framework follows the casual language modeling method to generate both vokens and texts in the text space jointly, and during the training phase, the developers append the vokens to the position of the ground truth images, and train the model to predict vokens within text generation.

Mapping Voken Features for Image Generation

After generating the text space, the framework aligns the hidden output state with the text conditional feature space of the text to image generation model. The framework also supports a feature mapper module that includes a dual-layer MLP model, a learnable decoder feature sequence, and a four-layer encoder-decoder transformer model.

Image Generation with LDM or Latent Diffusion Model

To generate the required images in the denoising process, the framework uses the mapping features as a conditional input. The framework also employs a LDM or Latent Diffusion Model for guidance, as during the training phase, the ground truth image is first converted into a latent feature using a pre-trained VAE following which, the developers obtain the latent noise feature by adding some noise.

The comprehensive approach deployed by the MiniGPT-5 framework allows developers to have a coherent understanding, and generation of both visual and textual elements, using specialized tokens, leveraging the capabilities of pretrained models, and using innovative training techniques.

MiniGPT-5 : Training and Results

When working on the MiniGPT-5 framework, developers observed that training on a limited interleaved text-and-image dataset directly can result in images with diminished quality, and misalignment given the significant domain shift between the image & text domains. To mitigate this issue, developers adopted two distinct training strategies,

  1. Encompassing the incorporation of classifier-free guidance techniques that boosts the effectiveness of generative tokens during the diffusion process.
  2. The second strategy is further divided into two stages
    1. An initial pre-training stage that focuses primarily on aligning coarse features.
    2. A fine-tuning stage that facilitates feature learning.

CFG or Classifier Free Guidance

The idea to first leverage CFG for multimodal generation came as a result of an attempt to enhance consistency & logic between the generated images & texts, and the CFG is introduced during the text to image diffusion process. This method observes that by training on both unconditional and conditional generation with conditioning dropout, the generative model can achieve enhanced conditional results.

Two-Stage Training Strategy

Given the significant domain shift observed between text-image generation, and pure text generation, the MiniGPT-5 framework uses a two-stage strategy for training

  1. Unimodal Alignment Stage or UAS,
  2. Multimodal Learning Stage or MLS.

Initially, the framework aligns the image generation features with the voken feature in single text-image pair datasets where each data sample contains only one text, and only one image, and the text is usually the image caption. In this stage, the framework allows the LLM to generate vokens by utilizing captions as LLM inputs.

Once the UAS has executed successfully, the model can generate images for single text descriptions, but struggles with interleaved language and vision generation including text-image pairs, and complicated reasoning is required for image and text generation. To tackle this hurdle, the developers have further fine tuned the MiniGPT-5 framework using PEFT parameters by interleaved vision-and-language datasets like VIST. During this stage, the framework constructs three different tasks from the dataset

  1. Text Only Generation : Generates the related text given the next image.
  2. Image Only Generation : Generates the related image given the next text.
  3. Multimodal Generation : Generates text image pairs using the given context.

MiniGPT-5 : Benchmarks and Results

To evaluate its performance in multimodal generation comprehensively, the MiniGPT-5 development team compares its performance with other prominent baseline models including Divter, GILL, and the Fine Tuned Unimodal Generation Model, and the comparison is demonstrated in the table below.

The MiniGPT-5 framework understands that the multimodal output might be meaningful as per the context, yet it might differ from the ground reality which is the primary reason why the MiniGPT-5 framework also incorporates human inputs to evaluate & assess the performance of the model. Overall, the effectiveness of the MiniGPT-5 framework for multimodal tasks is measured using three perspectives.

  1. Language Continuity : assessing whether the generated content aligns with the provided context seamlessly.
  2. Image Quality : assessing or evaluating the relevance & clarity of the image generated.
  3. Multimodal Coherence : to determine whether the combined text image output is in sync with the initial context.

VIST Final Step Evaluation

In the first stage of experiments, the MiniGPT-5 framework aims to generate the corresponding images, and the table below summarizes the results obtained from this setting.

As it can be seen, the MiniGPT-5 framework in all the three settings can outperform the fine-tuned SD2 framework, thus highlighting the effectiveness of the MiniGPT-5 pipeline.

The figure above compares the performance of the MiniGPT-5 framework with the fine-tuned MiniGPT-4 framework on the S-BERT, Rouge-L and Meteor performance metrics. The results indicate that the use of generative vokens does not affect the performance of the framework negatively when performing multimodal comprehension tasks. The results also demonstrate that the MiniGPT-5 framework is capable of utilizing long-horizontal multimodal input prompts across a wide array of data to generate high-quality & coherent images without compromising the ability of the original model for multimodal comprehension.

The table above compares the performance of three frameworks on 5,000 samples for multimodal generation from the aspects of Multimodal Coherence, Image Quality, and Language Continuity. As it can be observed, the MiniGPT-5 framework outperforms the other two baseline models by more than 70% cases. On the other hand, the table below demonstrates the performance of the MiniGPT-5 framework on the CC3M validation dataset for the generation of single images. Thanks to data limitations, developers found a gap for voken alignment when used with Stable Diffusion. Despite this limitation, the MiniGPT-5 framework outperforms the current state of the art baseline GILL framework across all metrics.

Conclusion

In this article, we have talked about MiniGPT-5, an interleaved language & vision generating algorithm technique that introduces the concept of “generative vokens” in an attempt to harness the capabilities of LLMs to generate multimodal data y aligning the large language model with a text to image generation model that is pre-trained. We have talked about the essential components & the overall architecture of the MiniGPT-5 framework along with the results that indicate substantial improvements in performance & efficiency when compared with the current baseline & state of the art models. MiniGPT-5 aspires to set a new benchmark in the multimodal content & data generation domain, and aims to resolve the challenges faced by previous models when trying to solve the same problem.

With AI, organizations are now seeing software developers as great collaborators

devteam-gettyimages-1460840795

The popular perception of software developers for decades has been that of brainy and somewhat introverted types who do their best work alone. However, research suggests today's software professionals are actually extraverted, preferring to work as actively as possible within broad teams and with end users. What's more, with artificial intelligence (AI) sweeping through IT shops, opportunities for higher-level advisory roles will only accelerate.

Generative AI will open up development processes to their businesses just as profoundly as methodologies such as Agile and DevOps, KPMG predicts. "In terms of how corporations develop and maintain software, it will prompt changes as big as, and likely even more impactful than, those created by Agile development methods, which enable rapid responses to changing software requirements and customer feedback."

Also: AI is great at coding, but there are some massive caveats

For starters, by automatically generating and testing code written in any language and running on any platform, developers will be freed up to move from project to project, and thus expand their agency across the breadth of their enterprises. "Rapidly onboard large groups of developers to accelerate new features or major changes to software," KPMG analysts advise. "These developers would be productive quickly and require less guidance from existing developers."
Developers themselves also see potential for wider collaboration with their business and technology counterparts, according to a survey of 500 developers by code-hosting platform GitHub. "Developers thrive in collaborative environments," writes Inbal Shani, chief product officer at GitHub. The bottom line is that "developers want to upskill, design solutions, get feedback from end users, and be evaluated on their communication skills."

More than four out of five developers expect AI coding tools will make their team more collaborative. Most also believe collaboration and communication should be just as important as code quality in terms of performance measures, yet only 33% report that their companies use collaboration and communication as a performance metric.
The survey shows that developers work with an average of 21 other developers on a typical project, and 52% report working with other teams daily or weekly. They rank regular touchpoints as the most important factor for effective collaboration. Yet developers also say they spend too much time on builds and tests, and current performance metrics do not adequately represent the contributions they make to their organizations.

Also: Meet the post-AI developer: More creative, more business-focused
Shani believes developer experience should be just as much of a priority to organizations as customer experience and user experience. The best path to code quality is through a productive developer experience that is built on collaboration across the board.
"Too many pings and messages can affect flow, but there's still a need to stay in touch," she observes. "In our survey, developers say effective collaboration results in improved test coverage and faster, cleaner, more secure code writing — which are best practices for any development team. This shows that when developers work effectively with others, they believe they build better and more secure software."
AI now plays a role in freeing up developer time and resources to pursue greater collaboration, the GitHub survey finds. Industry leaders concur that AI — in particular, generative AI — has the potential to elevate developer roles within their enterprises to that of advisors and business advocates. "As generative AI tools become more commonplace, we expect demand for IT professionals to shift from a builder role to a facilitator role," says Patrick Stokes, executive VP and general manager for Salesforce Platform.

Also: Top programming languages and topics: Here's what developers want to learn about
The automated development and deployment of software made possible through AI "has expanded the remit of conventional IT pros, agrees Rajesh Kumar R., CIO at LTIMindtree. "The hyper-automated environment has freed up the bandwidth of IT pros, enabling them to actively engage in mindful innovation and invention, solve complex business problems swiftly, and enhance the usability of software, rather than spending time on repetitive tasks," he says.

The CIO adds, "In its current form, generative AI stands to enhance developer productivity as it builds codes on demand for simpler and proven algorithms, increases code quality in test cases, and improves maintainability as it documents the code."

Developments in generative AI "represent a massive step forward in this journey because almost anyone can ask an AI to produce a functioning program," says Stokes. "Instead of spending hours writing that code, they can spend that time testing it, securing it, and tweaking its interfaces to satisfy its users best. The outcome is higher quality apps in much less time produced by people who will inevitably be even closer to the end-user experience."

Developer

7 Platforms for Getting High Paying Data Science Jobs

7 Platforms for Getting High Paying Data Science Jobs
Image by Author

If you are a recent graduate or recently laid off then this blog post is for you. These 7 platforms offer some of the highly lucrative jobs in data science. You can get full time, part time, contract, or gig jobs just by creating a profile and adding your achievements. I know, we are living in uncertain times and it is getting hard to land your dream job, but you have to start somewhere.

In this blog, we will discuss the top 7 platforms that can help you find high paying data science roles. Whether you are looking to break into the field or advance your career, these sites connect job seekers with top companies hiring data scientists. With the demand for data skills continuing to grow, these platforms give you access to remote, freelance, and traditional data science opportunities.

1. LinkedIn

I am a big fan of LinkedIn. It is one of the best platforms for data scientists to find high-paying roles and get recognized for their skills. On LinkedIn, you can search thousands of data science jobs at leading companies in all industries. LinkedIn uses powerful algorithms to recommend jobs that match your profile, skills, and interests.

7 Platforms for Getting High Paying Data Science Jobs
Image from LinkedIn Jobs

You can easily apply for positions and contact recruiters directly. LinkedIn also allows you to expand your professional network and connect with leaders in the data science field. By building out your profile with key skills, projects, and recommendations, recruiters can quickly identify you as a top data science candidate. The platform also makes it simple to stay up-to-date on the latest data science trends and best practices through curated content.

2. Wellfound

Wellfound (formerly AngelList Talent) is a great platform for data scientists to find high-paying remote jobs at startups and top tech companies. By creating a profile on Wellfound, you can access and apply for data science roles across fast-growing startups.

7 Platforms for Getting High Paying Data Science Jobs
Image from Wellfound

The platform is highly responsive — you can often expect to hear back from hiring managers within a couple days of applying. Wellfound makes it easy to identify startups you believe in and want to contribute your data skills to. It has proven to be an excellent source of data science job opportunities within the startup scene.

I've had positive experiences using Wellfound while searching for startup jobs myself. The platform matches your profile and data science expertise with relevant open positions at the most exciting young companies.

3. Toptal

My experience with Toptal was amazing — it's a great source for high-paying freelance and contract data science jobs. Toptal maintains a network of the top 3% of data professionals in the world, so you feel like an elite talent when working through them. You can get hired as a freelance data scientist, data engineer, machine learning engineer and more.

7 Platforms for Getting High Paying Data Science Jobs
Image from Toptal

The pay rates for projects are excellent. Toptal clients are top companies, startups and organizations so the work is challenging and rewarding. Being part of the exclusive Toptal network is also great for your ego and self-image as a data science professional. You go through a rigorous screening process to demonstrate your skills, which makes you feel accomplished when you are accepted.

4. Upwork

Upwork is a great freelance platform for data scientists to find high-paying, flexible jobs. Like Fiverr, Upwork connects you with clients looking for project-based work, hourly engagements, long-term contracts, and even potential full-time roles.

7 Platforms for Getting High Paying Data Science Jobs
Image from Upwork

The key is standing out and working on your Upwork portfolio. Make sure to highlight specific data science capabilities, tools you have experience with, and achievements from past projects.

Check Upwork frequently for new data science job postings that are a good match for your abilities. When applying, emphasize how you can provide value to potential clients.

5. Kolabtree

Kolabtree is a specialized freelancing platform for scientists and industry experts. While it takes time to build up your reputation, Kolabtree can be a source of lucrative data science jobs. I spent about 2 hours thoroughly completing my profile with all of my academic and professional credentials, achievements, and areas of expertise. This caught the attention of potential clients right away.

7 Platforms for Getting High Paying Data Science Jobs
Image from Kolabtree

Many projects on Kolabtree have fixed prices, but there is room to negotiate rates based on your capabilities and experience. In the beginning, I got some lower paying data projects as I was establishing myself on the platform. But by sticking with Kolabtree and successfully completing a number of projects, I began getting contacted for better paid contracts. The platform is very responsive — clients can easily reach out to you directly for potential projects matching your skills.

6. Indeed

Indeed is a leading job search platform that can be a great source of high-paying local data science roles, especially outside of North America. By creating a profile and selectively applying, Indeed connects you with thousands of data science opportunities in your area. It's important not to spam applications — carefully review and only apply to jobs that are a strong match for your abilities.

7 Platforms for Getting High Paying Data Science Jobs
Image from Indeed

Indeed is known for having many entry level and junior data science listings, but with the right experience you can find more advanced and senior level positions as well. For those living in South Asia or other markets outside North America, Indeed tends to be the go-to platform for local job hunting. Make sure to check Indeed regularly for new high-paying data science openings in your city or country.

7. Amazon Jobs

Major tech companies like Amazon offer lucrative data science roles that aren't always publicly listed on sites like LinkedIn and Indeed. To access these exclusive openings, you should check Amazon.jobs and other tech company job boards. These internal sites showcase hundreds of data science positions across various global offices.

7 Platforms for Getting High Paying Data Science Jobs
Image from Amazon Jobs

The majority of data jobs at top tech firms are only found on their internal career sites. Amazon is constantly hiring for data scientists across divisions like Alexa, AWS, retail, operations and more. These jobs tend to be very well-paid at all levels. Keep an eye on Amazon.jobs if you want to land a data science job at leading tech brand with strong compensation and benefits.

Conclusion

Landing a high paying job in data science can transform your career. While the competition for top roles is fierce, leveraging online platforms gives you access to incredible opportunities. The 7 high paying platforms allow you to connect with employers and clients looking for your specialized data skills. Whether searching for full-time employment or high-rate freelance work, build up your presence across these platforms.

Craft compelling profiles highlighting your data science credentials, expertise, and achievements. Be selective in applying and showcase how your capabilities can add value.

And keep trying. You will eventually land a job. While waiting for the job offer, keep learning new skills and working on portfolio projects.

Check out my latest blog to learn more:

  • 5 Portfolio Projects for Final Year Data Science Students
  • How to Ace Data Scientist Professional Certificate Exam
  • 5 Mistakes I Made While Switching to Data Science Career

Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master's degree in Technology Management and a bachelor's degree in Telecommunication Engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.

More On This Topic

  • The High Paying Side Hustles for Data Scientists
  • KDnuggets™ News 22:n04, Jan 26: The High Paying Side Hustles for Data…
  • 7 High Paying Side Hustles for Data Scientists
  • 6 Highest Paying Companies for Data Scientists
  • How to Get Data Science Interviews: Finding Jobs, Reaching Gatekeepers, and…
  • High-Fidelity Synthetic Data for Data Engineers and Data Scientists Alike

Graph of Thoughts: A New Paradigm for Elaborate Problem-Solving in Large Language Models

Graph of Thoughts: A New Paradigm for Elaborate Problem-Solving in Large Language Models
Key Takeaways

  • Graph of Thoughts (GoT) is a novel framework designed to enhance the prompting capabilities of Large Language Models (LLMs) for complex problem-solving tasks.
  • GoT surpasses existing paradigms like Chain-of-Thought (CoT) and Tree of Thoughts (ToT) by representing the information generated by an LLM as a graph, allowing for more flexible and efficient reasoning.
  • The framework has shown significant improvements in task performance, including a 62% increase in sorting quality and a cost reduction of over 31% compared to Tree of Thoughts.

This work brings the LLM reasoning closer to human thinking or brain mechanisms such as recurrence, both of which form complex networks.

Introduction

The burgeoning landscape of artificial intelligence has given rise to increasingly sophisticated Large Language Models (LLMs) capable of a wide range of tasks. Yet, one of the ongoing challenges is improving these models' ability to solve elaborate problems efficiently. Enter Graph of Thoughts (GoT), a framework hoping to take a giant leap in this direction. GoT advances the prompting capabilities of LLMs by structuring the information they generate into a graph, thereby enabling a more intricate and flexible form of reasoning.

While existing paradigms like Chain-of-Thought (CoT) and Tree of Thoughts (ToT) have contributed to the structured output and hierarchical reasoning in LLMs, they often operate within a linear or tree-like constraint. This limitation can sometimes hinder the model's ability to handle complex problem-solving tasks that require multi-dimensional reasoning and the ability to combine disparate pieces of information. Graph of Thoughts addresses this gap by introducing a graph-based structure for managing "LLM thoughts." This allows for an unprecedented level of flexibility in how information is stored, accessed, and manipulated within the model. With GoT, developers and researchers can fine-tune the prompting strategy to navigate this graph effectively, enabling LLMs to solve intricate problems in a more human-like manner.

Understanding Graph of Thoughts

Graph of Thoughts operates on a simple yet powerful concept: it models the information produced by an LLM as a graph where each vertex represents a unit of information, often referred to as "LLM thoughts." The edges between these vertices signify the dependencies or relationships between different units of thought. This graph-based approach allows for:

  • Combining arbitrary LLM thoughts into harmonious outcomes
  • Refining the essence of complex networks of thoughts
  • Strengthening thoughts with the use of feedback loops

In comparison to existing paradigms like CoT and ToT, GoT offers a more flexible and efficient way to manage and manipulate the information generated by LLMs.

Graph of Thoughts process compared
Figure 1: Comparison of Graph of Thoughts (GoT) to other prompting strategies (Image from paper)
Implementing Graph of Thoughts

To implement GoT, developers need to represent the problem-solving process as a graph, where each node or vertex represents a thought or a piece of information. Then, the relationships or dependencies between these thoughts are mapped as edges in the graph. This mapping allows for various operations like merging nodes to create more complex thoughts, or applying transformations to enhance the existing thoughts.

One of the standout features of GoT is its extensibility, allowing it to adapt to a variety of tasks and domains. Unlike more rigid structures, the graph-based representation in GoT can be dynamically altered during the problem-solving process. This means that as an LLM generates new thoughts or gains additional insights, these can be seamlessly incorporated into the existing graph without requiring a complete overhaul.

Moreover, GoT enables the implementation of feedback loops, where the model can revisit and refine its earlier thoughts based on newly acquired information. This dynamic and iterative process serves to significantly enhance the quality of the model's output, making it a particularly powerful tool for complex tasks that require ongoing refinement and adaptation.

Conclusion

The introduction of GoT may mark a significant advancement in the field of LLMs and their application in complex problem-solving tasks. By adopting a graph-based approach to represent and manipulate the information generated by LLMs, GoT offers a more flexible and efficient form of reasoning. Its success in improving task performance and reducing computational costs makes it a promising framework for future research and applications. Developers and researchers should explore this new paradigm in order to attempt to unlock the full problem-solving potential of their LLMs and improve their prompting.

Matthew Mayo (@mattmayo13) holds a Master's degree in computer science and a graduate diploma in data mining. As Editor-in-Chief of KDnuggets, Matthew aims to make complex data science concepts accessible. His professional interests include natural language processing, machine learning algorithms, and exploring emerging AI. He is driven by a mission to democratize knowledge in the data science community. Matthew has been coding since he was 6 years old.

More On This Topic

  • Top Open Source Large Language Models
  • More Free Courses on Large Language Models
  • Learn About Large Language Models
  • Introducing Healthcare-Specific Large Language Models from John Snow Labs
  • What are Large Language Models and How Do They Work?
  • AI: Large Language & Visual Models

Peak XV Unveils Ninth Cohort, 77% AI and Deep Tech Startups

Peak XV Unveils Ninth Cohort, 77% AI and Deep Tech Startups

India and Southeast Asia-focused VC fund, Peak XV Partners, has unveiled its ninth cohort, Surge 09, with a strong focus on AI and deep tech startups. The latest batch, comprising 13 startups, reflects the growing international interest in these sectors and aims to address the perceived lack of depth in India’s AI startup landscape.

Ten of the 13 startups in the Surge 09 cohort specialise in AI and deeptech, emphasising the fund’s commitment to emerging technologies. This comes as the global startup landscape increasingly gravitates toward AI, with more than half of Y Combinator’s most recent batch concentrating on AI with 35% AI startups.

Shailendra Singh, Managing Director of Peak XV, expressed optimism, stating, “Would I like to see more AI startups? The answer is yes. But do we see zero? No. We see some pretty interesting things that we have picked and invested in.” He also highlighted the evolving landscape in India, emphasising the country’s potential to foster greater expertise and innovation.

The 13 startups are Dozer, Elivaas, Ethereal Machines, Horizon Quantum Computing, InCore, Mercu, Mindgrove, Neurowyzr, Newtrace, Pix.ai, Relevance AI, ZeroK, and another stealth startup.

The Surge 09 cohort features startups addressing various issues, from early brain decline detection to quantum computing and sustainable hydrogen production. These startups are characterised by founders with PhDs and international experience, demonstrating the shift towards fundamental innovation in India.

Peak XV’s Surge program, set to complete five years in early 2024, has become a major player in early-stage investment in India and Southeast Asia. It has supported over 140 startups that have collectively raised more than $2 billion in follow-on funding.

Surge offers a unique model with a measured portfolio size for each batch, allowing founders to benefit from extensive guidance on product strategy and mission. Startups under Surge’s wing receive up to $3 million in seed funding and access to a broad array of resources.

The growing competition among venture firms in India to establish their own Surge-like programs is seen as a positive development for the ecosystem.

In addition to India and Southeast Asia, Peak XV is actively evaluating Australian startups, signalling a broader international expansion. Shailendra Singh highlighted the venture firm’s expertise in helping companies go global and indicated that Australian software startups are increasingly interested in building global firms.

The post Peak XV Unveils Ninth Cohort, 77% AI and Deep Tech Startups appeared first on Analytics India Magazine.

India Sees A Surge in Semiconductor Jobs

Recently, US-based semiconductor firm Micron announced that it is planning to set up a semiconductor fabrication industry in Sanand, India. Reports show that the company has started recruiting from local campuses.

Over thirty students have already received job offers at the company. Experts anticipate that approximately 150 engineering graduates from Gujarat will be employed in the emerging semiconductor sector by the end of the year.

Recent graduates in electronics and communications have been offered annual packages ranging from INR 15 lakh to INR 20 lakh. This is a positive development, especially considering the high demand for skilled workers in the semiconductor industry.

Micron’s Sanand facility stands as a significant semiconductor venture in India, backed by an estimated investment of INR 22,500 crore. The company is set to create one of the country’s largest assembly, testing, marking, and packaging (ATMP) plants at Sanand GIDC, projecting direct employment opportunities for 5,000 individuals and indirect employment for approximately 15,000 professionals.

Bright Future for Fresh Graduates

Chip manufacturing firm Tower Semiconductor, headquartered in Israel, has also expressed renewed interest in India’s chip incentive scheme. They are contemplating the establishment of a semiconductor fabrication plant within the country. This development will further amplify the demand for freshers in the semiconductor industry.

Moreover, Tata Group is set to invest INR 200 crore in establishing a semiconductor testing and packing unit in Narasapura, Kolar district, approximately 65 km from Bengaluru. As per a statement from the office of Karnataka’s Industries Minister on September 16, Tata Semiconductor Assembly and Test Private Limited will be generating 155 employment opportunities.

Similarly, Jaya Jagadish, country head of AMD India, and chairperson of the Semicon Talent Building Committee (TBC) recently said that as India strives to establish itself as a semiconductor manufacturing hub, the industry will create demand for 12 lakh jobs across the sector as manufacturing evolves and design functions solidify further.

The demand for skilled professionals in roles such as engineers, operators, and technicians is critical. She explained that they analyzed the growth and demand of India’s semiconductor industry, identifying a total requirement of approximately 1.2 million workers across various sectors. AMD recently announced the inauguration of its new design center campus in Bengaluru by the end of this year. Over the next five years, the company plans to generate 3,000 new engineering positions.

Furthermore, Lam Research also plans to train 60,000 Indian engineers using Semiverse Solution for Semiconductor Education.

The above mentioned developments aligns with the commitment made by Indian Union Minister Rajeev Chandrasekhar, aiming to produce a minimum of 85,000 global semiconductor talents within the next two years.

Chandrasekhar said that they see their entities, enterprises and engineers playing a deep, significant and decisive role in how the future of semiconductor design and manufacturing would be shaped. He said that the government is committed to being a catalyst for the success of such ventures.

Similarly, Indian Prime Minister Narendra Modi earlier said that the government has identified more than 300 colleges where semiconductor courses will be available and India will have more than 1 lakh semicon design engineers in the next five years.

Smartphone Manufacturing in India

In addition to the semiconductor industry, major smartphone companies such as Google and Apple have recently unveiled their intentions to commence phone manufacturing operations in India.

Apple has initiated iPhone production in India in collaboration with Foxconn, located in the southern state of Tamil Nadu. According to Chandrashekhar, Apple’s presence in India has resulted in the creation of over one lakh new direct manufacturing jobs in the past two years.

Likewise, Rick Osterloh, senior vice president of devices and services Google, revealed plans for local manufacturing of Pixel 8 and Pixel 8 Pro, unveiled recently, with the first devices set to launch in 2024. The company intends to collaborate with both domestic and international manufacturers for smartphone production, although specific names were not disclosed.

These advancements bode well for India’s semiconductor industry. With a skilled workforce and a growing domestic market for semiconductor products, India has the potential to make a significant impact on the global semiconductor market.

The post India Sees A Surge in Semiconductor Jobs appeared first on Analytics India Magazine.

The Growth Behind LLM-based Autonomous Agents

The Growth Behind LLM-based Autonomous Agents
Image by Editor | DALL-E 3

A lot has happened in the year 2023. We’ve seen the growth and emergence of Large Language Models (LLMs) and their particular use as fundamental controllers for autonomous agents. We’ve seen it for ourselves, many people have adopted these autonomous agents, and integrated them into organizations and more companies are interested in LLMs.

And yes, they have been successful.

But don’t you want to know more? Of course, you do.

Survey on LLM-based Autonomous Agents

Researchers from Gaoling School of Artificial Intelligence, Renmin University of China have come together to perform a comprehensive survey on LLM-based autonomous agents, in which they deliver a systematic review of the field of LLM-based autonomous agents from a holistic perspective.

The researchers delve into the construction of LLM-based autonomous agents aswell as a comprehensive overview of the diverse applications in a variety of fields, such as social science, natural science, and engineering.

So let’s get into it.

Below is an image of the growth trend in the field of LLM-based autonomous agents, through the number of published papers from January 2021 to August 2023.

As you can see, in the space of 2 years, LLMs have achieved notable successes, showing the wider public that AI applications have the potential to attain human-like intelligence. Comprehensive training datasets and a substantial number of model parameters work hand in hand in order to attain this.

So it seems like there is a lot of funding and research going into this field, therefore it is imperative to provide a systematic summary of the rapidly developing field to comprehensively understand the intricacies behind it and the benefits it will bring to inspire future research.

This is what this research team from the Gaoling School of Artificial Intelligence are doing.

The Growth Behind LLM-based Autonomous Agents
Image by LLM-Agent-Survey

Architecture Design of LLM-based Autonomous Agent

The whole aim behind LLM-based autonomous agents is that they have the ability to perform diverse tasks as if they have human-like capabilities. For this to be achievable, you need to look further into the architecture design of LLM-based autonomous agents:

  1. Which architecture should be designed to better use LLMs
  2. How to enable the agent to acquire capabilities for accomplishing specific tasks

As part of the systematic review, the researchers understood that LLMs need to fulfill specific roles and autonomously learn from the environment in order to evolve themselves like humans. This is where design rational agent architectures come into play.

The researchers have proposed a unified framework to summarize the number of developed modules to enhance LLMs:

  • Profile — identify the role of the agent
  • Memory — place the agent into a dynamic environment and enable it to recall past behaviors
  • Planning — place the agent into a dynamic environment and plan future actions.
  • Action — translating the agent’s decisions into specific outputs

The profiling module has a direct impact on the memory and planning modules, which all together these three modules influence the action module.

The Growth Behind LLM-based Autonomous Agents
Image by LLM-Agent-Survey

To delve into each module in depth, have a read of the paper: A Survey on Large Language Model-based Autonomous Agents.

In this paper, you can have a deeper look into the applications of LLM-based autonomous agents and proposed evaluation strategies in three distinct areas: social science, natural science, and engineering. LLM-based autonomous agents have shown significant potential to influence multiple domains, therefore, understanding how these applications are evaluated and the strategies used is important.

The Growth Behind LLM-based Autonomous Agents
Image by Survey LLM-based Autonomous Agents

As part of the research process, they also have an interactive table that contains more comprehensive papers related to LLM-based Agents.

Wrapping it up

As we can see, more and more people are peeling the skin back when it comes to LLMs. More people want to know what it’s really about, the architecture, the evaluation strategies and how it will impact our future. Is this to help build more trust around LLMs and AI applications in general or are we going to learn the truth about them?

Nisha Arya is a Data Scientist and Freelance Technical Writer. She is particularly interested in providing Data Science career advice or tutorials and theory based knowledge around Data Science. She also wishes to explore the different ways Artificial Intelligence is/can benefit the longevity of human life. A keen learner, seeking to broaden her tech knowledge and writing skills, whilst helping guide others.

More On This Topic

  • AgentGPT: Autonomous AI Agents in your Browser
  • Why You Need To Know About Autonomous AI Agents
  • Web LLM: Bring LLM Chatbots to the Browser
  • How DeepMind Trains Agents to Play Any Game Without Intervention
  • Baby AGI: The Birth of a Fully Autonomous AI
  • How to Better Leverage Data Science for Business Growth

Apple to Bring Generative AI to Siri

Apple to Bring Generative AI to Siri

In a major response to the recent AI frenzy in the technology industry, Apple is preparing to enter the generative AI landscape, and not just machine learning. The tech giant, which had been relatively passive amid the AI boom, is now gearing up to develop generative AI features across its entire range of devices, including iOS, Siri, and other apps.

Apple’s CEO Tim Cook, now asserts that the company has been working on generative AI technology for several years. Despite this assertion, it’s clear that Apple was caught off guard by the sudden AI fever in the industry and has been scrambling to catch up since late last year. Internally, there is a sense of anxiety and recognition of this significant delay in embracing generative AI technology.

Apple’s initial reluctance became evident as other tech giants, including Google and Microsoft introduced generative AI versions of their search engines, capable of generating human-like responses to user queries.

Microsoft also updated its Windows apps with Copilot, and Amazon.com Inc. enhanced Alexa with AI capabilities. Meanwhile, Apple’s only noteworthy AI release during this time was an improved auto-correct system in iOS 17.

Read: It’s High Time Apple Bought Stability AI

As previously reported, Apple developed its own large language model named Ajax and introduced an internal chatbot called “Apple GPT” for testing. The critical challenge now is to evaluate whether this technology can compete with existing offerings and how Apple can effectively integrate it into its products.

Apple’s generative AI initiative is being led by senior vice presidents John Giannandrea and Craig Federighi, who are referred to as the “executive sponsors” of the project. Eddy Cue, the head of services at Apple, is also involved, and the trio is set to invest around $1 billion annually in this endeavour.

Giannandrea’s team is responsible for developing the underlying technology for a new AI system. They are also working on a significant overhaul of Siri to make it smarter, with the possibility of releasing an improved version as early as next year. Federighi’s software engineering group is incorporating AI into the next version of iOS, with an emphasis on using a large language model to enhance the capabilities of Siri and the Messages app.

Apple’s software engineering teams are exploring the integration of generative AI into development tools like Xcode, which could aid app developers in writing applications more efficiently. Additionally, Cue’s team is working on adding AI features to as many Apple apps as possible, including Apple Music and productivity apps.

One ongoing debate within Apple revolves around the deployment of generative AI, considering options such as on-device, cloud-based, or a hybrid approach. An on-device setup would prioritise speed and privacy, while a cloud-based approach would enable more advanced operations. The choice between these deployment methods will be critical as Apple aims to stay competitive in the rapidly evolving field of generative AI.

The post Apple to Bring Generative AI to Siri appeared first on Analytics India Magazine.