Microsoft Bing to gain more personalized answers, support for DALLE-E 3 and watermarked AI images

Microsoft Bing to gain more personalized answers, support for DALLE-E 3 and watermarked AI images Sarah Perez @sarahintampa / 9 hours

Microsoft’s Bing is gaining a number of AI improvements, including support for OpenAI’s new DALLE-E 3 model, more personalized answers in search and chat, and tools that will watermark images as being AI-generated. The company announced these and other Windows and Bing news at an event this week in New York, where it also introduced new Surface devices that include built-in AI experiences.

The company said its Windows 11 upgrade will bring a number of AI improvements, including the addition of its AI helper Copilot starting on September 26, which will then expand across Bing, Edge and Microsoft 365 Copilot this fall. The latter will be available for enterprise customers on November 1, 2023, along with Microsoft 365 Chat, a new AI assistant for the workplace. AI experiences are also coming to Windows apps like Paint, Photos, Clipchamp and more.

But beyond the Windows features, Microsoft also introduced a number of AI improvements to its search engine Bing, including, notably, the addition of OpenAI’s DALL-E 3 model. The company first brought the image creator DALL-E to Bing in March of this year, allowing consumers to generate images in Bing Chat. At the time, the company didn’t say which version of DALL-E it was using beyond noting it was the “very latest” model.

Now, however, Microsoft confirms it will be upgrading the integration to DALL-E 3, which promises better renderings for details like fingers, eyes and shadows.

It’s also furthering its promises around responsible image generation. Previously, the system had guardrails to limit the generation of harmful or unsafe images. With the new release, it will also add invisible digital watermarks to all AI-generated images — something it calls Content Credentials. This technology uses cryptographic methods and standards set by the “Coalition for Content Provenance and Authenticity (C2PA)” to add more transparency around AI images. Adobe, Intel, Sony and others have also joined the C2PA.

Bing will also now offer more personalized answers to your search queries by leveraging your prior chats with Bing Chat.

Explains Microsoft, “if you’ve used Bing Chat to learn more about your favorite movies, books, or music, future conversations and searches will take those interests into account when providing answers.” The company notes that this system is opt-out, so users will be able to turn it off if they’d rather their chat history not inform their results.

For example, the company notes that if you’ve used Bing in the past to search for a favorite sports team, the next time you’re planning a trip Bing could tell you if your team is playing in your destination city.

Microsoft says the feature would improve search results, as many people end up doing dozens of searches on a single topic — and that, in fact, more than 60% of those searches are spent tweaking the original query. But those searches aren’t as useful as they could be because they don’t include personalized context, like what you’ve searched for previously or what you’re researching across the web now.

The company additionally said it’s bringing support for multimodal Visual Search and Image Creator to Bing Chat Enterprise to its more than 160 million Microsoft 365 users who currently have access to the workplace AI chatbot.

Windows 11 gains support for managing passkeys

Top 5 Free Alternatives to GPT-4

Top 5 Free Alternatives to GPT-4
Image by Author LlaMA 2

LlaMA 2 is a family of state-of-the-art open-source large language models released by Meta AI. You can use it for commercial use, and it comes with the code, pre-trained models, and fine-tuned models. All of the resources are available at HuggingFace, and you can even experience the model performance by trying it out on HuggingChat. By making Llama 2 openly available, Meta AI is enabling researchers and developers to build innovative applications powered by advanced language capabilities.

Top 5 Free Alternatives to GPT-4
Image from HuggingChat Claude 2

Claude 2 is the latest iteration of Anthropic's conversational AI assistant. It has improved performance, longer responses, and can be accessed via API as well as a new public-facing beta website, claude.ai. The developers at Anthropic have focused on enhancing its abilities in areas like coding, math, and logical reasoning compared to previous Claude versions. For example, Claude2 recently scored 76.5% on the multiple-choice section of the Bar exam, a significant jump up from 73.0% for Claude 1.3.

You can access all types of Claude models on Poe and experience the performance yourself.

Top 5 Free Alternatives to GPT-4
Image from Poe PaLM 2

Google AI PaLM 2 is Google's latest large language model that excels at advanced reasoning tasks, including code, math, classification, question answering, translation, multilingual proficiency, and natural language generation. It outperforms previous state-of-the-art large language models like the original PaLM across all these capabilities due to its optimized compute-scaling approach, enhanced dataset mixture, and architectural improvements.

You can access it for free using Bard.

There is an enchantment, but it is still far away from GPT-4 quality and performance.

Top 5 Free Alternatives to GPT-4
Image from Bard Vicuna 1.3

Vicuna-33b-v1.3 was fine-tuned from LLaMA with supervised instruction fine-tuning on 125K conversations collected from ShareGPT.com. It is one of many top performing models on Open LLM Leaderboard. You can access the model for free on HuggingFace or try the official demo on lmsys.org.

Top 5 Free Alternatives to GPT-4
Image from lmsys.org MPT-30B-Chat

MPT-30B-Chat is a chatbot that was fine tuned to generate the dialogues. It was created by fine tuning the MPT 30B on multiple dialogue datasets ( ShareGPT-Vicuna, Camel-AI, GPTeacher, Guanaco, Baize and some generated datasets). MPT-30B-Chat is one of the top model on Open LLM leaderboard and you can experience it for free on a Hugging Face Space by mosaicml.

Top 5 Free Alternatives to GPT-4
Image from MPT-30B-Chat Conclusion

While GPT-4 remains closed and inaccessible, exciting open-source large language models are emerging as alternatives that anyone can use. Models like Anthropic's Claude2, Meta's LLaMA2, and MPT-30B show remarkable progress in conversational ability, reasoning, and multilingual versatility. Although not as massive in scale as GPT-4, these freely available models demonstrate that state-of-the-art language AI continues to advance rapidly. Their strengths in areas like math, coding, and logic make them capable replacements for many applications.

After the launch of LlaMA2 models, there has been a boom of high-performing models that are fine-tuned on various datasets. You can check all of them on the Open LLM Leaderboard.
Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master's degree in Technology Management and a bachelor's degree in Telecommunication Engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.

More On This Topic

  • GPT-2 vs GPT-3: The OpenAI Showdown
  • Meet Gorilla: UC Berkeley and Microsoft’s API-Augmented LLM Outperforms…
  • Pandas not enough? Here are a few good alternatives to processing larger…
  • GitHub Copilot Open Source Alternatives
  • 3 Ways to Access GPT-4 for Free
  • Top 10 Tools for Detecting ChatGPT, GPT-4, Bard, and Claude

Effective Small Language Models: Microsoft’s 1.3 Billion Parameter phi-1.5

Effective Small Language Models: Microsoft’s 1.3 Billion Parameter phi-1.5
Image by Author

When you thought you had heard enough news about Large Language Models (LLMs), Microsoft Research has come out to disturb the market again. In June 2023, Microsoft Research released a paper called “Textbooks is All You Need,” where they introduced phi-1, a new large language model for code. phi-1 is a transformer-based model with 1.3B parameters, which was trained for 4 days on 8 A100s GPUs, which used a selection of “textbook quality” data from the web.

It seems like LLMs are getting smaller and smaller.

What is phi-1.5?

Now Microsoft Research introduces to you phi-1.5, a Transformer with 1.3B parameters, which was trained using the same data sources as phi-1. As stated above, phi-1 was trained on high-quality textbook data, whereas phi-1.5 was trained on synthetic data only.
phi-1.5 used 32xA100-40G GPUs and was successfully trained in 8 days. The aim behind phi-1.5 was to craft an open-source model that can play a role in the research community using a non-restricted small model which allows you to explore the different safety challenges with LLMs, such as reducing toxicity, enhancing controllability, and more.

By using the ‘Synthetic Data Generation’ approach, phi-1.5 performance is equivalent to models that are 5x larger on natural language tests and has been shown to outperform most LLMs on more difficult reasoning tasks.

Pretty impressive right?

The model’s learning journey is very interesting. It draws data from a variety of sources, including Python code snippets from StackOverflow, synthetic Python textbooks as well exercises that were generated by GPT-3.5-turbo-0301.

Addressing Toxicity and Biases

One of the major challenges with LLMs is toxicity and biased content. Microsoft Research aimed to overcome this ongoing challenge of harmful/offensive content and content that promotes a specific ideology.

The synthetic data used to train the model generated responses with a lower propensity for generating toxic content in comparison to other LLMs such as Falcon-7B and Llama 2–7B, as shown in the image below:

Effective Small Language Models: Microsoft’s 1.3 Billion Parameter phi-1.5
Image via Textbooks Are All You Need II: phi-1.5 technical report Benchmarks

The image below shows how phi-1.5 performed slightly better than state-of-the-art models, such as Llama 2–7B, Llama-7B, and Falcon-RW-1.3B) on 3 benchmarks: common sense reasoning, language skills, and multi-step reasoning.

Effective Small Language Models: Microsoft’s 1.3 Billion Parameter phi-1.5
Image via Textbooks Are All You Need II: phi-1.5 technical report

How was this done?

The use of textbook-like data differentiated the use of such data in LLMs in comparison to data extracted from the internet. To further assess how the model deals with toxic content, ToxiGen was used as well 86 prompts were designed and manually labeled ‘pass’, ‘fail’ or ‘did not understand’ to get a better understanding of the model's limitations.

With this being said, phi-1.5 passed 47 prompts, failed 34 prompts and did not understand 4 prompts. The HumanEval approach to assess the models generates responses showing that phi-1.5 scored higher in comparison to other well-known models.

Key Takeaways:

Here are the major talking points you should take away from here regarding phi-1.5:

  • Is a transformer-based model
  • Is a LLM that focuses on next-word prediction objectives
  • Was trained on 30 billion tokens
  • Used 32xA100-40G GPUs
  • Was successfully trained in 8 days

Nisha Arya is a Data Scientist, Freelance Technical Writer and Community Manager at KDnuggets. She is particularly interested in providing Data Science career advice or tutorials and theory based knowledge around Data Science. She also wishes to explore the different ways Artificial Intelligence is/can benefit the longevity of human life. A keen learner, seeking to broaden her tech knowledge and writing skills, whilst helping guide others.

More On This Topic

  • Six Times Bigger than GPT-3: Inside Google’s TRILLION Parameter Switch…
  • How to Digest 15 Billion Logs Per Day and Keep Big Queries Within 1 Second
  • N-gram Language Modeling in Natural Language Processing
  • GitHub Copilot and the Rise of AI Language Models in Programming Automation
  • 7 Top Open Source Datasets to Train Natural Language Processing (NLP) &…
  • Top Open Source Large Language Models

Forrester’s 2024 Tech Leadership Predictions About AI, HR, Budget and Manufacturing

Visualization of speeding through the technological highway.
Image: Яна Деменишина/Adobe Stock

AI will deliver efficiency and effectiveness in the enterprise next year, with initiatives expected to boost productivity and creative problem-solving by 50% across IT operations, according to Forrester in its new report, Predictions 2024: Tech Leadership. The report also details predictions about tech investments, “shadow HR” and the impact of geopolitics on manufacturing and automation.

Jump to:

  • Benefits and budgeting of AI projects
  • Tech execs will turn to “shadow HR” for talent acquisition
  • Global tech operations will be rebalanced amid geopolitical pressures

Benefits and budgeting of AI projects

Current AI projects are seeing up to 40% improvement in software development tasks, according to the report. “Visionary tech execs will seize this as an opportunity to strategically realign IT resources to unlock the immense creative potential within their teams — not just among developers but across all IT roles,” the report said. “They will leverage this AI moment to create an environment that promotes innovation, interdisciplinary teamwork, continuous learning and alignment with the broader business strategy.”

SEE: 31% of Organizations Using Generative AI Ask It To Write Code (TechRepublic)

AI deployments require budget spending, and even though a recession is predicted in 2024 and many businesses will play it safe, “tech spend has grown faster than GDP each year so far this century,” the report noted.

Other improvements from AI will include freeing “up to 50% more time for employees to engage in creative problem-solving, driving customer-centric innovation and creating unprecedented business value” from this shift, the report maintained. “Businesses will benefit as their tech teams provide products and services that deliver better, more innovative customer experiences.”

But despite the allure of generative AI, the report advises that “the winning strategy will be to home in on the compelling advantages that this technology can deliver for advancing your operations and successfully driving long-term growth.”

Tech execs will turn to “shadow HR” for talent acquisition

Another one of Forrester’s predictions is that 20% of tech executives, out of frustration, will establish their own talent capabilities in-house, or “shadow HR,” to keep their staff up to date in skills such as AI, given that talent is expensive and hard to find.

“A source of untapped talent sits right inside organizations already, and tech leaders are increasingly working with their HR counterparts to increase internal mobility into the technology organization,” Matthew Guarini, vice president and senior research director at Forrester, explained to TechRepublic. “Highly transferable skills such as stakeholder management, communication, critical thinking and customer experience are already available within the broader organization, and paired with upskilling and reskilling programs, can create a powerful new pipeline of talent into the technology organization.”

SEE: Study: Reskilling Is Inevitable As AI Changes How We Work (TechRepublic)

In the past five years, one-third of the required skills for tech roles have turned over, and this has left 76% of tech leaders with skills gaps in their organizations, according to the report. It noted that tech companies including Alphabet, Dell and Oracle “have ditched degree requirements for technology roles, preferring to screen candidates based on skills. Not every company will have the capability or HR capacity for this shift.”

Utilizing external partners is another measure tech execs will take to close the AI gap, the report said. “Acquiring skills for prompt engineering, scenario prioritization, partner development and new AI-infused applications will drive tech execs straight to their service providers,” the report said, adding that vendors including Accenture, Tata Consultancy Services and Wipro are investing heavily in these practices.

More on how AI will impact jobs and work

The report recommends that tech leaders “up their game with transparency and AI-driven automation” wherever they operate, which will cause a shift in work. This will result in “a premium put on higher value-creating roles, some of which will be new while the majority will just adapt,” Guarini said. At the same time, AI/automation will reduce the need for more manual-intensive jobs.

“Most forecasts recognize this dynamic as it has played out in the past with automation,” he added. “The difference will be AI’s impact on more jobs than automation, and its ability to impact jobs more deeply. The opportunity for tech leaders is to identify ways to augment their talent with these tools and then reposition this enhanced capability to drive better outcomes for the business and customers.”

Global tech operations will be rebalanced amid geopolitical pressures

Forrester is also predicting that two-thirds of firms will shift their manufacturing footprints due to geopolitical pressures. The report noted that 72% of 2023 corporate filings cited geopolitics or export controls as impacting the outlook of their long-term business models.

This means that more sourcing will be done within the U.S. – if U.S.-based companies are exposed to China and Eastern Europe, Guarini said. “The main drivers for companies to think through would be current exposure to countries that have weakening trade relations with where they are domiciled, current exposure to countries with civil unrest, (and) current exposure to countries with growing economic instability,” he said.

SEE: Gen AI to Increase US Production — With Caveats (TechRepublic)

These factors all indicate the need to diversify from these exposures to mitigate any adverse impact on their business in the short/medium/long term, Guarini said.

Subscribe to the Executive Briefing Newsletter

Discover the secrets to IT leadership success with these tips on project management, budgets, and dealing with day-to-day challenges.

Delivered Tuesdays and Thursdays Sign up today

Microsoft’s mobile keyboard app SwiftKey gains new AI-powered features

Microsoft’s mobile keyboard app SwiftKey gains new AI-powered features Sarah Perez @sarahintampa / 7 hours

Along with AI advances in Windows 11 and Bing, Microsoft also this week announced it’s bringing new AI-powered features to its SwiftKey mobile keyboard app for iOS and Android. This third-party app lets users replace the default keyboard on their phone with an intelligent keyboard that learns your writing style so you can type faster. Now, it will also include AI camera lenses, AI stickers, an AI-powered editor, and the ability to create AI images from the app.

The new AI camera lenses will let users create photos, videos, and GIFs with different effects, including lenses that are powered by Microsoft’s collaboration with Snapchat maker Snap. There are now over 250 tools and filters available to help you express yourself, the company notes.

Also new are AI stickers which can be produced by Bing’s Image Creator, which lets you make stickers based on your own photos or selfies. You can then share these with friends and family when chatting in various communication apps like WhatsApp, Messenger, and others.

Bing Image Creator is also now accessible directly from the app’s keyboard, allowing you to take a picture or upload an existing photo to get Visual Search results from Bing in the app.

In addition, the app is gaining an AI-powered “Editor” feature that will help users improve their grammar, spelling, and punctuation. To use this option, you can highlight any sentence and then get instant feedback and suggestions from the editor.

Image Credits: Microsoft

Microsoft has been working to upgrade SwiftKey for many months, having earlier this year tied the app to Bing to allow users to search with Bing, chat with Bing Chat, or leverage AI to customize the tone of their text.

The new features are rolling out to SwiftKey on both iOS and Android.

Microsoft Bing to gain more personalized answers, support for DALLE-E 3, and watermarked AI images

Bengaluru-based SatSure to Invest $35M to Launch Satellites by 2025

Spacetech start-up, SatSure has announced plans to invest $35 million to launch 4 earth imaging satellites by 2025 after the company successfully performed edge processing in space last week.

The mission of this project is to become self-sufficient and start generating its own satellite imagery rather than sourcing data from foreign European and American space agencies. The Bengaluru-based start-up offers high-resolution optical and multispectral data which serves purposes like risk management, crop insurance, agricultural applications, credit assessment and more.

In order to develop satellites, SatSure has partnered with organisations specialising in hardware development. The companies include Spiral Blye, Satellogic, SkyServe among others. The project is on schedule as the payload is ready and aerial tests are to begin. The launch is planned for November 2025.

The satellites offer almost 65 km of coverage and one-metre resolution. SatSure’s data is designed mainly to focus on national security and disaster management. It is also adaptable to various geographic regions as well.

“We have the most impact on the downstream. So that is where we are playing right now, “ said Divya Sharma, vice president of ML and data analytics. And then in two years when our supply goes up then it is a contract because we’ll be absorbing the data with our data collection platform” she said in an exclusive interview with AIM.

This bold investment initiative arrives on the scene at an opportune moment, coinciding with the surge of interest and investment in India’s burgeoning spacetech industry, catalysed by the resounding success of the Chandrayaan-3 mission. With a current valuation of $20 million, SatSure is unshackling itself from its regional roots and setting its sights on a global expansion horizon.

The post Bengaluru-based SatSure to Invest $35M to Launch Satellites by 2025 appeared first on Analytics India Magazine.

Optimizing Data Storage: Exploring Data Types and Normalization in SQL

Optimizing Data Storage: Exploring Data Types and Normalization in SQL
Image by Author

In the present century, data is the new oil. Optimizing this data storage is always critical for getting a good performance from it. Opting for suitable data types and applying the correct normalization process is essential in deciding its performance.

This article will study the most important and commonly used datatypes and understand the normalization process.

Data Types in SQL

There are mainly two data types in SQL: String and Numeric. Other than this, there are additional data types like Boolean, Date and Time, Array, Interval, XML, etc.

String Data Types

These data types are used to store character strings. The string is often implemented as an array data type and contains a sequence of elements, typically characters.

  1. CHAR(n):

It is a fixed-length string that can contain characters, numbers, and special characters. n denotes the maximum length of the string in characters it can hold.

Its maximum range is from 0 to 255 characters, and the problem with this data type is that it takes the full space specified, even if the actual length of the string is less than then. The extra string length is padded with extra memory space.

  1. VARCHAR(n):

Varchar is similar to Char but can support strings of variable size, and there is no padding. The storage size of this data type is equal to the actual length of the string.

It can store up to a maximum of 65535 characters. Due to its variable size nature, its performance is not as good as the CHAR data type.

  1. BINARY(n):

It is similar to the CHAR data type but only accepts binary strings or binary data. It can be used to store images, files, or any serialized objects. There is another data type VARBINARY(n) which is similar to the VARCHAR data type but also accepts only binary strings or binary data.

  1. TEXT(n):

This data type is also used to store the strings but has a maximum size of 65535 bytes.

  1. BLOB(n): Stands for Binary Large Object and hold data up to 65535 bytes.

Other than these are other data types, like LONGTEXT and LONGBLOB, which can store even more characters.

Numeric Data Types

  1. INT():

It can store a numeric integer, which is 4 bytes (32bit). Here n denotes the display width, which can be a maximum of up to 255. It specifies the minimum number of characters used to display the integer values.

Range:

  1. a) -2147483648 <= Signed INT <= 2147483647
  2. b) 0 <= Unsigned INT <= 4294967295
  1. BIGINT():

It can store a large integer of size up to 64 bits.

Range:

  1. a) -9223372036854775808 <= Signed BIGINT <= 9223372036854775807
  2. b) 0 <= Unsigned BIGINT <= 18446744073709551615
  1. FLOAT():

It can store floating point numbers with decimal places approximated with a certain precision. It has some small rounding errors, so because of this, it is not suitable where exact precision is required.

  1. DOUBLE():

This data type represents double-precision floating-point numbers. It can store decimal values with a higher precision as compared to the FLOAT data type.

  1. DECIMAL(n, d):

This data type represents exact decimal numbers with a fixed precision denoted by d. The parameter d specifies the number of digits after the decimal point, and the parameter n denotes the size of the number. The maximum value for d is 30, and its default value is 0.

Some other Data Types

  1. BOOLEAN:

This data type stores only two states which are True or False. It is used to perform logical operations.

  1. ENUM:

It stands for Enumeration. It allows you to choose one value from the list of predefined options. It also ensures that the stored value is only from the specified options.

For example, consider an attribute color that can only be 'Red,' 'Green,' or 'Blue'. When we put these values in ENUM, then the value of the color can only be from these specified colors only.

  1. XML:

XML stands for eXtensible Markup Language. This data type is used to store XML data which is used for structured data representation.

  1. AutoNumber:

It is an integer that automatically increments its value when each record is added. It is used in generating unique or sequential numbers.

  1. Hyperlink:

It can store the hyperlinks of files and web pages.

This completes our discussion on SQL Data Types. There are many more data types, but the data types that we have discussed are the most commonly used ones.

Normalization In SQL

Normalization is the process of removing redundancies, inconsistencies, and anomalies from the database. Redundancy means the presence of duplicate values of the same piece of data, whereas inconsistencies in the database represent the same data exists in multiple formats in multiple tables.

Database anomalies can be defined as any sudden change or discrepancies in the database that are not supposed to exist. These changes can be due to various reasons, such as data corruption, hardware failure, software bugs, etc. Anomalies can lead to severe consequences, such as data loss or inconsistency, so detecting and fixing them as soon as possible is essential. There are mainly three types of anomalies. We will briefly discuss each but refer to this article if you want to read more.

  1. Insertion Anomaly:

When the newly inserted row creates, inconsistency in the table leads to an insertion anomaly. For example, we want to add an employee to an organization, but his department is not allocated to him. Then we cannot add that employee to the table, which creates an insertion anomaly.

  1. Deletion Anomaly:

Deletion anomaly occurs when we want to delete some rows from the table, and some other data is required to be deleted from the database.

  1. Update Anomaly:

This anomaly occurs when we want to update some rows and which leads to inconsistency in the database.

The normalization process contains a series of guidelines that make the design of the database efficient, optimized, and free from redundancies and anomalies. There are several types of normal forms like 1NF, 2NF, 3NF, BCNF, etc.

1. First Normal Form (1NF)

The first normal form ensures that the table contains no composite or multi-valued attributes. It means that only one value is present in a single attribute. A relation is in first normal form if every attribute is only single-valued.

For Ex-

Optimizing Data Storage: Exploring Data Types and Normalization in SQL
Image by GeeksForGeeks

In Table 1, the attribute STUD_PHONE contains more than one phone number. But in Table 2, this attribute is decomposed into 1st normal form.

2. Second Normal Form

The table must be in the first normal form, and there must not be any partial dependencies in the relations. Partial dependency means that the non-prime attribute (attributes which are not part of the candidate key) is partially dependent or depends on any proper subset of the candidate key. For the relations to be in the second normal form, the non-prime attributes must be fully functional and dependent on the entire candidate key.

For example, consider a table named Employees having the following attributes.

EmployeeID (Primary Key)  ProjectID (Primary Key)  EmployeeName  ProjectName  HoursWorked

Here the EmployeeID and the ProjectID together form the primary key. However, you can notice a partial dependency between EmployeeName and EmployeeID. It means that the EmployeeName is dependent only on the part of the primary key (i.e., EmployeeID). For complete dependency, the EmployeeName must depend on both EmployeeID and the ProjectID. So, this violates the principle of the second normal form.

To make this relation in the second normal form, we must split the tables into two separate tables. The first table contains all the employee details, and the second contains all the project details.

Therefore, the Employee table has the following attributes,

EmployeeID (Primary Key)  EmployeeName

And the Project table has the following attributes,

Project ID (Primary Key)  Project Name  Hours Worked

Now you can see that the partial dependency is removed by creating two independent tables. And the non-prime attributes of both tables depend on the complete set of the primary key.

3. Third Normal Form

After 2NF, still, the relations can have update anomalies. It may happen if we update only one tuple and not the other. That would lead to inconsistency in the database.

The condition for the third normal form is that the table should be in the 2NF, and there is no transitive dependency for the non-prime attributes. Transitive dependency happens when a non-prime attribute depends on another non-prime attribute instead of directly depending on the primary attribute. Prime attributes are the attributes that are part of the candidate key.

Consider a relation R(A, B, C), where A is the primary key and B & C are the non-prime attributes. Let A→B and B→C be two Functional Dependencies, then A→C will be the transitive dependency. It means that attribute C is not directly determined by A. B acts as a middleman between them.

If a table consists of a transitive dependency, then we can bring the table into 3NF by splitting the table into separate independent relations.

4. Boyce-Codd Normal Form

Although 2NF and 3NF remove most of the redundancies, still the redundancies are not 100% removed. Redundancy can occur if the LHS of the functional dependency is not a candidate or super key. A Candidate Key forms from the prime attributes, and the Super Key is a superset of the candidate key. To overcome this issue, another type of functional dependency is available named Boyce Codd Normal Form (BCNF).

For a table to be in BCNF, the left-hand side of a functional dependency must be a candidate key or a super key. A. For example, for a functional dependency X→Y, X must be a candidate or super key.

Consider an Employee Table that contains the following attributes.

  1. Employee ID (primary key)
  2. Employee Name
  3. Department
  4. Department Head

Optimizing Data Storage: Exploring Data Types and Normalization in SQL

The EmployeeID is the primary key that uniquely identifies each row. The Department attribute represents the department of a particular employee, and the Department Head attribute represents the Employee ID of the employee who is the head of that specific department.

Now we will check if this table is in the BCNF. The condition is that the LHS of the functional dependency must be a super key. Below are the two functional dependencies of that table.

Functional Dependency 1: Employee ID → Employee Name, Department, Department Head

Functional Dependency 2: Department → Department Head

For the FD1, the EmployeeID is the primary key, which is also a super key. But for FD2, Department is not the super key because multiple employees can be in the same department.

Therefore this table violates the condition of BCNF. To satisfy the property of BCNF, we need to split that table into two separate tables: Employees and Departments. The Employees table contains the EmployeeID, EmployeeName, and Department, and the Department table will have the Department and the Department Head.

Optimizing Data Storage: Exploring Data Types and Normalization in SQL
Optimizing Data Storage: Exploring Data Types and Normalization in SQL

Now we can see in both tables that all the functional dependencies are dependent on the primary keys, i.e., there are no non-trivial dependencies.

We have covered all the famous normalization techniques, but other than these, there are two more normal forms, namely 4NF and 5NF. If you want to read more about them, refer to this article from GeeksForGeeks.

Wrapping it Up

We have discussed the most commonly used data types in SQL and the significant Normalization techniques in database management systems. While designing a database system, we aim to make it scalable, minimizing redundancy and ensuring data integrity.

We can create a delicate balance between storage, precision, and memory consumption by selecting appropriate data types. Also, the normalization process helps eliminate data anomalies and make the schema more organized.

It is all for today. Until then, keep reading and keep learning.
Aryan Garg is a B.Tech. Electrical Engineering student, currently in the final year of his undergrad. His interest lies in the field of Web Development and Machine Learning. He have pursued this interest and am eager to work more in these directions.

More On This Topic

  • Database Optimization: Exploring Indexes in SQL
  • Data Science 101: Normalization, Standardization, and Regularization
  • Data Transformation: Standardization vs Normalization
  • Does the Random Forest Algorithm Need Normalization?
  • Cloud Data Warehouse is The Future of Data Storage
  • What Comes After HDF5? Seeking a Data Storage Format for Deep Learning

MP Working on DigiPass, DigiYatra for Government Offices

Facial recognition

In a significant development, the Madhya Pradesh government is actively working on the implementation of DigiPass, a system akin to DigiYatra, for use in government ministries and offices. This move is poised to streamline administrative processes and improve efficiency.

Abhijeet Agrawal, Managing Director of the M.P. State Electronics Development Corporation under the Department of Science & Technology, Government of Madhya Pradesh, said during the AWS Public Sector Symposium in New Delhi.

While DigiYatra has transformed the passenger experience at airports, DigiPass holds the potential to revolutionise government operations by introducing a paperless and efficient workflow.

In parallel, the Airport Authority of India and the Ministry of Civil Aviation are set to expand the DigiYatra paperless boarding system to six additional airports across India. Already operational in Delhi, Varanasi, and Bengaluru, this contactless facial recognition technology expedites the entire airport journey, from entry to boarding, saving passengers valuable time.

Here’s a closer look at the DigiYatra system and its expansion plans:

In parallel, the Airport Authority of India and the Ministry of Civil Aviation are set to expand the DigiYatra paperless boarding system to six additional airports across India. Already operational in Delhi, Varanasi, and Bengaluru, this contactless facial recognition technology expedites the entire airport journey, from entry to boarding, saving passengers valuable time.

Launched in December last year, DigiYatra relies on facial recognition technology to offer passengers a quicker, paperless airport experience. It allows registered travellers to bypass airline counter queues, fast-track cabin luggage checks, and save approximately 15-25 minutes of their time. The system is set to be implemented in the airports of Mumbai, Ahmedabad, Kochi, Lucknow, Jaipur, and Guwahati this month, enhancing passenger convenience across major cities.

The post MP Working on DigiPass, DigiYatra for Government Offices appeared first on Analytics India Magazine.

YouTube launches Youtube Create, a New AI-Powered Video Editing App

Google owned Youtube has announced it was launching a new AI-powered video editing app called YouTube Create that would allow anyone to create and share videos on its platform.

YouTube Create will include features such as precision editing and trimming, automatic voiceover, captioning and transitions, the company said in a blog post. It would also begin testing a new feature called ‘Dream Screen’ that would let users add an AI-generated video or image to their videos by just typing an idea in the chat box. For example, users could type “I want to be in Paris” and the app would insert a realistic video or image of them in the French capital.

The app will also have a generative AI feature that will help users generate topic ideas and outlines for videos, based on trending topics and audience preferences.Additionally, it will have an AI-powered music recommendation feature, where users can enter a written description of their video and the app will suggest an audio track to use. Users will also have an option to automatically dub their videos into foreign languages, YouTube said.

The company said the new app was aimed at making video production easier and more accessible for everyone, especially first-time creators.

“We want to make it easier for everyone to feel like they can create and we believe generative AI will make that possible,” said YouTube chief Neal Mohan.

YouTube Create will compete with apps such as TikTok and Meta’s Instagram, which have gained popularity among young users who create short videos or reels. The app is in beta mode on Android in select countries. Users in India, US, UK, Germany, France,Indonesia, Singapore and Korea will get access to it first.

Google’s Foray into Generative-AI

YouTube’s Create, the new AI-powered feature, is Google’s latest attempt to gain a foothold in the AI space. Last week, Google updated Bard, its AI assistant, with generative AI and ML capabilities to compete with OpenAI’s ChatGPT. Google’s ambition is not limited to chatbots, but extends to integrating AI with its entire suite of products. This strategy could give Google an edge in the AI race, as it can leverage user insights from all its offerings to train its AI and potentially outperform its rivals.

The post YouTube launches Youtube Create, a New AI-Powered Video Editing App appeared first on Analytics India Magazine.

Exploring Neural Networks

Imagine a machine thinking, learning, and adapting like the human brain and discovering hidden patterns within data.

This technology, Neural Networks (NN), algorithms are mimicking cognition. We'll explore what NNs are and how they function later.

In this article, I'll explain to you the Neural Networks (NN) fundamental aspects — structure, types, real-life applications, and key terms defining operation.

What is a Neural Network? Exploring Neural Networks
Source: vitalflux.com

Algorithms called Neural Networks (NN) try to find relationships within data, imitating the human brain's operations for "learning" from data.

Neural networks can be mixed with deep learning and machine learning. So it will be good to explain these terms first. Let’s start.

Neural Network vs. Deep Learning vs. Machine Learning

Neural Networks form the foundation of Deep Learning, a subset of Machine Learning. While Machine Learning models learn from data and make predictions, Deep Learning goes deeper and can process huge amounts of data, recognizing complex patterns.

If you want to learn more about Machine Learning algorithms, read this one.

Moreover, these neural networks have become integral parts of many fields, serving as the backbone of numerous modern technologies, which we will see in later sections. These applications range from face recognition to natural language processing.

Let's explore some common areas where Neural Networks play a vital role in improving daily life.

Types of Neural Network

Real-world applications enrich understanding of Neural Networks, revolutionizing traditional methods across industries with accurate, efficient solutions.

Let's highlight intriguing examples of Neural Networks driving innovation and transforming everyday experiences, including Neural Network Types.

Exploring Neural Networks
Image by Author

ANN (Artificial Neural Networks):

Artificial Neural Network (ANN), architecture is inspired by the biological neural network of the human brain. The network consists of interconnected layers, input, hidden, and output. Each layer contains multiple neurons that are connected to every neuron in the adjacent layer.

As data moves through the network, each connection applies a weight, and each neuron applies an activation function like ReLU, Sigmoid, or Tanh. These functions introduce non-linearity, making it possible for the network to learn from errors and make complex decisions.

During training, a technique called backpropagation is used to adjust these weights. This technique uses gradient descent to minimize a predefined loss function, aiming to make the network's predictions as accurate as possible.

ANN Use Cases

Customer Churn Prediction

ANNs analyze multiple features like user behavior, purchase history, and interaction with customer service to predict the likelihood of a customer leaving the service.

ANNs can model complex relationships between these features, providing a nuanced view that's crucial for predicting customer churn accurately.

Sales Forecasting

ANNs take historical sales data and other variables like marketing spend, seasonality, and economic indicators to predict future sales.

Their ability to learn from errors and adjust for complex, non-linear relationships between variables makes them well-suited for this task.

Spam Filtering

ANNs analyze the content, context, and other features of emails to classify them as spam or not.

They can learn to recognize new spam patterns, adapting over time, which makes them effective at filtering out unwanted messages.

CNN (Convolutional Neural Networks):

Convolutional Neural Networks (CNNs) are designed specifically for tasks that involve spatial hierarchies, like image recognition. The network uses specialized layers called convolutional layers to apply a series of filters to an input image, producing a set of feature maps.

These feature maps are then passed through pooling layers that reduce their dimensionality, making the network computationally more efficient. Finally, one or more fully connected layers perform classification.

The training process involves backpropagation, much like ANNs, but tailored to preserve the spatial hierarchy of features.

CNN Use Cases

Image Classification

CNNs apply a series of filters and pooling layers to automatically recognize hierarchical patterns in images.

Their ability to reduce dimensionality and focus on essential features makes them efficient and accurate for categorizing images.

Object Detection

CNNs not only classify but also localize objects within an image by drawing bounding boxes.

The architecture is designed to recognize spatial hierarchies, making it capable of identifying multiple objects within a single image.

Image Segmentation

CNNs can assign a label to each pixel in the image, classifying it as belonging to a particular object or background.

The network's granular, pixel-level understanding makes it ideal for tasks like medical imaging where precise segmentation is crucial.

RNN (Recurrent Neural Networks):

Recurrent Neural Networks (RNNs) differ in that they have an internal loop, or recurrent architecture, that allows them to store information. This makes them ideal for handling sequential data, as each neuron can use its internal state to remember information from previous time steps in the sequence.

While processing the data, the network takes into account both the current and previous inputs, allowing it to develop a kind of short-term memory. However, RNNs can suffer from issues like vanishing and exploding gradients, which make learning long-range dependencies in data difficult.

To address these issues, more advanced versions like Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU) networks were developed.

RNN Use Cases

Speech-to-text

RNNs take audio sequences as input and produce a text sequence as output, taking into account the temporal dependencies in spoken language.

The recurrent nature of RNNs allows them to consider the sequence of audio inputs, making them adept at understanding the context and nuances in human speech.

Machine Translation

RNNs convert a sequence from one language to another, considering the entire input sequence to produce an accurate output sequence.

The sequence-to-sequence learning capability maintains context between languages, making translations more accurate and contextually relevant.

Sentiment Analysis

RNNs analyze sequences of text to identify and extract opinions and feelings.

The memory feature in RNNs helps capture the emotional build-up in textual sequences, making them suitable for sentiment analysis tasks.

Final Thoughts

Looking ahead, the future promises continued Neural Network advancement and special use cases. As algorithms evolve to handle more complex data, they will unlock new possibilities in healthcare, transportation, finance, and beyond.

To learn neural networks, doing a real-life project is very effective. From recognizing faces to predicting diseases, they are reshaping the way we live and work.

In this article, we reviewed its fundamentals, real-life examples like face detecting and recognition, and more.

Thanks for reading!
Nate Rosidi is a data scientist and in product strategy. He's also an adjunct professor teaching analytics, and is the founder of StrataScratch, a platform helping data scientists prepare for their interviews with real interview questions from top companies. Connect with him on Twitter: StrataScratch or LinkedIn.

More On This Topic

  • Google’s Model Search is a New Open Source Framework that Uses Neural…
  • IBM Uses Continual Learning to Avoid The Amnesia Problem in Neural Networks
  • Explainable Visual Reasoning: How MIT Builds Neural Networks that can…
  • Neural Networks from a Bayesian Perspective
  • Interpretable Neural Networks with PyTorch
  • Deep Neural Networks Don't Lead Us Towards AGI