AI Cloud Platform Gradient Partners With MongoDB

Gradient AI, a startup where developers can build and customise AI applications in the cloud using large language models, announced its partnership with database giant MongoDB. The partnership is said to strengthen the collaborative ties between the two companies helping Gradient AI users to build applications using the data platform.

Developing custom AI solutions that leverage both fine-tuning and Retrieval Augmented Generation (RAG) can pose challenges. However, individuals can build robust and cost-effective solutions with the assistance of Gradient and top frameworks such as MongoDB and LlamaIndex.

Chatbot With Multi-Integration

Financial chatbots can be created with MongoDB, where RAG can be used to reduce hallucination, MongoDB for Vector Store and LlamaIndex for RAG Orchestration. For instance, LlamaIndex can be used for uploading a PSF (point string function) into a raw string. Through Gradient’s web API’s, the models can be fine tuned, generate completions and embeddings from one platform. Connections can be made with Atlas Vector Search, and LlamaIndex can be used to bring together all the elements.

Source: GradientAI

Last month, Gradient came out of stealth mode with a $10 million funding led by Wing VC, Mango Capital, Tokyo Black and other VC platforms. The company was started with the focus to help enterprises unlock the full benefits of language models which require a reliable way to add private, proprietary data to them.

Gradient users are not required to initiate the training of LLMs from the beginning as the platform provides a selection of open-source LLMs, including Meta’s Llama 2, that users can fine-tune according to their requirements. Furthermore, Gradient offers pre-trained models designed for specific use cases such as data reconciliation, context gathering, and others that cater to various industries.

The post AI Cloud Platform Gradient Partners With MongoDB appeared first on Analytics India Magazine.

Data-driven, AI-powered supply chain part 3: Imagining the Future – Supply chain 5.0

Header-for-Part-3-1
Data-driven, AI-powered Supply Chain

The immediate and pressing need for ‘digitizing’ your supply-chain

  • According to a 2020 CGE report, companies which commit to digitizing their supply chains can expect to save up to 50 percent on supply-chain costs; besides, a 20% reduction in procurement costs, apart from an increase in revenue of over 10% — resulting from incremental competitive advantage. ( Source: Center for Global Enterprise (CGE) in partnership with CREATe.org)
  • World-class supply chain organizations can save up to 45 % in procurement costs through digital transformation. (Source: Hackett Group)

One may conclude: ‘Digitizing’ the supply-chain has become a survival necessity for companies to stay competitive. Apart from a substantial jump in the efficiency-effectiveness, the customer-experience, and upside to revenues, companies can expect a huge-huge cost-saving…

A Look at the Future: Components of Data-driven (Digital) Supply-chain

Note

AnyData-driven Supply-chain, or AI-driven Supply-chain will only be as good as the ‘quality, granularity, and richness’ of the underlying “data” supporting the decisions.So the most critical component of any “Digital Supply-chain” is the data, and the underlying Information Supply-chain supporting the Physical goods Supply-chain.

A good strategy consultant, among other things, is expected to begin by describing the end goal… So, before we attempt to create a roadmap for building a data-driven supply chain, let us first try to envision what it might look like.

Defining the “ideal” end-state: Featuring an AI-driven Digital Supply-chain of future

The supply chains of the future will be powered by self-managing, automated technologies that can collect and analyze data in real time. Few examples below:

  • #Cloud-based SCM tools & AI-driven algorithms for supply-chain planning and optimization
  • #Granular geo-tagged data from extended value-chain covering supplier’s suppliers; to customers’ customers
  • #Location-data (Location-stamp on every transaction)
  • #Robotics in factories and warehouses: Robotic Automatic Storage and Retrieval systems
  • #Self-driving trucks with GPS
  • #IOT and Radio-tags embedded into into every SKU & every device in the Supply chain
  • #Internet Based AI driven Bots (Listening Posts) for continuously scanning the market (internet) for better sources, better materials & better pricing & for quickly gathering & processing risk data (Geopolitical / supply-related)
  • #AI-driven automated processes wherever possible
  • #Data-driven decision-making across the value-chain.
DDSC-Figure-9-Imaging-the-Future-AI-Driven-Digital-Supply-chain-1
Figure 9: Imaging the Future: Data-driven, AI-powered Digital Supply-chain

1. Data Collection: This is the most important part of the ‘Digital Supply-chain’. The future supply chains are expected to extensively use IOT, Radio tags, GPS, etc. for Automated Data Acquisition (event-driven, or period-driven in rare cases).

  • # Enterprise data enhanced by seamlessly integrated geo-tagged micro-market data (perhaps made available on subscription mode by a third party)
  • # All transactional data without exception will be geo-tagged & instantly distributed to cloud-based Apps & Data lake-houses.
  • # Every Truck would have a GPS its location will be tracked 24×7. Every “Distribution Package” will carry a Radio-tag and its movement will be tracked through supply-chain.
  • # Every purchase, every order fulfillment, every job progress, every project will be tracked with time & location-stamped data.
  • # Every product & Every Job would carry a Radio-tag, and its movement through the workflow-stages would be automatically tracked.
  • # Every Store & Warehouse would have completely automated, AI-driven — Automated Storage and Retrieval Systems (ASRS), that automatically load and unload the unitized-palletized loads, while keeping track of the locations (Bin Number) for every Part-No & every SKU.
  • # All equipment & devices without exception would be IOT/GPS enabled and would automatically generate and distribute data into multiple analytics engines for control & continuous improvement.
  • # All workers may carry id-tags with a GPS, or their mobiles will be tracked through a custom-corporate app.
  • # All intra-office communication on mobiles through short-messaging apps & e-mails will be time-stamped and location-stamped.
  • # Real-time Batch Tracking & Product Traceability data through Radio-tags & Unique-ids for each SKU
  • # All communications with customers and vendors will be time and location-stamped.
  • # The internal Enterprise-data will be complimented/supplemented with external-data, from sources such as Census data, Market data, etc., — all sorted, normalized, and geo-tagged, before uploading into Data Lake-house.

2. Transportation:

  • # Every Truck would be Self-driven, GPS enabled, and will have the facilities for sending emergency distress signals.
  • # Digital Twin (SCADA equivalent system to control Transportation network) to carry drill-down information of every package, every SKU carried by every Truck, down to the level of ‘serial number’ of every SKU.
  • # AI-controlled Self-driving Drones for delivering products — especially for the last-mile delivery.

3. Warehousing & Supply-chain Operations: Completely Robotized — Automated AI-driven –

  • # Robotized picking and packing, AI-driven automated storage and retrieval system (ASRS)
  • # Automated Instant-Inventory taking at the press of a button using radiotags on every box, every pallet.

4. Customer / Vendor Collaboration

  • # Generative AI-based Chatbots to communicate
  • # Generative AI-based IVRS & Robotic calls for voice communication
  • # Generative AI-based Social Media Campaigns

5. Risk Management :

  • # Risk data to be collected using AI-driven Bot Listeners for picking up “Risk Signals”, “Mavericks”, or any “unusual activity”, or “unusual communications” — both from internal-enterprise-data & from the other external-data-sources including Public & Private databases, Social Media, and the Internet.
  • # AI-driven Risk Reporting, Early-Warning-Signals
  • # AI-driven automated ‘course-correction’ and ‘communications’ with vendors, customers, distributors, and employees.
  • # The Lifed-items (SKUs with an expiry-date) to be picked and packed based on shorted expiry Batches first.

6. Product Traceability, Batch Tracking & Adverse-event Reporting

  • # Real-time identification (Batch No & Serial No.) & Location tracking of product through Radio-tags & Unique-ids for each SKU — which maps to product-traceability-data on the cloud
  • # Traceability data includes: the origin of materials and parts, their processing history, and their distribution, as stated in ISO 9000:2005 (2005) & Additionally, data connected to the past, use, or location of a product.
  • # Easy & Instantaneous Batch-Recall in case of any AERAdverse Event Reported (as in the case of Food or Pharmaceuticals, or even consumer-durables like automobiles ) with reference to a particular batch-number, would automatically trigger an alert along with a “Do not Sell / Do not Use Warning” — not just for the particular batch-number, but also for all the batches which used materials & parts from the same input-batches from same suppliers. The alert would automatically replicate itself through the entire supply-chain (distribution network) no matter where the particular ‘risky’ batches are located.

7. Digital Twins:

  • # To be Mobile based: To be operated from “Anywhere & Anytime” for Planning-monitoring-correction & control
  • # Digital Twins to map 100% of the extended-supply-chain, all stages, including the controllable Vendor’s Supply-chain & Customer’s Supply-chain.
  • # Control Tower & AI to Co-Pilot decision making: AI-driven control towerto help managers make decisions & manage the overall performance of the supply-chain. Managers handle extreme exceptions alone.

8. Global Sourcing & Continuous improvement

  • # Continuous & Automated Scanning of the Internet for better Suppliers, Materials & better prices — using Bots
  • # Use of Generative AI for Price-Negotiations with customers.
  • # AI-driven BOTs (Listening-posts) on the Internet, and Social-Media for picking-up signals on Threats (such as Adverse Events and geo-political risks), or Opportunities (better materials, better prices, better sources

Building a AI-driven Digital Supply-chain

I have extensively written on data-driven decision-making, data-to-decision lifecycle, and creating a roadmap and a business case for building a data-driven organization earlier. ( Please refer to my book Big Data for Big Decisions: Building a Data-driven Organization, P. Taylor & Francis, UK, 2022 & Articles on Data Science Central).

As I have argued: fancy advanced analytics platforms, and AI apart, the key to building a Data-driven organization is to focus on the key decisions — the 10% of the organizational decisions that account for 90% of business outcomes.

So for building a data-driven supply-chain, the following key questions must be asked and answered:

1. What are the key “supply-chain decisions” that account for over 80–90% of Supply-chain efficiency?

2. What is the data that supports the identified key decisions?

Note: The bigger and more primal issue most organizations face is: that they simply do not have the data that supports (the) key decisions within the organization…not with the requisite granularity or the quality. Needless to say, building a data-driven supply chain in the absence of the data is a futile and vain exercise, which nevertheless many CDOs are attempting to.

To quickly summarize, I advocate the following steps for building a data-driven supply-chain

Building A Data-Driven Organization Building A Digital Supply-Chain
Start from decisions. List all Organizational decisions, along with metrics like “Value-at-stake”, and the “frequency Create a Master-list of Process-Constraints & List decisions associated with each process constraint.Each process constraint at each workflow-stage represents a set-of decisions? (There can be many process-constraints, but only a few significant ones)The same / similar process-constraints may be recurring at multiple workflow-stages.. Note the Frequency of recurrence. Working out “value-at-stake” for each decision: Can we estimate the ‘loss of throughput due to bottlenecks, at each stage?
Identify the 10% of the decisions that account for 90% of business outcomes List the bottlenecksfrom the biggest to smallest? And the 10% of the bottlenecks that account for 90% of throughput loss? Create a List of process constraints associated with the 10% of ‘critical’ bottlenecks. Note: All workflow stages are important. The absence of a current bottleneck does not diminish the possibility of a future bottleneck.
Create a Roadmap & Business case Create a roadmap for digitizing the supply chainby prioritizing the process constraints associated with the 10% of bottlenecks responsible for 80-90% of throughput loss.
Note: Digitizing the key process constraints at all workflow stages is important… The criticality of a bottleneck (& associated process constraints) dictates its priority.
Data behind the Decisions Create & Map the Information Supply-chain & the following as needed: Services Supply-chain & Cashflow Supply-chain Map the flow of information (Data + Insights)’ required at each workflow-stage Map the flow of information (Data + insights) required for ‘easing-out’ the process-constraints and the possible bottlenecks.
Map the Information supply-chain in-its entirety – (from data creation, to distribution, consumption, archiving, retrieval and repurposing)
Cross-map the Information Supply-chain with Physical-goods Supply-chain – (clearly indicating the cause & effect relationship between the ‘information availability’ and its ‘impact on the process-constraint’.)
Data that you need vs. data you have Check: If you have all the data that you need to support the Supply-chain decisions
For the Delta-Data: Create & implement a Data sourcing strategy
Building Analytics to improve the quality of decisions Build Analytics & AI solutions to improve the quality of Supply-chain decisions
Build a Digital Twin: A remote replica of that provides the status & the control mechanism for the entire supply-chain (similar to SCADA systems). Typically you can change the process parameters using ‘digital twin’.
Table: Roadmap for Building a Data-driven, AI-Powered Supply chain

Read Earlier: Data-driven Supply Chain Part-2: Theory of Constraints & The Concept Of Information Supply Chain.

Read Earlier: Data-driven Supply Chain Part-1: Roadmap for Building an AI-Powered, Data-driven Supply Chain

OpenAI Hires Google TPU Lead to Head Hardware Division

OpenAI has made a strategic move by appointing Richard Ho as the head of its hardware division. Ho’s arrival marks a significant shift in the company’s focus on hardware optimisation and co-design, targeting data center networks, racks, and buildings.

Ho’s appointment coincides with OpenAI’s expansion plans, as the company is actively seeking talent for a deep learning hardware/software co-design engineer position. The job listing outlines a focus on optimising hardware for workloads, evaluating accelerators, and influencing the roadmaps for data center networks, racks, and buildings in collaboration with partners.

OpenAI’s commitment to hardware optimisation has seen parallel developments with Microsoft’s Azure cloud. The company is part of a significant $13 billion investment primarily in cloud credits, with plans to contribute insights to Microsoft’s AI chip development. This relationship underscores OpenAI’s pivotal role in shaping the future of AI hardware.

Rumours have been circulating about OpenAI contemplating its proprietary chip hardware and exploring potential acquisitions. CEO Sam Altman has been pursuing fundraising for a separate chip company, hinting at a deeper exploration into hardware innovation within the organisation.

Before joining OpenAI, Ho spearheaded the chip engineering division at Lightmatter, a photonic computing company. However, his illustrious career spans over nine years at Google, where he held a prominent role in the development of Google’s Cloud Tensor Processing Units (TPU) as the senior director of engineering.

The TPU chip family at Google plays a pivotal role in training and inference for large language models within Google’s ecosystem and as a part of their cloud services. Ho’s tenure at Google also involved co-inventing innovative methods using machine learning for chip architecture design, showcasing his expertise in the field.

“I aim to build teams that efficiently build novel chips attacking the hardest problems with creativity and flair,” remarked Ho in his LinkedIn profile. However, he declined to provide further details on his new role at OpenAI.

Amidst these pivotal developments, Altman’s temporary departure and reinstatement by the OpenAI board have triggered discussions about the company’s future trajectory. Ho was among the employees who expressed intentions to leave the company during Altman’s absence.

Coinciding with Ho’s arrival, OpenAI also welcomed Google’s Todd Underwood, who was entrusted with leading a new Site Reliability Engineering team focused on research and training workloads.

As OpenAI solidifies its hardware division under Ho’s leadership, the company is poised for significant strides in shaping the hardware landscape of artificial intelligence, bridging innovation and functionality in the pursuit of groundbreaking AI solutions.

The post OpenAI Hires Google TPU Lead to Head Hardware Division appeared first on Analytics India Magazine.

11 Python Magic Methods Every Programmer Should Know

11 Python Magic Methods Every Programmer Should Know
Image by Author

In Python, magic methods help you emulate the behavior of built-in functions in your Python classes. These methods have leading and trailing double underscores (__), and hence are also called dunder methods.

These magic methods also help you implement operator overloading in Python. You’ve probably seen examples of this. Like using the multiplication operator * with two integers gives the product. While using it with a string and an integer k gives the string repeated k times:

 >>> 3 * 4  12  >>> 'code' * 3  'codecodecode'

In this article, we’ll explore magic methods in Python by creating a simple two-dimensional vector Vector2D class.

We’ll start with methods you’re likely familiar with and gradually build up to more helpful magic methods.

Let's start writing some magic methods!

1. __init__

Consider the following Vector2D class:

class Vector2D:      pass

Once you create a class and instantiate an object, you can add attributes like so: obj_name.attribute_name = value.

However, instead of manually adding attributes to every instance that you create (not interesting at all, of course!), you need a way to initialize these attributes when you instantiate an object.

To do so you can define the __init__ method. Let's define the define the __init__ method for our Vector2D class:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y    v = Vector2D(3, 5)

2. __repr__

When you try to inspect or print out the object you instantiated, you’ll see that you don't get any helpful information.

v = Vector2D(3, 5)  print(v)
Output >>> <__main__.Vector2D object at 0x7d2fcfaf0ac0>

This is why you should add a representation string, a string representation of the object. To do so, add a __repr__ method like so:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"    v = Vector2D(3, 5)  print(v)
Output >>> Vector2D(x=3, y=5)

The __repr__ should include all the attributes and information needed to create an instance of the class. The __repr__ method is typically used for the purpose of debugging.

3. __str__

The __str__ is also used to add a string representation of the object. In general, the __str__ method is used to provide info to the end users of the class.

Let's add a __str__ method to our class:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __str__(self):          return f"Vector2D(x={self.x}, y={self.y})"    v = Vector2D(3, 5)  print(v)
Output >>> Vector2D(x=3, y=5)

If there is no implementation of __str__, it falls back to __repr__. So for every class that you create, you should—at the minimum—add a __repr__ method.

4. __eq__

Next, let's add a method to check for equality of any two objects of the Vector2D class. Two vector objects are equal if they have identical x and y coordinates.

Now create two Vector2D objects with equal values for both x and y and compare them for equality:

v1 = Vector2D(3, 5)  v2 = Vector2D(3, 5)  print(v1 == v2)

The result is False. Because by default the comparison checks for equality of the object IDs in memory.

Output >>> False

Let’s add the __eq__ method to check for equality:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __eq__(self, other):          return self.x == other.x and self.y == other.y

The equality checks should now work as expected:

v1 = Vector2D(3, 5)  v2 = Vector2D(3, 5)  print(v1 == v2)
Output >>> True 

5. __len__

Python’s built-in len() function helps you compute the length of built-in iterables. Let’s say, for a vector, length should return the number of elements that the vector contains.

So let’s add a __len__ method for the Vector2D class:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __len__(self):          return 2    v = Vector2D(3, 5)  print(len(v))

All objects of the Vector2D class are of length 2:

Output >>> 2

6. __add__

Now let's think of common operations we’d perform on vectors. Let's add magic methods to add and subtract any two vectors.

If you directly try to add two vector objects, you’ll run into errors. So you should add an __add__ method:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __add__(self, other):          return Vector2D(self.x + other.x, self.y + other.y)

You can now add any two vectors like so:

v1 = Vector2D(3, 5)  v2 = Vector2D(1, 2)  result = v1 + v2  print(result)
Output >>> Vector2D(x=4, y=7)

7. __sub__

Next, let’s add a __sub__ method to calculate the difference between any two objects of the Vector2D class:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __sub__(self, other):          return Vector2D(self.x - other.x, self.y - other.y)
v1 = Vector2D(3, 5)  v2 = Vector2D(1, 2)  result = v1 - v2  print(result)
Output >>> Vector2D(x=2, y=3)

8. __mul__

We can also define a __mul__ method to define multiplication between objects.

Let's implement let's handle

  • Scalar multiplication: the multiplication of a vector by scalar and
  • Inner product: the dot product of two vectors
class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __mul__(self, other):          # Scalar multiplication          if isinstance(other, (int, float)):              return Vector2D(self.x * other, self.y * other)          # Dot product          elif isinstance(other, Vector2D):              return self.x * other.x + self.y * other.y          else:              raise TypeError("Unsupported operand type for *")

Now we’ll take a couple of examples to see the __mul__ method in action.

v1 = Vector2D(3, 5)  v2 = Vector2D(1, 2)    # Scalar multiplication  result1 = v1 * 2  print(result1)    # Dot product  result2 = v1 * v2  print(result2)
Output >>>    Vector2D(x=6, y=10)  13

9. __getitem__

The __getitem__ magic method allows you to index into the objects and access attributes or slice of attributes using the familiar square-bracket [ ] syntax.

For an object v of the Vector2D class:

  • v[0]: x coordinate
  • v[1]: y coordinate

If you try accessing by index, you’ll run into errors:

v = Vector2D(3, 5)  print(v[0],v[1])
---------------------------------------------------------------------------    TypeError                             	Traceback (most recent call last)     in ()  ----> 1 print(v[0],v[1])    TypeError: 'Vector2D' object is not subscriptable

Let’s implement the __getitem__ method:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __getitem__(self, key):          if key == 0:              return self.x          elif key == 1:              return self.y          else:              raise IndexError("Index out of range")

You can now access the elements using their indexes as shown:

v = Vector2D(3, 5)  print(v[0])    print(v[1])
Output >>>    3  5

10. __call__

With an implementation of the __call__ method, you can call objects as if they were functions.

In the Vector2D class, we can implement a __call__ to scale a vector by a given factor:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y   	       def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __call__(self, scalar):          return Vector2D(self.x * scalar, self.y * scalar)

So if you now call 3, you’ll get the vector scaled by factor of 3:

v = Vector2D(3, 5)  result = v(3)  print(result)
Output >>> Vector2D(x=9, y=15)

11. __getattr__

The __getattr__ method is used to get the values of specific attributes of the objects.

For this example, we can add a __getattr__ dunder method that gets called to compute the magnitude (L2-norm) of the vector:

class Vector2D:      def __init__(self, x, y):          self.x = x          self.y = y        def __repr__(self):          return f"Vector2D(x={self.x}, y={self.y})"        def __getattr__(self, name):          if name == "magnitude":              return (self.x ** 2 + self.y ** 2) ** 0.5          else:              raise AttributeError(f"'Vector2D' object has no attribute '{name}'")

Let’s verify if this works as expected:

v = Vector2D(3, 4)  print(v.magnitude)
Output >>> 5.0

Conclusion

That's all for this tutorial! I hope you learned how to add magic methods to your class to emulate the behavior of built-in functions.

We’ve covered some of the most useful magic methods. But this is not this is not an exhaustive list. To further your understanding, create a Python class of your choice and add magic methods depending on the functionality required. Keep coding!

Bala Priya C is a developer and technical writer from India. She likes working at the intersection of math, programming, data science, and content creation. Her areas of interest and expertise include DevOps, data science, and natural language processing. She enjoys reading, writing, coding, and coffee! Currently, she's working on learning and sharing her knowledge with the developer community by authoring tutorials, how-to guides, opinion pieces, and more.

More On This Topic

  • Python f-Strings Magic: 5 Game-Changing Tricks Every Coder Needs to Know
  • 25 Github Repositories Every Python Developer Should Know
  • The 6 Python Machine Learning Tools Every Data Scientist Should Know About
  • Three R Libraries Every Data Scientist Should Know (Even if You Use Python)
  • KDnuggets News, May 25: The 6 Python Machine Learning Tools Every…
  • 12 Docker Commands Every Data Scientist Should Know

Elections 2024: How AI will fool voters if we don’t do something now

capitolgettyimages-1773573761

Back in September, I investigated what generative AI would do to our elections. One of the bigger surprises was how much of an effect generative AI is likely to have on local elections.

Also: How Microsoft plans to protect elections from deepfakes

After reading the article, Prof. Robert Crossler reached out to discuss the issues surrounding elections in our digital-centric society. Crossler, an associate professor of information systems at Washington State University, is eminently qualified to discuss this area. His recent article in Government Information Quarterly looks at measuring the effect of political alignment, platforms, and fake news consumption on voter concern for election processes.

Crossler's work has been funded by the National Science Foundation and the Department of Defense. He served as president of the AIS Special Interest Group on Information Security and Privacy from 2019 to 2020. He was also awarded the 2013 Information Systems Society's Design Science Award for his work on information privacy.

Crossler and I had a fascinating discussion over email. What follows is the contents of our discussion. My questions are in bold font. The only edits are for punctuation.

ZDNET: In what ways can generative AI be leveraged to influence voters or election outcomes, and what are the potential risks associated with its use in this context?

Robert Crossler: Generative AI can be leveraged to further polarize voters. If there is an issue that one party wants to convince another person of, then generative AI can be used to create a narrative that supports this argument. The potential risk is that voters will believe something that is not true, but feel strongly that it is true because they have seen very realistic material that supports the viewpoint that they were led to believe.

Generative AI has the potential to target communications specifically at people based on easily gathered knowledge that people put in the public sphere. This is a technique that we are already seeing utilized for social-engineering attacks from a cybersecurity perspective.

With generative AI, the tools could be used to customize political messaging based on what the generative AI can easily determine are the interests and motivating factors of voters. Being able to do this at scale has the potential to take what have been targets made based on demographics to targets based on a much more granular understanding of the user.

Also: How to use Bing Chat (and how it's different from ChatGPT)

ZDNET: What are the ethical considerations regarding use of generative AI in the electoral process, and how should policymakers and the public approach these concerns?

RC: I wish I knew the answer to this question. But this is the exact question that should be asked and answered. Ethically, I believe that technology should not be used to distort the truth, or better yet make up alternative truths. However, utilizing policy to establish and then enforce ethical considerations in a changing technology is a difficult issue to get correct.

While it is possible to get politicians to agree on an ethical standard when it comes to the use of generative AI, I'm not sure how this applies the further you get from politicians. For example, Political Action Groups, how would this apply to them?

How would limiting their use of this technology impede on their First Amendment rights? I'm not a legal scholar, but anytime the government tries to tell people what they can and can't say, I start to get concerned.

Also: Google to require political ads to reveal if they're AI-generated

ZDNET: Can you share any findings or insights from your research that have implications for the broader understanding of technology's role in the democratic process?

RC: One of the things I found in my research is the role that social media, especially Facebook, has in spreading fake news. As a result, I don't use Facebook and I put little credence to what I read in other social media contexts. I would encourage others to not get drawn into a particular way of thinking based on what they are seeing in social media.

This is especially important given the algorithms that social media uses to engage users. These algorithms are written to show you more of what you engage with and less of what you don't. The result is an echo chamber of sorts, where the same information is confirmed but contrary information isn't necessarily shared.

Also: Facebook bans political campaigns from using its new AI-powered ad tools

ZDNET: What is the role of foreign actors in fake news and generative AI?

RC: Microsoft released information that supported the argument that China was using AI in social media to influence US voters. China and Russia are using generative AI deepfakes to control the narrative regarding the war in Ukraine. We're seeing AI used to try and control narratives around the conflict in Hamas and Israel.

Given what Microsoft reported and how generative AI is being used by foreign actors in relation to the war narrative, it seems reasonable to conclude that this will amplify as we approach Election Day 2024.

ZDNET: How big a threat are foreign actors to our election process? Are they specifically using generative AI to increase their influence or effectiveness?

RC: I think foreign actors pose a great threat to our election process. Foreign actors have always wanted certain people in power that would be more aligned with their agenda. Propaganda has always been a part of the process. Generative AI provides a tool that allows for more targeted and specific messaging that could likely prove to have an even greater influence than traditional methods of propaganda.

A related issue that I think is important to consider as we look at the election process and the potential power of foreign actors (but even domestic) is an unwillingness for political contenders to engage in debate. If the politicians won't put themselves into the public square to engage in conversation, then people will look elsewhere for information. The further people move from the primary source of information, the more likely it is that the information they are being provided is being manipulated by generative AI tools.

Also: AI boom will amplify social problems if we don't act now, says AI ethicist

ZDNET: How do you personally differentiate between legitimate political communication and potentially manipulative content generated by AI systems during election campaigns?

RC: I have started to do two things due to the last election and the disinformation during that election cycle. The first thing I do is triangulate the information that I am receiving. By that, I mean I want to consume information from multiple sources that perhaps have different biases in their reporting. The better I can triangulate information, the more confident I feel in the truth of that information.

The second thing I like to do, which enables me to accomplish the first, is to purposefully not allow myself to form an opinion when I initially learn about a new issue. This is especially important when it seems that the information is different than what I already knew, or potentially politically important. Waiting to form that opinion allows me to give enough time to triangulate with additional information that can confirm (or not) the original information.

ZDNET: Are there particular types of generative AI tech that are more prevalent or influential in the electoral context?

RC: I see this from a lens of media richness, with richer context having greater influence. For example, a video is richer than something that is text-only. What makes the world of generative AI much more effective is the ability for various types of generative content to be utilized together.

While a video may be a rather rich medium to communicate with people and can be very convincing, when there is a created text to accompany the video, it helps to triangulate for someone the information that is trying to be conveyed.

Also: Real-time deepfake detection: How Intel Labs uses AI to fight misinformation

ZDNET: How can the public be better informed about the use of generative AI in elections, and what role do media and educational institutions play in promoting awareness?

RC: The use of generative AI was thrust upon us approximately a year ago, in November 2022, and is still radically changing. I would encourage the public to be paranoid and seek out interactions with people in person over believing what they are seeing online.

I believe this is a time when the media can step up and be skeptical about everything and use their investigative tools to reveal truth. The importance of getting the story right should be more important than getting the story first. As media outlets build a reputation for shedding light on the truth, regardless of a bias they might bring, it will provide an outlet that voters can rely upon. I would love to see a media outlet embrace this as their approach to reporting news. Being first could very well lead to being wrong and losing credibility.

Educational institutions should utilize this as an area to encourage students to critically think. Critically evaluating news as we enter a new presidential election cycle will help create a more informed electorate. I'd also like to see more civil conversations emerge that allow people with differing views to discuss what they are seeing with others and to do so in a public square. The key to this is civil conversation — and if we can find a way to provide that environment, it can only help to clarify the truths in this election process.

Also: Most Americans think AI threatens humanity, according to a poll

ZDNET: How do you see the intersection of technology, democracy, and elections evolving in the coming years, and what potential challenges or opportunities do you anticipate?

RC: I only see our ability to discern the truth becoming more difficult as technology advances. This includes democracy and elections. Part of me wonders if we begin to step away from technology in order to discern the truth. This would allow for less input to be provided into our assessment of what is going on.

However, technology has allowed those without a voice to reveal truths of what is going on in their world. We have seen this throughout the world, that with social media especially, people are able to share what others may want to suppress. As a result, changes have been made.

The biggest challenge that needs to be addressed, and maybe it is addressed by those who own the generative AI technology, is to somehow inform the world when something is created with that technology. Without knowing what is created with this technology, it is going to be increasingly difficult for humans to be able to discern fact from fiction.

ZDNET: How will generative AI continue to influence the electoral process beyond the 2024 elections?

RC: I believe we will see some of the most creative uses of generative AI during the 2024 election cycle. After it is complete, these techniques will likely be enhanced and made better. I wish I knew the path to preventing this or what it will look like.

I hope this becomes a conversation that our politicians engage in, as I believe this is an area where the parties can come together. The manipulative power of generative AI has the potential to equally affect everyone. As such, by bringing together the best and brightest minds, I am hopeful that there is the ability to harness that which is good with these technologies and limit the influence of that which is harmful.

Also: Don't get scammed by fake ChatGPT apps: Here's what to look out for

ZDNET: What is your biggest fear about generative AI and elections? Explain why that is your biggest concern.

RC: My biggest fear about generative AI is that the choices of who represents America will be regretted after the fact. There isn't a mechanism I am aware of that allows Americans to easily undo the choices made on Election Day. If one voice is able to find a way to manipulate and control the outcome of an election, and then through governing it becomes apparent what happened, how do the people respond?

Why is this my fear? The observed influence and ability to target people in the voting process through social media has been highly effective. Generative AI has the potential to amplify this influence in incredible ways. Is there a point where that level of influence causes people to respond in unanticipated ways, throwing resulting in protests?

We've been seeing this occur over the last four years and it is happening in Europe. Are tensions getting high enough that things would explode further? Sometimes it feels that way.

ZDNET: Are there any positive uses for generative AI in American elections?

RC: Generative AI has great potential to increase efficiencies in the election process. From an administrative perspective, it can make the processing of communications much more effective. I would imagine that generative AI could be used to help a candidate prepare for interviews by having a smart agent ask them questions in preparation for a debate. There are a lot of ways that this could be used to make things better.

One skill I recently learned about was a way to use generative AI to take a sensational news story and reduce it to a boring news story. I look forward to using this during the election season to remove the sensationalism that often accompanies political news.

What do you think?

"Removing the sensationalism." That seems like a good place to end. I would like to thank Robert for taking his time and joining me in a fascinating conversation. If you have any questions, comments, or opinions about what we've discussed, please leave them in the comments below.

You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter on Substack, and follow me on Twitter at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, and on YouTube at YouTube.com/DavidGewirtzTV.

Artificial Intelligence

Faridabad-based Netweb Now a Manufacturing Partner for the NVIDIA Grace CPU Superchips

Netweb Technologies India Limited (Netweb), the leading Indian OEM in high-end computing, today announced that it is now a manufacturing partner for the NVIDIA Grace CPU Superchip and GH200 Grace Hopper Superchip MGX server designs.

Netweb will build and produce more than ten server variations under its Tyrone range of AI systems meant for a wide range of AI and high-performance computing/supercomputing applications.

Netweb’s AI systems with NVIDIA MGX will give a boost to the country’s ‘Make in India’ mission. At the same time, the local manufacturing of systems will build a local ecosystem to better address the demands around AI and accelerated computing applications of both government and private enterprises.

With NVIDIA MGX, a modular reference design, Netweb’s AI systems will target complex workloads of HPC, data science, large language models, edge computing, enterprise AI, and design and simulation. The product range will also support handling a wide range of simultaneous workloads such as AI training, inference, and 5G on a single system. At the same time, the designs ensure seamless upgrades for upcoming hardware generations.

“The success of generative AI and other related technology is directly correlated to the backend infrastructure and capabilities, so I believe India’s story on generative AI has only just begun. The initiative of Make-in-India by NVIDIA to support the PMO vision is a great beginning. It will bring locally manufactured cutting-edge technologies at par with global standards,” Sanjay Lodha, Chairman and Managing Director of Netweb, said.

The post Faridabad-based Netweb Now a Manufacturing Partner for the NVIDIA Grace CPU Superchips appeared first on Analytics India Magazine.

Will Hyderabad Dethrone Bangalore in the IT Hub Race?

Recently, Mahindra Group chairman Anand Mahindra claimed that Google’s decision to build its biggest office outside of the US in Hyderabad was a geopolitical move.

While it’s good news for the Indian tech ecosystem, what is odd is that Google opted for Hyderabad over Bangalore, which is the undisputed IT capital of India — our own ‘Silicon Valley’.

Mahindra’s perspective holds merit, but Google’s choice adds an interesting dimension to the narrative. Bangalore emerged as the country’s largest IT hub due to a confluence of factors, including the presence of top-tier educational institutions, the right infrastructure, favourable government policies and pleasant weather conditions.

In fact, the establishment of Electronic City in the 1970s marked the beginning of Bangalore’s journey as an IT destination. Today, Karnataka has a matured IT industry and is home to over 5,500 IT/ITES companies, most of which have a strong presence in Bangalore.

However, Hyderabad, over the years, has emerged as a challenger to Bangalore’s IT status. The city’s IT industry has grown tremendously over the years. In the financial year 2022-23, Telangana recorded an impressive 31.44 percent growth in IT exports and an increase of 16.29 percent in job creation in the IT sector.

Interestingly, last year, Hyderabad contributed to a third of the 4,50,000 IT jobs created in the country, relatively higher than the 1.46 lakh jobs created by Bangalore, according to a Nasscom report.

Favourable business conditions and progressive policies

Over the years, the government of Telangana, under the leadership of Chief Minister K Chandrashekar Rao, has implemented several proactive programmes and legislation to foster a business-friendly environment, acknowledging the significant potential of the IT sector.

Much of the credit of the monumental rise of Hyderabad as a challenger to Bangalore has to do with the radical policies introduced by Rao’s administration over the years.

The Telangana IT Policy, initiated in 2016, offers companies a wide array of incentives, such as tax reductions, along with infrastructure support and investment opportunities. The progressive policy is one of the reasons the state has successfully attracted numerous leading global tech companies.

“What has also helped the state in retaining big-ticket companies is the focus and clarity in the government as far as its industrial policy is concerned. Another factor is the continuity in policy in the past 20 years. Hyderabad and Telangana benefitted from policy continuity. Stability in the policy regime attracts investment,” Madan Pillutla, the dean of the Indian School of Business, told the Print.

Similarly, talking about Kaynes SemiCon’s decision to invest INR2800 crore to set up a OSAT and compound semiconductor facility at Kongara Kalan near Hyderabad, CEO Raghu Panicker said, “It’s the ecosystem. The speed at which Telangana works, quick approvals, and clearances were impressive.”

Favourite destination for GCCs

In the first half of 2023, Hyderabad witnessed the establishment of 11 new Global Capacity Centres (GCC), per the report titled ‘India GCC Trends – Half Yearly Analysis H1CY2023. Companies such as BlackBerry, CyberArk, Storable, and Align Technology established GCCs in Hyderabad in the first half of 2023.

An equivalent number of GCCs were also set up in Bangalore, accompanied by five in Pune and four in Mumbai. Over the years, Hyderabad has become the preferred choice for numerous multinational corporations to establish their offices. The fact that Google has opted to construct its largest office outside the US in Hyderabad rather than Bangalore speaks volumes about Hyderabad’s growing clout.

Notably, e-commerce giant Amazon’s largest campus in the world is also located in Hyderabad. “The second-largest campuses of Apple, Meta, Qualcomm, Micron, Novartis, Medtronic, Uber, Salesforce and many more have also been set up in Hyderabad in the last 9 years,” Rao posted on X.

Furthermore, while traditionally known as a net exporter of talent, Hyderabad is now experiencing a notable increase in the influx of talent. This shift could be attributed to the growing tech ecosystem in the state.

The choice of numerous MNCs to establish a presence in Hyderabad is likely influenced by the city’s rich talent pool. With premier educational institutions such as the International Institute of Information Technology, Hyderabad (IIITH), Indian Institute of Technology Hyderabad (IIT-H), and Birla Institute of Technology and Science (BITS), the city provides a skilled workforce.

Moreover, the state government has undertaken numerous skilling initiatives to ensure that the state fosters a skilled and employable workforce, driving economic growth and development.

Robust Infrastructure

MNC’s desire to set up camp in Hyderabad has to do with its robust infrastructure along with favourable business conditions. The city boasts an advanced metro system, well-designed roads and elevated connectivity and accessibility. These improvements contribute to Hyderabad’s appeal for businesses and residents, fostering commercial and residential real estate growth.

While speaking to the Indian American diaspora in San Jose on March 23 during a week-long tour of the US, Telangana IT and Industries Minister K T Rama Rao asserted that the quality of infrastructure in Hyderabad surpassed that of Bengaluru by twofolds.

“There is inertia in Bengaluru; come and invest in Hyderabad, which is a dynamic city. The cost of living is low too,” declared KTR.

“About 40 per cent of the people working in IT in Bengaluru are Telugu, and given an opportunity, they would come back to Hyderabad or Visakhapatnam or other cities in the two Telugu states.”

Moreover, according to a recent report, ​​Hyderabad has surpassed Bengaluru in the demand for office space, underscoring its prominence as a premier global business destination.

Earlier in 2020, Hyderabad was ranked as the world’s most dynamic city amongst 130 cities across the globe in the ‘City Momentum Index’ (CMI) published by JLL. Interestingly, Bangalore came 2nd, and other Indian cities in the Top 10 include Chennai at 5th and Delhi at 6th.

The Index highlights key factors contributing to the success of the world’s most dynamic cities, emphasising talent attraction and robust innovation economies. It also sheds light on the challenges these cities encounter while managing rapid growth and sustaining positive momentum in the long run.

Growing startup ecosystem

While Bangalore still holds the bragging rights and is home to the highest number of tech startups in the country, Hyderabad is also rapidly growing as a startup hub in India. The latest data shows that 5,062 startups have come up in the capital city of Karnataka since 2020. On the contrary, Hyderabad’s startup ecosystem now proudly boasts 4,369 tech startups.

Of late, investors’ attention has also shifted from major cities, such as Bangalore, Mumbai, and New Delhi, to Hyderabad. T-Hub, India’s largest incubation centre, has significantly nurtured entrepreneurship and innovation in Hyderabad. Incorporated in 2015, it has provided support, mentorship, and networking opportunities to startups, connecting them with influential stakeholders in the ecosystem.

“At T-Hub, we have nurtured over 3,500 startups, 150 mentors and 350 corporate partners,” Mahankali Srinivas Rao, CEO of T-Hub, said.

“The number of employees and employers who are looking at Hyderabad is growing. I sincerely believe that this is just the beginning,” Rao said earlier this year, while addressing leaders of the Hyderabad Software Enterprises Association (HYSEA).

Hence, Hyderabad’s growing IT sector, superior infrastructure, rising startup ecosystem, and talent influx position it as a dynamic and competitive alternative to Bangalore, challenging the traditional IT hierarchy in India. While Bangalore will continue to be the IT hub of India, whether Hyderabad can dethrone Bangalore as the coveted ‘IT capital’ remains to be seen.

The post Will Hyderabad Dethrone Bangalore in the IT Hub Race? appeared first on Analytics India Magazine.

How Big Data Is Saving Lives in Real Time: IoV Data Analytics Helps Prevent Accidents

How Big Data Is Saving Lives in Real Time: IoV Data Analytics Helps Prevent Accidents
Photo by Roberto Nickson

Internet of Vehicles, or IoV, is the product of the marriage between the automotive industry and IoT. IoV data is expected to get larger and larger, especially with electric vehicles being the new growth engine of the auto market. The question is: Is your data platform ready for that? This post shows you what an OLAP solution for IoV looks like.

What is special about IoV data?

The idea of IoV is intuitive: to create a network so vehicles can share information with each other or with urban infrastructure. What‘s often under-explained is the network within each vehicle itself. On each car, there is something called Controller Area Network (CAN) that works as the communication center for the electronic control systems. For a car traveling on the road, the CAN is the guarantee of its safety and functionality, because it is responsible for:

  • Vehicle system monitoring: The CAN is the pulse of the vehicle system. For example, sensors send the temperature, pressure, or position they detect to the CAN; controllers issue commands (like adjusting the valve or the drive motor) to the executor via the CAN.
  • Real-time feedback: Via the CAN, sensors send the speed, steering angle, and brake status to the controllers, which make timely adjustments to the car to ensure safety.
  • Data sharing and coordination: The CAN allows for data exchange (such as status and commands) between various devices, so the whole system can be more performant and efficient.
  • Network management and troubleshooting: The CAN keeps an eye on devices and components in the system. It recognizes, configures, and monitors the devices for maintenance and troubleshooting.

With the CAN being that busy, you can imagine the data size that is traveling through the CAN every day. In the case of this post, we are talking about a car manufacturer who connects 4 million cars together and has to process 100 billion pieces of CAN data every day.

IoV data processing

To turn this huge data size into valuable information that guides product development, production, and sales is the juicy part. Like most data analytic workloads, this comes down to data writing and computation, which are also where challenges exist:

  • Data writing at scale: Sensors are everywhere in a car: doors, seats, brake lights… Plus, many sensors collect more than one signal. The 4 million cars add up to a data throughput of millions of TPS, which means dozens of terabytes every day. With increasing car sales, that number is still growing.
  • Real-time analysis: This is perhaps the best manifestation of "time is life". Car manufacturers collect the real-time data from their vehicles to identify potential malfunctions, and fix them before any damage happens.
  • Low-cost computation and storage: It's hard to talk about huge data size without mentioning its costs. Low cost makes big data processing sustainable.

From Apache Hive to Apache Doris: a transition to real-time analysis

Like Rome, a real-time data processing platform is not built in a day. The car manufacturer used to rely on the combination of a batch analytic engine (Apache Hive) and some streaming frameworks and engines (Apache Flink, Apache Kafka) to gain near real-time data analysis performance. They didn't realize they needed real-time that bad until real-time was a problem.

Near Real-Time Data Analysis Platform

This is what used to work for them:

How Big Data Is Saving Lives in Real Time: IoV Data Analytics Helps Prevent Accidents
Data from the CAN and vehicle sensors are uploaded via 4G network to the cloud gateway, which writes the data into Kafka. Then, Flink processes this data and forwards it to Hive. Going through several data warehousing layers in Hive, the aggregated data is exported to MySQL. At the end, Hive and MySQL provide data to the application layer for data analysis, dashboarding, etc.

Since Hive is primarily designed for batch processing rather than real-time analytics, you can tell the mismatch of it in this use case.

  • Data writing: With such a huge data size, the data ingestion time from Flink into Hive was noticeably long. In addition, Hive only supports data updating at the granularity of partitions, which is not enough for some cases.
  • Data analysis: The Hive-based analytic solution delivers high query latency, which is a multi-factor issue. Firstly, Hive was slower than expected when handling large tables with 1 billion rows. Secondly, within Hive, data is extracted from one layer to another by the execution of Spark SQL, which could take a while. Thirdly, as Hive needs to work with MySQL to serve all needs from the application side, data transfer between Hive and MySQL also adds to the query latency.

Real-Time Data Analysis Platform

This is what happens when they add a real-time analytic engine to the picture:

How Big Data Is Saving Lives in Real Time: IoV Data Analytics Helps Prevent Accidents
Compared to the old Hive-based platform, this new one is more efficient in three ways:

  • Data writing: Data ingestion into Apache Doris is quick and easy, without complicated configurations and the introduction of extra components. It supports a variety of data ingestion methods. For example, in this case, data is written from Kafka into Doris via Stream Load, and from Hive into Doris via Broker Load.
  • Data analysis: To showcase the query speed of Apache Doris by example, it can return a 10-million-row result set within seconds in a cross-table join query. Also, it can work as a unified query gateway with its quick access to external data (Hive, MySQL, Iceberg, etc.), so analysts don't have to juggle between multiple components.
  • Computation and storage costs: Apache Doris uses the Z-Standard algorithm that can bring a 3~5 times higher data compression ratio. That's how it helps reduce costs in data computation and storage. Moreover, the compression can be done solely in Doris so it won't consume resources from Flink.

A good real-time analytic solution not only stresses data processing speed, it also considers all the way along your data pipeline and smoothens every step of it. Here are two examples:

1. The arrangement of CAN data

In Kafka, CAN data was arranged by the dimension of CAN ID. However, for the sake of data analysis, analysts had to compare signals from various vehicles, which meant to concatenate data of different CAN ID into a flat table and align it by timestamp. From that flat table, they could derive different tables for different analytic purposes. Such transformation was implemented using Spark SQL, which was time-consuming in the old Hive-based architecture, and the SQL statements are high-maintenance. Moreover, the data was updated by batch on a daily basis, which meant they could only get data from a day ago.

In Apache Doris, all they need is to build the tables with the Aggregate Key model, specify VIN (Vehicle Identification Number) and timestamp as the Aggregate Key, and define other data fields by REPLACE_IF_NOT_NULL. With Doris, they don't have to take care of the SQL statements or the flat table, but are able to extract real-time insights from real-time data.

How Big Data Is Saving Lives in Real Time: IoV Data Analytics Helps Prevent Accidents

3. DTC data query

Of all CAN data, DTC (Diagnostic Trouble Code) deserves high attention and separate storage, because it tells you what's wrong with a car. Each day, the manufacturer receives around 1 billion pieces of DTC. To capture life-saving information from the DTC, data engineers need to relate the DTC data to a DTC configuration table in MySQL.

What they used to do was to write the DTC data into Kafka every day, process it in Flink, and store the results in Hive. In this way, the DTC data and the DTC configuration table were stored in two different components. That caused a dilemma: a 1-billion-row DTC table was hard to write into MySQL, while querying from Hive was slow. As the DTC configuration table was also constantly updated, engineers could only import a version of it into Hive on a regular basis. That meant they didn't always get to relate the DTC data to the latest DTC configurations.

As is mentioned, Apache Doris can work as a unified query gateway. This is supported by its Multi-Catalog feature. They import their DTC data from Hive into Doris, and then they create a MySQL Catalog in Doris to map to the DTC configuration table in MySQL. When all this is done, they can simply join the two tables within Doris and get real-time query response.

How Big Data Is Saving Lives in Real Time: IoV Data Analytics Helps Prevent Accidents Conclusion

This is an actual real-time analytic solution for IoV. It is designed for data at really large scale, and it is now supporting a car manufacturer who receives 10 billion rows of new data every day in improving driving safety and experience.

Building a data platform to suit your use case is not easy, I hope this post helps you in building your own analytic solution.

Zaki Lu is a former product manager at Baidu and now DevRel for the Apache Doris open source community.

More On This Topic

  • Using Data Science to Predict and Prevent Real World Problems
  • How a Polytechnic Helps You Make the Tech-Business Connection
  • 90% of Today's Code is Written to Prevent Failure, and That's a Problem
  • How to Use Kafka Connect to Create an Open Source Data Pipeline for…
  • Real-Time Histogram Plots on Unbounded Data
  • The Architecture Behind DeepMind’s Model for Near Real Time Weather…

Pika, which is building AI tools to generate and edit videos, raises $55M

Pika, which is building AI tools to generate and edit videos, raises $55M Kyle Wiggers 10 hours

The generative AI hype hasn’t died down yet.

Case in point, Pika, a startup creating an AI-powered platform to edit and generate videos from captions and still images, today announced that it raised $55 million in a funding round led by Lightspeed Venture Partners with participation from Homebrew, Conviction Capital, SV Angel, Ben’s Bites and notable angel investors including Quora founder Adam D’Angelo, ex-GitHub CEO Nat Friedman and Giphy co-founder Alex Chung.

The fresh tranche comes just six months after Pika emerged from stealth and coincides with the early access launch of what Pika’s calling “Pika 1.0,” a new suite of videography tools that introduces a generative AI model capable of editing videos in a range of styles, like “3D animation,” “anime” and “cinematic.”

Pika Labs

Image Credits: Pika

“Video is at the heart of entertainment, yet the process of making high-quality videos to date is still complicated and resource-intensive,” Pika writes in a blog post published on its website this morning. “When we started Pika six months ago, we wanted to push the boundaries of technology and design a future interface of video making that is effortless and accessible to everyone. Since then, we’re proud to have grown the Pika community to half a million users, who are generating millions of videos per week.”

Pika was co-founded by Demi Guo and Chenlin Meng, both former Ph.D. students in Stanford’s Artificial Intelligence Lab. Prior to studying at Stanford, Guo worked as an engineer at Meta’s AI research division, while Meng co-authored a number of AI research papers including several pertaining to generative AI.

Pika Labs

Image Credits: Pika

Pika competes against generative AI video tools and models from the likes of Runway and Stability AI. But with Pika 1.0, Pika’s looking to up its game with several differentiating features.

For example, Pika 1.0 ships with a tool that can extend the length of existing videos or transform them into different styles, like “live action” to “animated” — or expand the canvas or aspect ratio of a video. Another module edits video content using AI, like changing someone’s clothing or even adding another character.

We’ll have to put those capabilities to the test once Pika 1.0 becomes widely available. But for what it’s worth, Lightspeed — which is also an investor in Stability AI — has confidence in the platform — even as tech giants like Google and Meta telegraph that they, too, are working on generative AI tools for video.

Pika Labs

Image Credits: Pika

“Just as other new AI products have done for text and images, professional-quality video creation will also become democratized by generative AI. We believe Pika will lead that transformation,” Lightspeed’s Michael Mignano said in a press release. “Given such an impressive technical foundation, rooted in an early passion for creativity, the Pika team seems destined to change how we all share our stories visually. At Lightspeed, we couldn’t be more excited to support their mission to allow anyone to bring their creative vision to life through video, and we’re thrilled to be investing alongside other amazing investors at the forefront of AI.”

Pika’s rapid growth is reflective of the continued, strong demand for generative AI of all flavors — from tools like Midjourney and DALL-E 3 to ChatGPT.

In a recent report, IDC projected that generative AI investments will rise from $16 billion this year to a whopping $143 billion in 2027. While generative AI accounts for just 9% of overall AI spending in 2023, the firm expects that’ll increase to 28% within five years.

The spending might just be justified. A recent poll — albeit focused only on users from the U.K. — found that Gen Z is embracing generative AI, with four in five (79%) teenagers aged 13-17 reporting having used generative AI tools, apps and services including ChatGPT and Snapchat’s My AI.

Then again, Gen Zers aren’t necessarily paying for generative AI. And enterprise customers, which have the biggest coffers to spend on it, are encountering hurdles deploying some forms of the tech.

O’Reilly’s 2023 generative AI in the enterprise report reveals that many corporate AI adopters (26%) are still in the early stages of piloting generative AI, and are deeply worried about the potential challenges present and future surrounding the tech — including unexpected outcomes, security, safety, fairness, bias and privacy. The difficulty of finding business use cases and concerns about legal issues (like who owns the copyright over AI-generated output) are holding generative AI back, the report implies, as are badly-conceived and poorly-implemented AI solutions.

AMD Eyes Big Wins with MI300X for AI Workloads 

AMD Eyes Big Wins with MI300X for AI Workloads

AMD predicts that MI300X will make AMD earn $2 billion revenue in 2024 because of the strong pull from the customers. The processor is set to release next week at the AMD Advancing AI event on December 6.

“We are seeing a tremendous pull for MI300X,” Mark Papermaster, CTO of AMD, told AIM, talking about how companies such as Lamini, Moreh, and Databricks are already announcing their success stories after using MI250x for training their generative AI models. “In fact, Lisa Su (CEO) has already said that MI300X is going to be the fastest growing product in the history of AMD,” added Papermaster.

Another reason for the success of AMD’s hardware is the success of its software stack – ROCm (AMD Radeon Open Compute), which Papermaster highlighted has reached high performance compute production level. Moreover, he added that the next version of ROCm, the 6.0, would be released soon, and would be production level ready for AI workloads.

“We are bringing a leadership product”

Microsoft announced at Ignite 2023 that it would be using MI300X for their AI workloads. Interestingly, at the same conference Microsoft also announced that it would be using the NVIDIA GH200 superchips for the same purpose. “Competition is good. It brings innovation,” said Papermaster. “It brings pricing that ensures value for the customers and spurs the industry forward.”

“We have not only brought competition, but are also bringing in a leadership product in inference applications,” he further added about MI300X. The upcoming GPU will be recognised for its leadership in both training and inference applications, solidifying AMD’s position in the AI hardware landscape.

“MI300X is a very high performance GPU which is capable of being scaled out to very large cluster sizes,” he further highlighted how AI is evolving and the company is also focusing on edge computing on smaller sizes within laptops.

AMD highlighted the widespread adoption of Ryzen AI, the inaugural dedicated AI accelerator available on an x86 processor. With over 50 systems now equipped with Ryzen 7000 Series processors and Ryzen AI, millions of AMD AI PCs are available in the market.

“Historically, if you go back in time, most AI inference seems done to the CPU. And still today, at least we fully support inferencing on AMD CPUs. But particularly generative AI applications are in much higher demand for computing capability. And so when you look at the generative AI, training and inferencing it requires acceleration, and that is where we’re bringing in competition with AMD instinct roadmap GPU.”

“What I want to highlight is that AMD has a broader strategy for AI than just GPUs. This is what makes us different from NVIDIA,” Gilles Garcia, senior director business lead, data center communication group at AMD, told AIM.

AMD believes that most of the AI workloads can be handled by CPUs alone. “CPUs are best for handling most of the problems with current edge processing such as thermal management, cost effectiveness, and reducing the footprint by working on edge,” said Garcia.

The startup and open source community in India

Focusing on the recent acquisition of Nod.ai, a software stack company that is now helping AMD, Papermaster said that AMD is always looking at the startup community in India. “For India, we have a strong design presence here. India will, of course, be central for our AI product development, hardware and software product development efforts.”

Highlighting how Hugging Face also started using AMD GPUs for testing, he added that with the explosion of the AI ecosystem, AMD is very much committed to an open ecosystem. “We are not proprietary or close,” he added.

“Rather than needing just one specific partnership to get access to AMD, we’ve differentiated because we’re very strongly committed to open source software and to open collaborations. I expect many such collaborations between AMD and India based on the strong adoption in India, and support of open source.”

To expand its research and engineering operations in India, AMD inaugurated its largest global design center in Bengaluru. The AMD Technostar R&D campus is a key component of the company’s $400 million investment in India over the next five years. The campus plans to accommodate around 3,000 new employees at AMD which will focus on R&D.

“We are very focused on workplace development here in India,” Papermaster continued. “Jaya Jagadish, our country head and senior vice president, was the leader of a government panel, which studied workforce development. She has made very specific recommendations to the government of India.”

Papermaster says that AMD has very innovative workforce development programmes, which the company is continually employing in India. “We have strong relationships with universities and we are also providing additional training to students. Then we bring them onto a very established internship programme,” he explained about how AMD has established an excellent pipeline for college graduate engineering in India.

The post AMD Eyes Big Wins with MI300X for AI Workloads appeared first on Analytics India Magazine.