I went to Microsoft to talk about AI. I’m still a little startled (but hopeful too)

Microsoft Silicon Valley office

I was excited, of course.

Earlier this year, Microsoft invited me to its Silicon Valley campus to contemplate AI, the generative kind.

So there I was in this enormous, airy building of the future, with hardly a human being in sight. I wafted into a very fetching theater – yes, they spelled my name wrong on the big screen — and there were several actual journalistic luminaries on stage too.

Also: The future of generative AI: Here's what technology analysts are saying

Oh, and a PR man from Microsoft, dressed all in black.

The specific subject was AI and journalism. Little did I know how much it would move my whole year. For here I was, perhaps for the first time, confronted with the three great horsepersons of AI — the Experimenters, the Fearful and the Capitulists.

Those last ones? Capitalists who have already capitulated to AI's allegedly all-encompassing powers.

Let's try something new

Some panelists mused that the seemingly sudden incursion of AI offered a fine chance to dabble with it.

At one news service, for example, they've experimented with an AI version of one of its sports editors.

Also: ZDNET looks back on tech in 2023, and looks ahead to 2024

The idea is to put some text into this AI-generated expert and let it/them speak. Yes, but could this bot give us a decent betting line on the Niners vs the Cowboys? You have to find a balance, said the news service's journalist.

But is there really any value added by being, well, so tech-clever?

This seemed to me to be one of the currently less-considered questions about AI. The year was all about generative AI. But, as with quite a few tech breakthroughs, how useful (and profitable) will it ultimately be? And precisely where and how?

So a strong impulse was to put it to the test in as many ways as possible. You never know what might emerge.

Oh, AI. Who is the fairest of them all? You Are.

But then there was the gentlefellow from another tech publication.

I know that sounds like the beginning of a joke, but it truly isn't. It's more of a mood enhancer.

Also: The future of work is more human than you'd think, say these business experts

You see, this particular gentlefellow is an enormous enthusiast for AI. Actually, 'enthusiast' doesn't quite cover it. He seemed more like the mesmerized subject of a despotic kingdom.

He offered the prediction that "generative AI will touch every piece of every part of the process of news from ideation, to headline generation to story editing."

Full disclosure: I had the idea for this column myself, my editor is a human being with infinite patience, and I'm writing with my own fingers and (what's left of my) mind.

That's the thing with predictions. They don't always come true. Just ask any AI sports editor.

I won't dwell on some of this nice man's more vivid admissions — they've been covered elsewhere.

For me, what was most affecting were these words: "Frankly, I think that the AI model is always more clever than me because it includes all of the written text throughout all of history."

I fear several things shot through my head at this very moment. No, of course 'How clever actually are you?' wasn't one of them.

Also: The promise and peril of AI at work in 2024, according to Deloitte's Tech Trends report

Not even when he said: "It feels like it is so much more talented than I'll ever be."

I met many people in college who'd read a lot of books. Seemingly every book that mattered. But would I have thought them clever or talented? Would I have imagined they'd, um, change the world? Not really. (Update: They didn't.)

Let AI be AI but you do you

I confess that night became my tech moment of the year.

Everywhere I went I seemed to meet either experimenters, fearers or, indeed, capitulists.

Both experimentation and fear are, of course, understandable. But to witness capitulism live on stage was a touch startling, and the feeling hasn't quite gone away.

Also: AI in 2023: A year of breakthroughs that left no human thing unchanged

I couldn't help offering a small retort to this committed AI capitulist.

"Maybe you should just believe in yourself more," I said. "I'm really concerned about you. Is it possible that because of your enthusiasm, you're already abdicating your own talents? You're actually maybe better than you think."

It kept on striking me throughout the year that for all the talk about the evil side of AI – and, because humans, there's plenty of evil – there's another side.

The true fascination of generative AI doesn't necessarily lie in what it's going to do to us, or even for us. It's what we can do with it.

It may actually show how clever we really are, not merely how clever we think we are.

If the internet taught us anything – yes, I know the jury is still out – it's that we derived enormous benefits and created new ways of talking and being — just as we endured awful changes of mood, behavior and hope.

Isn't that likely with AI too?

Also: Generative AI filled us with wonder in 2023 — but all magic comes with a price

2023 wasn't necessarily the year when AI began to take over our souls. It was merely a large new door opening, sucking us toward a blinding light.

I feel an appropriate level of fear having experimented with ChatGPT, with both hilarious and freakish results. (Watching a deepfake version of someone you do business with is truly a mind-twisting experience.)

But to be a capitulist strikes me as not merely defeatist, but just a little dull.

Here's an oddly optimistic thought: AI, this clever and talented thing, might even help us slow down a little and think a little more.

Also: Here's how to create your own custom chatbots using ChatGPT

Oh, and here's another optimistic thought: We've just learned that AI can get tired and lazy – perhaps it's embracing the human condition in the same way that we're embracing the AI condition.

After all, what can it do without our input? (Please don't answer that right now.)

We won't know for sure, of course until, say, 2032.

You may, of course, place your bets now with my personal AI Betbot service.

Featured

Top 7 Cybersecurity Threats for 2024

The rise and rapid adoption of new innovative technologies, such as generative artificial intelligence, no-code apps, automation and the Internet of Things, have dramatically changed the global cybersecurity and compliance landscape for every industry.

Cybercriminals are turning to new techniques, tools and software to launch attacks and create greater damage. As a result, the 2023 Cybersecurity Ventures Cybercrime Report predicts a rapid increase in damage costs associated with cybercrime — projected to cost $10.5 trillion globally in damages by the end of 2024. The report lists cost of data breaches, stolen funds, intellectual property theft, operational disruptions and post-attack recovery as the main expenses for organizations under this trend.

On the other hand, Google’s Cloud Cybersecurity Forecast 2024 report highlights the increased use of AI to scale malicious operations, nation-state-supported cybercriminal gangs, zero-day vulnerabilities and modern phishing as main attack vectors for the coming year.

To stay ahead of the curve, IT and security leaders should focus on layered security solutions and zero trust to keep their companies’ data safe from top cybersecurity threats like ransomware and phishing.

Learn about Kolide

Jump to:

  • 1. Ransomware
  • 2. OT-IT security
  • 3. Dark Web
  • 4. Malware as a service and hackers-for-hire
  • 5. Modern phishing
  • 6. IoT and Industrial IoT
  • 7. State-sponsored attacks
  • Staying vigilant in the evolving threat landscape

1. Ransomware

Ransomware — the breaching of business-critical systems and assets with the goal of encrypting them and holding them for ransom — will continue to plague organizations across all sectors in 2024. New and established cybercriminal groups will leverage ransomware as a service, making it easier than ever to launch sophisticated attacks. They will also employ evolving extortion tactics like double and triple extortion, pressuring victims through data leaks.

SEE: Here’s everything you need to know about ransomware.

As proven by the November 2023 ransomware attack on MeridianLink by ALPHV/BlackCat ransomware group, ransomware gangs are also willing to manipulate regulations. In that attack, BlackCat reported its own crime to put pressure on MeridianLink leveraging the new U.S. Securities and Exchange Commission law.

Healthcare, government and critical infrastructure will be particularly targeted by ransomware. Organizations must prioritize ransomware defense by updating systems, implementing robust backups, training employees and considering cyber insurance. More importantly, companies must ensure their security teams and experts have all the resources they need and are not working under unsustainable pressure.

2. OT-IT security

The convergence of operational technology and information technology in critical infrastructures, industrial facilities, public service providers and manufacturing plants is creating new vulnerabilities and attack opportunities for cybercriminals. Attacks on OT infrastructures via IT-compromised systems can disrupt operations, cause physical damage and risk public safety.

Notable 2023 OT-IT attacks include the late November ransomware attack on Ardent Health Services, which diverted ambulances and affected health emergency services across multiple U.S. states, and the attack on a water system in western Pennsylvania — claimed by an anti-Israeli Iranian cybercriminal group.

Organizations operating OT-IT systems must modernize legacy technology, deploy layered security, segment IT and OT networks, and implement robust access controls to prevent attacks.

3. Dark Web

The Dark Web, a hidden portion of the internet accessible only through specialized software and configurations, is a breeding ground for illegal activities. New trends on the dark web include the rise of organized criminal activity, characterized by the availability of:

  • No-code malware, which requires minimal technical expertise to use.
  • Plug-and-play kits, which are pre-configured tools for launching cyberattacks.
  • Dedicated customer support.

Additionally, fileless attacks, where attackers use stolen credentials purchased on the Dark Web to gain access to systems without leaving behind traditional malware traces, are one of the biggest trends to look out for. And zero-day brokers — cybercrime groups selling zero-day exploits on the Dark Web to multiple buyers — are becoming increasingly prevalent.

SEE: Here’s everything you need to know about the Dark Web.

In light of these evolving threats, it is crucial for organizations to consider actively monitoring the Dark Web through professional services. This proactive approach can provide valuable insights to help organizations mitigate the great number of different threats that come directly from the Dark Web.

4. Malware as a service and hackers-for-hire

The MaaS landscape has seen a dramatic increase in the availability of platforms and tools that broaden the range of accessible malware and attack functionalities. MaaS user interfaces have also become increasingly intuitive, incorporating tutorials and simplified processes, and diversified. They now cater to various budgets and needs that further lower the barrier to entry, while automation features have become increasingly prevalent.

Meanwhile, hackers-for-hire has become the norm, going even beyond the trend of effectively lowering the technical barriers of launching cyberattacks. This democratization of cybercrime is predicted to fuel a surge in both the number and sophistication of attacks in 2024. According to a Kaspersky report, 2024 will see more groups offering hack-for-hire services. Meanwhile,

SEE: A Kaspersky report reveals the top cyber threats for SMBs in 2023.

To navigate this evolving threat landscape, organizations must prioritize implementing strong layered security solutions capable of detecting and blocking malicious software before it can take root. By equipping employees with knowledge about MaaS and hackers-for-hire threats and social engineering tactics used to distribute malware, organizations can build a more resilient workforce. Regular data backups and encryption, coupled with a zero-trust security model, further bolster defenses by minimizing potential data loss and ensuring stringent access controls.

5. Modern phishing

Phishing attacks that leverage social engineering techniques and personalized messages to trick victims into revealing sensitive information or downloading or clicking on malicious files is evolving.

Traditional methods like mass-mailed generic messages are giving way to personalized and highly realistic attacks. Criminals use AI to automate campaigns and personalize messages with targeted details, generate convincing content like deep fakes and even automatically learn from successes.

To stay ahead, organizations must invest in tools that can detect AI-generated content, educate employees about these evolving threats, and run phishing simulations to identify the weak points in their organizations and secure workplaces.

6. IoT and Industrial IoT

IoT and Industrial IoT devices, with their growing ubiquity and often limited security, present an increasingly attractive target for cybercriminals. In 2023, attacks on IIoT devices saw a significant rise, with attackers leveraging vulnerabilities to launch distributed denial-of-service attacks, steal data and disrupt operations. These attacks evolved to include new techniques like exploiting supply chain vulnerabilities and compromising firmware updates, highlighting the need for enhanced security measures.

SEE: Top IIoT security risks.

To protect against these evolving threats in 2024, organizations must prioritize robust security practices throughout the entire IoT ecosystem. This includes implementing secure coding practices, regularly updating software and firmware, utilizing strong authentication protocols, and monitoring networks for suspicious activity.

Additionally, organizations need to consider adopting zero-trust security models and implementing segmentation strategies to isolate compromised devices and minimize attack impact.

7. State-sponsored attacks

Nation-state actors are increasingly using cyberattacks to achieve their political and strategic goals. These attacks can target critical infrastructure, steal sensitive information and disrupt essential services. 2023 saw an escalation of nation-state-supported cyber criminal activity linked to North Korea, seeking new mechanisms to fund weapon and government programs and navigate international sanctions; Russia, with hackers supporting the invasion of Ukraine and taking cyber warfare to international levels; and the new Israel-Hamas conflict.

Building strong relationships with government and law enforcement agencies and reporting security incidents is fundamental for organizations to mitigate state-backed threats.

2024 demands a proactive approach to thwarting state-sponsored attacks. Organizations need multilayered defenses, including sophisticated cybersecurity solutions, threat intelligence monitoring and robust incident response plans. By prioritizing comprehensive defense strategies and collaborating across sectors, organizations can better protect themselves from the evolving tactics of nation-state actors.

DOWNLOAD: These may be the top threats for 2024, but here are 50 cybersecurity threats to watch out for.

Staying vigilant in the evolving threat landscape

The cybersecurity landscape is constantly evolving, and threats are becoming more sophisticated. To mitigate modern cybersecurity and compliance threats, organizations must combine state-of-the-art technologies operating under holistic cybersecurity programs.

Strategies like zero-trust models are essential to strengthening companies’ security postures as they adapt efficiently and proactively to cybersecurity threats. Kolide — which sponsored this forward-looking report — offers a user- and device-level trust solution that empowers organizations with Okta to seamlessly deploy zero-trust access models and secure their environment and apps.

By staying vigilant and adapting to the changing threat landscape, organizations can protect themselves from cyberattacks and ensure the security of their data and systems.

Learn about Kolide

The challenges cloud migration and modernization solve for enterprises

Armor cloud storage sign with two up and down arrows in blue with reflection background. Cloud technology. 3d rendering – illustration.

With legacy systems, traditional data storage, and integration methods phasing out for many organizations, businesses are rapidly transitioning to cloud-based data to improve operational workflow and ensure a trustworthy data foundation. Cloud migration and modernization resolve numerous challenges commonly faced by modern companies. As more organizations adopt cloud initiatives for their business infrastructures, it is essential to understand and define the transformative process to maximize their distinctive value.

According to a recent report by Pluralsight, 70 percent of organizations say more than half of their infrastructure exists in the cloud, while 65 percent operate in multi-cloud environments. The market for cloud services and the benefits they deliver is quickly growing and evolving. Cloud applications remain among the most significant developments in data science innovation in recent years. Before business professionals can adopt and implement an initiative, they must know the differences between the two and their capabilities for maximizing business intelligence.

Transitioning to the cloud: Defining the market’s size and scope

The market value of public cloud spending on cloud-based technologies and digital transformations continues to skyrocket. Gartner predicts that worldwide end-user spending on public cloud services will grow 20.7 percent from 2022, totaling $591.8 billion by the end of 2023 (Table 1). Legacy data systems and traditional methods are on their way out as cloud-based infrastructures replace the data approaches of yesteryear. Achieving substantial business growth that accumulates value in today’s market requires a modern approach to data and analytics. This can only be acquired through a cloud-based digital transformation.

image-2
Table 1. Worldwide public cloud services end-user spending forecast in U.S. billions 2001-2023

With 94 percent of organizations already in the cloud, it is safe to assume that most enterprises have embraced the migration process to keep up with the latest developments. Cloud innovation is at an all-time high since businesses rely on these technologies and applications as the backbone for their current and future business strategies. With an ever-expanding cloud marketplace comes an increase in operational complexity. It is essential for organizations to have a firm grasp of what these cloud applications truly mean for the future of their businesses.

Research shows that approximately 75 percent of enterprises plan to invest in cloud technology platforms to facilitate innovation exchange. Additionally, around 47 percent of cloud decision-makers define digital transformation as “optimizing processes and becoming more operationally agile,” while another 40 percent claim it’s “improving customer experience.” Many organizations undergoing an enterprise-wide transition to cloud-centric data are still defining their strategy upon implementation. To reap the benefits of cloud developments while overcoming its many challenges, leaders need to maintain a comprehensive understanding of what transformation means for their organization and its specific business needs.

Benefits and challenges of migration and modernization

Implementing a cloud-based digital transformation is a tedious process that requires close examination and a defined strategy to ensure success. There isn’t much room for error, and it cannot deliver results if a business blindly initiates implementation. Cloud-based data and analytics initiatives can solve several issues plaguing many organizations, and, like most initiatives, it comes down to time and money. Cloud applications ensure faster deployment and incremental releases while reducing the costs of managing an on-premises infrastructure. Traditional data management applications requiring manual maintenance from business users can become problematic with an overabundance of users simultaneously. Delegating all management and maintenance duties to the cloud resolves the issue of having “too many cooks in the kitchen” while ensuring applications run smoothly and efficiently. Solving these common problems allows businesses to maximize cloud potential and reap valuable benefits.

Migrating or modernizing cloud data applications can create substantial advantages for an enterprise (Figure 1). Three of the most significant areas in which they provide benefits include:

Workload agility
The scalability and elasticity of the cloud allow business users to manage workload changes more easily while efficiently handling workload overtime changes. This reduces technical debt, eliminates maintenance costs, and improves system reliability.

Innovation and deployment
The cloud’s self-service nature facilitates quick and ingenious implementations by reducing time-to-market. Enterprises can roll out new features or releases on a fluid timeline utilizing the cloud’s progressive delivery methods.

Enhanced user experience
The cloud targets and resolves user interface shortcomings while quickly deploying the most effective features. It consistently enables current updates and reduces latencies or lag time for a faster user experience. Cloud-based data centers also allow enterprises to expand their global reach and improve accessibility for customers and users.

image-3
Figure 1: Cloud modernization outcomes

While cloud platforms can provide several business advantages for global organizations, there are also distinct challenges that, if overlooked, can potentially harm or slow down an enterprise. Legacy systems and applications collapse into one team with a move to a cloud-based modern organization. There is often strong internal resistance to cloud modernization programs as it causes an inherent culling of resources. Moving legacy workloads that process extensive amounts of data to the cloud can be both timely and costly due to the size and bandwidth involved. Data security is also at risk without experienced developers or administrators managing cloud deployment. If security best practices are not followed, security can be made worse through misconfigured databases or negligent vendors. Regarding budgetary constraints, some executives predict over 30 percent of cloud spending is wasted. An organization must maintain a well-defined plan or strategy when adopting a cloud architecture to make every penny count. Implementation without strategy is a recipe for disaster.

Key aspects and approaches of cloud modernization

Digital transformations require careful planning before they are put into action. It’s essential for IT and business leaders to know the key aspects of the modernization process, including:

Security compliance
Cloud providers offer several advanced security features and compliance certifications to help organizations adhere to strict privacy standards.

Serverless computing
Developers can put aside infrastructure worries with a serverless platform that allows them to build and deploy applications more seamlessly.

Containerization/orchestration
Applications and their many dependencies are packaged into containers to ensure consistent deployment. Tools such as Kubernetes automate and manage containerized applications.

Initiating a modernization strategy

While there are several available routes when selecting a modernization strategy, three significant strategies outweigh the rest when maximizing value and minimizing effort. These strategies include:

Process modernization
Modernize development and work operations to lower the total cost of ownership. This promotes collaboration between the development and operations teams when implementing automated processes with reduced errors.

Application modernization
Use virtualized workloads to modernize an application or framework through platform-as-a-service (PaaS) solutions. The cloud platform manages health, availability, and deployment, while the business only needs to provide codes and select configuration options.

Database modernization
Improving how data is stored, processed, and fed through a modernized database provides scalable and flexible solutions that will enhance the value and quality of centralized data.

Ensuring a successful cloud-based initiative

According to a recent Gartner report, 70 percent of enterprise workloads will be cloud-based by 2024, yet three out of four organizations do not have a fit-for-purpose cloud strategy. An organization’s cloud migration or modernization success relies entirely on its strategic approach. A more effective plan is for enterprises to establish their business goals and primary use cases before implementing a cloud platform. Ultimately, the endgame for a cloud modernization initiative is to grow the business and deliver significant value. This includes developing new and efficient workflows, reducing and/or optimizing tech costs, expanding products and/or services, and developing new ideas through data-centric decision-making.

Unfortunately, many businesses encounter obstacles when developing their strategic initiatives and implementing cloud-centric plans. Common pitfalls to avoid when creating a migration approach include:

Single vendor
To avoid putting all their eggs in a singular vendor basket, savvy enterprises employ a multi-cloud strategy. A cloud provider is easily replaceable, so it’s essential to leverage multiple automated tools and applications that support data discovery and optimization analysis.
Too much, too fast
While it may seem tempting to do everything simultaneously, it is vital for enterprises to start small when migrating to the cloud. By transitioning the workload into smaller portions, organizations are better aware of the benefits and the potential value while temporarily maintaining old technologies for a more seamless migration process. This reduces any uncertainties before going all in.
Lack of focus
Upon initial migration, it is crucial to prove the cloud’s worth by starting with the right use case(s) that can best benefit from cloud analytics and processing massive amounts of data.

Successfully creating a cloud-centric culture for an organization means establishing a clear path with a destination in sight. Business leaders support growth and innovation by encouraging employees to experiment, take risks, and embrace the change that comes with an adaptive cloud migration.
To optimize the chances of success and sidestep potential roadblocks, building a framework to structure a migration helps to plan and manage the process and creates accountability for every step of the process. While it’s essential to understand that all modernization strategies should be customized to the organization’s unique needs, here are the key phases of a successful strategy (Figure 2):

Assessment phase.
Analyze existing systems, architectures, infrastructures, and security. Develop budgets, bill of materials, and key performance indicators.
Design phase
Develop a deployment model, document the infrastructure architecture, select a vendor, plan the product, and build support within the organization.
Migration phase
Leverage different storage options to migrate infrastructure, applications, and data.
Operation phase
Manage workloads in the cloud including monitoring performance, managing resources, and maintaining security and compliance.
Optimization and ongoing support phase
Identify opportunities to improve efficiency and performance, reengineering as necessary for cost savings, improved business value, and other enhancements.

image-4
Figure 2: Cloud strategy phases

The future of the cloud: exploring innovative data science, ML, and AI possibilities

Cloud modernization is changing the way enterprises work and conduct business. As cloud technology evolves and develops, global enterprises will exponentially grow their business in new and innovative ways. Recent Microsoft research estimates that the value of the cloud computing market will reach $1,240.9 billion by the end of 2027. Cloud platforms deliver the flexibility and scalability necessary for optimizing analytical research to maximize profitability. Cloud technologies will continue to revolutionize the IT landscape, leading to new developments, discoveries, and strategic opportunities for organizations.

image-5
Figure 3: Cloud computing projected market size 2020-2030

The future of the cloud remains bright. Enterprises are reaping the benefits of the cloud and avoiding common problems plaguing many organizations. Business growth and innovation will continue at an augmented rate as data science, machine learning, and artificial intelligence specific developments and cloud applications thrive and evolve. Cloud initiatives reduce costs and ensure timely operational practices. Global organizations can leverage clean, accurate, and reliable data by implementing a modern cloud infrastructure, allowing employees to focus on other essential tasks and ensuring enterprises flourish, evolve, and expand quickly through streamlined operations.

About the Author:

Senthilkumar Thirunavukarasu is an IT professional with nearly 20 years of experience in data analytics, cloud computing, AI, and ML. He has successfully architected and built large-scale data integration solutions in the cloud for numerous enterprises across various industries, with a strong focus on setting up cloud-based infrastructure, building data engineering solutions, configuring security features, setting up data governance platforms and compliance standards. For more information, contact [email protected].

Rethinking Reproducibility As the New Frontier in AI Research

Reproducibility in ai research

Reproducibility, integral to reliable research, ensures consistent outcomes through experiment replication. In the domain of Artificial Intelligence (AI), where algorithms and models play a significant role, reproducibility becomes paramount. Its role in promoting transparency and trust among the scientific community is crucial. Replicating experiments and obtaining similar results not only validates methodologies but also strengthens the scientific knowledge base, contributing to the development of more reliable and efficient AI systems.

Recent advancements in AI emphasize the need for improved reproducibility due to the rapid pace of innovation and the complexity of AI models. In particular, the instances of irreproducible findings, such as in a review of 62 studies diagnosing COVID-19 with AI, emphasize the necessity to reevaluate practices and highlight the significance of transparency.

Moreover, the interdisciplinary nature of AI research, involving collaboration between computer scientists, statisticians, and domain experts, emphasizes the need for clear and well-documented methodologies. Thus, reproducibility becomes a shared responsibility among researchers to ensure that accurate findings are accessible to a diverse audience.

Examining the Reproducibility Challenges in AI Research

Addressing reproducibility challenges is crucial, especially in the face of recent instances of non-reproducible results in diverse domains like machine learning, including natural language processing and computer vision. This is also an indication of the difficulties researchers encounter when trying to replicate published findings with identical codes and datasets, hindering scientific progress and casting doubts on the capability and reliability of AI techniques.

Non-reproducible results have far-reaching consequences, eroding trust within the scientific community and hampering the widespread adoption of innovative AI methodologies. Moreover, this lack of reproducibility poses a threat to implementing AI systems in critical industries like healthcare, finance, and autonomous systems, leading to concerns regarding the reliability and generalizability of models.

Multiple factors contribute to the reproducibility crisis in AI research. For instance, the complex nature of modern AI models, combined with a deficiency in standardized evaluation practices and inadequate documentation, presents challenges in duplicating experimental setups. Researchers sometimes prioritize innovation over thorough documentation due to pressures to publish groundbreaking results. The interdisciplinary aspect of AI research further complicates the scenario, with differences in experimental practices and communication gaps among researchers from varied backgrounds impeding the replication of results.

Common Reproducibility Challenges in AI Research

In particular, the following reproducibility challenges are significant and require careful consideration to mitigate their adverse effects.

Algorithmic Complexity

Complex AI algorithms often have complex architectures and numerous hyperparameters. Effectively documenting and conveying the details of these models is a challenge that hinders transparency and validation of results.

Variability in Data Sources

Diverse datasets are crucial in AI research, but challenges arise due to differences in data sources and preprocessing methods. Replicating experiments becomes complex when these issues related to data are not thoroughly documented, affecting the reproducibility of results.

Inadequate Documentation

The dynamic nature of AI research environments, encompassing rapidly evolving software libraries and hardware configurations, adds an extra layer of complexity. Inadequate documentation of changes in the computing environment can lead to discrepancies in result replication.

Lack of Standardization

In addition, the absence of standardized practices for experimental design, evaluation metrics, and reporting worsens reproducibility challenges.

The Significance of Reproducibility in Scientific Research

At its core, reproducibility involves the ability to independently replicate and validate experimental results or findings reported in a study. This practice holds fundamental importance for several reasons.

Firstly, reproducibility promotes transparency within the scientific community. When researchers provide comprehensive documentation of their methodologies, including code, datasets, and experimental setups, it allows others to replicate the experiments and verify the reported outcomes. This transparency builds trust and confidence in the scientific process.

Likewise, in the context of machine learning, reproducibility becomes particularly vital as models progress from the development phase to operational deployment. ML teams encounter challenges associated with algorithm complexity, diverse datasets, and the dynamic nature of real-world applications. Reproducibility acts as a safeguard against errors and inconsistencies during this transition. By ensuring the replicability of experiments and results, reproducibility becomes a tool for validating the accuracy of research outcomes.

In addition, ML models trained on specific datasets and under particular conditions may exhibit varied performance when exposed to new data or deployed in different environments. The ability to reproduce results empowers ML teams to verify the robustness of their models, identify potential pitfalls, and enhance the generalizability of the developed algorithms.

Moreover, troubleshooting and debugging are facilitated by reproducibility. ML practitioners often encounter challenges when dealing with issues that arise during the transition of models from controlled research settings to real-world applications. Reproducible experiments serve as a clear benchmark for comparison, assisting teams in identifying discrepancies, tracing error origins, and incrementally enhancing model performance.

Best Practices for Achieving Reproducibility in AI Research

To achieve reproducibility in AI research, adherence to best practices is necessary to ensure the accuracy and reliability of presented and published results.

  • Thorough documentation is essential in this regard, encompassing the experimental process, data, algorithms, and training parameters.
  • Clear, concise, and well-organized documentation facilitates reproducibility.
  • Likewise, implementing quality assurance protocols, such as version control systems and automated testing frameworks, helps track changes, validate results, and enhance research reliability.
  • Open-source collaboration plays a vital role in fostering reproducibility. Leveraging open-source tools, sharing code, and contributing to the community strengthens reproducibility efforts. Embracing open-source libraries and frameworks fosters a collaborative environment.
  • Data separation, with a standardized methodology for splitting training and testing data, is crucial for reproducibility in AI research experiments.
  • Transparency holds immense importance. Researchers should openly share methodologies, data sources, and results. Making code and data available to other researchers enhances transparency and supports reproducibility.

Incorporating the above practices promotes trust within the AI research community. By ensuring experiments are well-documented, quality-assured, open-source, data-separated, and transparent, researchers contribute to the foundation of reproducibility, reinforcing the reliability of AI research outcomes.

The Bottom Line

In conclusion, emphasizing the significance of reproducibility in AI research is paramount for establishing the authenticity of research efforts. Transparency, particularly in response to recent instances of non-reproducible results, emerges as a critical aspect. The adoption of best practices, including detailed documentation, quality assurance, open-source collaboration, data separation, and transparency, plays a pivotal role in cultivating a culture of reproducibility.

Do companies have ethical guidelines for AI use? 56% of professionals are unsure, survey says

Scale on a block

Although AI has been around since the 1950s, it has seen tremendous growth within the past year. Tech giants have been implementing AI into their products and services, while individuals are using it to make their lives a little easier.

Deloitte surveyed companies and professionals in its second edition of the "State of Ethics and Trust in Technology" report, led by its Technology Trust Ethics practice. According to the report, 74% of companies have already begun testing generative AI, while 65% have begun to use it internally. The increasing awareness of AI's new capabilities has led to the pressing question of how organizations can use this technology ethically.

Also: The ethics of generative AI: How we can harness this powerful technology

Deloitte interviewed 26 specialists in various industries to gather information about how industry leaders are considering concerns about the ethical use of emerging technologies, including generative AI.

The company then tested hypotheses and delivered a 64-question survey to more than 1,700 businesses and technical professionals to gain further insights.

The report, by Beena Ammanath, managing director of Deloitte Consulting LLP and leader of Deloitte's Technology Trust Ethics practice, refers to emerging technologies as the following: Cognitive technologies (including general and generative AI and chatbots), digital reality, ambient experiences, autonomous vehicles, quantum computing, distributed ledger technology, and robotics.

According to the survey, 39% of survey respondents, consisting of business leaders and developers of emerging technologies, thought cognitive technologies had the most potential for social good, compared to 12% in digital reality, and 12% in ambient experiences.

Also: 5 essential traits that tomorrow's AI leader must have

However, 57% of survey respondents also thought that cognitive technologies had the greatest potential for serious ethical risk.

The most concerning statistic is that over half of the respondents (56%) said their "company does not have or are unsure if they have ethical principles guiding the use of generative AI."

Compared to Deloitte's report in 2022 about ethics and trust in emerging technologies, this year's report reveals that "organizations find themselves wrestling with new ethical issues posed by wide-scale adoption of this once-again new technology."

These issues are tied to concerns about how businesses and organizations are using these technologies.

Despite the many benefits of AI, 22% of respondents were concerned with data privacy while 14% cited transparency about how AI is trained with data to produce its outputs.

Also: Does your business need a chief AI officer?

Data poisoning as well as intellectual property and copyright were concerns that each consisted of 12% of survey respondents. Data poisoning is the "pollution" of data training sets by bad actors and can lead to inaccurate results produced by AI.

Deloitte's report also detailed the types of damage that survey respondents believe could arise when ethical violations are not taken seriously.

Reputational damage was the greatest source of concern coming from 38% of respondents, followed by human damage such as misdiagnoses or data privacy violations (27%), regulatory penalties like copyright infringement (17%), financial damage (9%), and employee dissatisfaction (9%).

These damages are evident in the several lawsuits that have already been filed due to privacy violations, copyright infringement, and other issues related to the unethical use of AI.

Also: AI and automation: Business leaders adopt small-scale solutions for greater impact

So how can companies ensure they using AI safely? Deloitte lists a multi-step approach to helping companies:

  • Exploration: Companies can begin by letting product owners, business leaders, and AI/ML practitioners explore generative AI through workshops to see how it could create value for their businesses. This way, companies can recognize the costs and benefits of incorporating AI into their businesses.
  • Foundational: Companies could buy or build AI platforms to implement generative AI into their businesses. Of the survey respondents, 30% of survey respondents' companies chose to use existing capabilities with major AI platforms. 8% of respondents created their own in-house AI platforms, while 5% decided not to use generative AI.
  • Governance: Creating standards and protocols for AI use could minimize the potentially harmful impacts of AI, so companies should determine what types of ethical principles they plan to uphold.
  • Trainings and education: Companies could mandate trainings that outline the ethical principles of using AI. In addition, technical trainings that educate employees about using a variety of LLMs could provide companies with more guidance about the ethical use of AI.
  • Pilots: Engineers and product leaders could run experiments on a variety of use cases to test proof of concepts and pilot programs and then eliminate aspects that are too risky.
  • Implementation: Companies should draft a plan for introducing a newly enhanced product into the market and assign accountability for product implementation and ownership. The company should also have a team of experts prepared to address any issues that may arise. Transparency is also crucial for this step, as companies should explain how user data is inputted into the model, how the model reaches its output, and how likely the model is to hallucinate.
  • Audit: According to one interviewee, companies will need to modify their policies depending on the risks of AI use. This could vary company by company, as not all organizations will incorporate AI for the same use case.

In considering the impact of generative AI on human workers, issues such as transparency and data privacy ranked above job displacement. Nevertheless, the report also mentioned that "49% said workers at their organization displaced by AI moved to different roles and retrained and upskilled." Furthermore, 11% were terminated, 13% were put in a different role without being retrained or upskilled, and 27% did not experience any job displacement at their organization from AI, according to Deloitte.

Also: Will AI hurt or help workers? It's complicated

"The sooner companies work together to identify the risks and establish governance up front, the better their ability may be to help generate stakeholder value, elevate their brands, create new markets, and contribute to building a more equitable world," said Ammanath.

Artificial Intelligence

The Best Data Science Resources, Bootcamp, and Courses to Learn Data Science in the New Year

Sponsored Content

The Best Data Science Resources, Bootcamp, and Courses to Learn Data Science in the New Year

Unlocking the secrets of data science is not just a skill; it's the powerhouse driving AI innovation and leveraging the colossal wave of big data.

As we stand on the threshold of a new year, the opportunity to delve into the dynamic world of data science has never been more compelling. We've partnered with Springboard, the leading data science bootcamp offering personalized 1:1 mentorship, dedicated career support, proven outcomes, and an unbeatable money-back job guarantee, to present a handpicked collection of resources to supercharge your data science journey in the coming year.

Best Resources

The following are the best resources to learn data science topics.

Kaggle

Kaggle offers machine learning competitions where teams compete for cash prizes to solve a specific task that requires AI. The site is also full of community-compiled datasets, projects, and courses that can help you take the next step in your data science journey.

Learn Python The Hard Way

Python is a foundational programming language for data science. The ability to import modules that help with data analysis, such as Pandas, helps Python users ingest and organize an impressive array of data. Learn Python The Hard Way is a book that’s offered for free that will take you from beginner to being able to confidently get out there and build things in Python.

R for Data Science

R vs. Python is a topic that comes up often regarding data science. Academic data practitioners especially value the ability to be able to work with R and R packages. The ideal answer is to be conversant in both R and Python, but that can be difficult, to say the least. But if you wanted an equivalent book resource to learn R, this online book, available for free, is the one to consult.

Codecademy

If you feel you’d rather learn by doing than parsing through a textbook, you’re in luck. There are plenty of platforms that offer that mode of learning. Codecademy has a pretty comprehensive library for Python linked here, and a bunch of skill paths that will take you into more advanced fields like deep learning, data wrangling, and building chatbots.

freeCodeCamp

freeCodeCamp is a not-for-profit organization that offers a curriculum for learning code. They also have numerous certifications that are geared toward data analysis, machine learning, and data science topics.

The Open Source Data Science Masters

The Open Source Data Science Masters offers a free and open-source version of a Master’s level curriculum for learning data science. It can take you from the very beginning to practicing towards your first data science job—it’s a good option for those who can persevere all the way through with self-study, though other options are available if more support or intervention is needed, such as courses and/or bootcamps.

Awesome Data Science

Awesome Data Science is a GitHub repository with compiled data science resources. In general, GitHub can be an awesome resource for looking at foundational data science libraries and resources from their code foundations. Most repositories will also have Readmes and documentation that can flesh out their optimal use.

Best Bootcamp

Bootcamps are a good solution if you’re looking for dedicated focus and support and backing on your path to a new career in data science.

Springboard (New Year’s Promo — $1,500 off)

Springboard offers a leading bootcamp solution for those looking to get into data science, backed by proven outcomes as well as dedicated 1:1 mentorship and career coaching. More than 90% of job-qualified individuals in the bootcamp got a job within 12 months — and the average salary increase is around $25,000. These outcomes are backed by the Springboard Job Guarantee, which gives individuals the promise of getting a job or their money back. Using the code kdnuggetsshine2024 will give anybody $1,500 off until January 2, 2024 at 11:59 pm PST.

Best Courses

Courses can be a good option for people looking for an in-between on a bootcamp and a series of free resources, giving you enough time to really focus on a topic, even if you might not have direct access to human, personalized support.

Datacamp

Datacamp offers courses in a variety of data science topics. They are a specialized platform dedicated to data science, sourcing instruction and material from data science experts in the industry.

Python for Data Science and Machine Learning (Udemy)

Udemy offers a variety of courses for learning on the topic. Since it’s a marketplace for courses, top instructors in a topic can quickly stand out. In this case, this comprehensive course covers both data science and machine learning.

More On This Topic

  • Best Resources to Learn Natural Language Processing in 2021
  • Which is Best: Data Science Bootcamp vs Degree vs Online Course
  • 6 Best Free Online Courses to Learn Python and Boost Your Career
  • 5 Highest-paid Languages to Learn This Year
  • Top 3 Free Resources to Learn Linear Algebra for Machine Learning
  • Top Free Resources To Learn ChatGPT

AMD and Lamini Come to the Rescue of AI Startups

Believe it or not, there is still a shortage of NVIDIA GPUs in the market. AI startups, along with Fortune 500 companies, have been struggling to onboard GPUs for training AI models, even as small as Llama 7B. This is where Lamini with its partnership with AMD has found its moat — LLM fine tuning on AMD using ROCm.

“If you are an AI startup blocked on GPUs, send me a note,” posted Gregory Diamos, co-founder of Lamini, on LinkedIn. “At Lamini, we have figured out how to use AMD GPUs, which gives us a relatively large supply compared to the rest of the market.”

Similarly, co-founder Sharon Zhou posted on X, “We just brought online a relatively enormous supply of GPUs Lamini. We’re allocating a significant subset to promising startups & open-source initiatives, building the future of LLMs.”

Lamini co-founders claimed that they have the most competitive perf/$ on the market right now, “because we figured out how to use AMD GPUs to get software parity with CUDA, trending beyond CUDA.”

Zhou, at AMD’s Advancing AI event, also discussed how they have been leveraging AMD hardware and software all this while, and proving that the open nature of the technology has been helping them fully own the technology. “We have reached beyond CUDA,” said Zhou. This was after the big reveal two months ago that Lamini was exclusively running on AMD GPUs for the past year.

What exactly has Lamini figured out?

In September, Lamini opened up its LLM Superstation for both, cloud and on-premise for its customers. The LLM Superstation is a finely tuned supercomputer that incorporates 128 AMD Instinct GPUs, utilising Lamini on the AMD ROCm open software ecosystem.

Read: AMD’s ROCm is Ready To Challenge NVIDIA’s CUDA

Lamini and ROCm have achieved significant maturity, facilitating the effective fine-tuning of expansive LLMs, including Meta AI’s Llama 2, which now Lamini customers can book for training LLMs, and making their AI models proprietary.

To ensure this, Lamini incorporates sophisticated optimisations tailored for enterprise LLMs, leveraging and extending PEFT (LoRA), RLHF, and toolformer. These optimisations enable data isolation across 4,266x models on a single server, accelerate model switching by 1.09 billion times, compress models by a factor of 32x, and seamlessly integrate LLMs with enterprise APIs without the need for hyperparameter search.

Furthermore, Lamini has integrated a range of novel optimisations designed to expedite LLMs while harnessing the distinctive capabilities of AMD’s MI platform. These enhancements empower the hosting of 200 billion parameter models on a singular server, accommodating 10,000 fine-tuned language models on a single server, managing 12,800 concurrent requests to a solitary server, and efficiently processing more than 3.5 million queries per day on a single node.

“What’s more, with Lamini, you can stop worrying about the 52-week lead time for NVIDIA H100s,” reads the Lamini blog. “Using Lamini exclusively, you can build your own enterprise LLMs and ship them into production on AMD Instinct GPUs.” The company claims the cost of using Lamini on AMD Instinct GPUs is 10 times lesser than AWS, without the wait time.

“We’ve deployed Lamini in our internal Kubernetes cluster with AMD Instinct GPUs, and are using fine tuning to create models that are trained on AMD code base across multiple components for specific developer tasks,” said Vamsi Boppana, SVP of AI at AMD.

Moreover, Diamos claims that MI250X runs bigger models than NVIDIA’s A100s. “We chose the Instinct MI250 as the foundation for Lamini because it runs the biggest models that our customers demand and integrates fine tuning optimisations. We use the large HBM capacity (128GB) on MI250 to run bigger models with lower software complexity than clusters of A100s.”

Companies love AMD GPUs

“Building LLMs should be easy. Every enterprise should be able to own LLM IP, just like they do for all their other software. We’re excited to partner with AMD because their GPUs unlock a huge opportunity for enterprises to get started with little to no lead time,” said Zhou. “The main reason for this is compute strategy.”

Read: AMD Loves Llama So Much

Apart from Lamini, and through Lamini, many companies are increasingly testing out AMD’s products. Now that MI300X is also out, companies such as Microsoft, Meta, and Oracle have already announced the integration of the GPUs into their data centres and providing their customers AMD compute. OpenAI is also integrating support for ROCm on its Triton compiler.

Databricks and Essential AI had also announced the use of MI250X for their services and said at the Advancing AI event that they are ready for MI300X.

All this is because of AMD’s performance and its open source approach. The training performance of the MI300X is exactly equal to the NVIDIA H100. But when it comes to inference, MI300X using Bloom 176B and Llama 2 70B offers 1.6X and 1.4X faster performance.

Zoho’s Sridhar Vembu also revealed at the Global AI Conclave that given the shortage of GPU supplies, the company has also been running its models on AMD. “AMD has a very competitive silicon now, so with that I think the supply situation should resolve then it’s a matter of getting the talent capital.”

The post AMD and Lamini Come to the Rescue of AI Startups appeared first on Analytics India Magazine.

Top 8 Vernacular Language Models Based on Llama 2 

Llama 2 is currently one of the sought-after open-source models globally, with 757K downloads on HuggingFace. Its free availability for research and commercial purposes make it attractive to a broad user base, including researchers, developers, and hobbyists.

It has gained popularity in India and other countries worldwide, where users have created ‘Local Llamas’ based on the local language of the region. In models like GPT-3.5, there is a challenge with ‘tokenisation,’ where Indic language text is not efficiently represented.

Compared to massive LLMs like GPT-4 or 3.5, Llama 2 requires less computational power and training data. This makes it more feasible to train and run on local hardware, even in areas with limited infrastructure.

Moreover, Llama 2 can be easily fine-tuned for specific tasks and domains using smaller datasets of local language text. This allows developers to tailor the model to their specific language and needs, improving its performance and relevance.

OpenHathi

Recently in India, Sarvam AI introduced OpenHathi-Hi-v0.1, marking the debut of the first Hindi LLM in the OpenHathi series. Built on a cost-effective platform, this model, extending from Llama2-7B, exhibits performance akin to GPT-3.5 for Indic languages.

OpenHathi, featuring a 48K-token extension of Llama2-7B’s tokenizer, undergoes a two-phase training process. The initial phase focuses on embedding alignment, aligning randomly initialized Hindi embeddings, followed by bilingual language modeling, teaching the model cross-lingual attention across tokens.

The model demonstrates robust performance across various Hindi tasks, comparable to, if not surpassing, GPT-3.5, while maintaining English proficiency. Sarvam AI’s evaluation includes non-academic, real-world tasks alongside standard Natural Language Generation (NLG) tasks. Evaluations against GPT-3.5 generation with GPT-4 as the judge revealed superior performance in Hindi, both in native and Romanised scripts.

Tamil Llama

Kaggle ML Engineer and Kaggle Master Abhinand Balachandra, recently introduced Tamil Llama, an Indic LLM engineered specifically to elevate the Tamil language domain. This AI model is built on top of Meta’s Llama 2.

It was trained with an additional 16,000 Tamil tokens, aiming to achieve superior text generation and comprehension in the Tamil language. This model serves as an extension of the LLaMA model and has been enhanced by incorporating extra Tamil tokens, utilizing the LoRA methodology for efficient training.

It presents four distinct variations: Tamil LLaMA 7B, 13B, 7B Instruct, and 14B Instruct. Throughout the training phase, the model’s vocabulary has been expanded to encompass 16,000 Tamil tokens, supplementing the original 32K tokens.

Telugu Llama

Ramsri Goutham Golla from Segmind.com is currently working on Telugu Llama. Meticulously trained on LLama 2, the popular language model, showcasing remarkable efficiency in token count for Telugu text.

Interestingly, Telugu language consumes less tokens as compared to English as per Golla. This breakthrough not only promises faster and cost-effective text generation in Telugu but also sets the stage for LLama 2 to excel in Indic languages, revolutionising the landscape of natural language processing.

Odia Generative AI

odia_llama2_7B_v1 created by Odia Generative AI is based on Llama2-7b and fine tuned with a 180k Odia instruction set. This set includes translated data from open-source resources and a purposefully crafted domain knowledge instruction set. The result is a model that effectively understands Odia instructions and generates responses, demonstrating its practical utility for the nuances of the Odia language.

SeaLLMs – Large Language Models for Southeast Asia

Alibaba Group Holding’s research division, Damo Academy, has introduced LLMs specifically designed for Southeast Asian languages. SeaLLMs are built upon the Llama-2 model and further advanced through continued pre-training with an extended vocabulary, specialized instruction and alignment tuning to better capture the intricacies of regional languages.

The Southeast Asia LLM (SeaLLM) was pre-trained on Vietnamese, Indonesian, Thai, Malay, Khmer, Lao, Tagalog, and Burmese data sets, and has outperformed other open-source models in linguistic and safety tasks. SeaLLMs exhibit remarkable proficiency in language understanding and generation tasks, presenting a formidable challenge to dominant models like ChatGPT-3.5, especially in Southeast Asian (SEA) languages.

VinaLLaMA: LLaMA-based Vietnamese Foundation Model

VinaLLaMA, a foundational LLM designed specifically for the Vietnamese language. VinaL- LaMA, built on top of LLaMA-2, represents a vital stride towards linguistic inclusivity in AI, adeptly addressing the syntactic and semantic intricacies of Vietnamese.

LLaMAntino: LLaMA 2 Models for Effective Text Generation in Italian Language

LLaMAntino is the Italian iteration of Llama 2, developed by researchers at the University of Bari Aldo Moro, Italy. Through the fine-tuning of pre-trained LLaMA 2 models with 13 and 7 billion parameters and using a substantial dataset of Italian text, they created the LLaMAntino-2-7b-hf-ITA and LLaMAntino-2-13b-hf-ITA models.

This model aims to provide Italian NLP researchers with a tool to tackle tasks such as information extraction and closed qa.

Dutch Llama v2 13B

It is a Dutch model based on Llama 2. This model is a fine-tuned version of BramVanroy/llama2-13b-ft-mc4_nl_cleaned_tiny on the BramVanroy/dutch_chat_datasets dataset on a context of 4096 tokens.

Depending on the input, the model can provide satisfactory outcomes, taking into account its 13B size and limited pretraining in Dutch. However, it’s essential to note that the model lacks human feedback during training and lacks safeguards, potentially leading to unexpected or offensive content based on the query.

The post Top 8 Vernacular Language Models Based on Llama 2 appeared first on Analytics India Magazine.

7 Reasons Why You’re Struggling to Land a Data Science Job

7 Reasons Why You're Struggling to Land a Data Science Job
Image by Editor

Tired of applying to data science roles and not hearing back from companies? Perhaps you managed to land a couple of interviews but weren't able to convert them to offers? Well, you’re not alone.

The job market is brutally competitive now. So just because it's difficult doesn't mean you're not good enough. That said, it's both important and helpful to take a step back and see how and where you can improve. And that’s exactly what this guide will help you with.

We’ll go over common reasons why aspiring data professionals like you struggle to make the cut. And how you can improve your chances of landing interviews and getting that job you want!

1. Your Coding Skills Are Rusty

It's a hard truth. So let's face it.

Say you’ve applied to a bunch of data science roles at companies that you’re interested in. And have been shortlisted for interviews.

Congratulations! You’re on the right track. The next goal is to convert the interview opportunity to a job offer. And the first step is to crack that coding interview.

You’ll first have a round of timed coding interviews—testing your problem-solving skills—followed by an SQL coding round.

But coding interviews are difficult to crack—even for experienced professionals. But consistent practice and spaced repetition can help you successfully crack these interviews.

Regularly practice coding interview questions on platforms like Leetcode and Hackerrank.

If you are looking for resources check out:

  • 5 Free Books to Help You Master Python
  • 7 Best Platforms to Practice SQL

Once you clear coding interviews, focus and prepare for technical rounds. Brush up your machine learning fundamentals. Also review your projects so you can explain their impact with confidence.

2. Your Résumé Does Not Stand Out

It is true that recruiters spend only a few seconds reviewing your resume and decide if it proceeds to the next phase or to the reject pile.

7 Reasons Why You're Struggling to Land a Data Science Job
Image by Author

So you should put in conscientious efforts to draft your resume. Be sure to tailor your resume based on the job specifications.

Here are a few resume tips:

  • Include relevant experiences and education sections.
  • List experiences and education and reverse chronological order.
  • Summarize experience in bulleted lists—quantifying impact and adding concise explanations.
  • Include a relevant projects section. Explain the projects in concise bullet points. Also include links to the projects.
  • Add a relevant skills section grouped by category like programming languages, tools and frameworks etc.

I’ll also suggest using a simple single-column layout that's easier to parse than complicated and fancy layouts.

3. Your Profile Is Either Too Specific Or Too Generic

When you’re applying to jobs, your resume and LinkedIn profile should be consistent without any conflicting details. And they should also be aligned with the experience and skill set that the role demands.

There are a couple of caveats you should avoid, though.

Your Profile Is Too Specific

Suppose you’re interested in medical imaging and computer vision. So almost all your projects are in computer vision. Such a profile may be a great fit for a computer vision engineer or a computer vision researcher role.

But what if you’re applying to a data scientist role at a FinTech company? Clearly, you don't stand out as a strong candidate.

Your Profile Is Too Generic

If you are an aspiring data scientist with strong SQL skills and experience building machine learning models, you can apply for the roles of data analyst and machine learning engineer as well.

But you don't want to make your resume/candidate profile look like you’re someone who wants to be a data analyst, a machine learning engineer, and a data scientist—all at once.

If you’re interested in all of these roles, have separate resumes for each.

It’s important to find a sweet middle ground that allows you to showcase your expertise and stand out as a potential candidate with a broad skill set that is aligned with the job’s requirements.

4. Your Projects Are Not Interesting Enough

Your projects help you gain a competitive edge over other candidates. So choose them wisely.

7 Reasons Why You're Struggling to Land a Data Science Job
Image by Author

Some aspiring data professionals put on their resume and portfolio certain projects which they shouldn't be. Yes, there are some beginner projects which are good for learning—but you should AVOID showcasing them in your portfolio.

Here are a few:

  • Titanic survival prediction
  • MNIST handwritten digit recognition
  • Classification using the iris dataset
  • Projects on the wine dataset

Just to name a few. These projects are too generic and basic to be able to land you an interview (let alone job offers).

So what are some interesting projects—especially if you are a beginner who is looking to break into this field?

Here are some beginner-level projects that would help you showcase your skills and emerge as a stronger candidate:

  • Customer segmentation
  • Loan default prediction
  • Market basket analysis
  • Customer churn prediction

Use real-world datasets to build your projects. This way you can showcase a lot of important skills: data collection, data cleaning, and exploratory data analysis besides model building.

Also include projects that are inspired by your interest. As I’d suggested in a previous pandas guide, try turning data from your interests and hobbies into interesting projects that will help you leave an impression on the interviewer.

5. Your Degree Stands in the Way

Another common road block aspiring data professionals face is their educational background. Breaking into data science can be especially difficult if you have majored in a field such as sociology, psychology, and the like.

While your skills—hard and soft skills—matter eventually, you should remember that you are competing with those who have an undergraduate or advanced degree in a related field.

So what can you do about this?

Look for ways to constantly upskill yourself. Remember, once you land your first data role, you can leverage your experience going forward.

Look for ways to work on relevant projects within your company. If your company has a dedicated data team, try to accept a small side project.

6. You’re Probably Not Learning in Public

Learning in public is super important, especially when you are trying to land your first job (and even after that, honestly).

I started writing online in late 2020. Since then, I’ve landed most of my opportunities through my work—tutorials and technical deep dives—that I published online.

So how and where do you start? Leverage social media platforms like LinkedIn and Twitter (X) to share your work with the community:

  • Built a project? Share it with your network. Ask for feedback. Improve.
  • Wrote a data science tutorial? Share it with your network.
  • Learned something new? Share it anyway.
  • Ran into an error that you eventually fixed? Yes, it’s worth sharing.

What you code on your laptop stays on your laptop. So be ready to put yourself out there and share what you build and learn.

Building a strong portfolio and online presence can be immensely helpful in the job search process. Because you never know which project or article might interest your future employer.

7. You Could Be More Proactive

Because of how competitive the job market is right now, you have to go beyond just applying to jobs—and start being more proactive.

7 Reasons Why You're Struggling to Land a Data Science Job
Image by Author

Here are a few simple steps that can help you make the difference:

  • Shortlist companies you're interested in.
  • Check for relevant openings.
  • Reach out to the recruiter with your resume and portfolio explaining why you would be a good fit for the role.
  • Connect with other professionals. Get into the habit of networking even when you have a stable full-time job.

Joining data science communities online can also be super helpful!

Wrapping Up

And that's a wrap. Here’s a quick review of what we’ve discussed:

  • Prepare for coding interviews. Practice on platforms like Leetcode and Hackerrank.
  • Tailor your profile to align with job requirements. But be consistent.
  • Put in efforts to work on your resume and project portfolio.
  • Start learning in public. Share what you build and learn.
  • Be proactive in networking with other professionals.

Good luck on your job search journey. I hope you land your data science role soon. What else would you add? Let us know in the comments.

Bala Priya C is a developer and technical writer from India. She likes working at the intersection of math, programming, data science, and content creation. Her areas of interest and expertise include DevOps, data science, and natural language processing. She enjoys reading, writing, coding, and coffee! Currently, she's working on learning and sharing her knowledge with the developer community by authoring tutorials, how-to guides, opinion pieces, and more.

More On This Topic

  • Unable to Land a Data Science Job? Here’s Why
  • KDnuggets™ News 22:n05, Feb 2: 7 Steps to Mastering Machine…
  • Data Science Projects That Will Land You The Job in 2022
  • A Data Science Portfolio That Will Land You The Job in 2022
  • 3 Data Science Projects Guaranteed to Land You That Job
  • A Data Science Portfolio That Will Land You The Job

Tech Mahindra Completes Project Indus, to Launch Soon 

Tech Mahindra’s outgoing chief, CP Gurnani, shared a Project Indus update on X, highlighting the development of an LLM specifically designed for Hindi and its 37 dialects.

Retiring after 19 years at Tech Mahindra he said, “I feel proud to share that the GenerativeAI project challenge I took up earlier has been successfully accomplished by our brilliant research team at Makers Lab.” Project Indus is in the beta testing phase within TechM. The pure Hindi LLM consists of 539 million parameters and 10 billion Hindi+ dialect tokens, as shared by Gurnani.

As I finish my run at @tech_mahindra, I feel proud to share that the #GenerativeAI project challenge I took up earlier has been successfully accomplished by our brilliant research team at Makers Lab.
We launched #ProjectIndus in beta testing phase within TechM. The pure Hindi… https://t.co/MjTsvWCGvk

— CP Gurnani (@C_P_Gurnani) December 20, 2023

“The model is probably the only one in the world that has all Hindi tokens and has been trained from the ground up. It will set the stage for the years to come as our prowess in deep tech,” he said.

“I now pass the baton over to Mohit Joshi, Nikhil Malhotra, and the team, as well as my talented Tech Mighties, who will take this one notch higher. Thank you for everything, team!” he added.

In October, Tech Mahindra announced its plan to release Project Indus by the end of December or early January. Introduced in August, the model is initially set to support 40 different Hindi dialects, with plans to add more languages and dialects in subsequent releases. Over the last two months, the 15-member Project Indus team has collected 1.2 terabytes of data in Hindi and related dialects.

The update by Gurnani comes in the backdrop of last week, which witnessed a slew of announcements from Indian companies and startups launching their large language models (LLMs). This includes Google-backed CoRover’s BharatGPT, Khosla Ventures-backed Sarvam.ai’sOpenHathi, Microsoft-backed Kissan AI’s Dhenu, and Ola’s Krutrim.

The post Tech Mahindra Completes Project Indus, to Launch Soon appeared first on Analytics India Magazine.