Skip to main navigation Skip to search Skip to main content

Threat Landscape of Adversarial Attacks on Generative AI and Large Language Models (LLMs): Exploring Different Types of Adversarial Attacks, Associated Risks, and Mitigation Strategies

  • Ishita Naik
  • , Dishita Naik
  • , Nitin Naik*
  • *Corresponding author for this work
  • Aston University
  • Birmingham City University

Research output: Chapter in Book/Published conference outputConference publication

Abstract

Generative Artificial Intelligence (Generative AI) has emerged as a transformative catalyst across disciplines and applications, fundamentally enhancing creativity, productivity, personalization, and problem-solving by creating novel, coherent, and contextually relevant content. Large Language Models (LLMs) are a specific type of generative AI that focuses mainly on understanding, generating, and manipulating human language. The proliferation of LLMs has amplified the threat of adversarial attacks that maliciously manipulate inputs or training data with the adversarial intent to exploit, compromise, or mislead LLMs. Unlike conventional cyberattacks that exploit commonly known software vulnerabilities or directly attack the IT infrastructure, adversarial attacks on LLMs exploit the statistical and linguistic patterns the LLMs have learned. These attack strategies coupled with the sheer scale of LLM deployments and diversity of inputs, exacerbates detection and mitigation challenges of ever-evolving adversarial attacks. Therefore, this paper will explore adversarial attacks on LLMs, and their associated risks and mitigations. Initially, it will explain what adversarial attacks on LLMs are and how they differ from conventional cyberattacks. Next, it will elucidate several types of adversarial attacks on LLMs to provide an in-depth overview of their nature and impact. Afterwards, it will review various risks associated with adversarial attacks on LLMs. Lastly, it will present various mitigation strategies for adversarial attacks on LLMs. This detailed analysis of adversarial attacks on LLMs, along with their associated risks and mitigation strategies, aims to provide in-depth insights into the security and safety challenges inherent to LLM deployment and usage.
Original languageEnglish
Title of host publicationContributions Presented at the International Conference on Computing, Communication, Cybersecurity & AI, July 10–11, 2025, Birmingham, UK
Subtitle of host publicationThe C3AI 2025
EditorsNitin Naik, Paul Jenkins, Shaligram Prajapat, Paul Grace
PublisherSpringer, Cham
Pages765-789
Number of pages25
ISBN (Electronic)9783032167910
ISBN (Print)9783032167903
DOIs
Publication statusPublished - 17 May 2026

Publication series

NameLecture Notes in Networks and Systems (LNNS)
PublisherSpringer Cham
Volume1811
ISSN (Print)2367-3370
ISSN (Electronic)2367-3389

Bibliographical note

Copyright © 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG. This version of the article has been accepted for publication, after peer review and is subject to Springer Nature’s AM terms of use [ https://www.springernature.com/gp/open-research/policies/accepted-manuscript-terms ] but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: https://doi.org/10.1007/978-3-032-16791-0

Keywords

  • AI model extraction attacks
  • AI model inversion attacks
  • AI model poisoning attacks
  • AI model theft attacks
  • AI models
  • Adversarial attacks
  • Backdoor poisoning attacks
  • Cyberattacks
  • Data poisoning attacks
  • Direct prompt injection attacks
  • Excessive agency attacks
  • Feature poisoning attacks
  • Generative AI
  • Indirect prompt injection attacks
  • Jailbreak attacks
  • LLMs
  • Label poisoning attacks
  • Large language models
  • Neural prompt-to-prompt attacks
  • Noise injection attacks
  • Prompt hijacking attacks
  • Prompt injection attacks
  • Prompt leaking attacks
  • RAG injection attacks
  • RAG poisoning attacks

Fingerprint

Dive into the research topics of 'Threat Landscape of Adversarial Attacks on Generative AI and Large Language Models (LLMs): Exploring Different Types of Adversarial Attacks, Associated Risks, and Mitigation Strategies'. Together they form a unique fingerprint.

Cite this