TY - GEN
T1 - Threat Landscape of Adversarial Attacks on Generative AI and Large Language Models (LLMs): Exploring Different Types of Adversarial Attacks, Associated Risks, and Mitigation Strategies
AU - Naik, Ishita
AU - Naik, Dishita
AU - Naik, Nitin
N1 - Copyright © 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG. This version of the article has been accepted for publication, after peer review and is subject to Springer Nature’s AM terms of use [ https://www.springernature.com/gp/open-research/policies/accepted-manuscript-terms ] but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: https://doi.org/10.1007/978-3-032-16791-0
PY - 2026/5/17
Y1 - 2026/5/17
N2 - Generative Artificial Intelligence (Generative AI) has emerged as a transformative catalyst across disciplines and applications, fundamentally enhancing creativity, productivity, personalization, and problem-solving by creating novel, coherent, and contextually relevant content. Large Language Models (LLMs) are a specific type of generative AI that focuses mainly on understanding, generating, and manipulating human language. The proliferation of LLMs has amplified the threat of adversarial attacks that maliciously manipulate inputs or training data with the adversarial intent to exploit, compromise, or mislead LLMs. Unlike conventional cyberattacks that exploit commonly known software vulnerabilities or directly attack the IT infrastructure, adversarial attacks on LLMs exploit the statistical and linguistic patterns the LLMs have learned. These attack strategies coupled with the sheer scale of LLM deployments and diversity of inputs, exacerbates detection and mitigation challenges of ever-evolving adversarial attacks. Therefore, this paper will explore adversarial attacks on LLMs, and their associated risks and mitigations. Initially, it will explain what adversarial attacks on LLMs are and how they differ from conventional cyberattacks. Next, it will elucidate several types of adversarial attacks on LLMs to provide an in-depth overview of their nature and impact. Afterwards, it will review various risks associated with adversarial attacks on LLMs. Lastly, it will present various mitigation strategies for adversarial attacks on LLMs. This detailed analysis of adversarial attacks on LLMs, along with their associated risks and mitigation strategies, aims to provide in-depth insights into the security and safety challenges inherent to LLM deployment and usage.
AB - Generative Artificial Intelligence (Generative AI) has emerged as a transformative catalyst across disciplines and applications, fundamentally enhancing creativity, productivity, personalization, and problem-solving by creating novel, coherent, and contextually relevant content. Large Language Models (LLMs) are a specific type of generative AI that focuses mainly on understanding, generating, and manipulating human language. The proliferation of LLMs has amplified the threat of adversarial attacks that maliciously manipulate inputs or training data with the adversarial intent to exploit, compromise, or mislead LLMs. Unlike conventional cyberattacks that exploit commonly known software vulnerabilities or directly attack the IT infrastructure, adversarial attacks on LLMs exploit the statistical and linguistic patterns the LLMs have learned. These attack strategies coupled with the sheer scale of LLM deployments and diversity of inputs, exacerbates detection and mitigation challenges of ever-evolving adversarial attacks. Therefore, this paper will explore adversarial attacks on LLMs, and their associated risks and mitigations. Initially, it will explain what adversarial attacks on LLMs are and how they differ from conventional cyberattacks. Next, it will elucidate several types of adversarial attacks on LLMs to provide an in-depth overview of their nature and impact. Afterwards, it will review various risks associated with adversarial attacks on LLMs. Lastly, it will present various mitigation strategies for adversarial attacks on LLMs. This detailed analysis of adversarial attacks on LLMs, along with their associated risks and mitigation strategies, aims to provide in-depth insights into the security and safety challenges inherent to LLM deployment and usage.
KW - AI model extraction attacks
KW - AI model inversion attacks
KW - AI model poisoning attacks
KW - AI model theft attacks
KW - AI models
KW - Adversarial attacks
KW - Backdoor poisoning attacks
KW - Cyberattacks
KW - Data poisoning attacks
KW - Direct prompt injection attacks
KW - Excessive agency attacks
KW - Feature poisoning attacks
KW - Generative AI
KW - Indirect prompt injection attacks
KW - Jailbreak attacks
KW - LLMs
KW - Label poisoning attacks
KW - Large language models
KW - Neural prompt-to-prompt attacks
KW - Noise injection attacks
KW - Prompt hijacking attacks
KW - Prompt injection attacks
KW - Prompt leaking attacks
KW - RAG injection attacks
KW - RAG poisoning attacks
UR - https://link.springer.com/chapter/10.1007/978-3-032-16791-0_38
UR - https://www.scopus.com/pages/publications/105040532063
U2 - 10.1007/978-3-032-16791-0_38
DO - 10.1007/978-3-032-16791-0_38
M3 - Conference publication
SN - 9783032167903
T3 - Lecture Notes in Networks and Systems (LNNS)
SP - 765
EP - 789
BT - Contributions Presented at the International Conference on Computing, Communication, Cybersecurity & AI, July 10–11, 2025, Birmingham, UK
A2 - Naik, Nitin
A2 - Jenkins, Paul
A2 - Prajapat, Shaligram
A2 - Grace, Paul
PB - Springer, Cham
ER -