TY - GEN
T1 - When Generative AI Gets Hacked: A Comprehensive Classification of Cyberattacks on Large Language Models (LLMs) and Their Mitigation Techniques
AU - Naik, Dishita
AU - Naik, Ishita
AU - Naik, Nitin
N1 - Copyright © 2026 The Author(s), under exclusive licence to Springer Nature Switzerland AG. 2024. This version of the article has been accepted for publication, after peer review and is subject to Springer Nature’s AM terms of use [ https://www.springernature.com/gp/open-research/policies/accepted-manuscript-terms ] but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: https://doi.org/10.1007/978-3-032-16791-0_5
PY - 2026/5/17
Y1 - 2026/5/17
N2 - Large Language Models (LLMs) have swiftly become prevalent in nearly every aspect of human life due to a combination of technological breakthroughs, practical usability, and rapid integration into everyday tools and workflows. Despite their remarkable capabilities, LLMs pose real challenges in their secure and safe development and deployment; and are vulnerable to various cyberattacks that can compromise their behaviour, outputs, security and performance. Understanding these vulnerabilities and potential cyberattacks on LLMs is essential for ensuring their secure and safe development and deployment. Numerous types of cyberattacks can be launched against LLMs, and there is currently no universally accepted classification system for these cyberattacks, as this remains an evolving area of research. This paper will provide a systematic and broad classification of LLM attacks into four major categories based on its four inherent and important components: input prompt, training data, underlying AI model and output; and these four categories of LLM attacks are: Input (Prompt) Related Cyberattacks, Data (Training) Related Cyberattacks, AI Model (Inference) Related Cyberattacks, and Output (Response) Related Cyberattacks. This paper will discuss all four aforementioned categories of cyberattacks on LLMs in detail including various types of cyberattacks in each category. Subsequently, it will discuss several risks associated with cyberattacks on LLMs. Finally, it will discuss several mitigation techniques for cyberattacks on LLMs. A rigorous examination and taxonomy of diverse cyberattacks targeting LLMs, alongside an analysis of associated risks and mitigation strategies, is poised to yield nuanced and actionable understanding regarding the security and safety landscape of LLMs. Through systematic classification and evaluation, such research will advance the field by illuminating various cyberattacks, vulnerabilities, risks, and defensive measures pertinent to LLM-based systems, thereby supporting more robust deployment and governance of these technologies in sensitive environments.
AB - Large Language Models (LLMs) have swiftly become prevalent in nearly every aspect of human life due to a combination of technological breakthroughs, practical usability, and rapid integration into everyday tools and workflows. Despite their remarkable capabilities, LLMs pose real challenges in their secure and safe development and deployment; and are vulnerable to various cyberattacks that can compromise their behaviour, outputs, security and performance. Understanding these vulnerabilities and potential cyberattacks on LLMs is essential for ensuring their secure and safe development and deployment. Numerous types of cyberattacks can be launched against LLMs, and there is currently no universally accepted classification system for these cyberattacks, as this remains an evolving area of research. This paper will provide a systematic and broad classification of LLM attacks into four major categories based on its four inherent and important components: input prompt, training data, underlying AI model and output; and these four categories of LLM attacks are: Input (Prompt) Related Cyberattacks, Data (Training) Related Cyberattacks, AI Model (Inference) Related Cyberattacks, and Output (Response) Related Cyberattacks. This paper will discuss all four aforementioned categories of cyberattacks on LLMs in detail including various types of cyberattacks in each category. Subsequently, it will discuss several risks associated with cyberattacks on LLMs. Finally, it will discuss several mitigation techniques for cyberattacks on LLMs. A rigorous examination and taxonomy of diverse cyberattacks targeting LLMs, alongside an analysis of associated risks and mitigation strategies, is poised to yield nuanced and actionable understanding regarding the security and safety landscape of LLMs. Through systematic classification and evaluation, such research will advance the field by illuminating various cyberattacks, vulnerabilities, risks, and defensive measures pertinent to LLM-based systems, thereby supporting more robust deployment and governance of these technologies in sensitive environments.
KW - Generative AI
KW - Large language models
KW - LLMs
KW - Cyberattacks on LLM
KW - Attacks on LLMs
UR - https://link.springer.com/chapter/10.1007/978-3-032-16791-0_5
UR - https://www.scopus.com/pages/publications/105040521321
U2 - 10.36227/techrxiv.176540281.10631689/v1
DO - 10.36227/techrxiv.176540281.10631689/v1
M3 - Conference publication
SN - 9783032167903
VL - 1811
T3 - Lecture Notes in Networks and Systems (LNNS)
SP - 97
EP - 130
BT - Contributions Presented at the International Conference on Computing, Communication, Cybersecurity & AI, July 10–11, 2025, Birmingham, UK
A2 - Naik, Nitin
A2 - Jenkins, Paul
A2 - Prajapat, Shaligram
A2 - Grace, Paul
PB - Springer, Cham
ER -