Skip to main navigation Skip to search Skip to main content

When Generative AI Gets Hacked: A Comprehensive Classification of Cyberattacks on Large Language Models (LLMs) and Their Mitigation Techniques

  • Birmingham City University
  • Aston University

Research output: Chapter in Book/Published conference outputConference publication

Abstract

Large Language Models (LLMs) have swiftly become prevalent in nearly every aspect of human life due to a combination of technological breakthroughs, practical usability, and rapid integration into everyday tools and workflows. Despite their remarkable capabilities, LLMs pose real challenges in their secure and safe development and deployment; and are vulnerable to various cyberattacks that can compromise their behaviour, outputs, security and performance. Understanding these vulnerabilities and potential cyberattacks on LLMs is essential for ensuring their secure and safe development and deployment. Numerous types of cyberattacks can be launched against LLMs, and there is currently no universally accepted classification system for these cyberattacks, as this remains an evolving area of research. This paper will provide a systematic and broad classification of LLM attacks into four major categories based on its four inherent and important components: input prompt, training data, underlying AI model and output; and these four categories of LLM attacks are: Input (Prompt) Related Cyberattacks, Data (Training) Related Cyberattacks, AI Model (Inference) Related Cyberattacks, and Output (Response) Related Cyberattacks. This paper will discuss all four aforementioned categories of cyberattacks on LLMs in detail including various types of cyberattacks in each category. Subsequently, it will discuss several risks associated with cyberattacks on LLMs. Finally, it will discuss several mitigation techniques for cyberattacks on LLMs. A rigorous examination and taxonomy of diverse cyberattacks targeting LLMs, alongside an analysis of associated risks and mitigation strategies, is poised to yield nuanced and actionable understanding regarding the security and safety landscape of LLMs. Through systematic classification and evaluation, such research will advance the field by illuminating various cyberattacks, vulnerabilities, risks, and defensive measures pertinent to LLM-based systems, thereby supporting more robust deployment and governance of these technologies in sensitive environments.
Original languageEnglish
Title of host publicationContributions Presented at the International Conference on Computing, Communication, Cybersecurity & AI, July 10–11, 2025, Birmingham, UK
Subtitle of host publicationThe C3AI 2025
EditorsNitin Naik, Paul Jenkins, Shaligram Prajapat, Paul Grace
PublisherSpringer, Cham
Pages97-130
Number of pages30
Volume1811
ISBN (Electronic)9783032167910
ISBN (Print)9783032167903
DOIs
Publication statusPublished - 17 May 2026

Publication series

NameLecture Notes in Networks and Systems (LNNS)
PublisherSpringer Cham
ISSN (Print)2367-3370
ISSN (Electronic)2367-3389

Bibliographical note

Copyright © 2026 The Author(s), under exclusive licence to Springer Nature Switzerland AG. 2024. This version of the article has been accepted for publication, after peer review and is subject to Springer Nature’s AM terms of use [ https://www.springernature.com/gp/open-research/policies/accepted-manuscript-terms ] but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: https://doi.org/10.1007/978-3-032-16791-0_5

Keywords

  • Generative AI
  • Large language models
  • LLMs
  • Cyberattacks on LLM
  • Attacks on LLMs

Fingerprint

Dive into the research topics of 'When Generative AI Gets Hacked: A Comprehensive Classification of Cyberattacks on Large Language Models (LLMs) and Their Mitigation Techniques'. Together they form a unique fingerprint.

Cite this