Securing Large Language Models Against Prompt Injection and Data Leakage Attacks

https://doi.org/10.5281/zenodo.19344116

Authors

  • Waleed Khan College of dupage 425 Fawell Blvd, Glen Ellyn, IL 60137, United States Author
  • Muhammad Akram Harper community college 1200 Algonquin Rd, Palatine, IL 60067, United States Author
  • Naseer Ahmad Department of Computer Science Lewis University USA Author
  • Amir Mohammad Delshadi New Mexico Highlands University, Las Vegas, MN, USA Author
  • Meher Sultana New Mexico Highlands University, Las Vegas, MN, USA Author
  • Muhammad Waleed Iqbal Department of Computer Science Comsats University Islamabad, Sahiwal Campus Author

Keywords:

, Prompt Injection Attacks, Data Leakage Prevention, AI Security, Adversarial Training, Natural Language Processing Security, Secure AI Systems

Abstract

Large Language Models (LLMs) have rapidly become integral to modern artificial intelligence applications, including conversational agents, software development tools, intelligent tutoring systems, and enterprise decision-support platforms. Their ability to process and generate human-like text has enabled widespread adoption across multiple industries. However, the increasing deployment of LLMs in real-world environments has introduced significant security and privacy risks, particularly in the form of prompt injection attacks and sensitive data leakage. Prompt injection attacks occur when maliciously crafted inputs manipulate the model into overriding system instructions, bypassing safety policies, or revealing hidden prompts and confidential information. Similarly, data leakage vulnerabilities may expose proprietary knowledge, training data, or sensitive organizational information through model responses. To address these challenges, this study proposes a comprehensive security architecture designed to enhance the robustness of LLM systems against adversarial prompts and information leakage. The proposed framework integrates multiple defense layers, including prompt validation, contextual anomaly detection, adversarial training, input sanitization, and output filtering mechanisms. Additionally, reinforcement learning–based safety policies are incorporated to guide model responses and enforce security constraints during inference. Experimental evaluation was conducted using a dataset containing both benign queries and adversarial prompts designed to exploit LLM vulnerabilities. Results demonstrate that the proposed framework significantly improves model resilience compared to baseline LLM implementations. Specifically, the prompt injection attack success rate decreased from 62% to 12%, while the data leakage rate dropped from 55% to 9%. Furthermore, the system maintained high response accuracy and improved overall reliability under adversarial conditions. These findings highlight the importance of integrating security-aware design principles into LLM deployment pipelines to ensure trustworthy, secure, and responsible AI systems.

Downloads

Download data is not yet available.

Downloads

Published

2026-03-11

How to Cite

Securing Large Language Models Against Prompt Injection and Data Leakage Attacks: https://doi.org/10.5281/zenodo.19344116. (2026). Annual Methodological Archive Research Review, 4(3), 112-129. https://amresearchjournal.com/index.php/Journal/article/view/1713

Similar Articles

121-130 of 1896

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)

1 2 3 > >>