Path-Aware Defense Against Prompt Injection Attacks in Retrieval-Augmented Large Language Models

Authors

  • Damien L. Perez School of Information Technology, University of Cincinnati, Cincinnati, OH, USA. Author
  • Iecac Rhodes Department of Computer Science, University of Central Florida, Orlando, FL, USA. Author

Keywords:

retrieval-augmented generation, prompt injection, adversarial robustness, path-aware defense, large language models, information retrieval, system security

Abstract

Retrieval-augmented generation (RAG) architectures have rapidly become a dominant paradigm for grounding large language model outputs in external knowledge, yet they introduce novel attack surfaces through prompt injection that threaten the integrity, confidentiality, and reliability of downstream applications. This paper presents a comprehensive system-level analysis of path-aware defense mechanisms designed to counter adversarial prompt injections in RAG pipelines. Departing from conventional input-filtering approaches, path-aware strategies operate on the provenance and trajectory of information as it moves from retrieval sources through fusion and generation stages, enabling fine-grained attribution, anomaly detection, and targeted intervention. The discussion examines structural trade-offs between defense granularity, latency, scalability, and model performance, and it locates path-aware methods within broader architectural patterns for robust AI infrastructure. The paper further addresses governance frameworks, fairness implications, operational resilience, and long-term sustainability, highlighting how path-level visibility transforms incident response, auditability, and regulatory compliance. By synthesizing insights from distributed systems, adversarial machine learning, and information retrieval, this work articulates a research roadmap for building resilient RAG ecosystems in which safety is not an afterthought but a property of the retrieval-generation pathway itself.

References

1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459-9474.

2. Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M. W. (2020). Retrieval augmented language model pre-training. In International Conference on Machine Learning (pp. 3929-3938). PMLR.

3. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610-623).

4. Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., ... & Oprea, A. (2021). Extracting training data from large language models. In 30th USENIX Security Symposium (pp. 2633-2650).

5. Wallace, E., Feng, S., Kandpal, N., Gardner, M., & Singh, S. (2019). Universal adversarial triggers for attacking and analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (pp. 2153-2162).

6. Perez, E., & Ribeiro, M. T. (2022). Ignore previous prompt: Attack techniques for language models. In NeurIPS ML Safety Workshop.

7. Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (pp. 79-90).

8. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., ... & Fedus, W. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.

9. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

10. Dinan, E., Humeau, S., Chintalapudi, B., & Weston, J. (2019). Build it break it fix it for dialogue safety: Robustness from adversarial human attack. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (pp. 4537-4546).

11. Papernot, N., McDaniel, P., Sinha, A., & Wellman, M. P. (2018). SoK: Security and privacy in machine learning. In 2018 IEEE European Symposium on Security and Privacy (EuroS&P) (pp. 399-414).

12. Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., & Shoham, Y. (2023). In-context retrieval-augmented language models. Transactions of the Association for Computational Linguistics, 11, 1316-1331.

13. C. Shi, S. Li, W. Lu, W. Wu, C. Wang, Z. Cheng, F. Shen, and T. Chua (2026)TraceRouter: robust safety for large foundation models via path-level intervention.arXiv preprint arXiv:2601.21900.

14. Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., ... & Sifre, L. (2022). Improving language models by retrieving from trillions of tokens. In International Conference on Machine Learning (pp. 2206-2240). PMLR.

15. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.

16. Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., ... & Yih, W. T. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 6769-6781).

17. Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H. T., ... & Le, Q. (2022). LaMDA: Language models for dialog applications. arXiv preprint arXiv:2201.08239.

18. Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., ... & Irving, G. (2019). Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593.

19. OpenAI (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.

20. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT (pp. 4171-4186).

21. Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.

Downloads

Published

2026-06-14

How to Cite

Path-Aware Defense Against Prompt Injection Attacks in Retrieval-Augmented Large Language Models. (2026). Journal of Advanced Artificial Intelligence Research, 1(1). https://www.jaair.org/index.php/home/article/view/132