A Multi-Agent Large Language Model Framework for Automated Scientific Literature Review and Knowledge Synthesis

Authors

  • Daniel R. Collins Department of Computer Science University of Massachusetts Lowell Lowell, MA, USA Author
  • Priya N. Raman School of Computing and Information University of Pittsburgh Pittsburgh, PA, USA Author

Abstract

The growth of scientific publications has made comprehensive literature review increasingly difficult for individual researchers and conventional search systems.  This paper presents ScHoL- ARSyNTH, a multi-agent large language model  (LLM) framework for automated scientific lit- erature review  and knowledge  synthesis.   The  framework  decomposes review generation  into coordinated  specialist  agents  for  query  planning,  paper  retrieval,  evidence  extraction,  claim verification,  thematic  clustering,  and  synthesis  writing.   Unlike  single-prompt  LLM  summa- rization,  the  proposed  design  maintains  explicit  evidence  provenance,  performs  cross-paper consistency checks, and separates exploratory retrieval from synthesis decisions.  We evaluate ScHoLARSyNTH on a curated benchmark of 4,200 papers from machine learning, biomedical informatics, and human–computer interaction, with 120 review tasks derived from survey arti- cle abstracts and section headings.  Experimental results show that ScHoLARSyNTH improves answer faithfulness from 0.76 to 0.89 and topical coverage from 0.71 to 0.84 compared with a strong retrieval-augmented generation baseline, while reducing unsupported claims by 43.8%. The system processes 100 candidate papers in 7.8 minutes on average, enabling interactive lit- erature scoping.   Ablation studies  demonstrate that verifier  and clustering  agents contribute most strongly to factual precision and structural coherence.  The findings indicate that multi- agent orchestration is a practical approach for transforming LLMs from generic summarizers into auditable research assistants for scientific knowledge synthesis.

References

[1] T. Brown et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, 2020.

[2] Gao, H., Zeng, W., Zhang, J., & Liang, Y. (2025, December). A large model API response quality prediction model based on least squares vector machine and SHAP interpretability analysis. In 2025 5th International Symposium on Artificial Intelligence and Big Data (AIBDF) (pp. 438-442). IEEE.

[3] H. Touvron et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023.

[4] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Ad- vances in Neural Information Processing Systems, 2020.

[5]Shih, K., Deng, Z., Chen, X., Zhang, Y., & Zhang, L. (2025, May). DST-GFN: A Dual-Stage Transformer Network with Gated Fusion for Pairwise User Preference Prediction in Dialogue Systems. In 2025 8th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE) (pp. 715-719). IEEE.

[6] Dou, Z., Cui, D., Yan, J., Wang, W., Chen, B., Wang, H., ... & Zhang, S. (2025). Dsadf: Thinking fast and slow for decision making. arXiv preprint arXiv:2505.08189.

[7] A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, 2017.

[8] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019.

[9] I. Beltagy, K. Lo, and A. Cohan, “SciBERT: A pretrained language model for scientific text,” in Proc. EMNLP-IJCNLP, 2019.

[10] Zhou, D. (2025, December). M-VP2: Microservice-Oriented Vulnerability Patch Planning-A Cost-Aware Approachusing Multi-Agent Reinforcement Learning. In 2025 5th International Conference on Computer, Internet of Things and Control Engineering (CITCE) (pp. 248-254). IEEE.

[11] N. Shinn et al., “Reflexion: Language agents with verbal reinforcement learning,” arXiv preprint arXiv:2303.11366, 2023.

[12] Q. Wu et al., “AutoGen: Enabling next-gen LLM applications via multi-agent conversation,” arXiv preprint arXiv:2308.08155, 2023.

[13] J. S. Park et al., “Generative agents: Interactive simulacra of human behavior,” in Proc. UIST, 2023.

[14]Wang, S., Feng, Y., & Fang, X. (2026). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving.

[15] V. Karpukhin et al., “Dense passage retrieval for open-domain question answering,” in Proc. EMNLP, 2020.

[16] G. Izacard and E. Grave, “Leveraging passage retrieval with generative models for open domain question answering,” in Proc. EACL, 2021.

[17] Fu, L., Chen, X., Gao, K., Huang, X., & Tong, K. (2025, October). Memory-Augmented Knowledge Fusion with Safety-Aware Decoding for Domain-Adaptive Question Answering. In 2025 6th International Conference on Machine Learning and Computer Application (ICMLCA) (pp. 1-6). IEEE.

[18] Shui, Y., Jin, R., Dou, Z., & Gao, Z. (2026). ProtoGuard-SL: Prototype Consistency Based Backdoor Defense for Vertical Split Learning. arXiv preprint arXiv:2604.03595.

[19] G. Li et al., “CAMEL: Communicative agents for mind exploration of large language model society,” arXiv preprint arXiv:2303.17760, 2023.

[20] S. Hong et al., “MetaGPT: Meta programming for a multi-agent collaborative framework,” in Proc. ICLR, 2024.

[21] C. Qian et al., “Communicative agents for software development,” arXiv preprint arXiv:2307.07924, 2023.

[22] G. Tsafnat et al., “Systematic review automation technologies,” Systematic Reviews, vol. 3, no. 1, 2014.

[23] B. C. Wallace, T. A. Trikalinos, J. Lau, C. Brodley, and C. H. Schmid, “Semi-automated screening of biomedical citations for systematic reviews,” BMC Bioinformatics, vol. 11, no. 1, 2010.

[24] A. Cohan et al., “A discourse-aware attention model for abstractive summarization of long documents,” in Proc. NAACL-HLT, 2018.

Downloads

Published

2026-06-16 — Updated on 2026-06-16

How to Cite

A Multi-Agent Large Language Model Framework for Automated Scientific Literature Review and Knowledge Synthesis. (2026). Journal of Advanced Artificial Intelligence Research, 5(1). https://www.jaair.org/index.php/home/article/view/103