A Multi-Agent Large Language Model Framework for Automated Scientific Literature Review and Knowledge Synthesis
Abstract
The growth of scientific publications has made comprehensive literature review increasingly difficult for individual researchers and conventional search systems. This paper presents ScHoL- ARSyNTH, a multi-agent large language model (LLM) framework for automated scientific lit- erature review and knowledge synthesis. The framework decomposes review generation into coordinated specialist agents for query planning, paper retrieval, evidence extraction, claim verification, thematic clustering, and synthesis writing. Unlike single-prompt LLM summa- rization, the proposed design maintains explicit evidence provenance, performs cross-paper consistency checks, and separates exploratory retrieval from synthesis decisions. We evaluate ScHoLARSyNTH on a curated benchmark of 4,200 papers from machine learning, biomedical informatics, and human–computer interaction, with 120 review tasks derived from survey arti- cle abstracts and section headings. Experimental results show that ScHoLARSyNTH improves answer faithfulness from 0.76 to 0.89 and topical coverage from 0.71 to 0.84 compared with a strong retrieval-augmented generation baseline, while reducing unsupported claims by 43.8%. The system processes 100 candidate papers in 7.8 minutes on average, enabling interactive lit- erature scoping. Ablation studies demonstrate that verifier and clustering agents contribute most strongly to factual precision and structural coherence. The findings indicate that multi- agent orchestration is a practical approach for transforming LLMs from generic summarizers into auditable research assistants for scientific knowledge synthesis.
References
[1] T. Brown et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, 2020.
[2] Gao, H., Zeng, W., Zhang, J., & Liang, Y. (2025, December). A large model API response quality prediction model based on least squares vector machine and SHAP interpretability analysis. In 2025 5th International Symposium on Artificial Intelligence and Big Data (AIBDF) (pp. 438-442). IEEE.
[3] H. Touvron et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023.
[4] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Ad- vances in Neural Information Processing Systems, 2020.
[5]Shih, K., Deng, Z., Chen, X., Zhang, Y., & Zhang, L. (2025, May). DST-GFN: A Dual-Stage Transformer Network with Gated Fusion for Pairwise User Preference Prediction in Dialogue Systems. In 2025 8th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE) (pp. 715-719). IEEE.
[6] Dou, Z., Cui, D., Yan, J., Wang, W., Chen, B., Wang, H., ... & Zhang, S. (2025). Dsadf: Thinking fast and slow for decision making. arXiv preprint arXiv:2505.08189.
[7] A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, 2017.
[8] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019.
[9] I. Beltagy, K. Lo, and A. Cohan, “SciBERT: A pretrained language model for scientific text,” in Proc. EMNLP-IJCNLP, 2019.
[10] Zhou, D. (2025, December). M-VP2: Microservice-Oriented Vulnerability Patch Planning-A Cost-Aware Approachusing Multi-Agent Reinforcement Learning. In 2025 5th International Conference on Computer, Internet of Things and Control Engineering (CITCE) (pp. 248-254). IEEE.
[11] N. Shinn et al., “Reflexion: Language agents with verbal reinforcement learning,” arXiv preprint arXiv:2303.11366, 2023.
[12] Q. Wu et al., “AutoGen: Enabling next-gen LLM applications via multi-agent conversation,” arXiv preprint arXiv:2308.08155, 2023.
[13] J. S. Park et al., “Generative agents: Interactive simulacra of human behavior,” in Proc. UIST, 2023.
[14]Wang, S., Feng, Y., & Fang, X. (2026). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving.
[15] V. Karpukhin et al., “Dense passage retrieval for open-domain question answering,” in Proc. EMNLP, 2020.
[16] G. Izacard and E. Grave, “Leveraging passage retrieval with generative models for open domain question answering,” in Proc. EACL, 2021.
[17] Fu, L., Chen, X., Gao, K., Huang, X., & Tong, K. (2025, October). Memory-Augmented Knowledge Fusion with Safety-Aware Decoding for Domain-Adaptive Question Answering. In 2025 6th International Conference on Machine Learning and Computer Application (ICMLCA) (pp. 1-6). IEEE.
[18] Shui, Y., Jin, R., Dou, Z., & Gao, Z. (2026). ProtoGuard-SL: Prototype Consistency Based Backdoor Defense for Vertical Split Learning. arXiv preprint arXiv:2604.03595.
[19] G. Li et al., “CAMEL: Communicative agents for mind exploration of large language model society,” arXiv preprint arXiv:2303.17760, 2023.
[20] S. Hong et al., “MetaGPT: Meta programming for a multi-agent collaborative framework,” in Proc. ICLR, 2024.
[21] C. Qian et al., “Communicative agents for software development,” arXiv preprint arXiv:2307.07924, 2023.
[22] G. Tsafnat et al., “Systematic review automation technologies,” Systematic Reviews, vol. 3, no. 1, 2014.
[23] B. C. Wallace, T. A. Trikalinos, J. Lau, C. Brodley, and C. H. Schmid, “Semi-automated screening of biomedical citations for systematic reviews,” BMC Bioinformatics, vol. 11, no. 1, 2010.
[24] A. Cohan et al., “A discourse-aware attention model for abstractive summarization of long documents,” in Proc. NAACL-HLT, 2018.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Advanced Artificial Intelligence Research

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.