Retrieval-Augmented Generation for Scientific Knowledge Discovery in Interdisciplinary Research Networks
Keywords:
Retrieval-Augmented Generation, scientific knowledge discovery, interdisciplinary research, large language models, knowledge graphs, system architecture, information retrieval, fairness, sustainability, socio-technical systemsAbstract
The accelerating pace of scientific discovery increasingly depends on the ability to integrate knowledge across disparate disciplines, yet traditional literature search and synthesis methods remain fragmented and labor-intensive. Retrieval-Augmented Generation (RAG) offers a transformative paradigm for scientific knowledge discovery by combining large language models with external knowledge bases, enabling context-aware, evidence-grounded responses that can bridge disciplinary boundaries. This paper presents a comprehensive examination of RAG systems within the context of interdisciplinary research networks, focusing on architectural design, structural trade-offs, governance challenges, and deployment infrastructure. We analyze the interplay between retrieval fidelity and generative coherence, the role of knowledge graph integration, and the socio-technical implications of scaling such systems across heterogeneous domains. Particular attention is given to issues of fairness, robustness, and sustainability, as well as the policy frameworks required to support responsible adoption. Through case illustrations from fields such as biomedicine, climate science, and materials engineering, we highlight both the potential and the current limitations of RAG for facilitating cross-domain synthesis. The paper concludes with forward-looking recommendations for system governance, benchmark development, and interdisciplinary collaboration that can guide future research and deployment of RAG-driven knowledge discovery platforms.
References
1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.
2. Chen, D., Fisch, A., Weston, J., & Bordes, A. (2017). Reading Wikipedia to answer open-domain questions. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 1870–1879.
3. Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., ... & Yih, W. (2020). Dense passage retrieval for open-domain question answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 6769–6781.
4. Guu, K., Lee, K., Tung, Z., Pasupat, P., & Liang, P. (2020). REALM: Retrieval augmented language model pre-training. Proceedings of the 37th International Conference on Machine Learning, 3929–3938.
5. Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
6. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., ... & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1–67.
7. Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2023). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. arXiv preprint arXiv:2310.11511.
8. Beltagy, I., Lo, K., & Cohan, A. (2019). SciBERT: A pretrained language model for scientific text. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 3615–3620.
9. Cohan, A., Feldman, S., Beltagy, I., Downey, D., & Weld, D. S. (2020). SPECTER: Document-level representation learning using citation-informed transformers. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2270–2282.
10. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 3982–3992.
11. Carbonell, J., & Goldstein, J. (1998). The use of MMR, diversity-based reranking for reordering documents and producing summaries. Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 335–336.
12. Johnson, J., Douze, M., & Jégou, H. (2019). Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547.
13. Thakur, N., Reimers, N., Rücklé, A., Srivatsa, A., & Gurevych, I. (2021). BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. arXiv preprint arXiv:2104.08663.
14. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
15. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623.
16. Wallace, E., Feng, S., Kandpal, N., Gardner, M., & Singh, S. (2019). Universal adversarial triggers for attacking and analyzing NLP. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2153–2162.
17. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645–3650.
18. Tshitoyan, V., Dagdelen, J., Weston, L., Dunn, A., Rong, Z., Kononova, O., ... & Jain, A. (2019). Unsupervised word embeddings capture latent knowledge from materials science literature. Nature, 571(7763), 95–98.
19. Jain, A., Ong, S. P., Hautier, G., Chen, W., Richards, W. D., Dacek, S., ... & Persson, K. A. (2013). Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. APL Materials, 1(1), 011002.
20. Wang, Y., Kordi, Y., Mishra, S., Sharma, A., & Mahajan, D. (2023). Interactive retrieval-augmented generation for scientific literature: A user study. arXiv preprint arXiv:2310.12345.
21. Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1(1), 1–13.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 Journal of Advanced Artificial Intelligence Research

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.