Personalized Drug Discovery Knowledge Retrieval Using Large Language Models and Biomedical Context Representation

Authors

  • Finn Carpenter Department of Computer Science, University of Houston, Houston, TX, USA. Author
  • Jingwenling Song Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO, USA. Author
  • Ole Bowman Department of Computer Science, University of Central Florida, Orlando, FL, USA. Author
  • Suraj Meaik Department of Computer Science, University of New Hampshire, Durham, NH, USA. Author

Keywords:

personalized retrieval; large language models; drug discovery; biomedical knowledge graphs; context representation; retrieval-augmented generation; fairness

Abstract

The accelerating growth of biomedical literature and multi-modal data has created an urgent need for intelligent retrieval systems that can navigate complex pharmacological spaces while adapting to the unique research contexts of individual scientists. This paper presents a system-level examination of personalized drug discovery knowledge retrieval that leverages large language models and rich biomedical context representations. Unlike generic search engines, the envisioned architecture fuses unstructured scientific text with structured knowledge graphs, molecular embeddings, and ontological frameworks to ground large language model reasoning in verifiable biomedical evidence. We analyze the structural trade-offs inherent in coupling retrieval-augmented generation pipelines with user profiling, focusing on how semantic context beyond surface-level keywords can be constructed from entity-linked biomedical concepts and relation-aware transformer encodings. The discussion extends to deployment-scale considerations, including latency-aware model serving, sustainability constraints of large-scale inference, and the maintenance of dynamic knowledge indices in the face of rapidly evolving biological databases. Furthermore, the paper addresses governance challenges related to algorithmic fairness in retrieved drug candidates, potential biases against underrepresented disease populations, and the explainability requirements for high-stakes pharmaceutical decision support. Policy implications are examined through the lens of regulatory oversight, intellectual property around AI-suggested compounds, and the need for auditable, transparent retrieval trajectories. By synthesizing perspectives from information retrieval, biomedical informatics, and socio-technical systems research, we offer a forward-looking framework that balances personalization depth with ethical and infrastructural robustness, charting a path toward responsible, context-aware knowledge discovery in drug development.

References

1. Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240.

2. Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., ... & Poon, H. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), 1–23.

3. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474).

4. Beltagy, I., Lo, K., & Cohan, A. (2019). SciBERT: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (pp. 3615–3620).

5. Zitnik, M., Agrawal, M., & Leskovec, J. (2018). Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics, 34(13), i457–i466.

6. Wishart, D. S., Feunang, Y. D., Guo, A. C., Lo, E. J., Marcu, A., Grant, J. R., ... & Wilson, M. (2018). DrugBank 5.0: a major update to the DrugBank database for 2018. Nucleic Acids Research, 46(D1), D1074–D1082.

7. Himmelstein, D. S., Lizee, A., Hessler, C., Brueggeman, L., Chen, S. L., Hadley, D., ... & Baranzini, S. E. (2017). Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife, 6, e26726.

8. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (pp. 4171–4186).

9. Teevan, J., Dumais, S. T., & Horvitz, E. (2005). Personalizing search via automated analysis of interests and activities. In Proceedings of the 28th Annual International ACM SIGIR Conference (pp. 449–456).

10. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35.

11. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623).

12. Bodenreider, O. (2004). The Unified Medical Language System (UMLS): integrating biomedical terminology. Nucleic Acids Research, 32(suppl_1), D267–D270.

13. Chen, D., Fisch, A., Weston, J., & Bordes, A. (2017). Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (pp. 1870–1879).

14. Yu, X. (2026, January). AI-Driven Personalization across Domains for Local Categorical Query Understanding and Context-Aware Retrieval. In Proceedings of the 2nd International Conference on Artificial Intelligence, Digital Media Technology and Social Computing (pp. 77-82).

15. Johnson, J., Douze, M., & Jégou, H. (2019). Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547.

16. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (pp. 3982–3992).

17. Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., ... & Yih, W.-t. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 6769–6781).

18. Alsentzer, E., Murphy, J. R., Boag, W., Weng, W.-H., Jin, D., Naumann, T., & McDermott, M. B. A. (2019). Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop (pp. 72–78).

19. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.

20. Tonekaboni, S., Joshi, S., McCradden, M. D., & Goldenberg, A. (2019). What clinicians want: Contextualizing explainable machine learning for clinical end use. In Proceedings of the 4th Machine Learning for Healthcare Conference (pp. 359–380).

Downloads

Published

2026-04-14

How to Cite

Personalized Drug Discovery Knowledge Retrieval Using Large Language Models and Biomedical Context Representation. (2026). Journal of Advanced Artificial Intelligence Research, 5(1). https://www.jaair.org/index.php/home/article/view/180