Explainable Federated Rule Mining for Privacy-Preserving Healthcare Risk Prediction under Heterogeneous Data Quality Conditions
Keywords:
Federated Learning, Rule Mining, Explainable AI, Healthcare Risk Prediction, Data Quality Heterogeneity, Privacy PreservationAbstract
The proliferation of electronic health records across fragmented healthcare systems has created unprecedented opportunities for data-driven risk prediction, yet institutional privacy regulations, data silos, and severe heterogeneity in data quality obstruct centralized analytical approaches. This paper presents a comprehensive system-level investigation of an explainable federated rule mining framework designed to operate under such conditions. We examine the architectural trade-offs necessary to integrate federated learning with interpretable rule discovery, ensuring that clinical decision support systems can ingest horizontally and vertically partitioned patient data without raw data exposure. Central to our analysis is the tension between global model expressiveness and local data characteristics shaped by varying completeness, timeliness, and measurement reliability. We discuss infrastructure requirements for orchestrating multi-party rule mining, including communication protocols, asynchronous aggregation strategies, and threshold-based rule pruning that account for site-specific quality metrics. The governance layer encompassing dynamic consent, differential privacy budgets, and auditable provenance trails is dissected in terms of legal compliance with regulations such as HIPAA and GDPR. Furthermore, we analyze how explainability is not merely a post hoc feature but an architectural requirement that shapes the rule representation, refinement cycles, and clinician-facing interfaces. A detailed examination of fairness and robustness reveals that heterogeneous data quality can amplify biases, requiring rule quality weighting mechanisms and community-based validation protocols. Deployment pathways, sustainability considerations within healthcare IT ecosystems, and cross-institutional policy frameworks are also addressed. The work contributes a holistic perspective that bridges technical federated rule mining techniques with operational, ethical, and regulatory realities, providing a blueprint for trustworthy and resilient predictive health systems.
References
1. McMahan, B., Moore, E., Ramage, D., Hampson, S., & y Arcas, B. A. (2017). Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (pp. 1273–1282). PMLR.
2. Rieke, N., Hancox, J., Li, W., Milletari, F., Roth, H. R., Albarqouni, S., ... & Cardoso, M. J. (2020). The future of digital health with federated learning. NPJ Digital Medicine, 3(1), 1–7.
3. Li, T., Sahu, A. K., Talwalkar, A., & Smith, V. (2020). Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, 37(3), 50–60.
4. Agrawal, R., & Srikant, R. (1994). Fast algorithms for mining association rules. In Proceedings of the 20th International Conference on Very Large Data Bases (pp. 487–499). Morgan Kaufmann.
5. Vellido, A. (2019). The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural Computing and Applications, 32, 18069–18083.
6. Weiskopf, N. G., & Weng, C. (2013). Methods and dimensions of electronic health record data quality assessment: Enabling reuse for clinical research. Journal of the American Medical Informatics Association, 20(1), 144–151.
7. Kahn, M. G., Callahan, T. J., Barnard, J., Bauck, A. E., Brown, J., Davidson, B. N., ... & Liaw, S. T. (2016). A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data. eGEMs, 4(1), 1244.
8. Wang, H., & Wu, Z. (2021). Heterogeneous data-aware federated learning. arXiv preprint arXiv:2105.00667.
9. Hashimoto, T., Srivastava, M., Namkoong, H., & Liang, P. (2018). Fairness without demographics in repeated loss minimization. In Proceedings of the 35th International Conference on Machine Learning (pp. 1929–1938). PMLR.
10. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., ... & Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210.
11. Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (pp. 4765–4774).
12. Dwork, C., & Roth, A. (2014). The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9(3–4), 211–407.
13. Hripcsak, G. (2019). Arden Syntax for medical logic modules. Journal of Biomedical Informatics, 93, 103161.
14. Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., & Zhang, L. (2016). Deep learning with differential privacy. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (pp. 308–318).
15. Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., ... & Papernot, N. (2021). Machine unlearning. In Proceedings of the IEEE Symposium on Security and Privacy (pp. 141–159).
16. U.S. Food and Drug Administration. (2019). Proposed regulatory framework for modifications to artificial intelligence/machine learning-based software as a medical device. FDA Discussion Paper.
17. Han, Z., Chen, W., & Han, Y. (2024, August). Data Quality-Driven Top-k Rule Discovery. In International Conference on Machine Learning, Cloud Computing and Intelligent Mining (pp. 24-38). Singapore: Springer Nature Singapore.
18. Ghassemi, M., Naumann, T., Schulam, P., Beam, A. L., Chen, I. Y., & Ranganath, R. (2020). A review of challenges and opportunities in machine learning for health. AMIA Summits on Translational Science Proceedings, 2020, 191–200.
19. Xiao, C., Choi, E., & Sun, J. (2018). Opportunities and challenges in developing deep learning models using electronic health records data: A systematic review. Journal of the American Medical Informatics Association, 25(10), 1419–1428.
20. Shiloach, M., Frencher, S. K., Steeger, J. E., Rowell, K. S., Bartzokis, K., Tomeh, M. G., ... & Hall, B. L. (2010). Toward robust information: Data quality and inter-rater reliability in the American College of Surgeons National Surgical Quality Improvement Program. Journal of the American College of Surgeons, 210(1), 6–16.
21. Johnson, A. E. W., Pollard, T. J., Shen, L., Lehman, L. H., Feng, M., Ghassemi, M., ... & Mark, R. G. (2016). MIMIC-III, a freely accessible critical care database. Scientific Data, 3, 160035.
22. Yang, Q., Liu, Y., Chen, T., & Tong, Y. (2019). Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology, 10(2), 1–19.
23. Chen, I. Y., Pierson, E., Rose, S., Joshi, S., Ferryman, K., & Ghassemi, M. (2021). Ethical machine learning in healthcare. Annual Review of Biomedical Data Science, 4, 123–144.
24. Cohen, I. G., Amarasingham, R., Shah, A., Xie, B., & Lo, B. (2014). The legal and ethical concerns that arise from using complex predictive analytics in health care. Health Affairs, 33(7), 1139–1147.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Advanced Artificial Intelligence Research

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.