Causality-Aware Federated Reinforcement Learning for Privacy-Preserving Budget Allocation in Social Commerce Advertising Ecosystems

Authors

  • Wesley Perez Department of Computer Science and Engineering, University of Nevada, Reno, Reno, NV, USA. Author
  • Aditya A. Malhotra School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, OR, USA. Author
  • Clifford Benson School of Information Technology, University of Cincinnati, Cincinnati, OH, USA. Author

Keywords:

federated reinforcement learning, causal inference, budget allocation, social commerce, differential privacy, advertising ecosystems, privacy-preserving machine learning, fairness, governance, cross-device attribution

Abstract

The rapid expansion of social commerce advertising ecosystems has created a critical need for budget allocation mechanisms that are both operationally efficient and privacy compliant. Traditional centralized approaches to budget optimization rely on aggregating granular user behavior data, raising significant privacy concerns and violating emerging regulatory frameworks. Meanwhile, conventional federated reinforcement learning methods, while addressing data decentralization, often fail to account for the structural causal relationships that drive advertising effectiveness in complex socio-technical systems. This paper proposes a causality-aware federated reinforcement learning framework designed specifically for privacy-preserving budget allocation across cross-device social commerce platforms. The architecture integrates three foundational components: a federated infrastructure that ensures raw user data never leaves local devices, a causal inference module that models the counterfactual impact of ad exposures on downstream conversions, and a multi-agent reinforcement learning layer that dynamically allocates budgets across creators, channels, and time periods under differential privacy constraints. We examine the structural trade-offs between statistical utility, privacy guarantees, and fairness across heterogeneous advertiser segments. The governance implications of such systems are analyzed, including issues of algorithmic transparency, accountability for budget decisions, and the prevention of discriminatory outcomes. Deployment considerations are discussed with respect to computational sustainability, communication overhead, and robustness to adversarial threats. We further situate the proposed framework within the broader policy landscape, including the implications of data minimization principles and zero-trust architectures. By synthesizing insights from causal inference, federated learning, and reinforcement learning, this work provides a comprehensive blueprint for next-generation advertising infrastructure that aligns economic incentives with ethical and legal mandates. The paper concludes with a forward-looking discussion on the integration of large language model-assisted policy generation and the potential for cultural bias in multi-regional deployments.

References

1. Dwork, C., & Roth, A. (2014). The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9(3-4), 211-407.

2. Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. Advances in Neural Information Processing Systems, 29.

3. Pearl, J. (2009). Causality: Models, reasoning, and inference (2nd ed.). Cambridge University Press.

4. McMahan, B., Moore, E., Ramage, D., Hampson, S., & y Arcas, B. A. (2017). Communication-efficient learning of deep networks from decentralized data. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, 1273-1282.

5. Shalev-Shwartz, S., & Ben-David, S. (2014). Understanding machine learning: From theory to algorithms. Cambridge University Press.

6. Zhou, D. (2026). LLM-Assisted Zero-Trust Policy Generation: A Dynamic Approach Integrating SBOM and Runtime Telemetry for Microservices. American Journal Of Big Data, 7(1), 212-228.

7. Bonawitz, K., Ivanov, V., Kreuter, B., Marcedone, A., McMahan, H. B., Patel, S., ... & Seth, K. (2017). Practical secure aggregation for privacy-preserving machine learning. Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 1175-1191.

8. Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., & Zhang, L. (2016). Deep learning with differential privacy. Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 308-318.

9. Huang, Z., Wang, J., Zhou, S., & Ren, K. (2022). Federated causal discovery. Proceedings of the AAAI Conference on Artificial Intelligence, 36(5), 5102-5110.

10. Dwork, C., McSherry, F., Nissim, K., & Smith, A. (2006). Calibrating noise to sensitivity in private data analysis. Theory of Cryptography Conference, 265-284.

11. Juels, A., & Ristenpart, T. (2014). Honey encryption: Security beyond the brute-force bound. Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, 399-410.

12. Ben-Sasson, E., Chiesa, A., Garman, C., Green, M., Miers, I., Tromer, E., & Virza, M. (2014). Zerocash: Decentralized anonymous payments from Bitcoin. Proceedings of the 2014 IEEE Symposium on Security and Privacy, 459-474.

13. Dai, R., Li, Y., & Zhang, K. (2023). Differentially private constraint-based causal discovery. Journal of Machine Learning Research, 24(189), 1-42.

14. Angrist, J. D., Imbens, G. W., & Rubin, D. B. (1996). Identification of causal effects using instrumental variables. Journal of the American Statistical Association, 91(434), 444-455.

15. Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., ... & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. International Conference on Machine Learning, 1928-1937.

16. Qi, J., Zhou, Q., & Kong, L. (2021). Federated reinforcement learning: A survey. ACM Computing Surveys, 54(9), 1-35.

17. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

18. Feldman, M., Friedler, S. A., Moeller, J., Scheidegger, C., & Venkatasubramanian, S. (2015). Certifying and removing disparate impact. Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 259-268.

19. Suresh, A. T., Yu, F. X., Kumar, S., & McMahan, H. B. (2017). Distributed mean estimation with limited communication. International Conference on Machine Learning, 3329-3337.

20. Nielsen, M. A., & Chuang, I. L. (2010). Quantum computation and quantum information: 10th anniversary edition. Cambridge University Press.

21. Blanchard, P., Guerraoui, R., & Stainer, J. (2017). Machine learning with adversaries: Byzantine tolerant gradient descent. Advances in Neural Information Processing Systems, 30.

22. Shi, C., Li, S., Guo, S., Xie, S., Wu, W., Dou, J., ... & Chua, T. S. (2025). Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation. arXiv preprint arXiv:2511.17282.

Downloads

Published

2026-06-12

How to Cite

Causality-Aware Federated Reinforcement Learning for Privacy-Preserving Budget Allocation in Social Commerce Advertising Ecosystems. (2026). Journal of Advanced Artificial Intelligence Research, 5(1). https://www.jaair.org/index.php/home/article/view/109