Reinforcement Learning Enabled Design of Electronic-State Engineering Strategies for Sustainable Oxygen Evolution Catalysis

Authors

  • Kekang Duan Department of Computer Science and Engineering, University at Buffalo, Buffalo, NY, USA. Author
  • Jack Battler Department of Computer Science, University of Central Florida, Orlando, FL, USA. Author
  • Bastian Ceaster School of Computing, Clemson University, Clemson, SC, USA. Author
  • Erthur Wignir Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO, USA. Author

Keywords:

reinforcement learning, oxygen evolution reaction, electronic-state engineering, sustainable catalysis, autonomous materials discovery, systems architecture

Abstract

The oxygen evolution reaction (OER) constitutes a critical bottleneck in sustainable energy conversion systems, yet the design of high-performance OER catalysts remains constrained by the vast and combinatorially complex landscape of electronic states that govern catalytic activity. Traditional approaches to electronic-state engineering rely on trial-and-error experimentation or density functional theory calculations that are ill-suited to navigate high-dimensional design spaces under operational constraints. This paper presents a systems-level framework for leveraging reinforcement learning (RL) to autonomously discover and optimize electronic-state engineering strategies for transition metal oxide catalysts. We conceptualize the catalyst design problem as a sequential decision-making process in which an RL agent modulates synthesis parameters, compositional levers, and defect chemistries to tune frontier electronic states such as metal-oxygen covalency, charge-transfer excitations, and correlated electron configurations. The architecture integrates a multi-fidelity simulation environment, active experimental validation loops, and a distributed computing infrastructure that balances exploration-exploitation trade-offs across heterogeneous computational resources. We provide a detailed examination of structural trade-offs in representing electronic state spaces, formulating reward signals that reconcile catalytic activity with long-term stability, and governing the socio-technical deployment of autonomous discovery platforms. Robustness considerations under epistemic uncertainty from both computational approximations and experimental variabilities are analyzed through the lens of domain randomization and ensemble policy methods. Furthermore, we address fairness and policy implications by proposing open-access benchmarking ecosystems and distributed manufacturing models that ensure globally equitable access to sustainably designed catalytic materials. This work reframes catalyst discovery as a large-scale adaptive system, offering governance principles and infrastructure blueprints that extend beyond OER to broader materials design challenges in the transition to a carbon-neutral energy infrastructure.

References

1. Seh, Z. W., Kibsgaard, J., Dickens, C. F., Chorkendorff, I., Nørskov, J. K., & Jaramillo, T. F. (2017). Combining theory and experiment in electrocatalysis: Insights into materials design. Science, 355(6321), eaad4998.

2. Hwang, J., Rao, R. R., Giordano, L., Katayama, Y., Yu, Y., & Shao-Horn, Y. (2017). Perovskites in catalysis and electrocatalysis. Science, 358(6364), 751–756.

3. Suntivich, J., May, K. J., Gasteiger, H. A., Goodenough, J. B., & Shao-Horn, Y. (2011). A perovskite oxide optimized for oxygen evolution catalysis from molecular orbital principles. Science, 334(6061), 1383–1385.

4. Jain, A., Ong, S. P., Hautier, G., Chen, W., Richards, W. D., Dacek, S., Cholia, S., Gunter, D., Skinner, D., Ceder, G., & Persson, K. A. (2013). Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. APL Materials, 1(1), 011002.

5. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

6. Grimaud, A., Diaz-Morales, O., Han, B., Hong, W. T., Lee, Y. L., Giordano, L., Stoerzinger, K. A., Koper, M. T. M., & Shao-Horn, Y. (2017). Activating lattice oxygen redox reactions in metal oxides to catalyse oxygen evolution. Nature Chemistry, 9(5), 457–465.

7. Peng, C. K., Lin, Y. C., Chiang, C. L., Qian, Z., Huang, Y. C., Dong, C. L., ... & Lin, Y. G. (2023). Zhang-Rice singlets state formed by two-step oxidation for triggering water oxidation under operando conditions. Nature Communications, 14(1), 529.

8. Zhou, Z., Li, X., & Zare, R. N. (2017). Optimizing chemical reactions with deep reinforcement learning. ACS Central Science, 3(12), 1337–1344.

9. Lookman, T., Balachandran, P. V., Xue, D., & Yuan, R. (2019). Active learning in materials science with emphasis on adaptive sampling using uncertainties for targeted design. npj Computational Materials, 5(1), 21.

10. Burger, B., Maffettone, P. M., Gusev, V. V., Aitchison, C. M., Bai, Y., Wang, X., Li, X., Alston, B. M., Li, B., Clowes, R., Rankin, N., Harris, B., Sprick, R. S., & Cooper, A. I. (2020). A mobile robotic chemist. Nature, 583(7815), 237–241.

11. Kandasamy, K., Dasarathy, G., Schneider, J., & Póczos, B. (2017). Multi-fidelity Bayesian optimisation with continuous approximations. Proceedings of the 34th International Conference on Machine Learning, 1799–1808.

12. Xie, T., & Grossman, J. C. (2018). Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties. Physical Review Letters, 120(14), 145301.

13. Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2016). Continuous control with deep reinforcement learning. Proceedings of the 4th International Conference on Learning Representations.

14. Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., & Stoica, I. (2018). Ray: A distributed framework for emerging AI applications. Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation, 561–577.

15. Hausknecht, M., & Stone, P. (2015). Deep recurrent Q-learning for partially observable MDPs. Proceedings of the 2015 AAAI Fall Symposium Series.

16. Roijers, D. M., Vamplew, P., Whiteson, S., & Dazeley, R. (2013). A survey of multi-objective sequential decision-making. Journal of Artificial Intelligence Research, 48, 67–113.

17. Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. Proceedings of the 16th International Conference on Machine Learning, 278–287.

18. Xu, M., Hu, J., & Panda, D. K. (2021). Accelerating reinforcement learning on high-performance computing systems. Proceedings of the 2021 IEEE International Parallel and Distributed Processing Symposium, 677–686.

19. Chiang, H. T. L., Hsu, J., Fiser, M., Tapia, L., & Faust, A. (2019). RL-RRT: Kinodynamic motion planning via learning reachability estimators from RL policies. IEEE Robotics and Automation Letters, 4(4), 4299–4306.

20. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J. W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., ... & Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1), 160018.

21. Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., & Abbeel, P. (2017). Domain randomization for transferring deep neural networks from simulation to the real world. Proceedings of the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, 23–30.

22. Stall, S., Yarmey, L., Cutcher-Gershenfeld, J., Hanson, B., Lehnert, K., Nosek, B., Parsons, M., Robinson, E., & Wyborn, L. (2019). Make scientific data FAIR. Nature, 570(7759), 27–29.

23. Zimmermann, J. B., Anastas, P. T., Erythropel, H. C., & Leitner, W. (2020). Designing for a green chemistry future. Science, 367(6476), 397–400.

24. National Academies of Sciences, Engineering, and Medicine. (2022). Automated research workflows for accelerated discovery: Closing the knowledge loop. The National Academies Press.

Downloads

Published

2026-07-03

How to Cite

Reinforcement Learning Enabled Design of Electronic-State Engineering Strategies for Sustainable Oxygen Evolution Catalysis. (2026). Journal of Advanced Artificial Intelligence Research, 1(1). https://www.jaair.org/index.php/home/article/view/152