Hierarchical Reinforcement Learning Framework for End-to-End Network Slice Lifecycle Management
Keywords:
network slicing, lifecycle management, hierarchical reinforcement learning, 5G and beyond, resource orchestration, quality of service, policy gradient, closed-loop automationAbstract
Network slicing has emerged as a cornerstone of 5G and beyond mobile systems, enabling the provisioning of customized logical networks over a shared physical infrastructure. The end-to-end lifecycle management of network slices—spanning preparation, commissioning, operation, and decommissioning—poses a complex sequential decision-making problem under uncertainty. Traditional orchestration approaches rely on heuristics and reactive control loops that struggle to meet dynamic quality-of-service requirements and resource efficiency goals across multiple administrative domains. This paper presents a hierarchical reinforcement learning framework for autonomous end-to-end slice lifecycle management. The proposed architecture decomposes the management plane into multiple temporal and functional abstraction levels, each handled by a reinforcement learning module operating on distinct timescales. A high-level meta-controller selects long-term slice objectives and policy options, intermediate controllers manage lifecycle state transitions and cross-domain coordination, while low-level controllers execute fine-grained resource allocation and fault mitigation actions. The framework leverages the options formalism and policy gradient methods to enable sample-efficient learning and scalable decision-making. We discuss architectural trade-offs, infrastructure implications, and system-level design choices that govern robustness, sustainability, and multi-tenant fairness. The paper further examines policy and governance dimensions, including trustworthiness, auditability, and regulatory compliance, arguing that hierarchical structure naturally accommodates human-intelligible control points and policy constraints. By integrating lifecycle-wide optimization with hierarchically abstracted control, the proposed framework addresses key shortcomings of flat reinforcement learning approaches and charts a path toward fully automated, zero-touch slice management in future communication networks.
References
1. Rost, P., Mannweiler, C., Michalopoulos, D. S., Sartori, C., Sciancalepore, V., Sastry, N., Holland, O., Tayade, S., Han, B., Bega, D., & Banchs, A. (2017). Network slicing to enable scalability and flexibility in 5G mobile networks. IEEE Communications Magazine, 55(5), 72–79.
2. Foukas, X., Patounas, G., Elmokashfi, A., & Marina, M. K. (2017). Network slicing in 5G: Survey and challenges. IEEE Communications Magazine, 55(5), 94–100.
3. 3GPP. (2020). 5G; Management and orchestration; Concepts, use cases and requirements (3GPP TS 28.530 version 16.2.0 Release 16). European Telecommunications Standards Institute.
4. ETSI. (2020). Zero-touch network and service management (ZSM); Reference architecture (ETSI GS ZSM 002 V1.1.1). European Telecommunications Standards Institute.
5. Li, R., Zhao, Z., Sun, Q., Chih-Lin, I., Yang, C., Chen, X., Zhao, M., & Zhang, H. (2019). Deep reinforcement learning for network slicing with heterogeneous resource requirements and time varying traffic. IEEE Transactions on Vehicular Technology, 68(12), 11496–11505.
6. Chen, Q., & Yu, F. R. (2020). Deep reinforcement learning for network slicing: A comprehensive survey. IEEE Communications Surveys & Tutorials, 23(1), 186–218.
7. Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
8. Sutton, R. S., Precup, D., & Singh, S. (1999). Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112(1–2), 181–211.
9. Kulkarni, T. D., Narasimhan, K., Saeedi, A., & Tenenbaum, J. (2016). Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation. Advances in Neural Information Processing Systems, 29, 3675–3683.
10. Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., & Kavukcuoglu, K. (2017). FeUdal networks for hierarchical reinforcement learning. Proceedings of the 34th International Conference on Machine Learning, 70, 3540–3549.
11. Huang, S., Li, J., & Li, X. (2021). Hierarchical reinforcement learning for network function virtualization resource allocation. IEEE Internet of Things Journal, 8(5), 3812–3822.
12. Li, Q. (2026). QoS Assurance Mechanism for 5G Network Slicing Based on the Deep Reinforcement Learning PPO Algorithm. arXiv preprint arXiv:2605.03345.
13. O-RAN Alliance. (2021). O-RAN Architecture Description v4.0. O-RAN Working Group 1.
14. Moubayed, A., Refaey, A., & Shami, A. (2020). Machine learning in network slicing—A survey. IEEE Access, 8, 70344–70367.
15. Garcia, J., & Fernandez, F. (2015). A comprehensive survey on safe reinforcement learning. Journal of Machine Learning Research, 16(1), 1437–1480.
16. Puiutta, E., & Veith, E. M. S. P. (2020). Explainable reinforcement learning: A survey. In A. Holzinger, P. Kieseberg, A. M. Tjoa, & E. Weippl (Eds.), International Cross-Domain Conference for Machine Learning and Knowledge Extraction (pp. 77–95). Springer.
17. Chen, D., Jiang, T., & Liu, Y. (2020). Fair resource allocation for network slices in 5G radio access networks with deep reinforcement learning. IEEE Transactions on Vehicular Technology, 69(9), 10197–10210.
18. Zhang, L., Afolabi, I., & Taleb, T. (2020). A survey on artificial intelligence for network slicing in beyond 5G: Techniques, open issues, and opportunities. IEEE Access, 8, 211014–211043.
19. Saad, W., Bennis, M., & Chen, M. (2020). A vision of 6G wireless systems: Applications, trends, technologies, and open research problems. IEEE Network, 34(3), 134–142.
20. Lu, Z., Lei, T., & Xu, J. (2022). Hierarchical reinforcement learning for intent-driven network slicing. IEEE Transactions on Cognitive Communications and Networking, 8(1), 180–193.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Advanced Artificial Intelligence Research

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.