Energy-Efficient Prompt Placement for Green Artificial Intelligence in Large-Scale NLP Systems
Keywords:
green artificial intelligence, prompt placement, energy-efficient NLP, large language models, selective prompt tuning, sustainable computing, system-level optimizationAbstract
The rapidly expanding computational demands of large-scale natural language processing systems have elevated energy consumption into a central challenge for sustainable artificial intelligence. While parameter-efficient fine-tuning methods such as prompt tuning and prefix tuning have substantially reduced the costs associated with adapting foundation models, the placement and activation of these prompts within inference pipelines remain under-explored from a system-level energy perspective. This paper presents a holistic examination of energy-efficient prompt placement as a lever for green AI in large-scale deployment environments. By framing prompt activation as a dynamic resource allocation problem, we analyze the structural trade-offs between inference latency, model quality, and carbon footprint across cloud, edge, and hybrid infrastructures. We explore architectural design patterns that support selective prompt insertion, caching strategies for reusable prompt states, and carbon-aware scheduling of prompt-dependent computations. The discussion extends to governance challenges surrounding fairness, transparency, and accountability when energy-saving policies differentially impact user populations or linguistic groups. Through cross-domain comparisons with adaptive computation mechanisms such as mixture-of-experts routing and early-exit architectures, we identify shared principles that can guide the development of energy-proportional prompt management frameworks. The paper concludes with forward-looking perspectives on how prompt placement strategies can be integrated into broader sustainability roadmaps for industrial-scale NLP systems, emphasizing the need for standardized energy reporting, open benchmarking databases, and multi-stakeholder policy coordination.
References
1. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645–3650). Association for Computational Linguistics.
2. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63.
3. Tay, Y., Dehghani, M., Bahri, D., & Metzler, D. (2020). Efficient transformers: A survey. ACM Computing Surveys, 53(6), Article 120.
4. Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., & Gelly, S. (2019). Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (pp. 2790–2799). PMLR.
5. Lester, B., Al-Rfou, R., & Constant, N. (2021). The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 3045–3059). Association for Computational Linguistics.
6. Li, X. L., & Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (pp. 4582–4597). Association for Computational Linguistics.
7. Zhu, W., & Tan, M. (2023, December). SPT: Learning to selectively insert prompts for better prompt tuning. In Proceedings of the 2023 conference on empirical methods in natural language processing (pp. 11862-11878).
8. Zaken, E. B., Ravfogel, S., & Goldberg, Y. (2022). BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (pp. 1–9). Association for Computational Linguistics.
9. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In Proceedings of the 10th International Conference on Learning Representations. OpenReview.
10. Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., & Dean, J. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In Proceedings of the 5th International Conference on Learning Representations. OpenReview.
11. Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
12. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A. G., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2704–2713). IEEE.
13. Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., & Liu, Q. (2020). TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 4163–4174). Association for Computational Linguistics.
14. Acun, B., Lee, K., Kazhamiaka, F., Maeng, K., Gupta, U., Chakkaravarthy, M., Brooks, D., & Wu, C.-J. (2023). Carbon Explorer: A holistic framework for carbon-efficient datacenter operations. In Proceedings of the 2nd Workshop on Sustainable Computer Systems (HotCarbon ’23). ACM.
15. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT ’19) (pp. 220–229). ACM.
16. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21) (pp. 610–623). ACM.
17. Jin, D., Jin, Z., Zhou, J. T., & Szolovits, P. (2020). Is BERT really robust? A strong baseline for natural language attack on text classification and entailment. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 34, pp. 8018–8025). AAAI Press.
18. Fowers, J., Ovtcharov, K., Papamichael, M. K., Massengill, T., Liu, M., Lo, D., Alkalay, S., Haselman, M., Adams, L., Ghandi, M., Heil, S., Cox, P., Shah, A., Bhardwaj, R., Caulfield, A. M., Chung, E. S., & Burger, D. (2018). Serving DNNs in real time at datacenter scale with Project Brainwave. In Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA ’18) (pp. 1–14). IEEE.
19. Jiang, Z., Lin, H., Zhong, Y., Huang, Q., Chen, Y., Zhang, Z., Peng, Y., Li, X., Xie, C., Nong, S., Jia, Y., He, S., Chen, H., Bai, Z., Hou, Q., Yan, S., Zhou, D., Sheng, Y., Jiang, Z., ... & Zheng, Y. (2024). MegaScale: Scaling large language model training to more than 10,000 GPUs. In Proceedings of the 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI ’24) (pp. 745–760). USENIX.
20. Rajbhandari, S., Rasley, J., Ruwase, O., & He, Y. (2020). ZeRO: Memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’20) (pp. 1–16). IEEE.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Advanced Artificial Intelligence Research

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.