Knowledge-Enhanced Text-to-Image Generation for Digital Cultural Heritage Preservation on Edge Computing Platforms
Keywords:
text-to-image generation, edge computing, digital cultural heritage, knowledge graphs, model adaptation, provenance, sustainabilityAbstract
The preservation of digital cultural heritage increasingly relies on generative artificial intelligence to reconstruct, visualize, and contextualize artifacts whose physical counterparts face deterioration, destruction, or inaccessibility. Text-to-image models offer a promising pathway for generating historically grounded visual content, yet generic models struggle with the domain specificity, contextual accuracy, and cultural sensitivity required by heritage institutions. This paper presents a system-level investigation of knowledge-enhanced text-to-image generation deployed on edge computing platforms, a paradigm that reconciles the need for powerful generative capabilities with the constraints of fieldwork, museums, and heritage sites. We examine how structured domain knowledge, including ontologies, curated metadata, and verified visual references, can be integrated into the generation pipeline through retrieval-augmented conditioning, prompting strategies, and lightweight adapter modules. The architectural discussion foregrounds edge-native deployment trade-offs, encompassing model compression, distributed inference, hardware heterogeneity, and the incorporation of emerging edge-AI accelerator ecosystems. Beyond technical architecture, we analyze governance frameworks for provenance, authenticity labeling, and bias mitigation, emphasizing the risk of erasing marginalized heritage narratives when generative models amplify training data imbalances. The paper further addresses robustness under degraded communication, energy sustainability at remote sites, and long-term digital accessibility. By synthesizing technical, organizational, and ethical dimensions, we argue that knowledge-enhanced generation on edge infrastructure constitutes a sustainable and culturally responsible foundation for next-generation digital heritage preservation systems. The analysis is grounded in contemporary research on multimodal learning, edge computing, and cultural informatics, offering forward-looking perspectives for system architects, heritage professionals, and policy makers.
References
1. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML).
2. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
3. Li, J., Li, D., Savarese, S., & Hoi, S. C. H. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML).
4. Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., & Taigman, Y. (2022). Make-A-Scene: Scene-based text-to-image generation with human knowledge. In Proceedings of the European Conference on Computer Vision (ECCV).
5. Fiorucci, M., Khoroshiltseva, M., Pontil, M., Traviglia, A., Del Bue, A., & James, S. (2020). Machine learning for cultural heritage: A survey. Pattern Recognition Letters, 133, 102-108.
6. Wang, X., Han, Y., Leung, V. C. M., Niyato, D., Yan, X., & Chen, X. (2022). Convergence of edge computing and deep learning: A comprehensive survey. IEEE Communications Surveys & Tutorials, 22(2), 869-904.
7. Xu, D., Zhu, Y., Choy, C. B., & Fei-Fei, L. (2017). Scene graph generation by iterative message passing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
8. Doerr, M. (2003). The CIDOC conceptual reference model: An ontological approach to semantic interoperability of cultural heritage information. In Proceedings of the Conference on Artificial Intelligence and Cultural Heritage.
9. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML).
10. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
11. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
12. Yao, J., Han, T., & Ansari, N. (2019). On mobile edge caching. IEEE Communications Surveys & Tutorials, 21(3), 2525-2553.
13. McMahan, B., Moore, E., Ramage, D., Hampson, S., & y Arcas, B. A. (2017). Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS).
14. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92.
15. UNESCO. (2015). Charter on the preservation of the digital heritage. Records of the General Conference, 38th session.
16. Chen, C., Wang, C., Li, Y., Wan, Z., Geng, M., Xiao, J., ... & Peng, Y. (2026). JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators. arXiv preprint arXiv:2606.28421.
17. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
18. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT).
19. Bolukbasi, T., Chang, K. W., Zou, J. Y., Saligrama, V., & Kalai, A. T. (2016). Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. In Advances in Neural Information Processing Systems (NIPS).
20. Hendrycks, D., & Dietterich, T. (2019). Benchmarking neural network robustness to common corruptions and perturbations. In Proceedings of the International Conference on Learning Representations (ICLR).
21. Lavoie, B. F., & Dempsey, L. (2004). Thirteen ways of looking at... digital preservation. D-Lib Magazine, 10(7/8).
22. Benzerki, M. A., Belkhir, L., & Elouatouat, N. (2020). Energy efficiency in edge AI: A review. IEEE Access, 8, 141224-141249.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Advanced Artificial Intelligence Research

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.