Adaptive Parameter-Efficient Fine-Tuning for Open-World Video Object Segmentation in Dynamic Scenes
Keywords:
parameter-efficient fine-tuning, open-world video object segmentation, dynamic scenes, adaptive systems, system design, robustness, governanceAbstract
The emergence of vision foundation models has radically transformed video object segmentation, yet their deployment in open-world dynamic scenes presents systemic challenges that extend far beyond model accuracy. This work systematically examines the design of adaptive parameter-efficient fine-tuning frameworks tailored to open-world video object segmentation under non-stationary conditions, emphasizing architectural decisions, infrastructure orchestration, robustness governance, and socio-technical policy dimensions. We argue that the current generation of parameter-efficient adapters, low-rank adaptors, and prompt-based tuning modules, while reducing the computational footprint of fine-tuning, lacks the architectural reflexivity needed to continuously reconfigure learnable parameters in response to distributional shifts, novel object categories, and resource-constrained edge environments. We propose a conceptual architecture in which a meta-controller monitors scene dynamics, uncertainty signals, and hardware telemetry to decide online which subsets of parameters should be updated, and at what granularity, without resorting to full-model retraining. Through this lens, the paper analyzes structural trade-offs involving latency, memory, energy consumption, and fairness, drawing on a broad body of evidence from video segmentation benchmarks, green AI studies, and algorithmic fairness research. Further, we explore the deployment ecosystems required to sustain such adaptive systems across geographies with heterogeneous regulatory regimes, addressing privacy, data sovereignty, and accountability gaps. The discussion extends to governance mechanisms that can audit adaptation trajectories and prevent discriminatory feedback loops in safety-critical applications. By synthesizing cross-domain insights from computer vision, systems engineering, and policy studies, the paper offers a roadmap for building robust, sustainable, and ethically-grounded segmentation infrastructures that can gracefully handle the unbounded variability of open-world dynamic scenes.
References
1. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollar, P., & Girshick, R. (2023). Segment anything. Proceedings of the IEEE/CVF International Conference on Computer Vision, 4015–4026.
2. Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., Radle, R., Rolland, C., Gustafson, L., & Feichtenhofer, C. (2024). SAM 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714.
3. Cheng, H. K., & Schwing, A. G. (2022). XMem: Long-term video object segmentation with an Atkinson-Shiffrin memory model. In European Conference on Computer Vision (pp. 640–658). Springer.
4. Yang, Z., Wei, Y., & Yang, Y. (2021). Associating objects with transformers for video object segmentation. Advances in Neural Information Processing Systems, 34, 2491–2502.
5. Yang, Z., & Yang, Y. (2022). Decoupling features in hierarchical propagation for video object segmentation. Advances in Neural Information Processing Systems, 35, 25202–25214.
6. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
7. Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., & Gelly, S. (2019). Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning (pp. 2790–2799). PMLR.
8. Li, X. L., & Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (pp. 4582–4597). Association for Computational Linguistics.
9. Ghiasi, G., Gu, X., Cui, Y., & Lin, T.-Y. (2022). Scaling open-vocabulary image segmentation with image-level labels. In European Conference on Computer Vision (pp. 540–557). Springer.
10. Wang, W., Feiszli, M., Wang, H., & Tran, D. (2021). Unidentified video objects: A benchmark for dense, open-world segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 10776–10785).
11. Athar, A., Luiten, J., Hermans, A., Ramanan, D., & Leibe, B. (2023). BURST: A benchmark for unifying object recognition, segmentation and tracking in video. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 1674–1683).
12. Qi, J., Gao, Y., Hu, Y., Wang, X., Liu, X., Bai, X., Belongie, S., Yuille, A., Torr, P. H. S., & Bai, S. (2022). Occluded video instance segmentation: A benchmark. International Journal of Computer Vision, 130(8), 2022–2039.
13. Dave, A., Khurana, T., Tokmakov, P., Schmid, C., & Ramanan, D. (2020). TAO: A large-scale benchmark for tracking any object. In European Conference on Computer Vision (pp. 436–454). Springer.
14. Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., & Girdhar, R. (2022). Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1290–1299).
15. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (pp. 8748–8763). PMLR.
16. Yang, L., Fan, Y., & Xu, N. (2019). Video instance segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 5188–5197).
17. Li, G., Yuan, H., Chen, S., Hu, Q., Wang, J., & Jiang, K. (2026). MFT: Memory-Aware Fine-Tuning of SAM2 for Efficient Long-Sequence Video Object Segmentation. IEEE Signal Processing Letters.
18. Karimi Mahabadi, R., Henderson, J., & Ruder, S. (2021). Compacter: Efficient low-rank hypercomplex adapter layers. In Advances in Neural Information Processing Systems, 34, 1022–1035.
19. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 59–68). ACM.
20. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63.
21. McMahan, B., Moore, E., Ramage, D., Hampson, S., & y Arcas, B. A. (2017). Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics (pp. 1273–1282). PMLR.
22. Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., & Wermter, S. (2019). Continual lifelong learning with neural networks: A review. Neural Networks, 113, 54–71.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Advanced Artificial Intelligence Research

This work is licensed under a Creative Commons Attribution 4.0 International License.
This article is published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.