Vision-Language Guided Identity Retrieval with Robust Appearance Disentanglement for Long-Term Person Tracking

Authors

  • Yuanhong Du Department of Electrical Engineering and Computer Science, University of Kansas, Lawrence, KS, USA. Author
  • Walid Stanley Department of Computer Science, University of New Hampshire, Durham, NH, USA. Author

Keywords:

person re-identification, vision-language models, appearance disentanglement, long-term tracking, system architecture, fairness, sustainable AI

Abstract

The sustained tracking of individuals across distributed camera networks over extended temporal windows remains an enduring systems challenge, as visual appearance can vary markedly due to clothing changes, illumination shifts, and occlusions. This paper offers a comprehensive system-level investigation into the convergence of vision-language guidance and robust appearance disentanglement for long-term person re-identification. We posit that integrating pre-trained vision-language models into the retrieval pipeline injects semantic reasoning about personal attributes that are resilient to superficial appearance transformations. A hybrid architecture is proposed in which multimodal fusion modules align visual embeddings with natural language descriptions of identity-relevant traits, while adversarial feature separation and semantic consistency regularization enforce identity-salient representations decoupled from transient clothing styles. The discussion moves beyond algorithmic design to encompass the governance, ethical deployment, and sustainability of such tracking infrastructures. We analyze structural trade-offs between model expressiveness and inference latency, energy-aware compression strategies, and privacy-preserving data flows. Drawing on smart city and public safety case illustrations, the paper examines how fairness audits and policy-aligned architectures can mitigate risks of demographic bias and erosion of civil liberties. The resulting framework provides a roadmap for designing person tracking systems that balance accuracy, robustness, accountability, and environmental responsibility in long-term operational contexts.

References

1. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (pp. 8748–8763). PMLR.

2. Li, J., Li, D., Xiong, C., & Hoi, S. (2022). BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Advances in Neural Information Processing Systems (Vol. 35, pp. 12888–12900).

3. Zheng, L., Shen, L., Tian, L., Wang, S., Wang, J., & Tian, Q. (2015). Scalable person re-identification: A benchmark. In Proceedings of the IEEE International Conference on Computer Vision (pp. 1116–1124).

4. Xiao, T., Li, S., Wang, B., Lin, L., & Wang, X. (2017). Joint detection and identification feature learning for person search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3415–3424).

5. Hermans, A., Beyer, L., & Leibe, B. (2017). In defense of the triplet loss for person re-identification. arXiv preprint, arXiv:1703.07737.

6. Ge, Y., Chen, D., & Li, H. (2020). Mutual mean-teaching: Pseudo label refinery for unsupervised domain adaptation on person re-identification. In International Conference on Learning Representations.

7. Ma, L., Jia, X., Sun, Q., Schiele, B., Tuytelaars, T., & Van Gool, L. (2019). Disentangled person image generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 12089–12098).

8. Eom, C., Lee, G., Lee, J., & Ham, B. (2021). Cloth-changing person re-identification using pose and shape. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 4220–4228).

9. Ding, Y., Wang, X., Yuan, H., Qu, M., & Jian, X. (2025). Decoupling feature-driven and multimodal fusion attention for clothing-changing person re-identification. Artificial Intelligence Review, 58(8), 241.

10. Liu, X., Song, M., Tao, D., Zhou, Z., Chen, C., & Bu, J. (2019). Video-based person re-identification with accumulative motion context. IEEE Transactions on Circuits and Systems for Video Technology, 29(9), 2776–2789.

11. Sun, P., Zhang, R., Jiang, Y., Kong, F., Xu, C., Zhan, W., Tomizuka, M., Li, Z., Yuan, Z., & Wang, C. (2019). MV-GAN: Multi-view generative adversarial network for two-stage cross-view person re-identification. In Proceedings of the IEEE International Conference on Computer Vision (pp. 5655–5664).

12. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (pp. 5998–6008).

13. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (pp. 4171–4186).

14. Huang, X., & Belongie, S. (2017). Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE International Conference on Computer Vision (pp. 1501–1510).

15. Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. In International Conference on Learning Representations.

16. Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the Conference on Fairness, Accountability and Transparency (pp. 77–91).

17. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 59–68).

18. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63.

19. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint, arXiv:2104.10350.

Downloads

Published

2026-06-14

How to Cite

Vision-Language Guided Identity Retrieval with Robust Appearance Disentanglement for Long-Term Person Tracking. (2026). Journal of Advanced Artificial Intelligence Research, 1(1). https://www.jaair.org/index.php/home/article/view/123