Multi-Granularity Attention Fusion for Video-Based Person Re-Identification in Dynamic Urban Scenes

Authors

  • Geapak Shukla Department of Computer Science and Engineering, University of Nevada, Reno, Reno, NV, USA. Author

Keywords:

person re-identification, multi-granularity attention, video analysis, urban surveillance, feature fusion, deep learning, system architecture, algorithmic fairness

Abstract

Video-based person re-identification in dynamic urban environments constitutes a critical enabler of intelligent urban systems, yet it poses formidable challenges due to uncontrolled variations in illumination, viewpoint, background clutter, occlusion, and human motion. This paper presents a systems-oriented analysis of multi-granularity attention fusion as a robust architectural paradigm for video-based person re-identification. Moving beyond purely model-centric treatments, the discussion frames re-identification within a complex socio-technical infrastructure that spans edge computing, privacy governance, and long-term sustainability. We examine how attention mechanisms operating at multiple spatial, temporal, channel, and part-based granularities can be systematically fused to improve re-identification accuracy while accommodating the constraints of real-world deployment. The analysis foregrounds structural trade-offs among computational cost, latency, modularity, and model interpretability, and it situates attention fusion design within broader debates on algorithmic fairness, accountability, and urban data governance. By integrating conceptual investigation, cross-domain comparison, and forward-looking policy reflection, this paper offers an interdisciplinary perspective on the design and deployment of attentive person re-identification systems in urban surveillance networks. The article ultimately argues that sustainable urban re-identification requires a holistic fusion not only of visual features but also of engineering robustness, ethical safeguards, and adaptive operational architectures.

References

1. Zheng, L., Yang, Y., & Hauptmann, A. G. (2016). Person re-identification: Past, present and future. arXiv preprint arXiv:1610.02984.

2. Ristani, E., Solera, F., Zou, R. S., Cucchiara, R., & Tomasi, C. (2016). Performance measures and a data set for multi-target, multi-camera tracking. In Computer Vision – ECCV 2016 Workshops (pp. 17-35). Springer.

3. Li, W., Zhao, R., Xiao, T., & Wang, X. (2014). DeepReID: Deep filter pairing neural network for person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 152-159).

4. Varior, R. R., Haloi, M., & Wang, G. (2016). Gated siamese convolutional neural network architecture for human re-identification. In European conference on computer vision (pp. 791-808). Springer.

5. Hermans, A., Beyer, L., & Leibe, B. (2017). In defense of the triplet loss for person re-identification. arXiv preprint arXiv:1703.07737.

6. Wang, G., Yuan, Y., Chen, X., Li, J., & Zhou, X. (2018). Learning discriminative features with multiple granularities for person re-identification. In Proceedings of the 26th ACM international conference on Multimedia (pp. 274-282).

7. Sun, Y., Zheng, L., Yang, Y., Tian, Q., & Wang, S. (2018). Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline). In Proceedings of the European conference on computer vision (ECCV) (pp. 480-496).

8. Chen, Y., Zhu, X., Zheng, W. S., & Lai, J. H. (2018). Person re-identification by camera correlation aware feature augmentation. IEEE transactions on pattern analysis and machine intelligence, 40(4), 961-974.

9. Liu, C. T., Wu, C. W., Wang, Y. C. F., & Chien, S. Y. (2019). Spatially and temporally efficient non-local attention network for video-based person re-identification. arXiv preprint arXiv:1908.01683.

10. Hou, R., Ma, B., Chang, H., Gu, X., Shan, S., & Chen, X. (2019). VRSTC: occlusion-free video person re-identification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 7183-7192).

11. Luo, H., Gu, Y., Liao, X., Lai, S., & Jiang, W. (2019). Bag of tricks and a strong baseline for deep person re-identification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops (pp. 0-0).

12. Ding, Y., Wang, X., Yuan, H., Qu, M., & Jian, X. (2025). Decoupling feature-driven and multimodal fusion attention for clothing-changing person re-identification. Artificial Intelligence Review, 58(8), 241.

13. Wang, Q., Yuan, Y., & Li, X. (2019). Learning video representations for person re-identification with temporal pooling. IEEE Transactions on Circuits and Systems for Video Technology, 30(7), 2151-2161.

14. Liu, H., Jie, Z., Jayashree, K., Qi, M., Jiang, J., Yan, S., & Feng, J. (2017). Video-based person re-identification with accumulative motion context. IEEE transactions on circuits and systems for video technology, 28(10), 2788-2802.

15. Li, S., Bak, S., Carr, P., & Wang, X. (2018). Diversity regularized spatiotemporal attention for video-based person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 369-378).

16. Su, C., Li, J., Zhang, S., Xing, J., Gao, W., & Tian, Q. (2017). Pose-driven deep convolutional model for person re-identification. In Proceedings of the IEEE international conference on computer vision (pp. 3960-3969).

17. Zheng, Z., Zheng, L., & Yang, Y. (2019). Pedestrian alignment network for large-scale person re-identification. IEEE Transactions on Circuits and Systems for Video Technology, 29(10), 3037-3045.

18. Xu, J., Zhao, R., Zhu, F., Wang, H., & Ouyang, W. (2018). Attention-aware compositional network for person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2119-2128).

19. Song, C., Huang, Y., Ouyang, W., & Wang, L. (2018). Mask-guided contrastive attention model for person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1179-1188).

20. Zhang, J., Wang, N., & Zhang, L. (2019). Multi-shot pedestrian re-identification via sequential decision making. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 11078-11087).

Downloads

Published

2026-05-17

How to Cite

Multi-Granularity Attention Fusion for Video-Based Person Re-Identification in Dynamic Urban Scenes. (2026). Journal of Advanced Artificial Intelligence Research, 1(1). https://www.jaair.org/index.php/home/article/view/114