Text-to-Person Retrieval with Disentangled Identity and Apparel Representations in Multimodal Vision Systems

Authors

  • Deepak Ganerjee Department of Computer Science, Colorado State University, Fort Collins, CO, USA. Author
  • Leif Barner Department of Computer Science and Engineering, University at Buffalo, Buffalo, NY, USA. Author
  • Karan Batra Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO, USA. Author

Keywords:

text-to-person retrieval, multimodal vision systems, disentangled representation, identity-apparel separation, person re-identification, large-scale infrastructure, fairness, system governance

Abstract

Text-to-person retrieval constitutes a rapidly expanding frontier in multimodal vision systems, enabling the search for individuals across large-scale image galleries using natural language queries. This paper provides a system-level examination of retrieval architectures that incorporate disentangled representations of identity and apparel, a design imperative for robust performance under clothing variations and temporal drift in appearance. Rather than focusing on narrowly defined model components, the discussion adopts an interdisciplinary lens spanning infrastructure design, cross-modal alignment, governance, fairness, sustainability, and operational deployment. The analysis traces the trajectory from monolithic person re-identification pipelines toward modular, semantically partitioned systems in which identity and transient visual attributes are explicitly separated through architectural constraints and training objectives. We explore the structural trade-offs introduced by disentanglement, including the coordination of multiple feature streams, the management of cross-modal attention, and the delicate equilibrium between apparel suppression and retention of complementary contextual cues. Deployment considerations such as distributed inference, model compression, and energy efficiency are examined alongside the sociotechnical implications of person search at scale. Bias amplification through apparel-based shortcuts, privacy risks, and the need for context-sensitive governance frameworks are treated as integral to system design rather than as afterthoughts. By surveying the state of the art and projecting forward-looking design principles, the paper articulates a comprehensive systems research agenda for next-generation multimodal retrieval platforms that balance technical performance with societal accountability.

References

1. Zheng, L., Shen, L., Tian, L., Wang, S., Wang, J., & Tian, Q. (2015). Scalable person re-identification: A benchmark. In Proceedings of the IEEE International Conference on Computer Vision (pp. 1116–1124).

2. Ahmed, E., Jones, M., & Marks, T. K. (2015). An improved deep learning architecture for person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3908–3916).

3. Li, S., Xiao, T., Li, H., Zhou, B., Yue, D., & Wang, X. (2017). Person search with natural language description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1970–1979).

4. Faghri, F., Fleet, D. J., Kiros, J. R., & Fidler, S. (2018). VSE++: Improving visual-semantic embeddings with hard negatives. In British Machine Vision Conference.

5. Ye, M., Shen, J., Lin, G., Xiang, T., Shao, L., & Hoi, S. C. H. (2021). Deep learning for person re-identification: A survey and outlook. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6), 2872–2893.

6. Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., ... & Lerchner, A. (2017). beta-VAE: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations.

7. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.

8. Lu, J., Batra, D., Parikh, D., & Lee, S. (2019). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems.

9. He, S., Luo, H., Wang, P., Wang, F., Li, H., & Jiang, W. (2021). Transreid: Transformer-based object re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 15013–15022).

10. Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., ... & Ng, A. Y. (2012). Large scale distributed deep networks. In Advances in Neural Information Processing Systems (pp. 1223–1231).

11. Lu, Y., Wu, Y., Liu, B., Zhang, T., Li, B., Chu, Q., & Yu, N. (2018). Cross-modality person re-identification with shared-specific feature transfer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1339–1347).

12. Yang, Q., Wu, A., & Zheng, W. S. (2021). Person re-identification by contour sketch under moderate clothing change. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(6), 2029–2046.

13. Torralba, A., & Efros, A. A. (2011). Unbiased look at dataset bias. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1521–1528).

14. Goodfellow, I. J., Shlens, J., & Szegedy, C. (2015). Explaining and harnessing adversarial examples. In International Conference on Learning Representations.

15. Ding, Y., Wang, X., Yuan, H., Qu, M., & Jian, X. (2025). Decoupling feature-driven and multimodal fusion attention for clothing-changing person re-identification. Artificial Intelligence Review, 58(8), 241.

16. Lee, K. H., Chen, X., Hua, G., Hu, H., & He, X. (2018). Stacked cross attention for image-text matching. In Proceedings of the European Conference on Computer Vision (pp. 201–216).

17. Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., & Zhang, L. (2016). Deep learning with differential privacy. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (pp. 308–318).

18. Pu, N., Chen, W., Liu, Y., Bakker, E. M., & Lew, M. S. (2021). Lifelong person re-identification via adaptive knowledge accumulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7901–7910).

19. Zhang, Y., Lu, Z., & Wang, S. (2020). Text-based person search with progressive multi-granularity learning. IEEE Access, 8, 40552–40561.

20. Zhang, X., Luo, H., Fan, X., Xiang, W., Sun, Y., Xiao, Q., ... & Sun, J. (2017). Alignedreid: Surpassing human-level performance in person re-identification. arXiv preprint arXiv:1711.08184.

21. Schroff, F., Kalenichenko, D., & Philbin, J. (2015). Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 815–823).

22. Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., ... & Vayena, E. (2018). AI4People—An ethical framework for a good AI society: Opportunities, risks, principles, and recommendations. Minds and Machines, 28(4), 689–707.

23. Dehghani, M., Zamani, H., Severyn, A., Kamps, J., & Croft, W. B. (2017). Neural ranking models with weak supervision. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 65–74).

24. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186).

25. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

Downloads

Published

2026-05-25

How to Cite

Text-to-Person Retrieval with Disentangled Identity and Apparel Representations in Multimodal Vision Systems. (2026). Journal of Advanced Artificial Intelligence Research, 1(1). https://www.jaair.org/index.php/home/article/view/120