35. Person re-identification across multi-camera systems using vision language models // IEEE FLLM 2025
Amirgaliyev B., Mussabek M., Bekmukhanbet A.
Person re-identification across multi-camera systems using vision language models // 2025 3rd International Conference on Foundation and Large Language Models (FLLM). — Vienna, Austria, 2025. — Pp. 1247–1253. — DOI: 10.1109/FLLM67465.2025.11391222.
Abstract: Intelligent solutions that improve public safety while handling enormous data streams are desperately needed, as evidenced by the explosive growth of surveillance systems. Large-scale deployment issues, appearance variations, and occlusion can all lead to errors in traditional camera tracking by saved recordings. In order to overcome these drawbacks, we suggest a deep learning-based framework that combines person description module and re-identification(Re-ID), which allows for strong identity association across several cameras. Our system uses detection, tracking modules to locate a person in the video and Re-ID to identify people in a variety of scenarios. Person description module utilizes vision language model(VLM) to describe target person. After testing this system with different VLMs the most suitable was Gemma3:4B, which achieves higher lexical overlap with the reference captions, reflected in its BLEU score of 0.0768 and ROUGE-L of 0.3283 and semantic similarity of 0.5887. These elements work together to improve reliability in crowd monitoring, anomaly detection, and multi-camera tracking.
Link / DOI: https://doi.org/10.1109/FLLM67465.2025.11391222
Отправить комментарий