Abstract:
Multimodal object re-identification aims to integrate complementary information from different modals to achieve the retrieval of the same target across non-overlapping surveillances. The primary challenge in this process lies in mitigating the cross-modal discrepancies. To address this issue, a hierarchical token selection transformer (HTSTrans) is proposed. HTSTrans consists of a multi-modal shared feature learning module (MSFLM), a hierarchical token selection module (HTSM), and a cross-modal alignment constraint module (CACM). MSFLM employs a weight-shared vision transformer (ViT) to learn shared multi-modal features, establishing preliminary feature consistency across different modals and providing a unified representation space for subsequent fine-grained feature learning. HTSM utilizes a hierarchically-aware progressive window mechanism to select crucial tokens. Through token fusion, it extracts modal-specific crucial detailed features, followed by multi-modal feature fusion to reduce cross-modal differences. CACM applies an inter-modal transmission constraint (ITC) and a cycle consistency constraint (CCC) to reduce feature discrepancies of the same identity in different modal spaces and enhance the discriminability of modal-invariant features. Experiments demonstrate that HTSTrans effectively improves the performance of multi-modal object re-identification.