基于层级化Token选择Transformer的多模态目标重识别方法

Multimodal object re-identification based on hierarchical token selection Transformer

  • 摘要: 多模态目标重识别旨在融合不同模态的互补信息,在视域不重叠的监控场景中实现对同一目标的检索,主要挑战在于缓解跨模态差异。为此,提出一种层级化token选择Transformer(HTSTrans)。HTSTrans由多模态共享特征学习模块(MSFLM)、层级化token选择模块(HTSM)和跨模态对齐约束模块(CACM)构成。MSFLM使用权重共享的视觉Transformer(ViT)学习多模态共享特征,在不同模态间初步建立特征一致性,为后续更精细化的特征学习提供统一的表示空间。HTSM使用层级感知的递进式窗口选择关键token,基于token融合实现对模态特定关键细节特征的提取,再通过多模态特征融合降低跨模态差异。CACM使用模态间传输约束(ITC)和循环一致性约束(CCC)降低同一目标身份在不同模态空间中的特征差异,并提升模态不变性特征的判别性。实验结果表明,HTSTrans可有效提升多模态目标重识别性能。

     

    Abstract: Multimodal object re-identification aims to integrate complementary information from different modals to achieve the retrieval of the same target across non-overlapping surveillances. The primary challenge in this process lies in mitigating the cross-modal discrepancies. To address this issue, a hierarchical token selection transformer (HTSTrans) is proposed. HTSTrans consists of a multi-modal shared feature learning module (MSFLM), a hierarchical token selection module (HTSM), and a cross-modal alignment constraint module (CACM). MSFLM employs a weight-shared vision transformer (ViT) to learn shared multi-modal features, establishing preliminary feature consistency across different modals and providing a unified representation space for subsequent fine-grained feature learning. HTSM utilizes a hierarchically-aware progressive window mechanism to select crucial tokens. Through token fusion, it extracts modal-specific crucial detailed features, followed by multi-modal feature fusion to reduce cross-modal differences. CACM applies an inter-modal transmission constraint (ITC) and a cycle consistency constraint (CCC) to reduce feature discrepancies of the same identity in different modal spaces and enhance the discriminability of modal-invariant features. Experiments demonstrate that HTSTrans effectively improves the performance of multi-modal object re-identification.

     

/

返回文章
返回