Opinion

FMMT: Infrared-Visible Fusion Boosts Multi-Camera Multi-Target Tracking

Daily briefingLucas MeyerMay 12, 2026· 3,970 views

New algorithm fuses infrared and visible video for robust multi-camera multi-target tracking in low-light conditions, with a new benchmark dataset.

Researchers have developed a multi-modal tracking algorithm, FMMT, that fuses infrared and visible light video to improve multi-camera multi-target tracking (MCMT) under challenging lighting conditions. The method uses a deep neural network to adaptively combine features from aligned visible and infrared camera pairs, then applies a global association Transformer to link targets across cameras in a single step. This global paradigm avoids error accumulation typical of two-step approaches and enhances robustness in low-light or nighttime scenes.

Key takeaways

  • FMMT dynamically weights infrared and visible features per channel, leveraging thermal contours in darkness and texture/color in normal light.
  • The authors introduce M3Track, the first multi-modal MCMT dataset, with 20 scenes, 100k image pairs, and 1.129 million annotated targets captured by moving drones.
  • On M3Track, FMMT achieves 61.7 CVMA and 70.3 CVIDF1, significantly outperforming comparative methods, with notable gains in nighttime scenarios.
  • The framework uses CenterNet for detection and DLA34 backbones, but these components are replaceable, offering flexibility for industrial integration.

For manufacturing automation, the approach is relevant for robust visual tracking in warehouses, outdoor logistics, or security perimeters where lighting varies. The fusion strategy and dataset provide a practical blueprint for systems needing reliable multi-camera tracking under low illumination, potentially improving autonomous mobile robot coordination and surveillance accuracy.

Source: 《机器人》期刊 (robot.sia.cn) · Published 2026-05-12 · “基于红外与可见光的多相机、多目标跟踪”