EDBT 2026 Demo / reviewers in the wild / expert
Andong Lu
dblp:245/2878
· DBLP profile ↗
19ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-0902-2260ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling the Power of Multi-Modal Template Update in RGBT TrackingabstractTemplate update is essential for improving the adaptability of tracking algorithms to target appearance variations. While previous methods have leveraged the spatio-temporal complementarity of multi-modal templates for RGBT tracking, a comprehensive analysis of the template update mechanism remains underexplored. In this work, we propose a novel prototype-based framework that decomposes the multi-modal template update process from the perspective of prototype learning into four key components: multi-modal prototype, prototype integration, prototype evaluation, and prototype update algorithm. Our findings highlight that the multi-modal prototype is the most critical factor in enhancing tracking adaptability to appearance variations, leading to more robust target representations. While prototype integration is less crucial when the target representation is already robust, it still contributes to learning a more discriminative representation. Additionally, the accuracy of template updates is strongly influenced by prototype evaluation, which controls the accuracy of the update process. Finally, the prototype update algorithm, which determines when and how template updates occur, is key to maintaining tracking robustness. Building on these insights, we introduce the Multi-modal Prototype RGBT Tracker (MPTrack), which adapts dynamically to appearance variations through prototype learning. MPTrack combines a fixed template from the first frame with both modality-shared and modality-specific templates, forming a robust multi-modal prototype representation. It incorporates a prototype evaluation module that guides updates based on template reliability, and an adaptive update algorithm to manage templates effectively. Additionally, a prototype-guided cross-modal integration module enhances the discriminative power of multi-modal relation modeling. Experimental results on five challenging RGBT tracking benchmarks demonstrate that MPTrack consistently outperforms state-of-the-art methods, setting new performance records. The experimental data and source code will be made publicly available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code. Lei Liu 0049, Chenglong Li 0002, Andong Lu, Yabin Zhu, Shoufei Han, Xinye Cai, Changhe Li |
IEEE Trans. Image Process. | 3 |
| 2026 | Pixel-Level RGBT Fusion Tracking via Heterogeneous Multi-Expert Distillation and Decoupled Representation LearningabstractPixel-level fusion is widely considered a lightweight yet limited strategy in RGB-Thermal (RGBT) tracking due to its shallow representational capacity. However, its actual limitations and potential remain largely unexplored. We systematically analyze fusion location, modality alignment, and tracking performance, revealing that despite lower modality gaps than feature-level fusion, pixel-level fusion lacks task-relevant discrimination, restricting its effectiveness. In this paper, we propose the Task-driven Pixel-level Fusion tracker (TPF), which preserves the efficiency of early fusion while enhancing discriminative capacity. Central to TPF is a lightweight pixel fusion adapter that ensures real-time image fusion with only 14.3KB extra parameters over the baseline at inference. To enhance its limited representational capacity, we propose a task-driven progressive learning framework consisting of two key stages. First, a heterogeneous multi-expert distillation scheme adaptively transfers image fusion knowledge from diverse models under tracking-guided evaluation, mitigating the generalization limitations of single-teacher distillation across varied tracking scenarios. Second, to overcome limited task discrimination caused by sparse, target-focused tracking supervision, we propose a decoupled representation learning strategy that offers dense, complementary guidance to improve target-background separation and fusion quality. A nearest-neighbor dynamic template update further enhances robustness to appearance changes. Extensive experiments on four RGBT tracking benchmarks show that TPF achieves competitive accuracy and speed, outperforming both feature-level and existing pixel-level fusion methods, offering new insights into efficient RGBT tracking. Andong Lu, Yuanzhi Guo, Kunpeng Wang 0005, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Cross-modulated Attention Transformer for RGBT TrackingabstractExisting Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and search-template correlation. Nevertheless, the independent search-template correlation calculations are prone to be affected by low-quality data, which might result in contradictory and ambiguous correlation weights. It not only limits the intra-modal feature representation, but also harms the robustness of cross-attention for multi-modal feature interaction and search-template correlation computation. To address these issues, we propose a novel approach called Cross-modulated Attention Transformer (CAFormer), which innovatively integrates inter-modality interaction into the search-template correlation computation within typical attention mechanism, for RGBT tracking. In particular, we first independently generate correlation maps for each modality and feed them into the designed correlation modulated enhancement module, which can modify inaccurate correlation weights by seeking the consensus between modalities. Such kind of design unifies self-attention and cross-attention schemes, which not only alleviates inaccurate attention weight computation in self-attention but also eliminates redundant computation introduced by extra cross-attention scheme. In addition, we design a collaborative token elimination strategy to further improve tracking inference efficiency and accuracy. Experiments on five public RGBT tracking benchmarks show the outstanding performance of the proposed CAFormer against state-of-the-art methods. Yun Xiao 0003, Jiacong Zhao, Andong Lu, Chenglong Li 0002, Yin Lin, Cong Liu 0006 |
AAAI | 3 |
| 2025 | RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion MambaabstractExisting RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue, this paper presents a novel All-layer multimodal Interaction Network, named AINet, which performs efficient and effective feature interactions of all modalities and layers in a progressive fusion Mamba, for robust RGBT tracking. Even though modality features in different layers are known to contain different cues, it is always challenging to build multimodal interactions in each layer due to struggling in balancing interaction capabilities and efficiency. Meanwhile, considering that the feature discrepancy between RGB and thermal modalities reflects their complementary information to some extent, we design a Difference-based Fusion Mamba (DFM) to achieve enhanced fusion of different modalities with linear complexity. When interacting with features from all layers, a huge number of token sequences (3840 tokens in this work) are involved and the computational burden is thus large. To handle this problem, we design an Order-dynamic Fusion Mamba (OFM) to execute efficient and effective feature interactions of all layers by dynamically adjusting the scan order of different layers in Mamba. Extensive experiments on four public RGBT tracking datasets show that AINet achieves leading performance against existing state-of-the-art methods. We will release the code upon acceptance of the paper. Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
AAAI | 1 |
| 2025 | Efficient RGBT Tracking via Heterogeneous Hierarchical Knowledge DistillationabstractThe increasing demand for real-time RGBT (RGB and Thermal) tracking in applications such as video surveillance, autonomous driving, and robotic navigation underscores the need for lightweight and efficient tracking frameworks. However, existing approaches require dual-stream architectures for repeated feature extraction, incurring high costs, while complex interaction strategies further reduce efficiency, limiting the real-time performance of RGBT trackers. To address this, we propose a novel Heterogeneous Hierarchical Knowledge Distillation framework (H2KD) to enable a single-stream RGBT tracker that maintains high efficiency while delivering performance comparable to existing dual-stream trackers. In particular, H2KD takes an existing dual-stream network tracker as the teacher and builds a simple single-stream network tracker as the student by concatenating the inputs. To inherit powerful representation of the heterogeneous teacher network, we expand the channel dimensions of single-stream networks to align with the fusion features of teacher network, and employ a hierarchical distillation strategy between their backbone networks. Moreover, H2KD also introduces architecture-independent prediction-level distillation between their prediction score maps to inherit the tracking capability of teachers more directly. Extensive experiments on three major RGBT tracking benchmarks and multiple dual-stream RGBT trackers demonstrate the effectiveness and generalization of the proposed method, which achieves competitive accuracy while achieving 120.2 FPS inference speed. Dengdi Sun, Chenglong Li 0002, Andong Lu |
ICME | 4 |
| 2025 | Modality-missing RGBT Tracking: Invertible Prompt Learning and High-quality Benchmarks
Andong Lu, Chenglong Li 0002, Jiacong Zhao, Jin Tang 0001, Bin Luo 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Efficient RGBT Tracking via Multi-Path Mamba Fusion NetworkabstractRGBT tracking aims to fully exploit the complementary advantages of visible and infrared modalities to achieve robust tracking, thus the design of multimodal fusion network is crucial. However, existing methods typically adopt CNNs or Transformer networks to construct the fusion network, which poses a challenge in achieving a balance between performance and efficiency. To overcome this issue, we introduce an innovative visual state space (VSS) model, represented by Mamba, for RGBT tracking. In particular, we design a novel multi-path Mamba fusion network that achieves robust multimodal fusion capability while maintaining a linear overhead. First, we design a multi-path Mamba layer to sufficiently fuse two modalities in both global and local perspectives. Second, to alleviate the issue of inadequate VSS modeling in the channel dimension, we introduce a simple yet effective channel swapping layer. Extensive experiments conducted on four public RGBT tracking datasets demonstrate that our method surpasses existing state-of-the-art trackers. Notably, our fusion method achieves higher tracking performance compared to the well-known Transformer-based fusion approach (TBSI), while also achieving 92.8% and 80.5% reductions in parameter count and computational cost, respectively. Fanghua Hong, Andong Lu, Lei Liu 0049, Qunjing Wang |
IEEE Signal Process. Lett. | 3 |
| 2025 | BR-MoE: Blind Multi-Modal Tracking With Route-Dynamic Mixture of Experts
Qingguo Meng, Andong Lu, Zhe Jin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Nighttime Person Re-Identification via Collaborative Enhancement Network With Multi-Domain LearningabstractPrevalent nighttime person re-identification (ReID) methods typically combine image relighting and ReID networks in a sequential manner. However, their performance (recognition accuracy) is limited by the quality of relighting images and insufficient collaboration between image relighting and ReID tasks. To handle these problems, we propose a novel Collaborative Enhancement Network called CENet, which performs the multilevel feature interactions in a parallel framework, for nighttime person ReID. In particular, the designed parallel structure of CENet can not only avoid the impact of the quality of relighting images on ReID performance, but also allow us to mine the collaborative relations between image relighting and person ReID tasks. To this end, we integrate the multilevel feature interactions in CENet, where we first share the Transformer encoder to build the low-level feature interaction, and then perform the feature distillation that transfers the high-level features from image relighting to ReID, thereby alleviating the severe image degradation issue caused by the nighttime scenario while avoiding the impact of relighting images. In addition, the sizes of existing real-world nighttime person ReID datasets are limited, and large-scale synthetic ones exhibit substantial domain gaps with real-world data. To leverage both small-scale real-world and large-scale synthetic training data, we develop a multi-domain learning algorithm, which alternately utilizes both kinds of data to reduce the inter-domain difference in training procedure. Extensive experiments on two real nighttime datasets,Night600andRGBNT201rgb, and a synthetic nighttime ReID dataset are conducted to validate the effectiveness of CENet. We release the code and synthetic dataset at: https://github.com/Alexadlu/CENet. Andong Lu, Chenglong Li 0002, Tianrui Zha, Xiaofeng Wang 0009, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | AFTER: Attention-Based Fusion Router for RGBT TrackingabstractMulti-modal feature fusion as a core investigative component of RGBT tracking emerges numerous fusion studies in recent years. However, existing RGBT tracking methods widely adopt fixed fusion structures to integrate multi-modal feature, which are hard to handle various challenges in dynamic scenarios. To address this problem, this work presents a novel Attention-based Fusion router called AFTER, which optimizes the fusion structure to adapt to the dynamic challenging scenarios, for robust RGBT tracking. In particular, we design a fusion structure space based on the hierarchical attention network, each attention-based fusion unit corresponding to a fusion operation and a combination of these attention units corresponding to a fusion structure. Through optimizing the combination of attention-based fusion units, we can dynamically select the fusion structure to adapt to various challenging scenarios. Unlike complex search of different structures in neural architecture search algorithms, we develop a dynamic routing algorithm, which equips each attention-based fusion unit with a router, to predict the combination weights for efficient optimization of the fusion structure. Extensive experiments on five mainstream RGBT tracking datasets demonstrate the superior performance of the proposed AFTER against state-of-the-art RGBT trackers. We release the code in https://github.com/Alexadlu/AFter. Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Duality-Gated Mutual Condition Network for RGBT TrackingabstractLow-quality modalities contain not only a lot of noisy information but also some discriminative features in RGB-Thermal (RGBT) tracking. However, the potentials of low-quality modalities are not well explored in existing RGBT tracking algorithms. In this work, we propose a novel duality-gated mutual condition network to fully exploit the discriminative information of all modalities while suppressing the effects of data noise. In specific, we design a mutual condition module, which takes the discriminative information of a modality as the condition to guide feature learning of target appearance in another modality. Such a module can effectively enhance target representations of all modalities even in the presence of low-quality modalities. To improve the quality of conditions and further reduce data noise, we propose a duality-gated mechanism and integrate it into the mutual condition module. To deal with the tracking failure caused by sudden camera motion, which often occurs in RGBT tracking, we design a resampling strategy based on optical flow. It does not increase much computational cost since we perform optical flow calculation only when the model prediction is unreliable and then execute resampling when the sudden camera motion is detected. Extensive experiments on four RGBT tracking benchmark datasets show that our method performs favorably against the state-of-the-art tracking algorithms. Andong Lu, Cun Qian, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Breaking Modality Gap in RGBT Tracking: Coupled Knowledge DistillationabstractModality gap between RGB and thermal infrared (TIR) images is a crucial issue but often overlooked in existing RGBT tracking methods. It can be observed that modality gap mainly lies in the image style difference. In this work, we propose a novel Coupled Knowledge Distillation framework called CKD, which pursues common styles of different modalities to break modality gap, for high performance RGBT tracking. In particular, we introduce two student networks and employ the style distillation loss to make their style features consistent as much as possible. Through alleviating the style difference of two student networks, we can break modality gap of different modalities well. However, the distillation of style features might harm to the content representations of two modalities in student networks. To handle this issue, we take original RGB and TIR networks as the teachers, and distill their content knowledge into two student networks respectively by the style-content orthogonal feature decoupling scheme. We couple the above two distillation processes in an online optimization framework to form new feature representations of RGB and thermal modalities without modality gap. In addition, we design a masked modeling strategy and a multi-modal candidate token elimination strategy into CKD to improve tracking robustness and efficiency respectively. Extensive experiments on five standard RGBT tracking datasets validate the effectiveness of the proposed method against state-of-the-art methods while achieving the fastest tracking speed of 96.4 FPS. Andong Lu, Jiacong Zhao, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001 |
ACM Multimedia | 1 |
| 2024 | Transformer RGBT Tracking With Spatio-Temporal Multimodal TokensabstractMany RGBT tracking researches primarily focus on modal fusion design, while overlooking the effective handling of target appearance changes. While some approaches have introduced historical frames or fuse and replace initial templates to incorporate temporal information, they have the risk of disrupting the original target appearance and accumulating errors over time. To alleviate these limitations, we propose a novel Transformer RGBT tracking approach, which mixes spatio-temporal multimodal tokens from the static multimodal templates and multimodal search regions in Transformer to handle target appearance changes, for robust RGBT tracking. We introduce independent dynamic template tokens to interact with the search region, embedding temporal information to address appearance changes, while also retaining the involvement of the initial static template tokens in the joint feature extraction process to ensure the preservation of the original reliable target appearance information that prevent deviations from the target appearance caused by traditional temporal updates. We also use attention mechanisms to enhance the target features of multimodal template tokens by incorporating supplementary modal cues, and make the multimodal search region tokens interact with multimodal dynamic template tokens via attention mechanisms, which facilitates the conveyance of multimodal-enhanced target change information. Our module is inserted into the transformer backbone network and inherits joint feature extraction, search-template matching, and cross-modal interaction. Extensive experiments on three RGBT benchmark datasets show that the proposed approach maintains competitive performance compared to other state-of-the-art tracking algorithms while running at 39.1 FPS. The project-related materials are available at:https://github.com/yinghaidada/STMT. Dengdi Sun, Yajie Pan, Andong Lu, Chenglong Li 0002, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Illumination Distillation Framework for Nighttime Person Re-Identification and a New BenchmarkabstractNighttime person Re-ID (person re-identification in the nighttime) is a very important and challenging task for visual surveillance but it has not been thoroughly investigated. Under the low illumination condition, the performance of person Re-ID methods usually sharply deteriorates. To address the low illumination challenge in nighttime person Re-ID, this paper proposes an Illumination Distillation Framework (IDF), which utilizes illumination enhancement and illumination distillation schemes to promote the learning of Re-ID models. Specifically, IDF consists of a master branch, an illumination enhancement branch, and an illumination distillation module. The master branch is used to extract the features from a nighttime image. The illumination enhancement branch first estimates an enhanced image from the nighttime image using a nonlinear curve mapping method and then extracts the enhanced features. However, nighttime and enhanced features usually contain data noise due to unstable lighting conditions and enhancement failures. To fully exploit the complementary benefits of nighttime and enhanced features while suppressing data noise, we propose an illumination distillation module. In particular, the illumination distillation module fuses the features from two branches through a bottleneck fusion model and then uses the fused features to guide the learning of both branches in a distillation manner. In addition, we build a real-world nighttime person Re-ID dataset, namedNight600, which contains 600 identities captured from different viewpoints and nighttime illumination conditions under complex outdoor environments. Experimental results demonstrate that our IDF can achieve state-of-the-art performance on two nighttime person Re-ID datasets (i.e.,Night600andKnight). We will release our code and dataset athttps://github.com/Alexadlu/IDF. Andong Lu, Zhang Zhang 0001, Yan Huang 0023, Yifan Zhang 0004, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Fusion Tree Network for RGBT TrackingabstractRGBT tracking is often affected by complex scenes (i.e., occlusions, scale changes, noisy background, etc). Existing works usually adopt a single-strategy RGBT tracking fusion scheme to handle modality fusion in all scenarios. However, due to the limitation of fusion model capacity, it is difficult to fully integrate the discriminative features between different modalities. To tackle this problem, we propose a Fusion Tree Network (FTNet), which provides a multi-strategy fusion model with high capacity to efficiently fuse different modalities. Specifically, we combine three kinds of attention modules (i.e., channel attention, spatial attention, and location attention) in a tree structure to achieve multi-path hybrid attention in the deeper convolutional stages of the object tracking network. Extensive experiments are performed on three RGBT tracking datasets, and the results show that our method achieves superior performance among state-of-the-art RGBT tracking models. Zhiyuan Cheng 0012, Andong Lu, Zhang Zhang 0001, Chenglong Li 0002, Liang Wang 0001 |
AVSS | 2 |
| 2022 | Dynamic Collaboration Convolution for Robust RGBT TrackingabstractLearning powerful representation of individual modality is critical for RGBT tracking. Recent works mainly focus on utilizing multiple convolutions to model feature representations of each modality. However, they usually leverage static convolutions to extract features, which are hard to handle complex input data. To deal with this problem, we propose a dynamic collaboration convolution, named DC-Conv, including a set of static convolutions and a weight-router module, for robust RGBT tracking. In specific, we set four static convolutions to each modality in every layer to model each modality, and design a weight-router module to fuse these static convolutions using learned dynamic weights. Such a dynamic weighting scheme makes the convolutions can be adapted to the variations of input data, and thus greatly improves the tracking performance. In addition, we propose an effective progressive learning algorithm to maximize the role of each convolution to make it capture discriminative representations. We evaluate our method on two public RGBT tracking benchmarks, and the results demonstrate the effectiveness of our tracker against state-of-the-art methods. Andong Lu, Chenglong Li 0002, Yan Huang 0023, Liang Wang 0001 |
ICPR | 2 |
| 2022 | Joint Token and Feature Alignment Framework for Text-Based Person SearchabstractText-based person search is a challenging crossmodal retrieval task. Existing works reduce the inter-modality and intra-class gaps by aligning local features extracted from image and text modalities, which easily lead to mismatching problems due to the lack of annotation information. Besides, it is sub-optimal to reduce two gaps simultaneously in the same feature space. This work proposes a novel joint token and feature alignment framework to reduce the inter-modality and intraclass gaps progressively. Specifically, we first build a dual-path feature learning network to extract features and conduct feature alignment to reduce the inter-modality gap. Second, we design a text generation module to generate token sequences using visual features, and then token alignment is performed to reduce the intra-class gap. Last, a fusion interaction module is introduced to further eliminate the modality heterogeneity using the strategy of multi-stage feature fusion. Extensive experiments on the CUHKPEDES dataset demonstrate the effectiveness of our model, which significantly outperforms previous state-of-the-art methods. Shangze Li, Andong Lu, Yan Huang 0008, Chenglong Li 0002, Liang Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | RGBT Tracking via Multi-Adapter Network with Hierarchical Divergence LossabstractRGBT tracking has attracted increasing attention since RGB and thermal infrared data have strong complementary advantages, which could make trackers all-day and all-weather work. Existing works usually focus on extracting modality-shared or modality-specific information, but the potentials of these two cues are not well explored and exploited in RGBT tracking. In this paper, we propose a novel multi-adapter network to jointly perform modality-shared, modality-specific and instance-aware target representation learning for RGBT tracking. To this end, we design three kinds of adapters within an end-to-end deep learning framework. In specific, we use the modified VGG-M as the generality adapter to extract the modality-shared target representations. To extract the modality-specific features while reducing the computational complexity, we design a modality adapter, which adds a small block to the generality adapter in each layer and each modality in a parallel manner. Such a design could learn multilevel modality-specific representations with a modest number of parameters as the vast majority of parameters are shared with the generality adapter. We also design instance adapter to capture the appearance properties and temporal variations of a certain target. Moreover, to enhance the shared and specific features, we employ the loss of multiple kernel maximum mean discrepancy to measure the distribution divergence of different modal features and integrate it into each layer for more robust representation learning. Extensive experiments on two RGBT tracking benchmark datasets demonstrate the outstanding performance of the proposed tracker against the state-of-the-art methods. Andong Lu, Chenglong Li 0002, Yuqing Yan, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Challenge-Aware RGBT Tracking
Chenglong Li 0002, Lei Liu 0049, Andong Lu, Jin Tang 0001 |
ECCV (22) | 3 |