VLDB 2026 Research / reviewers in the wild / expert
Hui Luo 0002
dblp:06/890-2
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-6698-5576ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosabstractMulti-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos, frequent viewpoint changes and complex UAV-ground relative motion dynamics pose significant challenges, which often lead to unstable affinity measurement and ambiguous association. Existing methods typically model motion and appearance cues separately, overlooking their spatio-temporal interplay and resulting in suboptimal tracking performance. In this work, we propose AMOT, which jointly exploits appearance and motion cues through two key components: an Appearance-Motion Consistency (AMC) matrix and a Motion-aware Track Continuation (MTC) module. Specifically, the AMC matrix computes bi-directional spatial consistency under the guidance of appearance features, enabling more reliable and context-aware identity association. The MTC module complements AMC by reactivating unmatched tracks through appearance-guided predictions that align with Kalman-based predictions, thereby reducing broken trajectories caused by missed detections. Extensive experiments on three UAV benchmarks, including VisDrone2019, UAVDT, and VT-MOT-UAV, demonstrate that our AMOT outperforms current state-of-the-art methods and generalizes well in a plug-and-play and training-free manner. Jianbo Ma 0001, Hui Luo 0002, Qi Chen 0014, Yuankai Qi, Yumei Sun, Amin Beheshti, Jianlin Zhang 0001, Ming-Hsuan Yang 0001 |
AAAI | 2 |
| 2026 | Multi-Stage Cross-Modality Feature Interaction for RGB-Thermal Multi-Object TrackingabstractRGB-Thermal multi-object tracking (RGB-T MOT) focuses on tracking multiple objects in complex scenarios, such as nighttime and low-light conditions, which is crucial for various applications, including video surveillance and unmanned aerial vehicle monitoring. Existing RGB-T MOT studies typically merge multi-source features in a single stage before the backbone network. These works fail to preserve fine-grained details and struggle with context-dependent fusion, resulting in suboptimal tracking performance. In this paper, we propose a novel multi-stage cross-modality spatial-temporal feature interaction network (MCTrack), which emphasizes two main aspects: temporal-aware learning and cross-modality information interaction. For temporal-aware learning, we introduce the temporal salient feature interaction (TSFI) module, which ensures trajectory continuity by capturing dynamic spatial variation of objects across RGB-T video pairs, enabling consistent object tracking over time. For cross-modality information interaction, we propose the bidirectional modality interaction (Bi-MI) module, which employs cross-modality transformers to extract complementary features from both RGB and thermal modalities at multiple stages of feature extraction, thereby improving tracking adaptability in diverse tracking scenarios. Additionally, we propose the cross-modality complementary mask (CCM) strategy, which applies non-overlapping random masks to feature maps to improve the robustness of cross-modality feature interaction. Extensive experiments and comparative analyses demonstrate that our MCTrack surpasses state-of-the-art trackers on both publicly available VT-MOT and UniRTL-MOT RGB-T datasets. The code is available at https://github.com/ydhcg-BoBo/RGB-T-MOT. Jianbo Ma 0001, Hui Luo 0002, Shuaicheng Niu, Peilin Zhao, Jianlin Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Test-Time Model Adaptation for Quantized Neural NetworksabstractQuantizing deep models prior to deployment is a widely adopted technique to speed up inference for various real-time applications, such as autonomous driving. However, quantized models often suffer from severe performance degradation in dynamic environments with potential domain shifts and this degradation is significantly more pronounced compared with their full-precision counterparts, as shown by our theoretical and empirical illustrations. To address the domain shift problem, test-time adaptation (TTA) has emerged as an effective solution by enabling models to learn adaptively from test data. Unfortunately, existing TTA methods are often impractical for quantized models as they typically rely on gradient backpropagation-an operation that is unsupported on quantized models due to vanishing gradients, as well as memory and latency constraints. In this paper, we focus on TTA for quantized models to improve their robustness and generalization ability efficiently. We propose a continual zeroth-order adaptation (ZOA) framework that enables efficient model adaptation using only two forward passes, eliminating the computational burden of existing methods. Moreover, we propose a domain knowledge management scheme to store and reuse different domain knowledge with negligible memory consumption, reducing the interference of different domain knowledge and fostering the knowledge accumulation during long-term adaptation. Experimental results on three classical architectures, including quantized transformer-based and CNN-based models, demonstrate the superiority of our methods for quantized model adaptation. On the quantized W6A6 ViT-B model, our ZOA is able to achieve a 5.0% improvement over the state-of-the-art FOA on ImageNet-C dataset. The source code is available at https://github.com/DengZeshuai/ZOA. Zeshuai Deng, Shuaicheng Niu, Hui Luo 0002, Shuhai Zhang, Wei Luo 0006, Mingkui Tan |
ACM Multimedia | 4 |
| 2025 | COLAFormer: Communicating local-global features with linear computational complexity
Zhengwei Miao, Hui Luo 0002, Meihui Li, Jianlin Zhang 0001 |
Pattern Recognit. | 2 |
| 2024 | Learning general features to bridge the cross-domain gaps in few-shot learning
Xiang Li 0212, Hui Luo 0002, Gaofan Zhou, Xiaoming Peng, Jianlin Zhang 0001, Meihui Li |
Knowl. Based Syst. | 2 |
| 2024 | Joint spatio-temporal modeling for visual tracking
Yumei Sun, Chuanming Tang, Hui Luo 0002, Xiaoming Peng, Jianlin Zhang 0001, Meihui Li |
Knowl. Based Syst. | 3 |
| 2024 | Improving Visual Representations of Masked Autoencoders With Artifacts SuppressionabstractRecently, Masked Autoencoders (MAE) have gained attention for their abilities to generate visual representations efficiently through pretext tasks. However, there has been little research evaluating the visual representations obtained by pre-trained MAE during the fine-tuning process. In this study, we address the gap by examining the attention maps within each block of the pre-trained MAE during the fine-tuning process. We observed artifacts in pre-trained models, which appear as significant responses in the attention maps of shallow blocks. These artifacts may negatively impact the transfer ability performance of MAE. To address this issue, we localize the cause of these artifacts to the asymmetry between the pre-training and fine-tuning processes. To suppress these artifacts, we propose a novel semantic masking strategy. This strategy aims to preserve complete and continuous semantic information within visible patches while maintaining randomness to facilitate robust representation learning. Experimental results demonstrate that the proposed masking strategy improves the performance of various downstream tasks while reducing artifacts. Specifically, we observed a 3.2% improvement in linear probing, a 0.5% enhancement in fine-tuning on Imagenet1K, and a 0.6% increase in semantic segmentation on ADE20K. Zhengwei Miao, Hui Luo 0002, Jianlin Zhang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Toward Compact and Robust Model Learning Under Dynamically Perturbed EnvironmentsabstractNetwork pruning has been widely studied to reduce the complexity of deep neural networks (DNNs) and hence speed up their inference. Unfortunately, most existing pruning methods ignore the changes in the model’s robustness before and after pruning, which makes pruned models vulnerable under dynamically perturbed environments (e.g., autonomous driving). Only a few works have explored the robustness of pruned models against adversarial attacks that significantly differ from perturbations in real-world scenarios. To bridge the gap between real-world applications and existing studies, in this work, we propose an adversarial pruning scheme, which automatically identifies and preserves robust channels to obtain robust pruned models that are suitable for practical deployment in dynamically perturbed environments. Specifically, to simulate real-world perturbations, we first employ multi-type adversarial attack samples and adversarial perturbation samples generated by an adversarial perturbation generator to create mixed noise samples. Then, we propose a plug-and-play feature scoring module and a novel contribution difference loss to evaluate the robustness of intermediate features dynamically. Next, to leverage robust intermediate features to identify robust channels, we have developed a simple but effective gating mechanism that evaluates the robustness of channels and preserves robust channels during training. Lastly, we compress the model in a layer-wise or block-wise manner. Compared to existing methods, our scheme enhances the robustness of the pruned model in a broader sense, making it better able to against dynamic perturbations in the real world. Extensive experimental results on well-known dataset benchmarks and popular network architectures demonstrate the effectiveness of our method. Hui Luo 0002, Zhuangwei Zhuang, Yuanqing Li 0001, Mingkui Tan, Cen Chen 0002, Jianlin Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Learning to Generate Diverse Data From a Temporal Perspective for Data-Free QuantizationabstractModel quantization is a prevalent method to compress and accelerate neural networks. Most existing quantization methods usually require access to real data to improve the performance of quantized models, which is often infeasible in some scenarios with privacy and security concerns. Recently, data-free quantization has been widely studied to solve the challenge of not having access to real data by generating synthetic data, among which generator-based data-free quantization is an important type. Previous generator-based methods focus on improving the performance of quantized models by optimizing the spatial distribution of synthetic data, while ignoring the study of changes in synthetic data from a temporal perspective. In this work, we reveal that generator-based data-free quantization methods usually suffer from the issue that synthetic data show homogeneity in the mid-to-late stages of the generation process due to the stagnation of the generator update, which hinders further improvement of the performance of the quantization model. To solve the above issue, we propose introducing the discrepancy between the full-precision and quantized models as new supervision information to update the generator. Specifically, we propose a simple yet effective adversarial Gaussian-margin loss, which promotes continuous updating of the generator by adding more supervision information to the generator when the discrepancy between the full-precision and quantized models is small, thereby generating heterogeneous synthetic data. Moreover, to mitigate the homogeneity of the synthetic data further, we augment the synthetic data with linear interpolation. Our proposed method can also promote the performance of other generator-based data-free quantization methods. Extensive experimental results show that our proposed method achieves superior performances for various settings on data-free quantization, especially in ultra-low-bit settings, such as 3-bit. Hui Luo 0002, Shuhai Zhang, Zhuangwei Zhuang, Jiajie Mai, Mingkui Tan, Jianlin Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |