EDBT 2026 Demo / reviewers in the wild / expert
Xiantao Hu
dblp:160/8016
· DBLP profile ↗
16ranked-venue papers
4as first author
15since 2021 · last 2026
0009-0007-1541-1717ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingabstractRGB-Thermal (RGBT) tracking aims to exploit visible and thermal infrared modalities for robust all-weather object tracking. However, existing RGBT trackers struggle to resolve modality discrepancies, which poses great challenges for robust feature representation. This limitation hinders effective cross-modal information propagation and fusion, which significantly reduces the tracking accuracy. To address this limitation, we propose a novel Contextual Aggregation with Deformable Alignment framework called CADTrack for RGBT Tracking. To be specific, we first deploy the Mamba-based Feature Interaction (MFI) that establishes efficient feature interaction via state space models. This interaction module can operate with linear complexity, reducing computational cost and improving feature discrimination. Then, we propose the Contextual Aggregation Module (CAM) that dynamically activates backbone layers through sparse gating based on the Mixture-of-Experts (MoE). This module can encode complementary contextual information from cross-layer features. Finally, we propose the Deformable Alignment Module (DAM) to integrate deformable sampling and temporal propagation, mitigating spatial misalignment and localization drift. With the above components, our CADTrack achieves robust and accurate tracking in complex scenarios. Extensive experiments on five RGBT tracking benchmarks verify the effectiveness of our proposed method. Hao Li 0101, Xiantao Hu, Wenning Hao, Dong Wang 0004, Huchuan Lu |
AAAI | 3 |
| 2026 | Motion-Aware Object Tracking via Motion and Geometry-Aware CuesabstractUnderstanding motion is essential for visual object tracking, especially in complex and dynamic scenarios. Yet, many existing methods rely on simplistic strategies such as template updates or temporal feature propagation, often overlooking the deeper modeling of motion information. To mitigate this limitation, we introduce a motion-aware spatio-temporal framework that enhances motion perception by explicitly matching motion patterns and modeling inter-frame motion relationships. Central to our design is a motion pattern dictionary, which encodes a diverse set of representative motion cues as learnable features. During tracking, features from the search region interact with the dictionary to retrieve the most relevant motion patterns, allowing the model to adapt to the current motion state. A dedicated decoder further incorporates temporal correlations to refine motion awareness. To complement motion modeling, we embed geometric cues into the search region features, which strengthens spatial perception, reduces ambiguity under occlusion, and improves foreground-background separation. Extensive evaluations on seven challenging benchmarks demonstrate the effectiveness of our design. In particular, MoDTrack_384 surpasses recent SOTA trackers on LaSOT by 1.2% in AUC, highlighting the benefits of motion pattern modeling and geometry-guided enhancement in mitigating tracking drift. Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Yufei Tan, Haiying Xia, Shuxiang Song 0001 |
AAAI | 4 |
| 2026 | Curriculum adaptation for one-stream RGB-T tracking
Xiantao Hu, Fansheng Zeng, Bineng Zhong 0001, Zhangyong Tang, Wenxuan Fang 0001, Jun Li 0027, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 1 |
| 2026 | Spiking pyramid wavelet transformation for high-efficient and low-energy image restoration
Chen Zhao 0002, Xiantao Hu, Rui Xie 0005, Jian Yang 0003, Ying Tai |
Pattern Recognit. | 2 |
| 2026 | Learning multi-scale spatial-frequency features for image denoising
Xu Zhao 0001, Chen Zhao 0002, Xiantao Hu, Hongliang Zhang 0002, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2026 | DVDPEC: Driving-Video Dehazing via Position Embedding-Based CodebookabstractDespite significant progress in real-world image dehazing, efficiently generating high-fidelity, haze-free videos (especially in driving scenarios) remains challenging. Existing methods generally extend image dehazing techniques to videos by employing pre-trained single image dehazing models for preprocessing followed by refinement stages. However, this disjointed two-stage process often leads to unrealistic textures and loss of detail, as it fails to leverage large amounts of high-quality images for prior learning and the subsequent refinement struggles to correct temporal inconsistencies across frames introduced in the first stage. To address these issues, we propose DVDPEC: a Driving Video Dehazing framework utilizing a Position Embedding-based (PE-based) Codebook and a novel Flow Selective Block (FSB). The PE-based codebook stores fine-grained, spatially aware textural information specific to driving videos and leverages implicit positional embeddings for precise, position-aware codebook matching. This enables accurate prior retrieval and improves dehazing results. The FSB aggregates information from adjacent frames by dynamically combining both image flow and prior flow, effectively mitigating flow estimation ambiguities caused by haze. It enhances information fusion across frames, leading to more coherent and visually appealing dehazed videos. Extensive experiments demonstrate that DVDPEC achieves state-of-the-art performance on real-world driving video dehazing tasks, significantly enhancing texture preservation and visual fidelity. Yu Zheng 0036, Wenxuan Fang 0001, Xiantao Hu, Junkai Fan, Jiangwei Weng, Jun Li 0027, Kai Zhang 0008, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Fast Adversarial Training With Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on VideosabstractAdversarial Training (AT) has been shown to significantly enhance adversarial robustness via a min-max optimization approach. However, its effectiveness in video recognition tasks is hampered by two main challenges. First, fast adversarial training for video models remains largely unexplored, which severely impedes its practical applications. Specifically, most video adversarial training methods are computationally costly, with long training times and high expenses. Second, existing methods struggle with the trade-off between clean accuracy and adversarial robustness. To address these challenges, we introduce Video Fast Adversarial Training with Weak-to-Strong consistency (VFAT-WS), the first fast adversarial training method for video data. Specifically, VFAT-WS incorporates the following key designs: First, it integrates a straightforward yet effective temporal frequency augmentation (TF-AUG), and its spatial-temporal enhanced form STF-AUG, along with Fast Gradient Sign Method (FGSM) to boost training efficiency and robustness. Second, it devises a weak-to-strong spatial-temporal consistency regularization, which seamlessly integrates the simple TF-AUG and the more complex STF-AUG. Leveraging the consistency regularization, it steers the learning process from simple to complex augmentations. Both of them work together to achieve a better trade-off between clean accuracy and robustness. Extensive experiments on UCF-101 and HMDB-51 with both CNN and Transformer-based models demonstrate that VFAT-WS achieves great improvements in adversarial robustness and corruption robustness, while accelerating training by nearly 490%. Songping Wang, Yueming Lyu, Xiantao Hu, Ziwen He, Wei Wang 0025, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Digital Twin-Enabled Mobility-Aware Cooperative Caching in Vehicular Edge ComputingabstractWith the advancement of vehicle-to-vehicle (V2V) ad hoc networks and wireless communication technologies, mobile edge caching has become a key enabler for enhancing network performance and user experience. However, traditional federated learning-based collaborative caching approaches in vehicular scenarios suffer from inadequate client selection mechanisms and limited prediction accuracy, which result in suboptimal cache hit ratios and increased content transmission latency. To address these challenges, we propose a Digital Twin-based Asynchronous Federated Learning-driven Predictive Edge Caching with Deep Reinforcement Learning (DAPR) framework. DAPR employs an intelligent client selection strategy based on asynchronous federated learning, which leverages mobility prediction and data quality assessment to avoid selecting highly mobile clients or clients with low-quality data, thereby significantly improving model convergence efficiency. In addition, we design a GRU-VAE prediction model that uses a Variational Autoencoder (VAE) to capture latent data distribution features and Gated Recurrent Units (GRUs) to model temporal dependencies, thereby substantially enhancing the accuracy of content request prediction. The predicted content popularities are then fed into a deep reinforcement learning-driven caching decision engine to dynamically optimize edge caching resource allocation. Extensive experiments demonstrate that DAPR achieves superior performance in terms of average reward, cache hit ratio, and transmission latency, thereby effectively improving the overall efficiency of vehicular edge caching systems. Zhenkui Shi, Chunpei Li, Mengkai Yan, Hongliang Zhang 0002, Xiantao Hu, Xianxian Li |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingabstractMultimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Xiantao Hu, Ying Tai, Xu Zhao 0001, Chen Zhao 0002, Zhenyu Zhang 0005, Jun Li 0027, Bineng Zhong 0001, Jian Yang 0003 |
AAAI | 1 |
| 2025 | Explicit Context Reasoning with Supervision for Visual TrackingabstractContextual reasoning with constraints is crucial for enhancing temporal consistency in cross-frame modeling for visual tracking. However, mainstream tracking algorithms typically associate context by merely stacking historical information without explicitly supervising the association process, making it difficult to effectively model the target's evolving dynamics. To alleviate this problem, we propose RSTrack, which explicitly models and supervises context reasoning via three core mechanisms. 1) Context Reasoning Mechanism : Constructs a target state reasoning pipeline, converting unconstrained contextual associations into a temporal reasoning process that predicts the current representation based on historical target states, thereby enhancing temporal consistency. 2) Forward Supervision Strategy : Utilizes true target features as anchors to constrain the reasoning pipeline, guiding the predicted output toward the true target distribution and suppressing drift in the context reasoning process. 3) Efficient State Modeling : Employs a compression-reconstruction mechanism to extract the core features of the target, removing redundant information across frames and preventing ineffective contextual associations. These three mechanisms collaborate to effectively alleviate the issue of contextual association divergence in traditional temporal modeling. Experimental results show that RSTrack achieves state-of-the-art performance on multiple benchmark datasets while maintaining real-time running speeds. Our code is available at https://github.com/GXNU-ZhongLab/RSTrack. Fansheng Zeng, Bineng Zhong 0001, Haiying Xia, Yufei Tan, Xiantao Hu, Liangtao Shi, Shuxiang Song 0001 |
ACM Multimedia | 5 |
| 2025 | Mamba Adapter: Efficient Multi-Modal Fusion for Vision-Language TrackingabstractUtilizing the high-level semantic information of language to compensate for the limitations of vision information is a highly regarded approach in single-object tracking. However, most existing vision-language (VL) trackers employ full-parameter fine-tuning, which can easily lead to catastrophic forgetting. Therefore, they fail to fully exploit the prior knowledge of pre-trained models from upstream tasks, resulting in unsatisfactory tracking performance. To alleviate the above problem, we propose a simple yet effective Vision-Language Tracking pipeline based on Mamba Adapter, named MAVLT, which adopts the idea of parameter-efficient fine-tuning (PEFT) to realize the interaction between vision-language modalities. This novel approach offers the following advantages: (1)The knowledge of the upstream pre-trained model is efficiently inherited by freezing its parameters. This ensures that the VL tracking framework only learns the modules for vision and language interaction, with a focus on the fusion between modalities. (2)The modal interaction between language and vision encoders is flexibly bridged in each encoder layer via proposed mamba adapter, enabling efficient interaction of visual and language information at multiple levels. Extensive experiments on five popular vision-language tracking benchmarks validate the effectiveness of the proposed MAVLT. Particularly, the MAVLT achieves 73.4% AUC score on the LaSOT benchmarks with only 0.18%(0.32M) of the total parameters updates. Code and models are available at https://github.com/GXNU-ZhongLab/MAVLT. Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Zhiyi Mo, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | SwimVG: Step-Wise Multimodal Fusion and Adaption for Visual GroundingabstractVisual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge separately by fully fine-tuning uni-modal pre-trained models, followed by a simple stack of visual-language transformers for multimodal fusion. However, these approaches not only limit adequate interaction between visual and linguistic contexts, but also incur significant computational costs. Therefore, to address these issues, we explore a step-wise multimodal fusion and adaption framework, namely SwimVG. Specifically, SwimVG proposes step-wise multimodal prompts (Swip) and cross-modal interactive adapters (CIA) for visual grounding, replacing the cumbersome transformer stacks for multimodal fusion. Swip can improve the alignment between the vision and language representations step by step, in a token-level fusion manner. In addition, weight-level CIA further promotes multimodal fusion by cross-modal interaction. Swip and CIA are both parameter-efficient paradigms, and they fuse the cross-modal features from shallow to deep layers gradually. Experimental results on four widely-used benchmarks demonstrate that SwimVG achieves remarkable abilities and considerable benefits in terms of efficiency. Liangtao Shi, Ting Liu 0018, Xiantao Hu, Yue Hu 0016, Quanjun Yin, Richang Hong |
IEEE Trans. Multim. | 3 |
| 2024 | Personalized federated learning based on feature fusionabstractFederated learning (FL) enables distributed clients to collaborate on training while storing their data locally to protect client privacy. However, due to data heterogeneity, including issues related to label distributions skew in heterogeneous scenarios, the resulting global model may not be suitable for all clients. In this work, we introduce a personalized federated learning method called pFedPM, which focuses on addressing this challenge of label distributions skew in heterogeneous scenarios. We replace traditional gradient uploading with feature uploading, and introduce a novel feature fusion scheme to learn personalized local model for clients. Specifically, the server receives feature information from clients, aggregates global features, and sends them back to the clients. Clients achieve personalization by fusing local and global features. Furthermore, we introduce a relation network as an additional decision layer, providing a non-linear learnable classifier to predict labels. Through the novel modeling techniques, our proposed method reduces communication costs and supports heterogeneous client models. Experimental results demonstrate that our approach outperforms recent FL methods on the MNIST, FEMNIST, and CIFAR-10 datasets while requiring less communication. Wolong Xing, Zhenkui Shi, Hongyan Peng, Xiantao Hu, Yaozong Zheng, Xianxian Li |
CSCWD | 4 |
| 2024 | Toward Modalities Correlation for RGB-T TrackingabstractRecently, RGB-T tracking methods have made significant progress, demonstrating remarkable capabilities in addressing the complexities of tracking tasks within demanding environments. However, these methods overlook instability of modal validity in real-world scenarios. This limits the model’s ability to understand the correlation between modalities, thereby hindering the model’s ability to fully leverage the synergistic effects of RGB and TIR. To address this challenge, we propose a novel RGB-T tracking model named MCTrack, from the perspective of leveraging correlation among modalities. First, during the feature extraction stage, we design a novel module based on channel matching modeling to construct bidirectional channel context information flow for two modalities. By leveraging information flow, specific modalities correlation information can be transmitted to two modes, augmenting the correlation between the two modes adaptively. Subsequently, after the feature extraction network, the features of each modality are decoded and transformed to generate more correlated feature representations. During this stage, we extract distinctive and collective features by leveraging the correlation among modalities. Then fusing these features and generated search region features specifically for localization. This aids the model in comprehending the correlation between RGB and TIR under complex scenarios, thereby enhancing its ability to capture and utilize key features. Based on extensive experiments conducted on four popular RGB-T tracking benchmarks, our model demonstrates superior performance, particularly showcasing impressive results on the LasHeR dataset with an achieved Precision of 71.6%. Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Transformer Tracking via Frequency FusionabstractTransformer has achieved impressive progress in visual tracking due to their capability of global modeling, which enables them to learn low-frequency features(i.e., high-level semantic information). However, it seems to overlook the high-frequency features(i.e., low-level texture and edge information) which are crucial to identify different intra-class object instances in the tracking task. To address this issue, we propose a transformer based tracker via frequency fusion perspective that investigated whether high-frequency and low-frequency features can be effectively combined to achieve robust tracking. Specifically, we design a simple yet effective two-stage fusion strategy and use an appropriate frequency fusion strategy in tracking process of each stage so as to make full use of frequency domain information. In the feature extraction stage, we use wavelet decomposition of high-frequency subbands to solve the performance loss caused by the transformer’s catastrophic forgetting of high-frequency information. In the prediction head stage, we use a variety of wavelet decomposition subbands to model the multi-frequency information. The two-stage fusion strategy makes our model extract more balanced and beneficial multi-frequency information, enabling it to effectively capture target texture information and local edge information while also being sensitive to global information. Extensive experiments on six challenging benchmarks (i.e., LaSOT$_{ext}$, UAV123, TNL2K, LaSOT, TrackingNet, and GOT-10k) demonstrates the superior performance of our tracker. Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Adaptive almost sure asymptotically stability for neutral-type neural networks with stochastic perturbation and Markovian switching
Liuwei Zhou, Zhijie Wang 0001, Xiantao Hu, Bo Chu, Wuneng Zhou |
Neurocomputing | 3 |