VLDB 2026 Research / reviewers in the wild / expert
Shilei Wang 0001
dblp:139/8291-1
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-0507-7256ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
Shilei Wang 0001, Pujian Lai, Jifeng Ning, Gong Cheng 0003 |
AAAI | 1 |
| 2026 | Mining representative tokens via transformer-based multi-modal interaction for RGB-T tracking
Pujian Lai, Shilei Wang 0001, Gong Cheng 0003 |
Pattern Recognit. | 3 |
| 2026 | Cross-alignment for efficient visual object tracking
Shilei Wang 0001, Mingjiang Liang, Shaoli Huang, Jifeng Ning, Gong Cheng 0003 |
Pattern Recognit. | 1 |
| 2026 | LTSTrack: Visual tracking with long-term temporal sequence
Zhaochuan Zeng, Shilei Wang 0001, Yidong Song, Zhenhua Wang 0003, Jifeng Ning |
Pattern Recognit. | 2 |
| 2026 | Exploring Pruning-Based Efficient Object Tracking via Hybrid Knowledge Distillation
Yidong Song, Shilei Wang 0001, Zhaochuan Zeng, Jikai Zheng, Zhenhua Wang 0003, Jifeng Ning |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Inter - Diffusion Generation Model of Speakers and Listeners for Effective CommunicationabstractFull-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of listeners in the interaction process and failing to fully explore the dynamic interaction between them. This paper innovatively proposes an Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication. For the first time, we integrate the full-body gestures of listeners into the generation framework. By devising a novel inter-diffusion mechanism, this model can accurately capture the complex interaction patterns between speakers and listeners during communication. In the model construction process, based on the advanced diffusion model architecture, we innovatively introduce interaction conditions and the GAN model to increase the denoising step size. As a result, when generating gesture sequences, the model can not only dynamically generate based on the speaker's speech information but also respond in realtime to the listener's feedback, enabling synergistic interaction between the two. Abundant experimental results demonstrate that compared with the current state-of-the-art gesture generation methods, the model we proposed has achieved remarkable improvements in the naturalness, coherence, and speech-gesture synchronization of the generated gestures. In the subjective evaluation experiments, users highly praised the generated interaction scenarios, believing that they are closer to real life human communication situations. Objective index evaluations also show that our model outperforms the baseline methods in multiple key indicators, providing more powerful support for effective communication. Jinhe Huang, Yongkang Cheng, Minghang Yu, Gaoge Han, Jinwei Li 0003, Jing Zhang 0164, Shilei Wang 0001, Xingjian Gu |
ICMR | 7 |
| 2025 | Multi-State Tracker: Enhancing Efficient Object Tracking via Multi-State Specialization and InteractionabstractEfficient trackers achieve faster runtime by reducing computational complexity and model parameters. However, this efficiency often compromises the expense of weakened feature representation capacity, thus limiting their ability to accurately capture target states using single-layer features. To overcome this limitation, we propose Multi-State Tracker (MST), which utilizes highly lightweight state-specific enhancement (SSE) to perform specialized enhancement on multi-state features produced by multi-state generation (MSG) and aggregates them in an interactive and adaptive manner using cross-state interaction (CSI). This design greatly enhances feature representation while incurring minimal computational overhead, leading to improved tracking robustness in complex environments. Specifically, the MSG generates multiple state representations at multiple stages during feature extraction, while SSE refines them to highlight target-specific features. The CSI module facilitates information exchange between these states and ensures the integration of complementary features. Notably, the introduced SSE and CSI modules adopt a highly lightweight hidden state adaptation-based state space duality (HSA-SSD) design, incurring only 0.1 GFLOPs in computation and 0.66 M in parameters. Experimental results demonstrate that MST outperforms all previous efficient trackers across multiple datasets, significantly improving tracking accuracy and robustness. In particular, it shows excellent runtime performance, with an AO score improvement of 4.5% over the previous SOTA efficient tracker HCAT on the GOT-10K dataset. The code is available at https://github.com/wsumel/MST. Shilei Wang 0001, Gong Cheng 0003, Pujian Lai, Junwei Han 0001 |
ACM Multimedia | 1 |
| 2024 | IoUNet++: Spatial cross-layer interaction-based bounding box regression for visual trackingabstractAbstract Accurate target prediction, especially bounding box estimation, is a key problem in visual tracking. Many recently proposed trackers adopt the refinement module called IoU predictor by designing a high‐level modulation vector to achieve bounding box estimation. However, due to the lack of spatial information that is important for precise box estimation, this simple one‐dimensional modulation vector has limited refinement representation capability. In this study, a novel IoU predictor (IoUNet++) is designed to achieve more accurate bounding box estimation by investigating spatial matching with a spatial cross‐layer interaction model. Rather than using a one‐dimensional modulation vector to generate representations of the candidate bounding box for overlap prediction, this paper first extracts and fuses multi‐level features of the target to generate template kernel with spatial description capability. Then, when aggregating the features of the template and the search region, the depthwise separable convolution correlation is adopted to preserve the spatial matching between the target feature and candidate feature, which makes their IoUNet++ network have better template representation and better fusion than the original network. The proposed IoUNet++ method with a plug‐and‐play style is applied to a series of strengthened trackers including DiMP++, SuperDiMP++ and SuperDIMP_AR++, which achieve consistent performance gain. Finally, experiments conducted on six popular tracking benchmarks show that their trackers outperformed the state‐of‐the‐art trackers with significantly fewer training epochs. Shilei Wang 0001, Baozhen Sun, Jifeng Ning |
IET Comput. Vis. | 1 |
| 2024 | Bidirectional Interaction of CNN and Transformer Feature for Visual TrackingabstractEmpowered by the sophisticated long-range dependency modeling ability of Transformer, tracking performance has seen a dynamic increase in recent years. Approaches in this vein leverage the Transformer feature to integrate the information of target and search regions while neglecting the superior local representation extracted by their CNN backbone. To address this, we introduce a BIdirectional inTeraction mechanism between CNN and Transformer features for visual tracking, termed BIT-Tracker, which admits a comprehensive fusion of local and global representations, and thus boosts tracking performance. The first ingredient of BIT-Tracker is an aggregation of multi-level Transformer features to achieve a better global modeling ability. In order to combine the merits of both local and global representations, our second ingredient performs a bi-directional interaction between CNN and Transformer features, where the interaction is achieved via either querying the CNN feature from the Transformer feature or querying the Transformer feature from the CNN feature. Afterwards, the outputs from both directions are fused to predict the temporal locations of targets. Extensive experiments demonstrate the effectiveness of the proposed feature aggregation and bi-directional interaction modules. Impressively, BIT-Tracker achieves leading performance on eight tracking benchmarks and outperforms SOTA results by salient margins. Code will be made available. Baozhen Sun, Zhenhua Wang 0003, Shilei Wang 0001, Yongkang Cheng, Jifeng Ning |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Modeling of Multiple Spatial-Temporal Relations for Robust Visual Object TrackingabstractRecently, one-stream trackers have achieved parallel feature extraction and relation modeling through the exploitation of Transformer-based architectures. This design greatly improves the performance of trackers. However, as one-stream trackers often overlook crucial tracking cues beyond the template, they prone to give unsatisfactory results against complex tracking scenarios. To tackle these challenges, we propose a multi-cue single-stream tracker, dubbed MCTrack here, which seamlessly integrates template information, historical trajectory, historical frame, and the search region for synchronized feature extraction and relation modeling. To achieve this, we employ two types of encoders to convert the template, historical frames, search region, and historical trajectory into tokens, which are then collectively fed into a Transformer architecture. To distill temporal and spatial cues, we introduce a novel adaptive update mechanism, which incorporates a thresholding component and a local multi-peak component to filter out less accurate and overly disturbed tracking cues. Empirically, MCTrack achieves leading performance on mainstream benchmark datasets, surpassing the most advanced SeqTrack by 2.0% in terms of the AO metric on GOT-10k. The code is available at https://github.com/wsumel/MCTrack. Shilei Wang 0001, Zhenhua Wang 0003, Qianqian Sun, Gong Cheng 0003, Jifeng Ning |
IEEE Trans. Image Process. | 1 |
| 2021 | Do We Really Need Frame-by-Frame Annotation Datasets for Object Tracking?abstractThere has been an increasing emphasis on building large-scale datasets as the driver of deep learning-based trackers' success. However, accurately annotating tracking data is highly labor-intensive and expensive, making it infeasible in real-world applications. In this study, we investigate the necessity of large-scale training data to ensure tracking algorithms' performance. To this end, we introduce a FAT (Few-Annotation Tracking) benchmark constructed by sampling one or a few frames per video from some existing tracking datasets. The proposed dataset can be used to evaluate the effectiveness of tracking algorithms considering data efficiency and new data augmentation approaches for object tracking. We further present AMMC (Augmentation by Mimicking Motion Change), a data augmentation strategy that enables learning high-performing trackers using small-scale datasets. AMMC first cuts out the tracked targets and performs a sequence of transformations to simulate the possible change by object motion. Then the transformed targets are pasted on the inpainted background images and further conjointly augmented to mimic variability caused by camera motion. Compared with standard augmentation methods, AMMC explicitly considers tracking data characteristics, which synthesizes more valid data for object tracking. We extensively evaluate our approach with two popular trackers on the FAT datasets. Experiments show that our method allows these trackers to even trained on a dataset requiring much less annotation to achieve comparable or even better performance to those on the full-annotation dataset. The results imply complete video annotation might not be necessary for object tracking if leveraging motion-driven data augmentations during training. Shaoli Huang, Shilei Wang 0001, Wei Liu 0007, Jifeng Ning |
ACM Multimedia | 3 |