VLDB 2026 Research / reviewers in the wild / expert
Xi Zhou 0001
dblp:42/5705-1
· DBLP profile ↗
21ranked-venue papers
0as first author
19since 2021 · last 2025
0000-0001-9943-5482ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 14 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MambaTrack: Exploiting Dual-Enhancement for Night UAV TrackingabstractNight unmanned aerial vehicle (UAV) tracking is impeded by the challenges of poor illumination, with previous daylight-optimized methods demonstrating suboptimal performance in low-light conditions, limiting the utility of UAV applications. To this end, we propose an efficient mamba-based tracker, leveraging dual enhancement techniques to boost night UAV tracking. The mamba-based low-light enhancer, equipped with an illumination estimator and a damage restorer, achieves global image enhancement while preserving the details and structure of low-light images. Additionally, we advance a cross-modal mamba network to achieve efficient interactive learning between vision and language modalities. Extensive experiments showcase that our method achieves advanced performance and exhibits significantly improved computation and memory efficiency. For instance, our method is 2.8× faster than CiteTracker and reduces 50.2% GPU memory. Our codes are available at https://github.com/983632847/Awesome-Multimodal-Object-Tracking. Chunhui Zhang 0001, Li Liu 0036, Xi Zhou 0001, Yanfeng Wang 0001 |
ICASSP | 4 |
| 2025 | Boosting Nighttime UAV Tracking via Self-prompting Autoregressive Learning and a New Benchmark
Chunhui Zhang 0001, Li Liu 0036, Xi Zhou 0001, Yanfeng Wang 0001 |
PRCV (16) | 4 |
| 2024 | WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale BenchmarkabstractUnderwater Object Tracking (UOT) is essential for identifying and tracking submerged objects in underwater videos, but existing datasets are limited in scale, diversity of target categories and scenarios covered, impeding the development of advanced tracking algorithms. To bridge this gap, we take the first step and introduce WebUOT-1M, \ie, the largest public UOT benchmark to date, sourced from complex and realistic underwater environments. It comprises 1.1 million frames across 1,500 video clips filtered from 408 target categories, largely surpassing previous UOT datasets, \eg, UVOT400. Through meticulous manual annotation and verification, we provide high-quality bounding boxes for underwater targets. Additionally, WebUOT-1M includes language prompts for video sequences, expanding its application areas, \eg, underwater vision-language tracking. Given that most existing trackers are designed for open-air conditions and perform poorly in underwater environments due to domain gaps, we propose a novel framework that uses omni-knowledge distillation to train a student Transformer model effectively. To the best of our knowledge, this framework is the first to effectively transfer open-air domain knowledge to the UOT model through knowledge distillation, as demonstrated by results on both existing UOT datasets and the newly proposed WebUOT-1M. We have thoroughly tested WebUOT-1M with 30 deep trackers, showcasing its potential as a benchmark for future UOT research. The complete dataset, along with codes and tracking results, are publicly accessible at \href{https://github.com/983632847/Awesome-Multimodal-Object-Tracking}{\color{magenta}{here}}. Chunhui Zhang 0001, Li Liu 0036, Guanjie Huang, Xi Zhou 0001, Yanfeng Wang 0001 |
NeurIPS | 5 |
| 2024 | High-compressed deepfake video detection with contrastive spatiotemporal distillation
Yizhe Zhu, Chunhui Zhang 0001, Jialin Gao, Xin Sun 0020, Zihan Rui, Xi Zhou 0001 |
Neurocomputing | 6 |
| 2024 | Point Spatio-Temporal Pyramid Network for Point Cloud Video UnderstandingabstractThe robustness to spatio-temporal sampling is significant for point cloud video understanding. Previous works overlook this issue and usually suffer notable performance drops when point densities and frame rates are changed. To remedy this, we propose a point spatio-temporal pyramid (PoST-Py) to improve the sampling robustness of point cloud video modeling. Specifically, we propose a pluggable PoST-Py to collect multi-scale feature maps from different layers of the backbone. Then, these features are integrated into a unified representation. This allows the model to capture multi-scale spatio-temporal information simultaneously. In addition, we employ the temporal cardinality difference to enhance the features to capture motion information. Extensive experiments show that PoST-Py achieves state-of-the-art performance, particularly with a notable improvement of over 2% under varying point sampling. This demonstrates the improved robustness of our method. The code is available athttps://github.com/JohnsonSign/PoST-Py. Longguang Wang, Yulan Guo, Xi Zhou 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Multi-Level Signal Fusion for Enhanced Weakly-Supervised Audio-Visual Video ParsingabstractThe weakly-supervised audio-visual video parsing (AVVP) task aims toparse a video into temporal events and predict their modality-specific categories. Current works primarily focus on refining training strategies and follow the framework fusing signals only at the segment level. However, they miss the point that video events, being composed of consecutive segments, require the integration of both local and global contexts to fully capture their essence. In this letter, we present theLocal-GlobalFusionNetwork (LGFNet), designed to facilitate multi-level interaction between audio and visual signals. Specifically, we create a two-dimensional map to generate multi-scale event proposals for both audio and visual modalities. Subsequently, we fuse audio and visual signals at both segment and event levels with a novel boundary-aware feature aggregation method, enabling the simultaneous capture of local and global information. To enhance the temporal alignment between the two modalities, we employ segment-level and event-level contrastive learning. In-depth experiments demonstrate the superiority of our LGFNet. Xin Sun 0020, Xi Zhou 0001 |
IEEE Signal Process. Lett. | 4 |
| 2023 | PointCMP: Contrastive Mask Prediction for Self-supervised Learning on Point Cloud VideosabstractSelf-supervised learning can extract representations of good quality from solely unlabeled data, which is ap-pealing for point cloud videos due to their high labelling cost. In this paper, we propose a contrastive mask prediction (PointCMP) framework for self-supervised learning on point cloud videos. Specifically, our PointCMP employs a two-branch structure to achieve simultaneous learning of both local and global spatiotemporal information. On top of this two-branch structure, a mutual similarity based augmentation module is developed to synthesize hard samples at the feature level. By masking dominant tokens and erasing principal channels, we generate hard samples to facilitate learning representations with better discrimi-nation and generalization performance. Extensive experiments show that our PointCMP achieves the state-of-the-art performance on benchmark datasets and outperforms existing full-supervised counterparts. Transfer learning results demonstrate the superiority of the learned representations across different datasets and tasks. Xiaoxiao Sheng, Longguang Wang, Yulan Guo, Xi Zhou 0001 |
CVPR | 6 |
| 2023 | Exploiting Multi-modal Fusion for Robust Face Representation Learning with Missing Modality
Yizhe Zhu, Xin Sun 0020, Xi Zhou 0001 |
ICANN (2) | 3 |
| 2023 | Audio-Driven Talking Head Video Generation with Diffusion ModelabstractSynthesizing high-fidelity talking head videos by fitting input audio sequences is a highly anticipated technique in many applications, such as digital humans, virtual video conferences, and human-computer interaction. Popular GAN-based methods aim to align speech audio with lip motions and head poses. However, existing methods are prone to training instability and even mode collapse, resulting in low-quality video generation. In this paper, we propose a novel audio-driven diffusion method for generating high-resolution realistic videos of talking heads with the help of the denoising diffusion model. Specifically, the face attribute disentanglement module is proposed to disentangle eye blinking and lip motion features, where the lip motion features are synchronized with audio features via the contrastive learning strategy, and the disentangled motion features are aligned well with the talking head. Furthermore, the denoising diffusion model takes the source image and the warped motion features as input to generate the high-resolution realistic talking head with diverse head poses. Extensive evaluations using multiple metrics demonstrate that our method outperforms the current techniques both qualitatively and quantitatively. Yizhe Zhu, Chunhui Zhang 0001, Xi Zhou 0001 |
ICASSP | 4 |
| 2023 | Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud VideosabstractRecently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually notoriously expensive. Moreover, training via one or only a few traditional tasks (e.g., classification) may be insufficient to learn subtle details of the spatio-temporal structure existing in point cloud videos. In this paper, we propose a Masked Spatio-Temporal Structure Prediction (MaST-Pre) method to capture the structure of point cloud videos without human annotations. MaST-Pre is based on spatio-temporal point-tube masking and consists of two self-supervised learning tasks. First, by reconstructing masked point tubes, our method is able to capture the appearance information of point cloud videos. Second, to learn motion, we propose a temporal cardinality difference prediction task that estimates the change in the number of points within a point tube. In this way, MaST-Pre is forced to model the spatial and temporal structure in point cloud videos. Extensive experiments on MSRAction-3D, NTU-RGBD, NvGesture, and SHREC’17 demonstrate the effectiveness of the proposed method. The code is available at https://github.com/JohnsonSign/MaST-Pre. Xiaoxiao Sheng, Hehe Fan, Longguang Wang, Yulan Guo, Xi Zhou 0001 |
ICCV | 8 |
| 2023 | AVForensics: Audio-driven Deepfake Video Detection with Masking Strategy in Self-supervisionabstractExisting cross-dataset deepfake detection approaches exploit mouth-related mismatches between the auditory and visual modalities in fake videos to enhance generalisation to unseen forgeries. However, such methods inevitably suffer performance degradation with limited or unaltered mouth motions, we argue that face forgery detection consistently benefits from using high-level cues across the whole face region. In this paper, we propose a two-phase audio-driven multi-modal transformer-based framework, termed AVForensics, to perform deepfake video content detection from an audio-visual matching view related to full face. In the first pre-training phase, we apply the novel uniform masking strategy to model global facial features and learn temporally dense video representations in a self-supervised cross-modal manner, by capturing the natural correspondence between the visual and auditory modalities regardless of large-scaled labelled data and heavy memory usage. Then we use these learned representations to fine-tune for the down-stream deepfake detection task in the second phase, which encourages the model to offer accurate predictions based on captured global facial movement features. Extensive experiments and visualizations on various public datasets demonstrate the superiority of our self-supervised pre-trained method for achieving generalisable and robust deepfake video detection. Yizhe Zhu, Jialin Gao, Xi Zhou 0001 |
ICMR | 3 |
| 2023 | All in One: Exploring Unified Vision-Language Tracking with Multi-Modal AlignmentabstractCurrent mainstream vision-language (VL) tracking framework consists of three parts,i.e., a visual feature extractor, a language feature extractor, and a fusion model. To pursue better performance, a natural modus operandi for VL tracking is employing customized and heavier unimodal encoders, and multi-modal fusion models. Albeit effective, existing VL trackers separate feature extraction and feature integration, resulting in extracted features that lack semantic guidance and have limited target-aware capability in complex scenarios, e.g., similar distractors and extreme illumination. In this work, inspired by the recent success of exploring foundation models with unified architecture for both natural language and computer vision tasks, we propose an All-in-One framework, which learns joint feature extraction and interaction by adopting a unified transformer backbone. Specifically, we mix raw vision and language signals to generate language-injected vision tokens, which we then concatenate before feeding into the unified backbone architecture. This approach achieves feature integration in a unified backbone, removing the need for carefully-designed fusion modules and resulting in a more effective and efficient VL tracking framework. To further improve the learning efficiency, we introduce a multi-modal alignment module based on cross-modal and intra-modal contrastive objectives, providing more reasonable representations for the unified All-in-One transformer backbone. Extensive experiments on five benchmarks, i.e., OTB99-L, TNL2K, LaSOT, LaSOTExt and WebUAV-3M, demonstrate the superiority of the proposed tracker against existing state-of-the-art (SOTA) methods on VL tracking. Codes will be available at https://github.com/983632847/All-in-One here. Chunhui Zhang 0001, Xin Sun 0020, Yiqian Yang, Li Liu 0036, Xi Zhou 0001, Yanfeng Wang 0001 |
ACM Multimedia | 6 |
| 2023 | Video Moment Retrieval via Comprehensive Relation-Aware NetworkabstractVideo moment retrieval aims to retrieve a target moment from an untrimmed video that semantically corresponds to the given language query. Existing methods commonly treat it as a regression task or a ranking task from the perspective of computer vision. Most of these works neglect comprehensive relations between video content and language context at a multi-granularity level and fail to efficiently model temporal relations among different video moments. In this paper, we formulate video moment retrieval into video reading comprehension by treating the input video as a text passage and language query as a question. To tackle the above impediments, we propose a Comprehensive Relation-aware Network (CRNet) to perceive comprehensive relations from extensive aspects. Specifically, we unite visual and textual features simultaneously at both clip-level and moment-level to thoroughly exploit inter-modality information, leading to a coarse-and-fine cross-modal interaction. Moreover, a background suppression module is introduced to restrain irrelevant background clips, meanwhile, a novel IoU attention mechanism and graph attention layer are efficiently devised to focus on the dependencies among highly-correlated video moments for the best choice selection. In-depth experiments on three public datasets TACoS, ActivityNet Captions, and Charades-STA demonstrate the superiority of our solution. Xin Sun 0020, Jialin Gao, Yizhe Zhu, Xi Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | HiCo: Hierarchical Contrastive Learning for Ultrasound Video Model Pretraining
Chunhui Zhang 0001, Yixiong Chen, Li Liu 0036, Xi Zhou 0001 |
ACCV (6) | 5 |
| 2022 | You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in VideosabstractMoment retrieval in videos is a challenging task that aims to retrieve the most relevant video moment in an untrimmed video given a sentence description. Previous methods tend to perform self-modal learning and cross-modal interaction in a coarse manner, which neglect fine-grained clues contained in video content, query context, and their alignment. To this end, we propose a novel Multi-Granularity Perception Network (MGPN) that perceives intra-modality and inter-modality information at a multi-granularity level. Specifically, we formulate moment retrieval as a multi-choice reading comprehension task and integrate human reading strategies into our framework. A coarse-grained feature encoder and a co-attention mechanism are utilized to obtain a preliminary perception of intra-modality and inter-modality information. Then a fine-grained feature encoder and a conditioned interaction module are introduced to enhance the initial perception inspired by how humans address reading comprehension problems. Moreover, to alleviate the huge computation burden of some existing methods, we further design an efficient choice comparison module and reduce the hidden size with imperceptible quality loss. Extensive experiments on Charades-STA, TACoS, and ActivityNet Captions datasets demonstrate that our solution outperforms existing state-of-the-art methods. Xin Sun 0020, Jialin Gao, Xi Zhou 0001 |
SIGIR | 5 |
| 2022 | Efficient Video Grounding With Which-Where Reading ComprehensionabstractVideo grounding aims at localizing the temporal moment related to the given language description, which is very helpful to many cross-modal content understanding applications like visual question answering and sentence-video search. Existing approaches usually directly regress the temporal boundaries of an event described by a query sentence in the video sequence. This direct regression manner often encounters a large decision space due to diverse target events and variable video durations, leading to inaccurate localization as well as inefficient grounding. This paper presents an efficient framework termed from which to where to facilitate video grounding. The core idea is imitating the reading comprehension process to gradually narrow the decision space, in what we decompose the direct regression into two steps. The “which” step first roughly selects a candidate area by evaluating which video segment in the predefined set is closest to the ground truth. To this end, we formulate this step into a multi-choice reading comprehension problem and propose a criterion to select the best-matched segment. In this way, the excessive decision space is effectively reduced. The “where” step aims to precisely regress the temporal boundary of the selected video segment from the shrunk decision space. We thus introduce a triple-span representation for each candidate video segment to use the regional context for better boundary regression. The “which” and “where” steps can be combined into a unified framework and learned end-to-end, leading to an efficient video grounding system. Extensive experiments on Charades-STA, ActivityNet-Captions, and TACoS benchmarks clearly demonstrate the effectiveness of our framework. Jialin Gao, Xin Sun 0020, Bernard Ghanem, Xi Zhou 0001, Shiming Ge |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Relation-aware Video Reading Comprehension for Temporal Language GroundingabstractTemporal language grounding in videos aims to localize the temporal span relevant to the given query sentence.Previous methods treat it either as a boundary regression task or a span extraction task.This paper will formulate temporal language grounding into video reading comprehension and propose a Relation-aware Network (RaNet) to address it.This framework aims to select a video moment choice from the predefined answer set with the aid of coarse-and-fine choice-query interaction and choice-choice relation construction.A choicequery interactor is proposed to match the visual and textual information simultaneously in sentence-moment and token-moment levels, leading to a coarse-and-fine cross-modal interaction.Moreover, a novel multi-choice relation constructor is introduced by leveraging graph convolution to capture the dependencies among video moment choices for the best choice selection.Extensive experiments on ActivityNet-Captions, TACoS, and Charades-STA demonstrate the effectiveness of our solution.Codes will be available at https: //github.com/Huntersxsx/RaNet. Jialin Gao, Xin Sun 0020, Mengmeng Xu 0006, Xi Zhou 0001, Bernard Ghanem |
EMNLP (1) | 4 |
| 2021 | Skeleton-Based Action Recognition With Focusing-Diffusion Graph Convolutional NetworksabstractGraph Convolutional Networks have been successfully applied in skeleton-based action recognition. The key is fully exploring the spatial-temporal context. This letter proposes a Focusing-Diffusion Graph Convolutional Network (FDGCN) to address this issue. Each skeleton frame is first decomposed into two opposite-direction graphs for subsequent focusing and diffusion processes. Next, the focusing process generates a spatial-level representation for each frame individually by an attention module. This representation is regarded as a supernode to aggregate the feature from each joint node in each frame for spatial context extraction. After generating supernodes for the entire sequence, a transformer encoder layer is proposed to capture the temporal context further. Finally, these supernodes pass the embedded spatial-temporal context back to the spatial joints through the diffusion graph in the diffusing process. Extensive experiments on the NTU RGB+D and Skeleton-Kinetics benchmarks demonstrate the effectiveness of our approach. Jialin Gao, Tong He 0002, Xi Zhou 0001, Shiming Ge |
IEEE Signal Process. Lett. | 3 |
| 2021 | Self-Guided Body Part Alignment With Relation Transformers for Occluded Person Re-IdentificationabstractPerson re-identification in the wild is often challenged by occlusion. Existing methods mainly rely on learned external cues like pose or parsing to ease occlusion distraction. This knowledge highly related to body semantics may introduce alignment effects, leading to additional requirements for dedicated training data and inference computation. We propose the Self-guided Body Part Alignment method that learns cue-free semantic-aligned local prediction for feature representations to avoid high-cost dependence on external cues. First, scale-wise global spatial attention is utilized to determine essential body parts automatically. A relation transformer network is then employed to predict semantic-aligned local parts, guided with anchored global information by constraint loss. Similarity metrics for all parts are merged with threshold conditions to filter invisible body parts comprehensively. Experimental results on occluded and holistic person reID benchmarks show the proposed method outperforms other cue-relied and cue-free methods. As far as we know, this is the first method that applies transformer networks on local predictions for occluded reID tasks. Guanshuo Wang, Jialin Gao, Xi Zhou 0001, Shiming Ge |
IEEE Signal Process. Lett. | 4 |
| 2020 | Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid NetworkabstractAccurate temporal action proposals play an important role in detecting actions from untrimmed videos. The existing approaches have difficulties in capturing global contextual information and simultaneously localizing actions with different durations. To this end, we propose a Relation-aware pyramid Network (RapNet) to generate highly accurate temporal action proposals. In RapNet, a novel relation-aware module is introduced to exploit bi-directional long-range relations between local features for context distilling. This embedded module enhances the RapNet in terms of its multi-granularity temporal proposal generation ability, given predefined anchor boxes. We further introduce a two-stage adjustment scheme to refine the proposal boundaries and measure their confidence in containing an action with snippet-level actionness. Extensive experiments on the challenging ActivityNet and THUMOS14 benchmarks demonstrate our RapNet generates superior accurate proposals over the existing state-of-the-art methods. Jialin Gao, Zhixiang Shi, Guanshuo Wang, Yufeng Yuan, Shiming Ge, Xi Zhou 0001 |
AAAI | 7 |
| 2020 | Receptive Multi-Granularity Representation for Person Re-IdentificationabstractA key for person re-identification is achieving consistent local details for discriminative representation across variable environments. Current stripe-based feature learning approaches have delivered impressive accuracy, but do not make a proper trade-off between diversity, locality, and robustness, which easily suffers from part semantic inconsistency for the conflict between rigid partition and misalignment. This paper proposes a receptive multi-granularity learning approach to facilitate stripe-based feature learning. This approach performs local partition on the intermediate representations to operate receptive region ranges, rather than current approaches on input images or output features, thus can enhance the representation of locality while remaining proper local association. Toward this end, the local partitions are adaptively pooled by using significance-balanced activations for uniform stripes. Random shifting augmentation is further introduced for a higher variance of person appearing regions within bounding boxes to ease misalignment. By twobranch network architecture, different scales of discriminative identity representation can be learned. In this way, our model can provide a more comprehensive and efficient feature representation without larger model storage costs. Extensive experiments on intra-dataset and cross-dataset evaluations demonstrate the effectiveness of the proposed approach. Especially, our approach achieves a state-of-the-art accuracy of 96.2%@Rank-1 or 90.0%@mAP on the challenging Market-1501 benchmark. Guanshuo Wang, Yufeng Yuan, Shiming Ge, Xi Zhou 0001 |
IEEE Trans. Image Process. | 5 |