Xiujun Shu

dblp:273/0053 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Why and How: Knowledge-Guided Learning for Cross-Spectral Image Patch Matching
abstract
Recently, cross-spectral image patch matching based on feature relation learning has attracted extensive attention. However, existing methods focus on mining richer feature relations by building complex relation extraction structures. Meanwhile, performance bottlenecks have gradually emerged. To address this, we make the first attempt to explore a stable and efficient bridge between descriptor learning and metric learning, and construct a Knowledge-Guided Learning Network (KGL-Net), which achieves significant performance improvements while abandoning complex network structures. Specifically, we find that there is feature extraction consistency between metric learning based on feature difference learning and descriptor learning based on Euclidean distance. This provides the foundation for bridge building. To ensure the stability and efficiency of the constructed bridge, on the one hand, we conduct an in-depth exploration of 20 combined network architectures. On the other hand, a feature-guided loss is constructed to achieve mutual guidance of features. In addition, unlike existing methods, we consider that the feature mapping ability of the metric branch should receive more attention. Therefore, a hard negative sample mining for metric learning (HNSM-M) strategy is constructed. To the best of our knowledge, this is the first time that hard negative sample mining for metric networks has been implemented and brings significant performance gains. Extensive experimental results show that our KGL-Net achieves SOTA performance in multiple cross-spectral image patch matching datasets. Our code is available at https://github.com/YuChuang1205/KGL-Net.
Chuang Yu 0003, Yunpeng Liu 0001, Jinmiao Zhao, Ao Chen 0003, Xiujun Shu, Bo Wang 0162, Zelin Shi, Xiangyu Yue 0001
IEEE Trans. Image Process.5
2025 Chi-Square Wavelet Graph Neural Networks for Heterogeneous Graph Anomaly Detection
abstract
Graph Anomaly Detection (GAD) in heterogeneous networks presents unique challenges due to node and edge heterogeneity. Existing Graph Neural Network (GNN) methods primarily focus on homogeneous GAD and thus fail to address three key issues: (C1) Capturing abnormal signal and rich semantics across diverse meta-paths; (C2) Retaining high-frequency content in HIN dimension alignment; and (C3) Learning effectively from difficult anomaly samples with class imbalance. To overcome these, we propose ChiGAD, a spectral GNN framework based on a novel Chi-Square filter, inspired by the wavelet effectiveness in diverse domains. Specifically, ChiGAD consists of: (1) Multi-Graph Chi-Square Filter, which captures anomalous information via applying dedicated Chi-Square filters to each meta-path graph; (2) Interactive Meta-Graph Convolution, which aligns features while preserving high-frequency information and incorporates heterogeneous messages by a unified Chi-Square Filter; and (3) Contribution-Informed Cross-Entropy Loss, which prioritizes difficult anomalies to address class imbalance. Extensive experiments on public and industrial datasets show that ChiGAD outperforms state-of-the-art models on multiple metrics. Additionally, its homogeneous variant, ChiGNN, excels on seven GAD datasets, validating the effectiveness of Chi-Square filters. Our code is available at https://github.com/HsipingLi/ChiGAD.
Xiping Li, Xiangyu Dong 0002, Xingyi Zhang 0003, Kun Xie 0010, Yuanhao Feng, Bo Wang 0162, Guilin Li 0001, Wuxiong Zeng, Xiujun Shu, Sibo Wang 0001
KDD (2)9
2025 Precise occlusion-aware and feature-level reconstruction for occluded person re-identification
Xiujun Shu, Hanjun Li 0002, Ruizhi Qiao, Weijian Ruan, Hanjing Su, Bo Wang 0162, Shouzhi Chen
Neurocomputing1
2023 Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer
abstract
Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge from a language model, while ignoring the rich semantic information inherent in image-text pairs. Instead, recently developed open-vocabulary (OV) based methods succeed in exploiting such information of image-text pairs in object detection, and achieve impressive performance. Inspired by the success of OV-based methods, we propose a novel open-vocabulary framework, named multi-modal knowledge transfer (MKT), for multi-label classification. Specifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model. To facilitate transferring the image-text matching ability of VLP model, knowledge distillation is employed to guarantee the consistency of image and label embeddings, along with prompt tuning to further update the label embeddings. To further enable the recognition of multiple objects, a simple but effective two-stream module is developed to capture both local and global features. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on public benchmark datasets.
Sunan He, Taian Guo, Tao Dai 0001, Ruizhi Qiao, Xiujun Shu, Bo Ren 0002, Shutao Xia
AAAI5
2023 Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in Movies
abstract
Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1) For different annotators, labeling highlight has uncertainty, which leads to inaccurate and time-consuming annotations. (2) Besides previous supervised or unsupervised settings, some existing video corpora can be useful, e.g., trailers, but they are often noisy and incomplete to cover the full highlights. In this work, we study a more practical and promising setting, i.e., reformulating high-light detection as “learning with noisy labels”. This setting does not require time-consuming manual annotations and can fully utilize existing abundant video corpora. First, based on movie trailers, we leverage scene segmentation to obtain complete shots, which are regarded as noisy labels. Then, we propose a Collaborative noisy Label Cleaner (CLC) framework to learn from noisy highlight moments. CLC consists of two modules: augmented cross-propagation (ACP) and multimodality cleaning (MMC). The former aims to exploit the closely related audio-visual signals and fuse them to learn unified multimodal representations. The latter aims to achieve cleaner highlight labels by observing the changes in losses among different modalities. To verify the effectiveness of CLC, we further collect a large-scale highlight dataset named MovieLights. Comprehensive experiments on MovieLights and YouTube Highlights datasets demonstrate the effectiveness of our approach. Code has been made available at: https://github.com/TencentYoutuResearch/HighlightDetection-CLC.
Bei Gan, Xiujun Shu, Ruizhi Qiao, Haoqian Wu, Hanjun Li 0002, Bo Ren 0002
CVPR2
2023 NewsNet: A Novel Dataset for Hierarchical Temporal Segmentation
abstract
Temporal video segmentation is the get-to- go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations can help compare the semantics in the corresponding scales, but lack a wider view of larger temporal spans, especially when the video is complex and structured. Therefore, we present two abstractive levels of temporal segmentations and study their hierarchy to the existing fine-grained levels. Accordingly, we collect NewsNet, the largest news video dataset consisting of 1,000 videos in over 900 hours, associated with several tasks for hierarchical temporal video segmentation. Each news video is a collection of stories on different topics, represented as aligned audio, visual, and textual data, along with extensive frame-wise annotations in four granularities. We assert that the study on NewsNet can advance the understanding of complex structured video and benefit more areas such as short-video creation, personalized advertisement, digital instruction, and education. Our dataset and code is publicly available at https://github.com/NewsNet-Benchmark/NewsNet.
Haoqian Wu, Mingchen Zhuge, Bing Li 0024, Ruizhi Qiao, Xiujun Shu, Bei Gan, Liangsheng Xu, Bo Ren 0002, Mengmeng Xu 0006, Wentian Zhang, Ramachandra Raghavendra, Chia-Wen Lin, Bernard Ghanem
CVPR7
2023 D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance Annotation
abstract
Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. Recently, weakly supervised methods still have a large performance gap compared to fully supervised ones, while the latter requires laborious timestamp annotations. In this study, we aim to reduce the annotation cost yet keep competitive performance for TSG task compared to fully supervised ones. To achieve this goal, we investigate a recently proposed glance-supervised temporal sentence grounding task, which requires only single frame annotation (referred to as glance annotation) for each query. Under this setup, we propose a Dynamic Gaussian prior based Grounding framework with Glance annotation (D3G), which consists of a Semantic Alignment Group Contrastive Learning module (SA-GCL) and a Dynamic Gaussian prior Adjustment module (DGA). Specifically, SA-GCL samples reliable positive moments from a 2D temporal map via jointly leveraging Gaussian prior and semantic consistency, which contributes to aligning the positive sentence-moment pairs in the joint embedding space. Moreover, to alleviate the annotation bias resulting from glance annotation and model complex queries consisting of multiple events, we propose the DGA module, which adjusts the distribution dynamically to approximate the ground truth of target moments. Extensive experiments on three challenging benchmarks verify the effectiveness of the proposed D3G. It outperforms the state-of-the-art weakly supervised methods by a large margin and narrows the performance gap compared to fully supervised methods. Code is available at https://github.com/solicucu/D3G.
Hanjun Li 0002, Xiujun Shu, Sunan He, Ruizhi Qiao, Taian Guo, Bei Gan, Xing Sun 0001
ICCV2
2023 MFGNet: Dynamic Modality-Aware Filter Generation for RGB-T Tracking
abstract
Many RGB-T trackers attempt to attain robust feature representation by utilizing an adaptive weighting scheme (or attention mechanism). Different from these works, we propose a new dynamic modality-aware filter generation module (named MFGNet) to boost the message communication between visible and thermal data by adaptively adjusting the convolutional kernels for various input images in practical tracking. Given the image pairs as input, we first encode their features with the backbone network. Then, we concatenate these feature maps and generate dynamic modality-aware filters with two independent networks. The visible and thermal filters will be used to conduct a dynamic convolutional operation on their corresponding input feature maps respectively. Inspired by residual connection, both the generated visible and thermal feature maps will be summarized with input feature maps. The augmented feature maps will be fed into the RoI align module to generate instance-level features for subsequent classification. To address issues caused by heavy occlusion, fast motion and out-of-view, we propose to conduct a joint local and global search by exploiting a new direction-aware target driven attention mechanism. The spatial and temporal recurrent neural network is used to capture the direction-aware context for accurate global attention prediction. Extensive experiments on three large-scale RGB-T tracking benchmark datasets validated the effectiveness of our proposed algorithm.
Xiao Wang 0014, Xiujun Shu, Shiliang Zhang, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001
IEEE Trans. Multim.2
2022 Hyperspherical Learning in Multi-Label Classification
Bo Ke, Yunquan Zhu, Mengtian Li 0002, Xiujun Shu, Ruizhi Qiao, Bo Ren 0002
ECCV (25)4
2022 Exploiting robust unsupervised video person re-identification
abstract
Abstract Unsupervised video person re‐identification (reID) methods usually depend on global‐level features. Many supervised reID methods employed local‐level features and achieved significant performance improvements. However, applying local‐level features to unsupervised methods may introduce an unstable performance. To improve the performance stability for unsupervised video reID, this paper introduces a general scheme fusing part models and unsupervised learning. In this scheme, the global‐level feature is divided into equal local‐level feature. A local‐aware module is employed to explore the potentials of local‐level feature for unsupervised learning. A global‐aware module is proposed to overcome the disadvantages of local‐level features. Features from these two modules are fused to form a robust feature representation for each input image. This feature representation has the advantages of local‐level feature without suffering from its disadvantages. Comprehensive experiments are conducted on three benchmarks, including PRID2011, iLIDS‐VID, and DukeMTMC‐VideoReID, and the results demonstrate that the proposed approach achieves state‐of‐the‐art performance. Extensive ablation studies demonstrate the effectiveness and robustness of proposed scheme, local‐aware module and global‐aware module. The code and generated features are available at https://github.com/deropty/uPMnet .
Xianghao Zang, Ge Li 0002, Wei Gao 0003, Xiujun Shu
IET Image Process.4
2022 Weakly-supervised anomaly detection in video surveillance via graph convolutional label noise cleaning
Nannan Li 0001, Jia-Xing Zhong, Xiujun Shu, Huiwen Guo
Neurocomputing3
2022 Temporal Weighting Appearance-Aligned Network for Nighttime Video Retrieval
abstract
Video-based person re-identification (ReID) aims at re-identifying video sequences of a specified person from videos captured by disjoint cameras. Existing datasets and works on this task all focus on daytime scenarios and cannot adapt well to the nighttime scenarios, which is also of significant importance for practical applications. In this paper, we contribute a new dataset for nighttime video-based ReID, termed NIVIR, which contains 800 identities with over 228,000 images. NIVIR contains video shots under various lighting conditions, different weathers, and complex scenarios, which is consistent with the real nighttime outdoor surveillance. Furthermore, we propose a temporal weighting appearance-aligned network (TWAN) for nighttime video-based ReID, which is composed of a correlation-based appearance-aligned module (CAM) and a temporal weighting module (TWM). Specifically, CAM is proposed to reconstruct the adjacent feature maps to guarantee the appearance alignment between the central frame and its adjacent frames. TWM is designed to evaluate the frame quality of a tracklet and generate temporal weights to enhance the video representation. Extensive experiments conducted on our new NIVIR dataset demonstrate that the proposed TWAN outperforms the state-of-the-art methods. We believe that our NIVIR dataset and the comprehensive attempts for solving the nighttime ReID problem will push forward the development of the ReID research community.
Weijian Ruan, Yiran Tao, Linjun Ruan, Xiujun Shu, Yu Qiao 0001
IEEE Signal Process. Lett.4
2022 Large-Scale Spatio-Temporal Person Re-Identification: Algorithms and Benchmark
abstract
Person re-identification (re-ID) in the scenario with large spatial and temporal spans has not been fully explored. This fact partially occurs because existing benchmark datasets were mainly collected with limited spatial and temporal ranges,e.g.,using videos recorded in a few days by cameras in a specific region of the campus. Such limited spatial and temporal ranges make it hard to simulate the difficulties of person re-ID in real scenarios. In this work, we contribute a novel Large-scale Spatio-Temporal (LaST) person re-ID dataset, including 10,862 identities with more than 228k images. Compared with existing datasets, LaST presents more challenging and high-diversity re-ID settings and significantly larger spatial and temporal ranges. For instance, each person can appear in different cities or countries, and in various time slots from day to evening, and in different seasons from spring to winter. To our best knowledge, LaST is a novel person re-ID dataset with the largest spatio-temporal ranges. Based on LaST, we verified its challenge by conducting a comprehensive performance evaluation of 14 re-ID algorithms. We further propose an easy-to-implement baseline that works well in such challenging re-ID settings. We also verified that models pre-trained on LaST can generalize well on existing datasets with short-term and cloth-changing scenarios. We expect LaST to inspire future works toward more realistic and challenging re-ID tasks. More information about the dataset is available athttps://github.com/shuxjweb/last.git.
Xiujun Shu, Xiao Wang 0014, Xianghao Zang, Shiliang Zhang, Yuanqi Chen, Ge Li 0002, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and Benchmark
abstract
Tracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared with traditional bounding box (BBox) based tracking, this setting guides object tracking with high-level semantic information, addresses the ambiguity of BBox, and links local and global search organically together. Those benefits may bring more flexible, robust and accurate tracking performance in practical scenarios. However, existing natural language initialized trackers are developed and compared on benchmark datasets proposed for tracking-by-BBox, which can’t reflect the true power of tracking-by-language. In this work, we propose a new benchmark specifically dedicated to the tracking-by-language, including a large scale dataset, strong and diverse baseline methods. Specifically, we collect 2k video sequences (contains a total of 1,244,340 frames, 663 words) and split 1300/700 for the train/testing respectively. We densely annotate one sentence in English and corresponding bounding boxes of the target object for each video. We also introduce two new challenges into TNL2K for the object tracking task, i.e., adversarial samples and modality switch. A strong baseline method based on an adaptive local-global-search scheme is proposed for future works to compare. We believe this benchmark will greatly boost related researches on natural language guided tracking.
Xiao Wang 0014, Xiujun Shu, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001
CVPR2
2021 Learning to disentangle scenes for person re-identification
Xianghao Zang, Ge Li 0002, Wei Gao 0003, Xiujun Shu
Image Vis. Comput.4
2021 Diverse part attentive network for video-based person re-identification
Xiujun Shu, Ge Li 0002, Longhui Wei, Jia-Xing Zhong, Xianghao Zang, Shiliang Zhang, Yaowei Wang 0001, Yongsheng Liang 0001, Qi Tian 0001
Pattern Recognit. Lett.1
2021 Semantic-Guided Pixel Sampling for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (re-ID) is a new rising research topic that aims at retrieving pedestrians whose clothes are changed. This task is quite challenging and has not been fully studied to date. Current works mainly focus on body shape or contour sketch, but they are not robust enough due to view and posture variations. The key to this task is to exploit cloth-irrelevant cues. This paper proposes a semantic-guided pixel sampling approach for the cloth-changing person re-ID task. We do not explicitly define which feature to extract but force the model to automatically learn cloth-irrelevant cues. Specifically, we firstly recognize the pedestrian's upper clothes and pants, then randomly change them by sampling pixels from other pedestrians. The changed samples retain the identity labels but exchange the pixels of clothes or pants among different pedestrians. Besides, we adopt a loss function to constrain the learned features to keep consistent before and after changes. In this way, the model is forced to learn cues that are irrelevant to upper clothes and pants. We conduct extensive experiments on the latest released PRCC dataset. Our method achieved 65.8% on Rank1 accuracy, which outperforms previous methods with a large margin. The code is available athttps://github.com/shuxjweb/pixel_sampling.git.
Xiujun Shu, Ge Li 0002, Xiao Wang 0014, Weijian Ruan, Qi Tian 0001
IEEE Signal Process. Lett.1
2020 Cellular Network Radio Propagation Modeling with Deep Convolutional Neural Networks
abstract
Radio propagation modeling and prediction is fundamental for modern cellular network planning and optimization. Conventional radio propagation models fall into two categories. Empirical models, based on coarse statistics, are simple and computationally efficient, but are inaccurate due to oversimplification. Deterministic models, such as ray tracing based on physical laws of wave propagation, are more accurate and site specific. But they have higher computational complexity and are inflexible to utilize site information other than traditional global information system (GIS) maps.
Xiujun Shu, Bingwen Zhang, Jie Ren 0013, Lizhou Zhou, Xin Chen 0062
KDD2