EDBT 2026 Demo / reviewers in the wild / expert
Xiuguo Bao
dblp:62/6224
· DBLP profile ↗
28ranked-venue papers
0as first author
16since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 since 2021Artificial intelligence and machine learning · 13 · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Label-informed knowledge integration: Advancing visual prompt for VLMs adaptation
Yunhong Wang 0001, Guodong Wang 0006, Yingjie Gao 0001, Xiuguo Bao, Di Huang 0001 |
Comput. Vis. Image Underst. | 6 |
| 2025 | Anomaly-aware self-supervised feature learning for weakly supervised video anomaly detection
Zhen Yang 0037, Guodong Wang 0006, Yuanfang Guo, Xiuguo Bao, Di Huang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2025 | Cross-Modal Contrastive Masked AutoEncoder for Compressed Video Pre-TrainingabstractIn this paper, we propose a novel Transformer based approach, namely Cross-modal Contrastive Masked AutoEncoder (C2MAE), to Self-Supervised Learning (SSL) on compressed videos. A unified Transformer encoder is employed to discover relationships of visual tokens from RGBs, motion vectors and residuals. A hybrid SSL framework is proposed, which combines the complementary advantages of Masked Image Modeling (MIM) and Contrastive Learning (CL) pretext tasks, for powerful representation learning. The MIM branch extends VideoMAE by a new Fine-Grained Motion-aware Masking (FGMM) strategy and a modified Multi-modal Reconstruction (MR) task, where FGMM computes motion saliency maps as motion priors to guide the masks so that it well fits for the data properties in the compressed domain and the MR task highlights the reconstruction of raw videos by joint representations from corresponding compressed videos in addition to that in each single modality. The CL branch introduces the Contrastive Cross-modal Learning (CCL) module, and the features from a compressed video clip and the ones from its raw video counterpart are compared instead of widely used augmented data. Due to these designs, C2MAE significantly enhances interactions across modalities to compensate the sparsity of I-frames and the coarse and noisy nature of P-frames, thus delivering much stronger pre-trained models. Extensive experiments are conducted on the UCF-101, HMDB-51 and Kinetics-400 benchmarks with state-of-the-art results reported, demonstrating its effectiveness. Jiaxin Chen 0002, Guohao Li 0010, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Rotation Has Two Sides: Evaluating Data Augmentation for Deep One-class ClassificationabstractOne-class classification (OCC) involves predicting whether a new data is normal or anomalous based solely on the data from a single class during training. Various attempts have been made to learn suitable representations for OCC within a self-supervised framework. Notably, discriminative methods that use geometric visual transformations, such as rotation, to generate pseudo-anomaly samples have exhibited impressive detection performance. Although rotation is commonly viewed as a distribution-shifting transformation and is widely used in the literature, the cause of its effectiveness remains a mystery. In this study, we are the first to make a surprising observation: there exists a strong linear relationship (Pearson's Correlation, $r > 0.9$) between the accuracy of rotation prediction and the performance of OCC. This suggests that a classifier that effectively distinguishes different rotations is more likely to excel in OCC, and vice versa. The root cause of this phenomenon can be attributed to the transformation bias in the dataset, where representations learned from transformations already present in the dataset tend to be less effective, making it essential to accurately estimate the transformation distribution before utilizing pretext tasks involving these transformations for reliable self-supervised representation learning. To the end, we propose a novel two-stage method to estimate the transformation distribution within the dataset. In the first stage, we learn general representations through standard contrastive pre-training. In the second stage, we select potentially semantics-preserving samples from the entire augmented dataset, which includes all rotations, by employing density matching with the provided reference distribution. By sorting samples based on semantics-preserving versus shifting transformations, we achieve improved performance on OCC benchmarks. Guodong Wang 0006, Yunhong Wang 0001, Xiuguo Bao, Di Huang 0001 |
ICLR | 3 |
| 2024 | LightRL-AD: A Lightweight Online Reinforcement Learning Approach for Autonomous Defense against Network AttacksabstractWith the rapid growth of the Internet, network structure has become increasingly complex, leading to more diverse and impactful network attacks. Traditional methods of detecting and defending against network attacks struggle with increasingly complex situations due to human decision-making processes. Recent research has started exploring autonomous defense mechanisms for network attacks within software-defined network (SDN) environments. However, these methods typically employ complex reinforcement learning techniques, making them challenging to implement in online deployment environments. In this paper, we propose LightRL-AD, a lightweight online reinforcement learning approach for autonomous defense against network attacks in SDN. LightRL-AD integrates a machine learning-based Intrusion Detection System (IDS), a reinforcement learning-based Intrusion Prevention System (IPS), and a Moving Target Defense (MTD) mechanism. The ML-based IDS classifies network flows into categories such as malicious or benign, while the RL-based IPS utilizes the SARSA algorithm to determine and execute appropriate defensive actions, ensuring robust network security. We employ specific hardware and software to establish a simulated SDN network for our experiments. And we implement LightRL-AD in the network and evaluate its performance. Experimental results demonstrate that LightRL-AD performs better to defend against slow-rate DDoS attacks autonomously. Fengyuan Shi 0005, Zhou Zhou 0007, Qingyun Liu 0001, Xiuguo Bao |
TrustCom | 8 |
| 2024 | Towards Video Anomaly Detection in the Real World: A Binarization Embedded Weakly-Supervised NetworkabstractIn this letter, we pioneer to propose a binarization embedded weakly-supervised video anomaly detection (BE-WSVAD) method by constructing a binarized GCN-based anomaly detection module. Compared to the existing weakly-supervised video anomaly detection (WS-VAD) methods, BE-WSVAD focuses on the detection efficiency, which is ignored by the existing literature yet vital in real applications. Specifically, to improve the detection performance of the binary anomaly detection module, we propose a binary network augmentation strategy in the training process. Due to the weakly supervision mechanism, the videos employed in the training process are usually lengthy, in which the lengthy-input dependencies tend to be exploited to improve the detection performance with extra memory consumption. Then, we propose the short-input inference modes, which can largely reduce the desired length of the input video. Experimental results demonstrate the superiority of our BE-WSVAD in terms of the memory and computational consumptions while giving comparable accuracies. Zhen Yang 0037, Yuanfang Guo, Junfu Wang, Di Huang 0001, Xiuguo Bao, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | AHIP: An Adaptive IP Hopping Method for Moving Target Defense to Thwart Network AttacksabstractIn a static network, attackers can easily launch network attacks on target hosts which have long-term constant IP addresses. In order to defend against attackers effectively, many defense approaches use IP hopping to dynamically transform IP configuration. However, these approaches usually focus on one type of network attacks, scanning attacks or Denial of Service (DoS) attacks, and cannot sense network situations. This paper proposes AHIP, an adaptive IP hopping method for moving target defense (MTD) to defend against different network attacks. We use a trained lightweight one-dimensional convolutional neural network (1D-CNN) detector to judge whether there are no attacks, scanning attacks or DoS attacks in the network, which can adaptively trigger corresponding IP hopping strategy. We use specific hardware and software to create the software defined network (SDN) environment for experiments. The experiments prove that AHIP performs better to thwart network attacks and has lower system overhead. Fengyuan Shi 0005, Zhou Zhou 0007, Qingyun Liu 0001, Xiuguo Bao |
CSCWD | 6 |
| 2023 | Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation for Anomaly DetectionabstractAnomaly detection (AD), aiming to find samples that deviate from the training distribution, is essential in safety-critical applications. Though recent self-supervised learning based attempts achieve promising results by creating virtual outliers, their training objectives are less faithful to AD which requires a concentrated inlier distribution as well as a dispersive outlier distribution. In this paper, we propose Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation (UniCon-HA), taking into account both the requirements above. Specifically, we explicitly encourage the concentration of inliers and the dispersion of virtual outliers via supervised and unsupervised contrastive losses, respectively. Considering that standard contrastive data augmentation for generating positive views may induce outliers, we additionally introduce a soft mechanism to re-weight each augmented inlier according to its deviation from the inlier distribution, to ensure a purified concentration. Moreover, to prompt a higher concentration, inspired by curriculum learning, we adopt an easy-to-hard hierarchical augmentation strategy and perform contrastive aggregation at different depths of the network based on the strengths of data augmentation. Our method is evaluated under three AD settings including unlabeled one-class, unlabeled multi-class, and labeled multi-class, demonstrating its consistent superiority over other competitors. Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001 |
ICCV | 5 |
| 2023 | Compressed Video Prompt TuningabstractCompressed videos offer a compelling alternative to raw videos, showing the possibility to significantly reduce the on-line computational and storage cost. However, current approaches to compressed video processing generally follow the resource-consuming pre-training and fine-tuning paradigm, which does not fully take advantage of such properties, making them not favorable enough for widespread applications. Inspired by recent successes of prompt tuning techniques in computer vision, this paper presents the first attempt to build a prompt based representation learning framework, which enables effective and efficient adaptation of pre-trained raw video models to compressed video understanding tasks. To this end, we propose a novel prompt tuning approach, namely Compressed Video Prompt Tuning (CVPT), emphatically dealing with the challenging issue caused by the inconsistency between pre-training and downstream data modalities. Specifically, CVPT replaces the learnable prompts with compressed modalities (\emph{e.g.} Motion Vectors and Residuals) by re-parameterizing them into conditional prompts followed by layer-wise refinement. The conditional prompts exhibit improved adaptability and generalizability to instances compared to conventional individual learnable ones, and the Residual prompts enhance the noisy motion cues in the Motion Vector prompts for further fusion with the visual cues from I-frames. Additionally, we design Selective Cross-modal Complementary Prompt (SCCP) blocks. After inserting them into the backbone, SCCP blocks leverage semantic relations across diverse levels and modalities to improve cross-modal interactions between prompts and input flows. Extensive evaluations on HMDB-51, UCF-101 and Something-Something v2 demonstrate that CVPT remarkably outperforms the state-of-the-art counterparts, delivering a much better balance between accuracy and efficiency. Jiaxin Chen 0002, Xiuguo Bao, Di Huang 0001 |
NeurIPS | 3 |
| 2023 | S$^{2}$-Net:Semantic and Saliency Attention Network for Person Re-IdentificationabstractPerson re-identification is still a challenging task when moving objects or another person occludes the probe person. Mainstream methods based on even partitioning apply an off-the-shelf human semantic parsing to highlight the non-collusion part. In this paper, we apply an attention branch to learn the human semantic partition to avoid misalignment introduced by even partitioning. In detail, we propose a semantic attention branch to learn 5 human semantic maps. We also note that some accessories or belongings, such as a hat, bag, may provide more informative clues to improve the person Re-ID. Human semantic parsing, however, usually treats non-human parts as distractions and discards them. To fetch the missing clues, we design a branch to capture the salient non-human parts. Finally, we merge the semantic and saliency attention to build an end-to-end network, named as S$^{2}$-Net. Specifically, to further improve Re-ID, we develop a trade-off weighting scheme between semantic and saliency attention and set the right weight with the actual scene. The extensive experiments show that S$^{2}$-Net gets the competitive performance. S$^{2}$-Net achieves 87.4% mAP on Market1501 and obtains 79.3%/56.1% rank-1/mAP on MSMT17 without semantic supervision. The source codes are available athttps://github.com/upgirlnana/S2Net. Xuena Ren, Dongming Zhang 0004, Xiuguo Bao, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Video Anomaly Detection by Solving Decoupled Spatio-Temporal Jigsaw Puzzles
Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001 |
ECCV (10) | 5 |
| 2022 | Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion EnhancementabstractCompressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues (e.g., motion vectors and residuals). However, this task severely suffers from the coarse and noisy dynamics and the insufficient fusion of the heterogeneous RGB and motion modalities. To address the two issues above, this paper proposes a novel framework, namely Attentive Cross-modal Interaction Network with Motion Enhancement (MEACI-Net). It follows the two-stream architecture, i.e. one for the RGB modality and the other for the motion modality. Particularly, the motion stream employs a multi-scale block embedded with a denoising module to enhance representation learning. The interaction between the two streams is then strengthened by introducing the Selective Motion Complement (SMC) and Cross-Modality Augment (CMA) modules, where SMC complements the RGB modality with spatio-temporally attentive local motion features and CMA further combines the two modalities with selective feature augmentation. Extensive experiments on the UCF-101, HMDB-51 and Kinetics-400 benchmarks demonstrate the effectiveness and efficiency of MEACI-Net. Jiaxin Chen 0002, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001 |
IJCAI | 4 |
| 2022 | Person Re-identification with a Cloth-Changing Aware TransformerabstractCloth-Changing person re-identification is a challenging problem for its huge intra-class variance caused by changing clothes. In this paper, we focus on learning cloth-insensitive features without any cues from external models. Specifically, we introduce a novel Cloth-Changing Aware Transformer(CCAT) to learn identity-relevant cues. The transformer encoder utilizes Cloth Information Embedding(CIE) to encode outfits information and automatically focuses on the identity-discriminative feature. The tokens from the transformer encoder are then fed to the Cloth-Insensitive Consistency Learning(CICL) decoder to learn cloth-irrelevant features from pairwise pedestrian images. Besides, we impose Intra-class Constraint Learning(ICL) to pull the global tokens from the encoder closer. In this way, the model learns the identity-relevant and cloth-insensitive features. Extensive experimental results on three cloth-changing person ReID datasets demonstrate that our proposed algorithm can extract highly robust feature representations of cloth-changing persons, and it outperforms the state-of-the-art cloth-changing person ReID approaches. Our method achieves 87.4% and 69.7% on R-1 accuracy on the latest released LTCC [28] and PRCC [40], which outperforms previous methods with a large margin. The code is available at CCAT. Xuena Ren, Xiuguo Bao |
IJCNN | 3 |
| 2022 | Double Granularity Relation Network with Self-criticism for Occluded Person Re-identification
Xuena Ren, Xiuguo Bao |
MMM (1) | 3 |
| 2021 | Multi-Scale Background Suppression Anomaly Detection In Surveillance VideosabstractVideo anomaly detection has been widely applied in various surveillance systems for public security. However, the existing weakly supervised video anomaly detection methods tend to ignore the interference of the background frames and possess limited ability to extract effective temporal information among the video snippets. In this paper, a multi-scale background suppression based anomaly detection (MSBSAD) method is proposed to suppress the interference of the background frames. We propose a multi-scale temporal convolution module to effectively extract more temporal information among the video snippets for the anomaly events with different durations. A modified hinge loss is constructed in the suppression branch to help our model to better differentiate the abnormal samples from the confusing samples. Experiments on UCF Crime demonstrate the superiority of our MS-BSAD method in the video anomaly detection task. Yuanfang Guo, Jinjie Wei, Xiuguo Bao, Di Huang 0001 |
ICIP | 4 |
| 2021 | Sequential dynamic event recommendation in event-based social networks: An upper confidence bound approach
Yuan Liang 0003, Chunlin Huang, Xiuguo Bao |
Inf. Sci. | 3 |
| 2020 | Semantic-Guided Shared Feature Alignment for Occluded Person Re-IDentificationabstractOccluded Person Re-ID is a challenging task under resolved. Instead of extracting features over the entire image which would easily cause mismatching, we propose Semantic-Guided Shared Feature Alignment (SGSFA) method to extract features focusing on the non-occluded parts. SGSFA parses human body regions through Semantic Guided (SG) branch and aligns regions through Spatial Feature Alignment (SFA) branch simultaneously, and gets enriched representations over the regions for Re-ID. Dynamic classification loss of spatial features and their dynamical sequential combinations in the training stage help facilitate feature diversity. During the matching stage, we use only the visible feature shared by probe and gallery with no extra cues. The experiment results show that SGSFA achieves rank-1 of 62.3% and 50.5% respectively for Occluded-DukeMTMC and P-DukeMTMC-reID, surpassing the state-of-the-art by a large margin. Xuena Ren, Xiuguo Bao |
ACML | 3 |
| 2020 | Towards Practical Compressed Video Action Recognition: A Temporal Enhanced Multi-Stream NetworkabstractCurrent compressed video action recognition methods are mainly based on complete data. However, in a real transmission scenario, the compressed video packets are usually disorderly received and even lost due to network jitters or congestion. To recognize actions in early phases with limited packets, e.g. for quickly forecasting possible potential risks, in this paper, we propose a Temporal Enhanced Multi-Stream Network (TEMSN) towards practical compressed video action recognition. First, we make use of three modalities in the compressed domain as complementary cues and build a multi-stream network to capture rich information from compressed video packets. Second, we design a temporal enhanced module based on an Encoder-Decoder structure, which is applied to each stream to infer missing packets, generating more accurate action dynamics. Thanks to the multiple modalities and their temporal enhancement, our approach better models actions with partial available compressed video packets. Experiments on the HMDB-51 and UCF-101 datasets validate its effectiveness and efficiency. Longteng Kong, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001, Yunhong Wang 0001 |
ICPR | 4 |
| 2019 | GeoCET: Accurate IP Geolocation via Constraint-Based Elliptical Trajectories
Xiuguo Bao, Yongzheng Zhang 0002, Huanhuan Yang |
CollaborateCom | 2 |
| 2019 | Context-adaptive matching for optical flow
Yueran Zu, Wenzhong Tang, Xiuguo Bao, Ke Gao 0012 |
Multim. Tools Appl. | 3 |
| 2018 | GeoBLR: Dynamic IP Geolocation Method Based on Bayesian Linear Regression
Xiuguo Bao, Yongzheng Zhang 0002 |
CollaborateCom | 2 |
| 2018 | Saliency guided fast interpolation for large displacement optical flowabstractThe optical flow estimation is still an open question nowadays. One of the bottlenecks of it is the interpolation speed. In this paper, a saliency guide fast interpolation method is proposed which is more than about 2 times faster than the traditional one. The method runs on CPU without any supervision or semantic segmentation information. To make it faster, a fast saliency detection method is introduced to separate the image into two parts. The non-saliency superpixels are interpolated faster with random search only. The salient superpixels are interpolated by propagation and random search. To keep it accurate, the relative initial movement is used to guide the search area when computing the affine model. A soft affine model evaluation is introduced to make the optical flow result more robust. Extensive experiments on challenging datasets MPI-Sintel and KITTI-15 show that our method is efficient and effective. Yueran Zu, Xiuguo Bao, Wenzhong Tang |
ICPR | 3 |
| 2018 | Deep Convolutional Neural Network with Scalogram for Audio Scene Modeling
Hangting Chen, Pengyuan Zhang, Haichuan Bai, Qingsheng Yuan, Xiuguo Bao, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2015 | A Well-Behaved TV Logo Recognition Method Using Heuristic Optimality Checked Basis Pursuit Denoising
Xiuguo Bao |
ICIG (1) | 3 |
| 2014 | Characterizing emotion entrainment in social mediaabstractThe sociological theory of entrainment accounts for the synchronization of human rhythmic modalities through social interactions: they coordinate in a variety of dimensions including linguistic styles, facial expressions, music pace, applause, and so on. Though highly relevant, emotion entrainment has received little attention to date. In addition, most previous studies on entrainment are done through small scale or controlled laboratory studies. In this paper, we investigate emotion entrainment in the context of online social media. To the best of our knowledge, this is the first time that emotion entrainment has been examined on a large scale, real world setting. For this purpose, we propose a framework that can model entrainment phenomenon and measure its effect. Our framework differentiates from previous research by its model-free essential and discerning in entrainment directions. These traits enable us to model entrainment dynamics under few assumptions, and distinguish emotion flow of entrainment. In our studies, we investigate entrainment patterns under different emotion states, i.e. positive, neutral and negative. We discover that entrainments under different emotions all follow a power law distribution. Besides, people are willing to entrain to others under positive emotion, and users with positive emotion are more likely to be entrained. By inspecting the interactions between entrainment and emotion, we reveal that entrainment has an effect of negotiating different emotion types toward an even distribution. Saike He, Xiaolong Zheng 0001, Xiuguo Bao, Hongyuan Ma, Daniel Dajun Zeng, Bo Xu 0002, Changliang Li, Hongwei Hao |
ASONAM | 3 |
| 2014 | PAAP: prefetch-aware admission policies for query results cache in web search enginesabstractCaching query results is an efficient technique for Web search engines. Admission policy can prevent infrequent queries from taking space of more frequent queries in the cache. In this paper we present two novel admission policies tailored for query results cache. These policies are based on query results prefetching information. We also propose a demote operation for the query results cache to improve the cache hit ratio. We then use a trace of over 5 million queries to evaluate our admission policies, as well as traditional policies. Experimental results show that our prefetch-aware admission policies can achieve hit ratios better than state-of-the-art admission policies. Hongyuan Ma, Wei Liu 0005, Bingjie Wei, Xiuguo Bao, Bin Wang 0004 |
SIGIR | 5 |
| 2012 | Acquiring netizen group's opinions for modeling food safety eventsabstractFood safety events are typical public security events that draw great public concern. In food safety events, millions of netizens pay close attention to the event, express their opinions online and thus influence the decisions of government or food producers. Modeling netizen groups, especially the dynamics of their opinions in these events, can help us understand the mechanism and evolvement of such events and provide valuable insights for social management. However, conventional computational models, such as agent-based models, are usually constructed manually. In this paper, we propose an approach to acquiring netizen group's opinions from online comments to facilitate the modeling of food safety events. We conduct experimental study on typical events happened in China and empirically evaluate the performance of our proposed approach. The results verify the effectiveness of the approach. Zhangwen Tan, Wenji Mao, Daniel Dajun Zeng, Xiuguo Bao |
ISI | 5 |
| 2011 | Graph-based multi-space semantic correlation propagation for video retrieval
Bailan Feng, Juan Cao 0001, Xiuguo Bao, Yongdong Zhang 0001, Shouxun Lin, Xiao-chun Yun |
Vis. Comput. | 3 |