EDBT 2026 Demo / reviewers in the wild / expert
Xiaoshuai Wu
dblp:299/2195
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | K-ProtoDiff: Key Prototypes-Guided Diffusion for Time Series GenerationabstractTime series generation is essential for advancing data-driven modeling and decision-making across a wide range of domains. However, existing approaches primarily focus on global patterns, often failing to capture local key patterns such as abrupt changes or anomalies. These key patterns are crucial for interpretability and operational decision making, as they frequently represent intervention points with significant real-world impact. To bridge this gap, we propose Key Prototypes-Guided Diffusion (K-ProtoDiff) for time series generation , a new model that learns the global data distribution while preserving localized key patterns critical for temporal dynamics. In K-ProtoDiff, we first derive time series prototype representations through adaptive self-supervised learning. Then, a key prototype assignment module is used to extract prototype weights, forming key prototype-aware representations that serve as conditional guidance for generation. During sampling, to further enhance the fidelity of key patterns during the denoising process, we propose Reflection Sampling (R-Sampling), a step-wise refinement strategy that encourages the reverse trajectory to better align with key prototype constraints. Experiments on nine real-world datasets demonstrate that K-ProtoDiff significantly outperforms state-of-the-art baselines in key pattern retention, achieving an average 77.6% improvement in key pattern preservation. Yuhang Duan, Lin Lin 0008, Xiaoshuai Wu |
AAAI | 3 |
| 2026 | Sim-to-Real: An Unsupervised Noise Layer for Screen-Camera Watermarking RobustnessabstractUnauthorized screen capturing and dissemination pose severe security threats such as data leakage and information theft. Several studies propose robust watermarking methods to track the copyright of Screen-Camera (SC) images, facilitating post-hoc certification against infringement. These techniques typically employ heuristic mathematical modeling or supervised neural network fitting as the noise layer, to enhance watermarking robustness against SC. However, both strategies cannot fundamentally achieve an effective approximation of SC noise. Mathematical simulation suffers from biased approximations due to the incomplete decomposition of the noise and the absence of interdependence among the noise components. Supervised networks require paired data to train the noise-fitting model, and it is difficult for the model to learn all the features of the noise. To address the above issues, we propose Simulation-to-Real (S2R). Specifically, an unsupervised noise layer employs unpaired data to learn the discrepancy between the modeled simulated noise distribution and the real-world SC noise distribution, rather than directly learning the mapping from sharp images to real-world images. Learning this transformation from simulation to reality is inherently simpler, as it primarily involves bridging the gap in noise distributions, instead of the complex task of reconstructing fine-grained image details. Extensive experimental results validate the efficacy of the proposed method, demonstrating superior watermark robustness and generalization compared to state-of-the-art methods. Xin Liao 0001, Baowei Wang, Han Fang 0004, Xiaoshuai Wu, Grace Guiling Wang |
AAAI | 5 |
| 2026 | Locate Core, Refine Path: A Training-Free Closed-Loop Paradigm for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) aims to dynamically segment a target object across video sequences based on a given natural language expression. Although recent decoupled methods surpass end-to-end models, they suffer from a critical trade-off: relying on fine-tuned large language models for key frame selection incurs high computational costs, whereas employing sparse sampling risks missing optimal reference frames. Furthermore, the subsequent mask propagation lacks self-correction mechanisms, leading to irreversible cumulative errors over time. In this work, we propose LCRP, a training-free framework driven by two novel components. The Hybrid Candidate Sampling (HCS) module significantly improves frame selection by integrating lightweight semantic parsing with CLIP-based filtering. Meanwhile, the Multi-Anchor Refinement (MAR) module utilizes anchor frames as dynamic checkpoints to detect temporal drift and execute localized re-propagation. Extensive experiments indicate that our proposed method achieves state-of-the-art results. Code is available at https://github.com/Crystal535/Locate-Core-Refine-Path. Jizhe Yu, Hao Zhang 0218, Xiya Bu, Yuhang Duan, Xiaoshuai Wu, Yu Liu 0035 |
ICMR | 5 |
| 2026 | Unifying Gradient Leakage Attacks Against Privacy-Protected Federated Learning in IoT NetworksabstractFederated learning (FL) is a transformative paradigm for the Internet of Things (IoT), enabling decentralized model training across distributed IoT devices while reducing reliance on centralized data collection. Crucially, FL cuts communication overhead, an essential benefit in bandwidth-limited IoT environments. However, repeated gradient exchanges between edge clients (e.g., sensors, mobile devices) and the central server expose vulnerabilities to gradient leakage attacks (GLAs), allowing adversaries to reconstruct private training data from shared gradients. While various gradient protection strategies, such as differential privacy, sparsification, and clipping, have been introduced to mitigate this risk, most existing GLAs are designed for specific protection schemes and fail under heterogeneous deployments. In this work, we propose a unified GLA framework that targets diverse gradient protection techniques in FL systems. Our approach tackles two core challenges: (i) aligning protected gradients with their raw counterparts to enable robust feature extraction, and (ii) identifying critical features for efficient and accurate data reconstruction.We introduce a Taylor-based gradient approximation method for alignment and design a feature reconstructor that enhances both performance and computational efficiency. Extensive experiments across various FL scenarios demonstrate the framework’s superior reconstruction capability under different protection schemes, emphasizing the need for robust privacy-preserving mechanisms in IoT networks. Hui Zhou 0014, Zheng Qin 0001, Peng Sun 0003, Yipeng Zou, Xiaoshuai Wu |
IEEE Internet Things J. | 6 |
| 2026 | Versatile and harmless deepfake proactive forensics via conditional watermarking
Xiaoshuai Wu, Xin Liao 0001, Jie Zhang 0004, Jinlin Guo |
Inf. Sci. | 1 |
| 2025 | Vision-Text Interaction with Orientation-Awareness for Referring Remote Sensing Image Segmentation
Xiaoshuai Wu, Yu Liu 0035, Kaiping Xu, Hao Zhang 0218, Jizhe Yu |
ICANN (2) | 1 |
| 2025 | PA2Net: Pyramid Attention Aggregation Network for Saliency Detection
Jizhe Yu, Yu Liu 0035, Xiaoshuai Wu, Kaiping Xu, Jiangquan Li |
MMM (3) | 3 |
| 2025 | Flexible Partial Screen-Shooting Watermarking With Provable RobustnessabstractScreen-shooting watermarking is an effective means of protecting screen content from unauthorized capture and illegal dissemination. However, existing methods are primarily designed for full-image capture, making them ineffective for partial screen-shooting prevalent in real-world scenarios. To address this limitation, we propose FPSMark, a flexible watermarking method tailored for partial screen-shooting that embeds consistent watermarks in multiple uniformly distributed cover blocks. Specifically, considering that robustness requirements vary according to the layout of each image, we model the mathematical relationship between the watermark block count and robustness, proving the flexibility of FPSMark in ensuring partial screen-shooting robustness. Moreover, partial screen-shooting disrupts watermark synchronization, posing challenges for precise watermark localization. To overcome this, we design an intrinsic signal localization network optimized with a hybrid loss. The localization network exploits the inherent distinctions between the watermark and non-watermark features, while the hybrid loss constrains the network at three dimensions: pixel-level, region-level, and sample-level. Experimental results demonstrate the superiority of FPSMark, showing robust performance across partial capture percentages. Its extraction accuracy exceeds 98% even with only half of the image captured, and it achieves 82% accuracy at a 40% capture ratio, whereas existing methods achieve only around 50% under the same conditions. Xin Liao 0001, Han Fang 0004, Jinlin Guo, Xiaoshuai Wu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Are Watermarks Bugs for Deepfake Detectors? Rethinking Proactive Forensics
Xiaoshuai Wu, Xin Liao 0001, Bo Ou, Zheng Qin 0001 |
IJCAI | 1 |
| 2024 | SafePaint: Anti-forensic Image Inpainting with Domain Adaptation
Dunyun Chen, Xin Liao 0001, Xiaoshuai Wu |
ACM Multimedia | 3 |
| 2024 | Screen-Shooting Resistant Watermarking With Grayscale Deviation SimulationabstractWith the prevalence of electronic devices in our daily lives, content leakages frequently occur, and to enable leakage tracing, screen-shooting resistant watermarking has attracted tremendous attention. However, current studies often overlook a thoughtful investigation of the cross-media screen-camera process and fail to consider the effect of grayscale deviation on the screen. In this paper, we proposescreen-shootingdistortionsimulation ($\bf {SSDS}$), which involves a grayscale deviation function for constructing a more practical noise layer. We divide SSDS into screen displaying and camera shooting. For screen displaying, different viewing angles result in grayscale deviation with distinct intensities, and we simulate the distortions by modeling the relative position of the viewing point and the screen plane. For camera shooting, a series of distortion functions are used to approximate the perturbations in the camera pipeline, including defocus blur, noise and JPEG compression. Furthermore, the gradient-guided encoder is designed to conduct the embedding in the texture region using a modification cost map. Experimental results show that our proposed watermarking framework outperforms the state-of-the-art methods in terms of robustness and visual quality. Yiyi Li, Xin Liao 0001, Xiaoshuai Wu |
IEEE Trans. Multim. | 3 |
| 2023 | SepMark: Deep Separable Watermarking for Unified Source Tracing and Deepfake DetectionabstractMalicious Deepfakes have led to a sharp conflict over distinguishing between genuine and forged faces. Although many countermeasures have been developed to detect Deepfakes ex-post, undoubtedly, passive forensics has not considered any preventive measures for the pristine face before foreseeable manipulations. To complete this forensics ecosystem, we thus put forward the proactive solution dubbed SepMark, which provides a unified framework for source tracing and Deepfake detection. SepMark originates from encoder-decoder-based deep watermarking but with two separable decoders. For the first time the deep separable watermarking, SepMark brings a new paradigm to the established study of deep watermarking, where a single encoder embeds one watermark elegantly, while two decoders can extract the watermark separately at different levels of robustness. The robust decoder termed Tracer that resists various distortions may have an overly high level of robustness, allowing the watermark to survive both before and after Deepfake. The semi-robust one termed Detector is selectively sensitive to malicious distortions, making the watermark disappear after Deepfake. Only SepMark comprising of Tracer and Detector can reliably trace the trusted source of the marked face and detect whether it has been altered since being marked; neither of the two alone can achieve this. Extensive experiments demonstrate the effectiveness of the proposed SepMark on typical Deepfakes, including face swapping, expression reenactment, and attribute editing. Code will be available at https://github.com/sh1newu/SepMark. Xiaoshuai Wu, Xin Liao 0001, Bo Ou |
ACM Multimedia | 1 |
| 2023 | FAMM: Facial Muscle Motions for Detecting Compressed Deepfake Videos Over Social NetworksabstractAs a face manipulation technique, the misuse of Deepfakes poses potential threats to the state, society, and individuals. Several countermeasures have been proposed to reduce the negative effects produced by Deepfakes. Current detection methods achieve satisfactory performance in dealing with uncompressed videos. However, videos are generally compressed when spread over social networks because of limited bandwidth and storage space, which generates compression artifacts and the detection performance inevitably decreases. Hence, how to effectively identify compressed Deepfake videos over social networks becomes a significant problem in video forensics. In this paper, we propose a facial-muscle-motions-based (FAMM) framework to solve the problem of compressed Deepfake video detection. Specifically, we first locate faces from consecutive frames and extract landmarks from the face images. Then, continuous facial landmarks are utilized to construct facial muscle motion features by modeling the five sensory and face regions. Finally, we fuse the diverse forensic knowledge using Dempster-Shafer theory and provide the final detection results. Furthermore, we demonstrate the effectiveness of FAMM through analyzing mutual information, compression procedure, and facial landmarks for compressed Deepfake videos. Theoretical analyses illustrate that compression does not affect facial muscle motion feature construction and the differences in designed features exist between the real and Deepfake videos. Extensive experimental results conclude that the proposed method outperforms the state-of-the-art methods in detecting compressed Deepfake videos. More importantly, FAMM achieves comparable detection performance on compressed videos that are over real-world social networks. Xin Liao 0001, Yumei Wang, Tianyi Wang 0006, Xiaoshuai Wu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Group Consensus for Heterogeneous Multiagent Systems With Time Delays Based on Frequency Domain ApproachabstractThis article investigates the group consensus problem for heterogeneous multiagent systems with time delays via pinning control. Under the designed control protocol, agents in the system could be grouped arbitrarily, and agents in the same subgroup could converge to a constraint position. Meanwhile, the kinetics of agents in the same subgroup could be same or different. Both the fixed and switching topologies are considered. Based on the frequency domain method and stability theory, sufficient conditions for the system achieve group consensus are derived. Finally, several numerical examples are presented to verify the performance of the control protocol. Fenglan Sun, Xiaoshuai Wu, Jürgen Kurths, Wei Zhu 0004 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Sign steganography revisited with robust domain selection
Xiaoshuai Wu, Yanli Chen 0002, Ming Xu 0001, Ning Zheng 0001, Xiangyang Luo 0001 |
Signal Process. | 1 |
| 2021 | Secure reversible data hiding in encrypted images based on adaptive prediction-error labeling
Xiaoshuai Wu, Ming Xu 0001, Ning Zheng 0001 |
Signal Process. | 1 |