VLDB 2026 Research / reviewers in the wild / expert
Shuwang Zhou
dblp:11/9937
· DBLP profile ↗
15ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0003-3471-7563ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Facial Privacy Protection for Remote PhotoplethysmographyabstractRemote photoplethysmography (rPPG) has emerged as a crucial technology for contactless health monitoring, providing a convenient and non-invasive method to measure physiological signals from skin videos. Because face videos are commonly used for rPPG measurements, privacy concerns arise due to the inherent sensitivity of facial biometric data. Concerns about privacy breaches in facial video recordings have hindered telemedicine advancements and limited the creation of large-scale medical datasets, restricting the development of rPPG-based technologies. Additionally, the necessity to transmit and store rPPG videos in these applications necessitates video compression as an indispensable step. However, existing facial privacy protection techniques and video compression methods tend to degrade the rPPG signal in videos. To address these challenges, this study proposes a straightforward yet effective face anonymization module-a plug-and-play component employing spatial pixel redistribution algorithms to achieve: 1) eliminating identifiable biometric features while preserving the physiological information; 2) facilitating video compression by a macroblock reassembly strategy based on chromaticity clustering. Experiments on three rPPG datasets illustrate that the proposed method preserves physiological information in anonymized videos while effectively facilitating video compression. Jieying Wang, Caifeng Shan, Shuwang Zhou, Minglei Shu |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | CorGPT: Coronary Angiography Imaging Analysis Using Large Medical Vision-Language Models
Baoqian Huang, Hongkuan Zhang, Shuwang Zhou |
ICIC (25) | 5 |
| 2025 | GRAIL: Guided Retrieval via Layer-Wise Discrepancy from Extrapolated Final Distributions
Hongkuan Zhang, Shuwang Zhou |
ICONIP (4) | 4 |
| 2025 | Disruptive Attacks on Face Swapping via Low-Frequency Perceptual PerturbationsabstractDeepfake technology, driven by Generative Adversarial Networks (GANs), poses significant risks to privacy and societal security. Existing detection methods are predominantly passive, focusing on post-event analysis without preventing attacks. To address this, we propose an active defense method based on low-frequency perceptual perturbations to disrupt face-swapping manipulation, reducing the performance and naturalness of generated content. Unlike prior approaches that used low-frequency perturbations to impact classification accuracy, our method directly targets the generative process of deepfake techniques.We combine frequency and spatial domain features to strengthen defenses. By introducing artifacts through low-frequency perturbations while preserving high-frequency details, we ensure the output remains visually plausible. Additionally, we design a complete architecture featuring an encoder, a perturbation generator, and a decoder, leveraging discrete wavelet transform (DWT) to extract low-frequency components and generate perturbations that disrupt facial manipulation models. Experiments on CelebA-HQ and LFW demonstrate significant reductions in face-swapping effectiveness, improved defense success rates, and preservation of visual quality. Mengxiao Huang, Minglei Shu, Shuwang Zhou, Zhaoyang Liu 0002 |
IJCNN | 3 |
| 2025 | An Effective GANs-Based Method for High-Information-Value Multimodal Wearable Sensor Data SynthesisabstractIn the field of Human Activity Recognition (HAR), current GANs-based approaches for sensor data synthesis frequently fail to account for the varying informational significance of data samples, often treating them with uniform importance. This leads to the generation of a substantial number of data samples with limited informational utility by the trained generator, which ultimately fail to contribute meaningfully to HAR tasks. This paper introduces an effective GANs-based sensor data synthesis method for generating more data samples with high information value. Firstly, we incorporate active learning mechanisms into the adversarial training architecture to guide a more comprehensive learning of the data distribution. Secondly, we enhance the model’s ability to learn temporal, spatial, and spatial-temporal features by combining convolutional, recurrent, and self-attention modules. Thirdly, we validate the proposed method on the public dataset using both quantitative and qualitative metrics. The experimental results demonstrate that the proposed method can synthesize high-information-value multimodal sensor data, which are more valuable for training downstream HAR models. Yang Gu 0001, Shuwang Zhou |
IJCNN | 3 |
| 2025 | SES-Net: Multi-dimensional Spot-Edge-Surface Network for Nuclei Segmentation
Congjian Lu, Shuwang Zhou, Ke Shan, Hongkuan Zhang, Zhaoyang Liu 0002 |
MMM (4) | 2 |
| 2025 | KEREM: Enhancing Reliability and Transparency in Medical QA through LLM and Knowledge Graph FusionabstractOpen-domain medical question answering (QA) systems face significant challenges in achieving accurate reasoning and transparent explanations. In this study, we propose KEREM (Knowledge graph Enhanced Reasoning with Explainable Modeling), a novel framework that integrates large language models (LLMs) with structured medical knowledge graphs (KGs) to enable deep multimodal joint reasoning. KEREM supports multi-hop inference and causal chain explanations through four key modules: input processing and entity alignment, knowledge subgraph construction, joint reasoning and path control, and answer generation with natural language explanations. We evaluate KEREM on two real-world medical QA benchmarks—CMCQA and ChatDoctor-5k—where it consistently outperforms existing baselines in terms of answer accuracy, reasoning transparency, and structural consistency. Furthermore, lightweight supervised fine-tuning demonstrates KEREM’s strong contextual transferability and significantly improves its generation quality in specialized clinical settings. These findings highlight KEREM’s effectiveness in generating accurate answers and causal explanations, establishing a solid foundation for trustworthy medical QA in high-stakes clinical domains. Shaojie Dong, Zhe Zhu, Pengyao Xu, Ke Shan, Shuwang Zhou |
SMC | 5 |
| 2025 | Physiological Information Preserving Video Compression for rPPGabstractRemote photoplethysmography (rPPG) has recently attracted much attention due to its non-contact measurement convenience and great potential in health care and computer vision applications. Early rPPG studies were mostly developed on self-collected uncompressed video data, which limited their application in scenarios that require long-distance real-time video transmission, and also hindered the generation of large-scale publicly available benchmark datasets. In recent years, with the popularization of high-definition video and the rise of telemedicine, the pressure of storage and real-time video transmission under limited bandwidth have made the compression of rPPG video inevitable. However, video compression can adversely affect rPPG measurements. This is due to the fact that conventional video compression algorithms are not specifically proposed to preserve physiological signals. Based on this, we propose a video compression scheme specifically designed for rPPG application. The proposed approach consists of three main strategies: 1) facial ROI-based computational resource reallocation; 2) rPPG signal preserving bit resource reallocation; and 3) temporal domain up- and down-sampling coding. UBFC-rPPG, ECG-Fitness, and a self-collected dataset are used to evaluate the performance of the proposed method. The results demonstrate that the proposed method can preserve almost all physiological information after compressing the original video to 1/60 of its original size. The proposed method is expected to promote the development of telemedicine and deep learning techniques relying on large-scale datasets in the field of rPPG measurement. Jieying Wang, Caifeng Shan, Shuwang Zhou, Minglei Shu |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | An Encoder-Decoder Based Approach for ECG Delineation
Xiangju Kong, Shuwang Zhou, Tianlei Gao, Zhe Zhu |
ICONIP (4) | 2 |
| 2024 | A Multi-Lead Electrocardiogram Signal Classification Method Based on Temporal and Multi-View Contrastive LearningabstractThe rise in wearable devices has led to the generation of a large amount of unlabeled electrocardiogram (ECG) data. Effectively utilising this data has been a challenge. One approach to address this issue is contrastive learning. However, most existing contrastive learning methods based on data augmentation primarily utilize the augmented electro-cardiogram (ECG) signals for comparison. These approaches have certain drawbacks. While the augmented ECG signals may help highlight certain features, they could also potentially mask important variations in the original signal, leading to the loss of crucial information. Moreover, solely relying on augmented data may lead to the model relying too heavily on specific variation patterns, overlooking the authentic features present in the original signal, resulting in poor generalization and robustness for downstream tasks. To address this issue, we propose a multi-lead electrocardiogram signal classification method based on temporal and multi-view contrastive learning. The method utilizes the time invariance of ECG signals and the multi-view information provided by the original ECG signals and augmented signals for contrastive learning. Through this approach, it not only addresses the limitations of relying solely on augmented signals but also leverages the prior knowledge that ECG signal categories remain stable over short periods of time to fully model the time context of ECG signals, capturing dynamic features. The integration of the temporal context comparison module and the multi-view comparison module significantly enhances the performance of downstream classification tasks. The outcomes of our experiments show that our method performs better on four datasets than previous approaches, which even exceeds supervised performance. When we pretrain on the SPH dataset and fine-tune on the PTB-XL dataset, our approach shows the best performance. Specifically, our method's AUROC exceeds that of the best baseline model by 4.9% and surpasses supervised performance by 1.3%. Luyao Li, Hui Liu 0046, Shuwang Zhou, Zhaoyang Liu 0002, Minglei Shu |
SMC | 3 |
| 2023 | Dynamic Facial Expression Recognition in Unconstrained Real-World Scenarios Leveraging Dempster-Shafer Evidence Theory
Tianyi Wang 0006, Shuwang Zhou, Minglei Shu |
ICANN (2) | 3 |
| 2023 | A lightweight 2-D CNN model with dual attention mechanism for heartbeat classification
Hongfu Xie, Shuwang Zhou, Tianlei Gao, Minglei Shu |
Appl. Intell. | 3 |
| 2023 | Curvilinear Structure Tracking Based on Dynamic Curvature-penalized Geodesics
Li Liu 0065, Shuwang Zhou, Minglei Shu, Laurent D. Cohen, Da Chen 0002 |
Pattern Recognit. | 3 |
| 2022 | Dynamic knowledge graph reasoning based on deep reinforcement learning
Hao Liu 0066, Shuwang Zhou, Changfang Chen, Tianlei Gao, Jiyong Xu, Minglei Shu |
Knowl. Based Syst. | 2 |
| 2010 | Constraint-based sensor network nodes particle swarm search localization algorithmabstractA constraint-based sensor network nodes particle swarm search localization algorithm (CPL) is presented. First of all, a constraint domain of an unknown node must be determined; Then the positions which meet specific criteria is searched out by particle swarm optimization algorithm and the searching results within the constraint domain are recorded; Finally, the unknown node's localization can be obtained by calculating the average recording results. As is shown in the experiment results, CPL has strong robustness, and comparing with normal schemes such as least square method (LS), CPL's positioning accuracy can improve 50% when the ranging error is 35%. Shuwang Zhou, Yinglong Wang 0001, Qiang Guo 0003, Nuo Wei |
CSCWD | 1 |