Zhaodong Sun

dblp:277/9723 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-0597-0765ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Radar-APLANC: Unsupervised Radar-based Heartbeat Sensing via Augmented Pseudo-Label and Noise Contrast
abstract
Frequency Modulated Continuous Wave (FMCW) radars can measure subtle chest wall oscillations to enable non-contact heartbeat sensing. However, traditional radar-based heartbeat sensing methods face performance degradation due to noise. Learning-based radar methods achieve better noise robustness but require costly labeled signals for supervised training. To overcome these limitations, we propose the first unsupervised framework for radar-based heartbeat sensing via Augmented Pseudo-Label and Noise Contrast (Radar-APLANC). We propose to use both the heartbeat range and noise range within the radar range matrix to construct the positive and negative samples, respectively, for improved noise robustness. Our Noise-Contrastive Triplet (NCT) loss only utilizes positive samples, negative samples, and pseudo-label signals generated by the traditional radar method, thereby avoiding dependence on expensive ground-truth physiological signals. We further design a pseudo-label augmentation approach featuring adaptive noise-aware label selection to improve pseudo-label signal quality. Extensive experiments on the Equipleth dataset and our collected radar dataset demonstrate that our unsupervised method achieves performance comparable to state-of-the-art supervised methods.
Zhaodong Sun, Xu Cheng 0003, Zuxian He
AAAI2
2026 PhysFlow: Frequency-Selective Flow Matching with Dual-Stream Expert Fusion for Remote Photoplethysmography
abstract
Remote photoplethysmography (rPPG) enables non-contact heart rate monitoring through facial videos, offering significant potential for health monitoring and telemedicine applications. Existing methods typically learn direct mappings from facial videos to rPPG signals. However, they often treat the rPPG signal holistically, overlooking the different contributions of different frequency bands, which may miss potential band-specific features and lead to inaccurate estimation and reduced robustness. To overcome this limitation, we propose PhysFlow, a novel Flow Matching framework for robust rPPG measurement. PhysFlow introduces a frequency-selective wavelet loss to emphasize physiologically important frequency bands while suppressing noise. The framework employs a highly extensible Dual-Stream Velocity Estimator to predict velocity fields from spatiotemporal and signal perspectives, integrated via an Expert Fusion mechanism. Additionally, a Wavelet Enhancement module is introduced to transform the intermediate flow state into multi-scale spectral-temporal features to facilitate velocity field prediction. Besides, classifier-free guidance is incorporated to enhance conditional control. We validate PhysFlow on four datasets, showing that our method significantly outperforms existing approaches in both intra-dataset and cross-dataset evaluations. The code is available at https://github.com/reimu996/PhysFlow/.
Youchen Luo, Zhaodong Sun, Huiyu Yang, Wenye Geng, Yuwei Chen 0005
ICMR2
2026 UDA-rPPG: Unsupervised Geometric-Physiological Domain Anchoring for Low-Light rPPG Measurement
abstract
Remote photoplethysmography (rPPG) is a critical technique for non-contact monitoring of human vital signs using facial video data. Most of the existing rPPG approaches, either supervised ones relying on ground-truth physiological signals or less constrained unsupervised ones, primarily address the problem of inaccurate physiological measurements under normal lighting conditions. However, few works focus on handling physiological measurements in extremely low-light scenarios. To this end, we propose an unsupervised geometric-physiological domain anchoring for low-light rPPG measurement (UDA-rPPG). Firstly, we develop a geometric anchoring video enhancement module (GAEM) that can enhance video brightness while preserving rPPG signals, achieving accurate geometric-domain face anchoring. Secondly, we introduce a low-light stable spatial-temporal network (LS-Phys), which focuses on high-frequency information to mitigate noise in low-light scenarios. Finally, a novel highest-peak priority learning strategy is presented to learn physiological-domain rPPG signal anchoring by emphasizing peak information, which enhances the robustness of rPPG measurements in low-light environments. Additionally, we construct a comprehensive low-light rPPG dataset (LRPD) that contains both visible and near-infrared videos under low-light scenarios. Extensive experiments demonstrate the superior performance of our approach over state-of-the-art unsupervised rPPG methods in different light conditions and verify the generalization of UDA-rPPG on cross-dataset testing. Our code and dataset are available at https://github.com/wwenmaositu/LS-rPPG-LRPD.
Xu Cheng 0003, Zhaodong Sun
IEEE Trans. Circuits Syst. Video Technol.4
2025 From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification
abstract
Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities, raising privacy and ownership concerns that make existing centralized training impractical for VI-ReID. To tackle these challenges, we propose L2RW, a benchmark that brings VI-ReID closer to real-world applications. The rationale of L2RW is that integrating decentralized training into VI-ReID can address privacy concerns in scenarios with limited data-sharing regulation. Specifically, we design protocols and corresponding algorithms for different privacy sensitivity levels. In our new benchmark, we ensure the model training is done in the conditions that: 1) data from each camera remains completely isolated, or 2) different data entities (e.g., data controllers of a certain region) can selectively share the data. In this way, we simulate scenarios with strict privacy constraints which is closer to real-world conditions. Intensive experiments with various server-side federated algorithms are conducted, showing the feasibility of decentralized VI-ReID training. Notably, when evaluated in unseen domains (i.e., new data entities), our L2RW, trained with isolated data (privacy-preserved), achieves performance comparable to SOTAs trained with shared data (privacy-unrestricted). We hope this work offers a novel research entry for deploying VI-ReID that fits real-world scenarios and can benefit the community.
Hao Yu 0015, Xu Cheng 0003, Haoyu Chen 0001, Zhaodong Sun, Guoying Zhao 0001
CVPR5
2025 FusionPhys: A Flexible Framework for Fusing Complementary Sensing Modalities in Remote Physiological Measurement
Chenhang Ying, Huiyu Yang, Jieyi Ge, Zhaodong Sun, Xu Cheng 0003, Kui Ren 0001
ICCV4
2025 Evidential Remote Physiological Measurement via Uncertainty-aware Fusion of Video and RF
abstract
Remote physiological measurement enables the capture of vital signals in a non-contact way, which offers significant potential for various applications. Monitoring these signals is achieved through video cameras or radio frequency (RF) sensors, with recent few methods attempting to fuse both sources to leverage complementary patterns for enhanced accuracy. However, these two modalities operate on distinct principles, where video-based methods detect subtle facial color changes from blood volume variations, while RF-based methods capture subtle body vibration due to heartbeats. In practical applications, they may encounter interference at different occasions. Treating these modalities as equally reliable in all situations can lead to suboptimal fusion. To address this issue, we propose an evidential video-RF fusion framework for robust remote physiological signal measurement. We design an uncertainty regression head for each uni-modality, which estimates uncertainty features together with the corresponding physiological signal in each branch. Then an evidential multi-modal fusion module is employed to dynamically fuse the two modalities according to their uncertainty. Extensive experiments carried on public and self-collected datasets show that the proposed method not only achieves superior fusion performance on easy data collected under well-controlled environment, it also generalizes well to unseen data which represents challenging practical conditions that one or both sensors are disturbed.
Jieyi Ge, Zhaodong Sun, Wei Peng 0009, Chenhang Ying, Yuwei Chen 0005, Kui Ren 0001
ACM Multimedia2
2025 Learning From Yourself to Others for Unsupervised Visible-Infrared Re-Identification
abstract
Unsupervised visible-infrared person re-identification (US-VI-ReID) aims to match unlabeled pedestrian images captured under varying lighting conditions. The key challenge lies in generating accurate pseudo-labels, alongside alleviating the significant modality gap between visible and infrared modalities. Existing methods mainly focus on mitigating the effects of noisy labels through loss functions during backward propagation. However, these noisy labels already influence the forward propagation, leading to incorrect cross-modality correspondences. To address this issue, we propose a Hierarchical Centrality Collaborative Learning (HCCL) framework for US-VI-ReID, which proactively identifies noisy labels during the forward propagation. The rationale behind HCCL is that intra-modality refinement serves as the foundation for establishing cross-modality correspondences, reflecting the principle of learning from yourself to others. For intra-modality learning, we propose a Closeness Centrality Selection (CCS), quantifying sample confidence using closeness centrality to identify noisy samples. By discarding the noisy samples during forward propagation, CCS mitigates their adverse effects and ensures identity-consistent representation learning. For cross-modality learning, a Hierarchical Consistency Matching (HCM) is proposed to establish local instance-level label associations by leveraging bidirectional consistency with the most reliable samples identified during intra-modality learning. These local associations are then propagated to guide the global cluster-level cross-modality correspondences. Extensive experiments demonstrate that our HCCL achieves competitive performance on mainstream datasets, even surpassing some supervised counterparts. Additionally, outstanding results on corrupted datasets verify its generalizability and robustness.
Wenhui Ji, Xu Cheng 0003, Zhaodong Sun, Guoying Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Biometric Authentication Based on Enhanced Remote Photoplethysmography Signal Morphology
abstract
Remote photoplethysmography (rPPG) is a non-contact method for measuring cardiac signals from facial videos, offering a convenient alternative to contact photoplethysmography (cPPG) obtained from contact sensors. Recent studies have shown that each individual possesses a unique cPPG signal morphology that can be utilized as a biometric identifier, which has inspired us to utilize the morphology of rPPG signals extracted from facial videos for person authentication. Since the facial appearance and rPPG are mixed in the facial videos, we first de-identify facial videos to remove facial appearance while preserving the rPPG information, which protects facial privacy and guarantees that only rPPG is used for authentication. The de-identified videos are fed into an rPPG model to get the rPPG signal morphology for authentication. In the first training stage, unsupervised rPPG training is performed to get coarse rPPG signals. In the second training stage, an rPPG-cPPG hybrid training is performed by incorporating external cPPG datasets to achieve rPPG biometric authentication and enhance rPPG signal morphology. Our approach needs only de-identified facial videos with subject IDs to train rPPG authentication models. The experimental results demonstrate that rPPG signal morphology hidden in facial videos can be used for biometric authentication. The code is available at https://github.com/zhaodongsun/rppg_biometrics.
Zhaodong Sun, Jukka Komulainen, Guoying Zhao 0001
IJCB1
2024 Contrast-Phys+: Unsupervised and Weakly-Supervised Video-Based Remote Physiological Measurement via Spatiotemporal Contrast
abstract
Video-based remote physiological measurement utilizes facial videos to measure the blood volume change signal, which is also called remote photoplethysmography (rPPG). Supervised methods for rPPG measurements have been shown to achieve good performance. However, the drawback of these methods is that they require facial videos with ground truth (GT) physiological signals, which are often costly and difficult to obtain. In this paper, we propose Contrast-Phys+, a method that can be trained in both unsupervised and weakly-supervised settings. We employ a 3DCNN model to generate multiple spatiotemporal rPPG signals and incorporate prior knowledge of rPPG into a contrastive loss function. We further incorporate the GT signals into contrastive learning to adapt to partial or misaligned labels. The contrastive loss encourages rPPG/GT signals from the same video to be grouped together, while pushing those from different videos apart. We evaluate our methods on five publicly available datasets that include both RGB and Near-infrared videos. Contrast-Phys+ outperforms the state-of-the-art supervised methods, even when using partially available or misaligned GT signals, or no labels at all. Additionally, we highlight the advantages of our methods in terms of computational efficiency, noise robustness, and generalization.
Zhaodong Sun
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Many birds, one stone: Medical image segmentation with multiple partially labeled datasets
Qing Liu 0003, Hailong Zeng, Zhaodong Sun, Guoying Zhao 0001, Yixiong Liang
Pattern Recognit.3
2022 Contrast-Phys: Unsupervised Video-Based Remote Physiological Measurement via Spatiotemporal Contrast
Zhaodong Sun
ECCV (12)1
2022 Privacy-Phys: Facial Video-Based Physiological Modification for Privacy Protection
abstract
The invisible remote photoplethysmography (rPPG) signals in facial videos can reveal the cardiac rhythm and physiological status. Recent studies show that rPPG is a non-contact way for emotion recognition, disease detection, and biometric identification, which means there is a potential privacy problem about physiological information leakage from facial videos. Therefore, it is essential to process facial videos to prevent rPPG extraction in privacy-sensitive situations such as online video meetings. In this letter, we propose Privacy-Phys, a novel method based on a pre-trained 3D convolutional neural network, to modify rPPG in facial videos for privacy protection. Our experimental results show that our approach can modify rPPG signals in facial videos more effectively and efficiently than the previous baseline. Our method can be applied to process facial videos in online video meetings or video-sharing platforms to prevent rPPG from being captured maliciously.
Zhaodong Sun
IEEE Signal Process. Lett.1
2022 Non-Contact Atrial Fibrillation Detection From Face Videos by Learning Systolic Peaks
abstract
OBJECTIVE: We propose a non-contact approach for atrial fibrillation (AF) detection from face videos. METHODS: Face videos, electrocardiography (ECG), and contact photoplethysmography (PPG) from 100 healthy subjects and 100 AF patients are recorded. Data recordings from healthy subjects are all labeled as healthy. Two cardiologists evaluated ECG recordings of patients and labeled each recording as AF, sinus rhythm (SR), or atrial flutter (AFL). We use the 3D convolutional neural network for remote PPG monitoring and propose a novel loss function (Wasserstein distance) to use the timing of systolic peaks from contact PPG as the label for our model training. Then a set of heart rate variability (HRV) features are calculated from the inter-beat intervals, and a support vector machine (SVM) classifier is trained with HRV features. RESULTS: Our proposed method can accurately extract systolic peaks from face videos for AF detection. The proposed method is trained with subject-independent 10-fold cross-validation with 30 s video clips and tested on two tasks. 1) Classification of healthy versus AF: the accuracy, sensitivity, and specificity are 96.00%, 95.36%, and 96.12%. 2) Classification of SR versus AF: the accuracy, sensitivity, and specificity are 95.23%, 98.53%, and 91.12%. In addition, we also demonstrate the feasibility of non-contact AFL detection. CONCLUSION: We achieve good performance of non-contact AF detection by learning systolic peaks. SIGNIFICANCE: non-contact AF detection can be used for self-screening of AF symptoms for suspectable populations at home or self-monitoring of AF recurrence after treatment for chronic patients.
Zhaodong Sun, Juhani Junttila, Mikko Tulppo, Tapio Seppänen
IEEE J. Biomed. Health Informatics1
2021 A Plug-and-Play Deep Image Prior
abstract
Deep image priors (DIP) offer a novel approach for the regularization that leverages the inductive bias of a deep convolutional architecture in inverse problems. However, the quality of DIP approaches often degrades when the number of iterations exceeds a certain threshold due to overfitting. To mitigate this effect, this work incorporates a plug-and-play prior scheme which can accommodate additional regularization steps within a DIP framework. Our modification is achieved using an augmented Lagrangian formulation of the problem, and is solved using an Alternating Direction Method of Multipliers (ADMM) variant, which can capture existing DIP approaches as a special case. We show experimentally that our ADMM-based DIP pairing outperforms competitive baselines in PSNR while exhibiting less overfitting.
Zhaodong Sun, Fabian Latorre, Thomas Sanchez, Volkan Cevher
ICASSP1