VLDB 2026 Research / reviewers in the wild / expert
Hao Lu 0009
dblp:72/5422-9
· DBLP profile ↗
27ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0002-2241-6598ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 15 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distributed resource orchestration in heterogeneous multi-cloud environments: A shadow-price based collaborative mechanism with individual rationality
Yang Song 0022, Hao Lu 0009, Xingwei Wang 0001, Min Huang 0001 |
Comput. Networks | 2 |
| 2026 | Align the GAP: Prior-Based Unified Multi-task Remote Physiological Measurement Framework For Domain Generalization and PersonalizationabstractAbstract Multi-source synsemantic domain generalization (MSSDG) for multi-task remote physiological measurement seeks to enhance the generalizability of these metrics and attracts increasing attention. However, challenges like partial labeling and environmental noise may disrupt task-specific accuracy. Meanwhile, given that real-time adaptation is necessary for personalized products, the test-time personalized adaptation (TTPA) after MSSDG is also worth exploring, while the gap between previous generalization and personalization methods is significant and hard to fuse. Thus, we proposed a unified framework for MSSD G and TTP A employing P riors ( GAP ) in biometrics and remote photoplethysmography (rPPG). We first disentangled information from face videos into invariant semantics, individual bias, and noise. Then, multiple modules incorporating priors and our observations were applied in different stages and for different facial information. Then, based on the different principles of achieving generalization and personalization, our framework could simultaneously address MSSDG and TTPA under multi-task remote physiological estimation with minimal adjustments. We expanded the MSSDG benchmark to the TTPA protocol on six publicly available datasets and introduced a new real-world driving dataset with complete labeling. Extensive experiments that validated our approach, and the codes along with the new dataset are in https://github.com/WJULYW/GAP . Jiyao Wang 0002, Xiao Yang 0025, Hao Lu 0009, Dengbo He, Kaishun Wu |
Int. J. Comput. Vis. | 3 |
| 2026 | Dynamic chunking-driven intelligent transmission mechanism for distributed systems
Enliang Lv, Xingwei Wang 0001, Bo Yi 0002, Hao Lu 0009, Min Huang 0001, Yue Kou, Keqin Li 0001 |
Knowl. Based Syst. | 4 |
| 2026 | Intelligent Cross-Domain Data Orchestration in Computing Power Networks: An Attention-Enhanced Multi-Agent Reinforcement Learning ApproachabstractComputing Power Networks (CPNs) integrate heterogeneous resources across the cloud–edge–end continuum to support wide-area distributed computational services, but the geographical separation of computation and data makes cross-domain data access a major bottleneck. Intelligent cross-domain data orchestration in CPNs is difficult because replica selection and end-to-end path planning must be jointly optimized under Service Level Agreement (SLA) and resource constraints, while each domain observes only partial congestion and resource information. This paper presents AE-MAAC, an attention-enhanced multi-agent reinforcement learning framework that formulates cross-domain data orchestration as a Multi-Agent Markov Decision Process (MMDP) with a hierarchical composite action space and constraint-aware masking under a centralized-training–decentralized-execution paradigm. An attention-based state representation captures heterogeneous cross-domain topology and resource information, an attention-enhanced centralized critic strengthens inter-domain credit assignment in large-scale settings, and parallel dual-policy actors together with a parallel experience ensemble and prioritized sampling improve training stability in large constrained action spaces. Extensive simulations across three CPN scales show that AE-MAAC achieves the highest average episode reward. In the representative 5×10 network, it reaches an average episode reward of 431.7 with a 94.2% request success rate and a 259.8 ms average end-to-end delay, while yielding a lower multi-objective cost than state-of-the-art RL baselines. Yan Wang 0146, Xingwei Wang 0001, Hao Lu 0009, Bo Yi 0002, Min Huang 0001, Yue Kou |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | Towards Generalizable Multi-Camera 3D Object Detection via Perspective RenderingabstractDetecting and localizing objects in 3D space using multiple cameras, known as Multi-Camera 3D Object Detection (MC3D-Det), has gained prominence with the advent of bird's-eye view (BEV) approaches. However, these methods often struggle with the serious domain gaps caused by various viewpoints and environments between the training and testing domains. To address this challenge, we propose a novel framework that aligns 3D detection with 2D camera plane results by perspective rendering, thus achieving consistent and accurate results when facing serious domain shifts. Our approach consists of two main steps in both source and target domains: 1) rendering diverse view maps from BEV features by leveraging implicit foreground volumes and 2) rectifying the perspective bias of these maps. This design promotes the learning of perspective- and context-independent features, crucial for accurate object detection across varying viewpoints, camera parameters, and environmental conditions. Notably, our model-agnostic approach preserves the original network structure without incurring additional inference costs, facilitating seamless integration across various models and simplifying deployment. Worth noting is that our approach achieves satisfactory results in real data when trained only with virtual datasets, eliminating the need for real scene annotations. Experimental results on both Domain Generalization (DG) and Unsupervised Domain Adaptation (UDA) demonstrate its effectiveness. Hao Lu 0009, Qing Lian, Dalong Du, Ying-Cong Chen |
AAAI | 1 |
| 2025 | Period-LLM: Extending the Periodic Capability of Multimodal Large Language ModelabstractPeriodic or quasi-periodic phenomena reveal intrinsic characteristics in various natural processes, such as weather patterns, movement behaviors, traffic flows, and biological signals. Given that these phenomena span multiple modalities, the capabilities of Multimodal Large Language Models (MLLMs) offer promising potential to effectively capture and understand their complex nature. However, current MLLMs struggle with periodic tasks due to limitations in: 1) lack of temporal modelling and 2) conflict between short and long periods. This paper introduces Period-LLM, a multimodal large language model designed to enhance the performance of periodic tasks across various modalities, and constructs a benchmark of various difficulty for evaluating the cross-modal periodic capabilities of large models. Specially, We adopt an "Easy to Hard Generalization" paradigm, starting with relatively simple text-based tasks and progressing to more complex visual and multimodal tasks, ensuring that the model gradually builds robust periodic reasoning capabilities. Additionally, we propose a "Resisting Logical Oblivion" optimization strategy to maintain periodic reasoning abilities during semantic alignment. Extensive experiments demonstrate the superiority of the proposed Period-LLM over existing MLLMs in periodic tasks. The code is available at https: //github.com/keke-nice/Period-LLM. Yuting Zhang 0008, Hao Lu 0009, Qingyong Hu, Yin Wang 0004, Kaishen Yuan, Xin Liu 0012, Kaishun Wu |
CVPR | 2 |
| 2025 | Rhythmguassian: Repurposing Generalizable Gaussian Model for Remote Physiological Measurement
Hao Lu 0009, Yuting Zhang 0008, Jiaqi Tang 0005, Wenhang Ge, Wei Wei 0008, Kaishun Wu, Ying-Cong Chen |
ICCV | 1 |
| 2025 | Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language ModelsabstractLarge Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate LLMs with an important representation. To effectively encode occupancy as input for the LLM and address the category imbalances associated with occupancy, we propose Motion Separation Variational Autoencoder (MS-VAE). This innovative approach utilizes prior knowledge to distinguish dynamic objects from static scenes before inputting them into a tailored Variational Autoencoder (VAE). This separation enhances the model's capacity to concentrate on dynamic trajectories while effectively reconstructing static scenes. The efficacy of Occ-LLM has been validated across key tasks, including 4D occupancy forecasting, self-ego planning, and occupancybased scene question answering. Comprehensive evaluations demonstrate that Occ-LLM significantly surpasses existing state-of-the-art methodologies, achieving gains of about 6% in Intersection over Union (IoU) and 4% in mean Intersection over Union (mIoU) for the task of 4D occupancy forecasting. These findings highlight the transformative potential of Occ-LLM in reshaping current paradigms within robotic and autonomous driving. Tianshuo Xu, Hao Lu 0009, Xu Yan 0005, Yingjie Cai, Ying-Cong Chen |
ICRA | 2 |
| 2025 | SeMi: When Imbalanced Semi-Supervised Learning Meets Mining Hard ExamplesabstractSemi-Supervised Learning (SSL) can leverage abundant unlabeled data to boost model performance. However, the class-imbalanced data distribution in real-world scenarios poses great challenges to SSL, resulting in performance degradation. Existing class-imbalanced semi-supervised learning (CISSL) methods mainly focus on rebalancing datasets but ignore the potential of using hard examples to enhance performance, making it difficult to fully harness the power of unlabeled data even with sophisticated algorithms. To address this issue, we propose a method that enhances the performance of Imbalanced Semi-Supervised Learning by Mining Hard Examples (SeMi). This method distinguishes the entropy differences among logits of hard and easy examples, thereby identifying hard examples and increasing the utility of unlabeled data, better addressing the imbalance problem in CISSL. In addition, we maintain a class-balanced memory bank with confidence decay for storing high-confidence embeddings to enhance the pseudo-labels' reliability. Although our method is simple, it is effective and seamlessly integrates with existing approaches. We perform comprehensive experiments on standard CISSL benchmarks and experimentally demonstrate that our proposed SeMi outperforms existing state-of-the-art methods on multiple benchmarks, especially in reversed scenarios, where our best result shows approximately a 54.8% improvement over the baseline methods. Our code is available at https://github.com/pywin/SeMi. Yin Wang 0004, Hao Lu 0009, Zhen Qin 0004, Hailiang Zhao, Guanjie Cheng, Xin Du 0002, Ge Su, Li Kuang, MengChu Zhou, Shuiguang Deng |
ACM Multimedia | 3 |
| 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous DrivingabstractLarge reconstruction model has remarkable progress, which can directly predict 3D or 4D representations for unseen scenes and objects. However, current work has not systematically explored the potential of large reconstruction models in the field of autonomous driving. To achieve this, we introduce the Large 4D Gaussian Reconstruction Model (DrivingRecon). With an elaborate and simple framework design, it not only ensures efficient and high-quality reconstruction, but also provides potential for downstream tasks. There are two core contributions: firstly, the Prune and Dilate Block (PD-Block) is proposed to prune redundant and overlapping Gaussian points and dilate Gaussian points for complex objects. Then, dynamic and static decoupling is tailored to better learn the temporary-consistent geometry across different time. Experimental results demonstrate that DrivingRecon significantly improves scene reconstruction quality compared to existing methods. Furthermore, we explore applications of DrivingRecon in model pre-training, vehicle type adaptation, and scene editing. Our code will be available. Hao Lu 0009, Tianshuo Xu, Wenzhao Zheng, Dalong Du, Masayoshi Tomizuka, Kurt Keutzer, Ying-Cong Chen |
NeurIPS | 1 |
| 2025 | PMMJC: A preference-based multi-stage matching-mechanism for JointCloud environments
Hao Lu 0009, Jianzhi Shi, Yang Song 0022, Xingwei Wang 0001, Bo Yi 0002, Yudi Cheng, Min Huang 0001, Sajal K. Das 0001 |
J. Netw. Comput. Appl. | 1 |
| 2025 | PhysMLE: Generalizable and Priors-Inclusive Multi-Task Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) has been widely applied to measure heart rate from face videos. To increase the generalizability of the algorithms, domain generalization (DG) attracted increasing attention in rPPG. However, when rPPG is extended to simultaneously measure more vital signs (e.g., respiration and blood oxygen saturation), achieving generalizability brings new challenges. Although partial features shared among different physiological signals can benefit multi-task learning, the sparse and imbalanced target label space brings the seesaw effect over task-specific feature learning. To resolve this problem, we designed an end-to-end Mixture of Low-rank Experts for multi-task remote Physiological measurement (PhysMLE), which is based on multiple low-rank experts with a novel router mechanism, thereby enabling the model to adeptly handle both specifications and correlations within tasks. Additionally, we introduced prior knowledge from physiology among tasks to overcome the imbalance of label space under real-world multi-task physiological measurement. For fair and comprehensive evaluations, this paper proposed a large-scale multi-task generalization benchmark, named Multi-Source Synsemantic Domain Generalization (MSSDG) protocol. Extensive experiments with MSSDG and intra-dataset have shown the effectiveness and efficiency of PhysMLE. In addition, a new dataset was collected and made publicly available to meet the needs of the MSSDG. The code and data are available at https://github.com/WJULYW/PhysMLE. Jiyao Wang 0002, Hao Lu 0009, Ange Wang, Xiao Yang 0025, Ying-Cong Chen, Dengbo He, Kaishun Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Advancing Generalizable Remote Physiological Measurement Through the Integration of Explicit and Implicit Prior KnowledgeabstractRemote photoplethysmography (rPPG) is a promising technology for capturing physiological signals from facial videos, with potential applications in medical health, affective computing, and biometric recognition. The demand for rPPG tasks has evolved from achieving high performance in intra-dataset testing to excelling in cross-dataset testing (i.e., domain generalization). However, most existing methods have overlooked the incorporation of prior knowledge specific to rPPG, leading to limited generalization capabilities. In this paper, we propose a novel framework that effectively integrates both explicit and implicit prior knowledge into the rPPG task. Specifically, we conduct a systematic analysis of noise sources (e.g., variations in cameras, lighting conditions, skin types, and motion) across different domains and embed this prior knowledge into the network design. Furthermore, we employ a two-branch network to disentangle physiological feature distributions from noise through implicit label correlation. Extensive experiments demonstrate that the proposed method not only surpasses state-of-the-art approaches in RGB cross-dataset evaluation but also exhibits strong generalization from RGB datasets to NIR datasets. The code is publicly available at https://github.com/keke-nice/Greip. Yuting Zhang 0008, Hao Lu 0009, Xin Liu 0012, Ying-Cong Chen, Kaishun Wu |
IEEE Trans. Image Process. | 2 |
| 2024 | Bi-TTA: Bidirectional Test-Time Adapter for Remote Physiological Measurement
Hao Lu 0009, Ying-Cong Chen |
ECCV (11) | 2 |
| 2024 | An Incremental Unified Framework for Small Defect Inspection
Jiaqi Tang 0005, Hao Lu 0009, Xiaogang Xu 0002, Ruizheng Wu, Sixing Hu, Tong Zhang 0001, Tsz Wa Cheng, Ming Ge, Ying-Cong Chen, Fugee Tsung |
ECCV (31) | 2 |
| 2024 | Backdoor Contrastive Learning via Bi-level Trigger OptimizationabstractContrastive Learning (CL) has attracted enormous attention due to its remarkable capability in unsupervised representation learning. However, recent works have revealed the vulnerability of CL to backdoor attacks: the feature extractor could be misled to embed backdoored data close to an attack target class, thus fooling the downstream predictor to misclassify it as the target. Existing attacks usually adopt a fixed trigger pattern and poison the training set with trigger-injected data, hoping for the feature extractor to learn the association between trigger and target class. However, we find that such fixed trigger design fails to effectively associate trigger-injected data with target class in the embedding space due to special CL mechanisms, leading to a limited attack success rate (ASR). This phenomenon motivates us to find a better backdoor trigger design tailored for CL framework. In this paper, we propose a bi-level optimization approach to achieve this goal, where the inner optimization simulates the CL dynamics of a surrogate victim, and the outer optimization enforces the backdoor trigger to stay close to the target throughout the surrogate CL procedure. Extensive experiments show that our attack can achieve a higher attack success rate (e.g., 99\% ASR on ImageNet-100) with a very low poisoning rate (1\%). Besides, our attack can effectively evade existing state-of-the-art defenses. Weiyu Sun, Hao Lu 0009, Ying-Cong Chen, Ting Wang 0006, Lu Lin 0001 |
ICLR | 3 |
| 2024 | rPPG-HiBa: Hierarchical Balanced Framework for Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) is a promising technique for non-contact physiological signal measurement. It has great potential applications in human health monitoring and emotion analysis. However, existing methods for the rPPG task ignore the long-tail phenomenon of physiological signal data, especially on multi-domain joint training. In addition, we find that the long-tail problem of the physiological label (phys-label) exists in different datasets, and the long-tail problem of some domain exists under the same phys-label. To tackle these problems, we propose a hierarchical balanced framework, to mitigate the bias caused by domain and phys-label imbalance. Specifically, we propose anti-spurious domain center learning tailored to learning domain-balanced embeddings space. Then, we adopt compact-aware continuity regularization to estimate phys-label-wise imbalances and construct continuity between embeddings. Extensive experiments demonstrate that our method outperforms the state-of-the-art in cross-dataset and intra-dataset settings. Our code is available at https://github.com/pywin/HiBa. Yin Wang 0004, Hao Lu 0009, Ying-Cong Chen, Li Kuang, MengChu Zhou, Shuiguang Deng |
ACM Multimedia | 2 |
| 2024 | HAWK: Learning to Understand Open-World Video AnomaliesabstractVideo Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios.
In this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk. Jiaqi Tang 0005, Hao Lu 0009, Ruizheng Wu, Xiaogang Xu 0002, Bin Guo 0001, Jiangbo Lu, Qifeng Chen 0001, Ying-Cong Chen |
NeurIPS | 2 |
| 2024 | Hierarchical Style-Aware Domain Generalization for Remote Physiological MeasurementabstractThe utilization of remote photoplethysmography (rPPG) technology has gained attention in recent years due to its ability to extract blood volume pulse (BVP) from facial videos, making it accessible for various applications such as health monitoring and emotional analysis. However, the BVP signal is susceptible to complex environmental changes or individual differences, causing existing methods to struggle in generalizing for unseen domains. This article addresses the domain shift problem in rPPG measurement and shows that most domain generalization methods fail to work well in this problem due to ambiguous instance-specific differences. To address this, the article proposes a novel approach called Hierarchical Style-aware Representation Disentangling (HSRD). HSRD improves generalization capacity by separating domain-invariant and instance-specific feature space during training, which increases the robustness of out-of-distribution samples during inference. This work presents state-of-the-art performance against several methods in both cross and intra-dataset settings. Jiyao Wang 0002, Hao Lu 0009, Ange Wang, Ying-Cong Chen, Dengbo He |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | ConDiff-rPPG: Robust Remote Physiological Measurement to Heterogeneous OcclusionsabstractRemote photoplethysmography (rPPG) is a contactless technique that facilitates the measurement of physiological signals and cardiac activities through facial video recordings. This approach holds tremendous potential for various applications. However, existing rPPG methods often did not account for different types of occlusions that commonly occur in real-world scenarios, such as temporary movement or actions of humans in videos or dust on camera. The failure to address these occlusions can compromise the accuracy of rPPG algorithms. To address this issue, we proposed a novel Condiff-rPPG to improve the robustness of rPPG measurement facing various occlusions. First, we compressed the damaged face video into a spatio-temporal representation with several types of masks. Second, the diffusion model was designed to recover the missing information with observed values as a condition. Moreover, a novel low-rank decomposition regularization was proposed to eliminate background noise and maximize informative features. ConDiff-rPPG ensured consistency in optimization goals during the training process. Through extensive experiments, including intra- and cross-dataset evaluations, as well as ablation tests, we demonstrated the robustness and generalization ability of our proposed model. Jiyao Wang 0002, Ximeng Wei, Hao Lu 0009, Ying-Cong Chen, Dengbo He |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | rPPG-MAE: Self-Supervised Pretraining With Masked Autoencoders for Remote Physiological MeasurementsabstractRemote photoplethysmography (rPPG) is an important technique for detecting human vital signs and has received extensive attention. For a long time, researchers have focused attention on supervised methods that rely on large amounts of labeled data. These methods are limited by their need for large amounts of data and the difficulty of acquiring ground truth physiological signals. To address these issues, several self-supervised methods based on contrastive learning have been proposed. However, they focus on contrastive learning between samples, which neglects inherent self-similar priors in physiological signals and seems to have a limited ability to cope with noise. In this paper, a linear self-supervised reconstruction task was designed for extracting the inherent self-similar priors in physiological signals. In addition, a specific noise-insensitive strategy was explored for reducing the interference of motion and illumination. The framework proposed in this paper, rPPG-MAE, demonstrates excellent performance even on the challenging VIPL-HR dataset. We also evaluate the proposed method on two public datasets, namely, PURE and UBFC-rPPG. The results show that our method not only outperforms existing self-supervised methods but also outperforms state-of-the-art (SOTA) supervised methods. One important observation is that the quality of the dataset appears to be more important than the size of the dataset used in self-supervised pretraining of the rPPG. The source code is available athttps://github.com/linuxsino/rPPG-MAE. Xin Liu 0012, Yuting Zhang 0008, Zitong Yu, Hao Lu 0009, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Self-Similarity Prior Distillation for Unsupervised Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) is a non-invasive technique that aims to capture subtle variations in facial pixels caused by changes in blood volume resulting from cardiac activities. Most existing unsupervised methods for rPPG tasks focus on the contrastive learning between samples while neglecting the inherent self-similarity prior in physiological signals. In this paper, we propose a Self-Similarity Prior Distillation (SSPD) framework for unsupervised rPPG estimation, which capitalizes on the intrinsic temporal self-similarity of cardiac activities. Specifically, we first introduce a physical-prior embedded augmentation technique to mitigate the effect of various types of noise. Then, we tailor a self-similarity-aware network to disentangle more reliable self-similar physiological features. Finally, we develop a hierarchical self-distillation paradigm for self-similarity-aware learning and rPPG signal decoupling. Comprehensive experiments demonstrate that the unsupervised SSPD framework achieves comparable or even superior performance compared to the state-of-the-art supervised methods. Meanwhile, SSPD has the lowest inference time and computation cost among end-to-end models. Weiyu Sun, Hao Lu 0009, Ying Chen 0006, Xiaolin Huang, Ying-Cong Chen |
IEEE Trans. Multim. | 3 |
| 2023 | Neuron Structure Modeling for Generalizable Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) technology has drawn increasing attention in recent years. It can extract Blood Volume Pulse (BVP) from facial videos, making many applications like health monitoring and emotional analysis more accessible. However, as the BVP signal is easily affected by environmental changes, existing methods struggle to generalize well for unseen domains. In this paper, we systematically address the domain shift problem in the rPPG measurement task. We show that most domain generalization methods do not work well in this problem, as domain labels are ambiguous in complicated environmental changes. In light of this, we propose a domain-label-free approach called NEuron STructure modeling (NEST). NEST improves the generalization capacity by maximizing the coverage of feature space during training, which reduces the chance for under-optimized feature activation during inference. Besides, NEST can also enrich and enhance domain invariant features across multi-domain. We create and benchmark a large-scale domain generalization protocol for the rPPG measurement task. Extensive experiments show that our approach outperforms the state-of-the-art methods on both cross-dataset and intra-dataset settings. The codes are available at https://github.com/LuPaoPao/NEST. Hao Lu 0009, Zitong Yu, Xuesong Niu, Ying-Cong Chen |
CVPR | 1 |
| 2023 | Resolve Domain Conflicts for Generalizable Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) technology has become increasingly popular due to its non-invasive monitoring of various physiological indicators, making it widely applicable in multimedia interaction, healthcare, and emotion analysis. Existing rPPG methods utilize multiple datasets for training to enhance the generalizability of models. However, they often overlook the underlying conflict issues in the rPPG field, such as (1) label conflict resulting from different phase delays between physiological signal labels and face videos at the instance level, and (2) attribute conflict stemming from distribution shifts caused by head movements, illumination changes, skin types, etc. To address this, we introduce the DOmain-HArmonious framework (DOHA). Specifically, we first propose a harmonious phase strategy to eliminate uncertain phase delays and preserve the temporal variation of physiological signals. Next, we design a harmonious hyperplane optimization that reduces irrelevant attribute shifts and encourages the model's optimization towards a global solution that fits more valid scenarios. Our experiments demonstrate that DOHA significantly improves the performance of existing methods under multiple protocols. Weiyu Sun, Hao Lu 0009, Ying Chen 0006, Xiaolin Huang, Ying-Cong Chen |
ACM Multimedia | 3 |
| 2021 | Dual-GAN: Joint BVP and Noise Modeling for Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) based physiological measurement has great application values in health monitoring, emotion analysis, etc. Existing methods mainly focus on how to enhance or extract the very weak blood volume pulse (BVP) signals from face videos, but seldom explicitly model the noises that dominate face video content. Thus, they may suffer from poor generalization ability in unseen scenarios. This paper proposes a novel adversarial learning approach for rPPG based physiological measurement by using Dual Generative Adversarial Networks (Dual-GAN) to model the BVP predictor and noise distribution jointly. The BVP-GAN aims to learn a noise-resistant mapping from input to ground-truth BVP, and the Noise-GAN aims to learn the noise distribution. The two GANs can promote each other’s capability, leading to improved feature disentanglement between BVP and noises. Besides, a plug-and-play block named ROI alignment and fusion (ROI-AF) block is proposed to alleviate the inconsistencies between different ROIs and exploit informative features from a wider receptive field in terms of ROIs. In comparison to state-of-the-art methods, our approach achieves better performance in heart rate, heart rate variability, and respiration frequency estimation from face videos. Hao Lu 0009, Hu Han 0001, Shaohua Kevin Zhou |
CVPR | 1 |
| 2021 | BVPNet: Video-to-BVP Signal Prediction for Remote Heart Rate EstimationabstractIn this paper, we propose a new method for remote photoplethysmography (rPPG) based heart rate (HR) estimation. In particular, our proposed method BVPNet is streamlined to predict the blood volume pulse (BVP) signals from face videos. Towards this, we firstly define ROIs based on facial landmarks and then extract the raw temporal signal from each ROI. Then the extracted signals are pre-processed via first-order difference and Butterworth filter and combined to form a Spatial-Temporal map (STMap). We then propose to revise U-Net, in order to predict BVP signals from the STMap. BVPNet takes into account both temporal and frequency domain losses in order to learn better than conventional models. Our experimental results suggest that our BVPNet outperforms the state-of-the-art methods on two publicly available datasets (MMSE-HR and VIPL-HR). Abhijit Das 0001, Hao Lu 0009, Hu Han 0001, Antitza Dantcheva, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2021 | NAS-HR: Neural architecture search for heart rate estimation from face videosabstractIn anticipation of its great potential application to natural human-computer interaction and health monitoring, heart-rate (HR) estimation based on remote photoplethysmography has recently attracted increasing research attention. Whereas the recent deep-learning-based HR estimation methods have achieved promising performance, their computational costs remain high, particularly in mobile-computing scenarios. We propose a neural architecture search approach for HR estimation to automatically search a lightweight network that can achieve even higher accuracy than a complex network while reducing the computational cost. First, we define the regions of interests based on face landmarks and then extract the raw temporal pulse signals from the R, G, and B channels in each ROI. Then, pulse-related signals are extracted using a plane-orthogonal-to-skin algorithm, which are combined with the R and G channel signals to create a spatial-temporal map. Finally, a differentiable architecture search approach is used for the network-structure search. Compared with the state-of-the-art methods on the public-domain VIPL-HR and PURE databases, our method achieves better HR estimation performance in terms of several evaluation metrics while requiring a much lower computational cost1. Hao Lu 0009, Hu Han 0001 |
Virtual Real. Intell. Hardw. | 1 |