Dongmin Huang

dblp:260/8478 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-2604-5878ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Generalized Camera-Based Contactless Seq2Seq Sleep Staging
Ming Xia 0003, Qiongyan Wang, Dongmin Huang, Hanrong Cheng, Peifen Chen, Mei Zi, Wenjin Wang 0002
IEEE Internet Things J.3
2026 Plantar Perfusion Imaging for Peripheral Arterial Disease Screening: A Proof-of-Concept Study
abstract
The diagnosis of peripheral artery disease (PAD) typically relies on specialized equipment such as ultrasound. The delayed PAD detection of these approaches may lead to amputation and even death. To achieve rapid and ubiquitous PAD screening, we propose a novel concept of camera-based plantar perfusion imaging (CPPI) for PAD diagnosis and severity classification. Specifically, we performed a simulation trial that used an RGB camera to record the plantar video of 20 subjects and a cuff with different pressures applied to the left leg to simulate different degrees of lower limb blockage. We generated the plantar perfusion maps using remote photoplethysmography imaging and proposed a multi-view perfusion (MVP) feature set to represent the perfusion maps for PAD classification. The experimental results show that the Pearson correlation coefficients between MVP and Doppler ultrasound (clinical reference) features were larger than 0.9. MVP feature combined with Support Vector Machine obtains 91.47% accuracy in distinguishing the normal and obstructed states, and 76.48% accuracy in differentiating four different degrees of vascular obstruction. The clinical benchmark demonstrated the potential of CPPI as a rapid, sensitive, and easy-to-use diagnostic tool for PAD, suitable for large-scale screening in home or community settings.
Ningbo Zhao, Dongmin Huang, Yonglong Ye, Zi Luo, Hongzhou Lu, Wenjin Wang 0002
IEEE J. Biomed. Health Informatics3
2026 Camera-Based Respiratory Imaging System for Monitoring Infant Thoracoabdominal Patterns of Respiration
abstract
Existing respiratory monitoring techniques primarily focus on respiratory rate measurement, neglecting the potential of using thoracoabdominal patterns of respiration for infant lung health assessment. To bridge this gap, we exploit the unique advantage of spatial redundancy of a camera sensor to analyze the infant thoracoabdominal respiratory motion. Specifically, we propose a camera-based respiratory imaging (CRI) system that utilizes optical flow to construct a spatio-temporal respiratory imager for comparing the infant chest and abdominal respiratory motion, and employs deep learning algorithms to identify infant abdominal, thoracoabdominal synchronous, and thoracoabdominal asynchronous patterns of respiration. To alleviate the challenges posed by limited clinical training data and subject variability, we introduce a novel multiple-expert contrastive learning (MECL) strategy to CRI. It enriches training samples by reversing and pairing different-class data, and promotes the representation consistency of same-class data through multi-expert collaborative optimization. Clinical validation involving 44 infants shows that MECL achieves 70% in sensitivity and 80.21% in specificity, which validates the feasibility of CRI for respiratory pattern recognition. This work investigates a novel video-based approach for assessing the infant thoracoabdominal patterns of respiration, revealing a new value stream of video health monitoring in neonatal care.
Dongmin Huang, Yongshen Zeng, Yingen Zhu, Xiaoyan Song, Liping Pan, Jie Yang 0083, Hongzhou Lu, Wenjin Wang 0002
IEEE J. Biomed. Health Informatics1
2025 Implementation of Voxel Selective Ellipse Normalization to Enhance Radar Respiration Estimation in Metallic Chamber
abstract
Compared to broader physical activities, detecting nuanced respiratory movements poses a significant challenge in indoor health monitoring systems. While respiratory activity can be conceptualized as periodic chest movements akin to mechanical vibrations, uncontrollable environmental factors often introduce noise into detected radar signals. The clutter in the field of view, especially metallic objects, such as hospital steel beds, degrades the performance of radar physiological monitoring: 1) amplifying noise of multipath effects and 2) misleading the informative localization module. In this article, we propose a preprocessing scheme of Search-Voxel Ellipse Normalization for respiratory detection system, including an ellipse normalization method combined with the fitting-cost voxel selection policy, to improve the respiration detection performance using MIMO frequency modulated continuous wave radar. This article provides an in-depth assessment of the designed system, including a metal-insulated room test, involving ten participants in different postures. The results show notable performance improvements of our proposed ENDTW-MVMD method, especially in lowering the mean absolute error from the best state-of-the-art 0.93–0.75 bpm and stabilization in voxel selection. The proposed approach is thoroughly evaluated against established methods across various dimensions, such as voxel selection, independent performance, frequency estimation, and ablation studies.
Yao Ge 0002, Yingen Zhu, Sidra Liaqat, Dongmin Huang, Liangyue Yu, Chengkai Tang, Muhammad Ali Imran 0001, Wenjin Wang 0002, Qammer H. Abbasi
IEEE Internet Things J.4
2025 Toward Camera-PRV-Based Early Warning in Hospital ICU: A Pilot Study
abstract
In the Intensive Care Unit (ICU), monitoring a patient’s heart rate variability (HRV) can provide vital information regarding the physiological status and autonomic nervous system, serving as a potential metric for detecting adverse events in ICU. Although camera-based remote photoplethysmography (rPPG) has shown to be effective for heart rate monitoring in clinical environments, it remains insufficient research on utilizing rPPG-derived pulse rate variability (PRV) for monitoring autonomic nervous system (ANS) changes of ICU patients. In addition, the agreement between HRV and PRV is controversial, especially for ICU patients with multiple concurrent diseases. This study aims to validate the accuracy of camera-PRV measurements by analyzing eight HRV/PRV parameters from subjects with different health conditions, further discussing its feasibility as an early-warning tool. The accuracy of rPPG-derived PRV was comparable to that of contact-PPG and has high correlations with Electrocardiogram. Combined with machine learning algorithms, the system was able to accurately differentiate between healthy subjects and ICU patients (accuracy = 0.86, F1 score = 0.83, sensitivity = 0.79, specificity = 0.79). Furthermore, the system demonstrated good reliability in classifying subjects with different health status and ages. However, it could not effectively correspond to the APACHE II score when classifying ICU patients with different severities based on HRV/PRV. Because the APACHE II scoring system evaluates the changes in the patient’s ANS as well as a variety of confounding factors, of which the HRV/PRV parameters represent only a portion. In summary, the camera-PRV system was able to provide the individual physiological tracking and coarse-level modeling of the group health status. It is expected to achieve a more comprehensive health status assessment by integrating more vital signs in the future.
Jincan Lou, Yifeng Tancheng, Dongmin Huang, Hongzhou Lu, Wenjin Wang 0002
IEEE Internet Things J.6
2025 Camera-Based Bi-Modal PPG-SCG: Sleep Privacy-Protected Contactless Vital Signs Monitoring
abstract
The monitoring of respiratory rate (RR), heart rate (HR), HR variability (HRV), and blood pressure (BP) during sleep allows for a comprehensive evaluation of sleep quality, facilitating the understanding and improvement of a person’s sleep health. Contactless physiological monitoring using cameras has gained popularity recently due to its convenient, infection-free, continuous, and versatile nature. However, the privacy concerns limit the application of camera-based solutions in sleep monitoring setups. This study proposes a novel hybrid setup that integrates camera-based seismocardiography (CamSCG) and photoplethysmography (CamPPG) for contactless measurement of RR, HR, and HRV during sleep while simultaneously estimating BP. For the proximal SCG, we employed camera-based laser speckle vibrometry to measure cardiac motions from the chest, and benchmarked it with a millimeter-wave radar (RFSCG). For the distal photoplethysmographic (PPG), a defocused camera was utilized to measure pulse signals from the facial skin while protecting privacy. In this setup, we analyzed the single-modality in measuring RR, HR, and HRV, and established two bi-modalities (CamSCG-CamPPG and RFSCG-CamPPG) to measure pulse transit time (PTT) features for BP calibration. The benchmark involving 19 subjects highlights the potential of camera-based bi-modal SCG-PPG for privacy-protected vital signs monitoring during sleep.
Yingen Zhu, Yao Ge 0002, Dongmin Huang, Pong C. Yuen, Fu Xiao 0001, Wenjin Wang 0002
IEEE Internet Things J.5
2025 Camera-Based Infant Suffocation Risk Detection Via Text-to-Image Generation for Guarding Sleep Safety
abstract
Current camera-based infant monitoring mainly focuses on physiological measurement, overlooking its important semantic analysis potential for detecting accidental suffocation caused by oronasal occlusion during sleep. However, developing a robust infant suffocation risk detection model typically requires substantial labeled data, which is very difficult to obtain in real-world scenarios. To address this, we utilized the text-to-image diffusion model to generate diverse infant images depicting oronasal occlusion and non-occlusion scenarios controlled by text prompts. To ease the process of labeling, self- and semi-supervised learning algorithms are leveraged to learn the semantic information from unlabeled data with the support of minimal labeled data to train different model architectures. To evaluate the feasibility of this solution, we conducted a clinical trial in the neonatology department, which collected video data from 22 infants under various oronasal occlusion scenarios using breathable covers (e.g. clinical tissue). The clinical evaluation shows that most models trained on 25,000 generated images achieved over 90% performance on metrics of accuracy, recall, and F1-score, outperforming conventional approaches that pre-train and fine-tune the model using over 90,000 labeled task-related online images. This demonstrates the feasibility of leveraging text-to-image generated data to achieve robust camera-based infant suffocation risk detection, so as to secure the sleep safety of infants. More importantly, it beacons the potential of using text-based large-scale model to solve the general issue of scarcity of human data in artificial intelligence-based healthcare or clinical applications.
Dongmin Huang, Chuchu Liao, Jingyun Mai, Xiaoxiao He, Liping Pan, Ming Xia 0003, Huailei Lai, Xuhui Yang, Zhenlang Lin, Wenjin Wang 0002
IEEE J. Biomed. Health Informatics1
2025 Prototype-Driven Hard-Sample Contrastive Learning for Camera-Based Respiratory Imaging Analysis
abstract
Respiratory spatial patterns describe the distribution and dynamics of lung conditions, and monitoring of their asymmetry or irregularities enables a more comprehensive assessment of respiratory functions. The feasibility of using the camera pixel array sensing combined with machine learning for analyzing respiratory spatial patterns was demonstrated, however, this approach faces challenges in patient generalization due to limited clinical data and individual respiratory variability. Data augmentation methods may address this by synthesizing new data, but they have a risk of destroying the symmetry or regular semantic information of respiratory patterns. To address this, we propose a prototype-driven hard-sample contrastive learning (PHCL) method tailored for camera-based respiratory imaging analysis. It first separates the samples into simple and hard-to-learn samples using prototypes and Gini-index distance measurement. Then it synthesizes a new feature by blending one-class simple samples and other-class hard samples to construct a transition boundary between different classes, so as to broaden the feature distribution. Then it employs contrastive learning to emphasize feature consistency between prototypes and hard same-class samples from different subjects to mitigate individual respiratory variability and refine class boundaries. Extensive experiments were conducted in the neonatal intensive care unit and the thoracic surgery department, where PHCL outperforms image augmentation and advanced feature augmentation methods by 1-10% in both accuracy and F1-score. Our work provides valuable insights into the analysis of asymmetric and irregular respiratory activities.
Dongmin Huang, Ming Xia 0003, Liping Pan, Qiqiong Wang, Xiaoyan Song, Xiaoting Tao, Kun Qiao, Hongzhou Lu, Wenjin Wang 0002
IEEE J. Biomed. Health Informatics1
2024 Camera-Based Respiratory Imaging for Intelligent Rehabilitation Assessment of Thoracic Surgery Patients
abstract
Camera-based respiration monitoring is currently focused on the continuous measurement of respiratory rate, overlooking its potential in lung health assessment. Inspired by auscultation and palpation that use respiratory symmetry to assess the lung rehabilitation of thoracic surgery patients, we exploit the advantage of spatial redundancy of a camera sensor to replicate this clinical routine. In particular, we propose a camera-based respiratory imaging (CRI) system that leverages optical flow and deep learning algorithms to analyze the symmetric/asymmetric patterns of chest respiratory motion, and classify the subject as health, left lesion, or right lesion. To mitigate the issues of sample scarcity and subject variance, we introduce a novel multiple-prototype contrastive model (MPCM) that uses the symmetric respiration hypothesis to generate more training data, and produces multiple deep prototypes to enhance the consistency of deep representation of samples from different subjects. The clinical validation involving 45 subjects demonstrates the feasibility of CRI for lung rehabilitation assessment, where MPCM achieves above 70% in the used evaluation indices (e.g. accuracy, sensitivity, and specificity). This study demonstrates a new value stream in video health monitoring that uses camera-based respiratory imaging for the lung rehabilitation assessment of thoracic patients after surgery.
Dongmin Huang, Xiaoting Tao, Yingen Zhu, Kun Qiao, Hongzhou Lu, Wenjin Wang 0002
IEEE Internet Things J.1
2024 Generalized Camera-Based Infant Sleep-Wake Monitoring in NICUs: A Multi-Center Clinical Trial
abstract
The infant sleep-wake behavior is an essential indicator of physiological and neurological system maturity, the circadian transition of which is important for evaluating the recovery of preterm infants from inadequate physiological function and cognitive disorders. Recently, camera-based infant sleep-wake monitoring has been investigated, but the challenges of generalization caused by variance in infants and clinical environments are not addressed for this application. In this paper, we conducted a multi-center clinical trial at four hospitals to improve the generalization of camera-based infant sleep-wake monitoring. Using the face videos of 64 term and 39 preterm infants recorded in NICUs, we proposed a novel sleep-wake classification strategy, called consistent deep representation constraint (CDRC), that forces the convolutional neural network (CNN) to make consistent predictions for the samples from different conditions but with the same label, to address the variances caused by infants and environments. The clinical validation shows that by using CDRC, all CNN backbones obtain over 85% accuracy, sensitivity, and specificity in both the cross-age and cross-environment experiments, improving the ones without CDRC by almost 15% in all metrics. This demonstrates that by improving the consistency of the deep representation of samples with the same state, we can significantly improve the generalization of infant sleep-wake classification.
Dongmin Huang, Dongfang Yu, Yongshen Zeng, Xiaoyan Song, Liping Pan, Junli He, Lirong Ren, Jie Yang 0083, Hongzhou Lu, Wenjin Wang 0002
IEEE J. Biomed. Health Informatics1
2024 Multi-Task Learning for Audio-Based Infant Cry Detection and Reasoning
abstract
Infant cry is a crucial indicator that offers valuable insights into their physical and mental conditions, such as hunger and pain. However, the scarcity of infant cry datasets hinders the model's generalization in real-life scenarios. The varying voiceprint characteristics among infants further exacerbate this challenge, deteriorating the model's performance on unseen infants. To this end, we propose a multi-task model for Infant Cry Detection and Reasoning (ICDR). It leverages datasets from two tasks to enrich data diversity and introduces an efficient attention module to achieve inter-task feature supplementarity. To mitigate the impact of subject differences, ICDR introduces an intra-task contrastive mixture of experts (CMoE) module that adaptively allocates experts to reduce subject variance and applies contrastive learning to enhance the representation consistency of samples from different infants in the same state. Extensive cross-subject experiments show that ICDR outperforms the state-of-the-art models in infant cry detection and reasoning, with an improvement of 2-9% in the F1-score. This demonstrates that multi-task learning effectively enhances the model's generalization ability by inter-task attention and intra-task CMoE.
Ming Xia 0003, Dongmin Huang, Wenjin Wang 0002
IEEE J. Biomed. Health Informatics2
2023 A Multi-center Clinical Trial for Camera-based Infant Sleep and Awake Detection in Neonatal Intensive Care Unit
abstract
Infants need adequate sleep to develop their brain and cardiovascular systems, especially for preterm infants in the Neonatal Intensive Care Unit. Camera-based infant monitoring is an emerging direction of research in video health monitoring. Intuitively, camera can easily identify the sleep-awake stage of infants by detecting the state of eyes, e.g. closed or opening. Thus in this paper, we propose to explore the unique advantage of camera-based facial analysis for sleep-awake detection, as a fundamental step toward infant sleep monitoring. A multi-center clinical trial was conducted to collect infant videos for investigating the feasibility of our proposal. A benchmark including four machine learning methods of SVM, KNN, MLP, and CNN (ResNet18) was set up to classify the sleep/awake stage of infants. To alleviate the overfitting issue caused by over-sampling of a sleeping infant, we propose to integrate ResNet18 with the contrastive learning strategy to strengthen the consistency of facial features learned from different infants. The clinical evaluation shows that all benchmarked methods obtained an accuracy above 75% while the proposed method achieved the best accuracy of 86%. This invokes further explorations of using facial/eye features of infants for sleep-awake staging, towards intelligent contactless sleep analysis of infants in combination with camera-based vital signs monitoring.
Yuya Yuan, Dongmin Huang, Lirong Ren, Xiaoyan Song, Liping Pan, Hongzhou Lu, Wenjin Wang 0002
HealthCom2
2023 A Contrastive Embedding-Based Domain Adaptation Method for Lung Sound Recognition in Children Community-Acquired Pneumonia
abstract
Lung sound analysis has been used for assessing the lung conditions of children with community-acquired pneumonia (CAP). However, the inconsistent data distribution, caused by variant CAP symptoms appearing at different ages of children, limits the generalization ability of most diagnostic models. The data scarcity will further exacerbate this problem. Therefore, we propose a contrastive embedding-based domain adaptation network (CEDANN) to eliminate individual differences and alleviate data scarcity for improving the generalization ability. It forces the embedding layer to learn the subject-independent but task-dependent features by the adversarial learning between the domain classifier and the task classifier. Multiple contrast input tuples are constructed by combining samples from different classes to increase input combinations to alleviate data scarcity. The proposed method is evaluated on a multi-center clinical dataset. The subject-independent experiments show that CEDANN improves the sensitivity from 45.93% to 59.06% and the specificity from 35.07% to 59.55% in identifying CAP-confirmed, symptomatic relief, and recovery children. It demonstrates the effectiveness of CEDANN in the diagnosis and prognosis of children CAP.
Dongmin Huang, Lingwei Wang, Hongzhou Lu, Wenjin Wang 0002
ICASSP1
2023 PSDCE: Physiological signal-based double chaotic encryption for instantaneous E-healthcare services
Dongmin Huang, Shengwen Fan, Kaining Han, Gwanggil Jeon, Joel J. P. C. Rodrigues
Future Gener. Comput. Syst.2
2021 Differences first in asymmetric brain: A bi-hemisphere discrepancy convolutional neural network for EEG emotion recognition
Dongmin Huang, Sentao Chen, Cheng Liu 0001, Lin Zheng 0003, Zhihang Tian, Dazhi Jiang
Neurocomputing1