Jilong Kuang

dblp:82/8314 · DBLP profile ↗
← Back
59ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0003-2942-9102ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 15 since 2021Human-computer interaction and ubiquitous computing · 9 · 3 since 2021Computer networks · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSystems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 MAD-Fusion: Modality-Aware Dynamic Fusion in Identification of Activities of Daily Living
abstract
Activities of daily living (ADL) identification with wearables has significant implications in healthy lifestyle management and offers an important sensor-based supervised learning research benchmark. Most ADL studies use single-modality (i.e., single sensor type like smartwatch only or earbuds only) data while multi-sensor data fusion studies using early-stage fusion of modalities from multi-sensors and multi-devices are emerging. To improve classification performance and model interpretability, leveraging early-stage fusion, late-stage fusion and individual modalities, we introduced novel modality-aware dynamic fusion (MAD-Fusion) models for multi-sensor data-fusion-based ADL identification. Based on early-stage fusion, we incorporated conformal prediction for uncertainty quantification, uncertainty late-stage fusion for cross-modality interpretability, and multi-modal strength-aware classification module. Trained on 36 independent subjects and tested on 4 independent subjects from Samsung ADL dataset with multi-sensor (accelerometers and gyroscopes) and multi-device (earbuds and smartwatch), MAD-Fusion not only achieved the state-of-the-art classification performance (accuracy: 0.9504, F1-score: 0.9142), but also enabled better interpretability of contributions and uncertainties from different modalities. The additional contribution of each building block is validated systematically. Furthermore, we validated MAD-Fusion’s superior performance on two public datasets in multi-sensor single-device settings (UCI-HAR and USC-HAD datasets). On all three datasets, MAD-Fusion manifested statistically significantly superiority comparing against baselines of single-modality, early-stage fusion and late-stage fusion (p< 0.001). To conclude, the novel MAD-Fusion models improve the classification performance, uncertainty quantification and interpretability for the ADL identification and can be applied to broader supervised learning areas requiring high performance, rigorous uncertainty quantification and model interpretability.
Xianghao Zhan, Ebrahim Nemati, Mohsin Y. Ahmed, Sharath Chandrashekhara, Jilong Kuang
IEEE Internet Things J.6
2025 BallistoBud: Heart Rate Variability Monitoring using Earbud Accelerometry for Stress Assessment
Mehrab Bin Morshed, David Jimmy Lin, Hao Zhou 0001, Wendy Berry Mendes, Jilong Kuang
CHI8
2025 Towards Real-Time Acute Stress Management via Integrated Biosensing and Neurostimulation: A Closed-Loop Earbud Platform
Li Zhu 0004, William Schuerman, Matthew K. Leonard, Wendy Berry Mendes, Wallace Ming Yip Wong, Jilong Kuang, Sharanya Arcot Desai
GLOBECOM7
2025 Optimizing Biomarkers from Earbud Ballistocardiogram: Calibration and Calibration-Free Algorithms for Accelerometer Axis Selection and Fusion
abstract
The earbud-based ballistocardiogram (BCG) assessment holds significant promise for monitoring diverse physiological signals, including stress, cardiac activity, and blood pressure. However, unlike traditional methods that measure the force component along the head-to-foot axis for enhanced BCG signal quality, ear-worn devices are prone to orientation misalignment, leading to significant variations in BCG morphology. To address this challenge, we propose two novel algorithms: one that employs sensor-to-body-segment calibration and another that applies a calibration-free, physiologically informed axis fusion method to enhance earbud-based BCG signal assessment. We evaluate the performance of these approaches against existing methods, focusing on heart rate variability (HRV) estimation and morphological feature extraction. Through a comprehensive investigation, we aim to identify optimal strategies for obtaining high-quality BCG signals using ear-worn devices.
Mehrab Bin Morshed, Holland Ernst, Li Zhu 0004, Jilong Kuang
ICASSP9
2025 Evaluation of Wearable Head BCG for PTT Measurement in Blood Pressure Intervention
abstract
This study evaluates the usability of wearable head ballistocardiography (BCG) in providing accurate pulse transit time (PTT) measurements during blood pressure (BP) interventions. Head BCG is a new technique enabling measurement of proximal aortic blood ejection from sensors placed at distal sites, which envisions PTT measurement from single integrated device for cuff-less BP estimation. However, due to its low signal-to-noise ratio and sensitivity to motion artifacts, accurate beat selection is crucial to ensure the integrity of PTT calculation. In this paper, using inertial measurement unit (IMU) sensors integrated in a prototype earbud, we investigate whether the wearable head BCG signal is aligned with the ground-truth proximal reference acquired from the synchronously-recorded impedance cardiography (ICG) signal, to assess the usability of head BCG as the proximal indicator for PTT measurement at different stages of BP intervention. Wearable BCG signals showed highest reliability during resting states, with 63% of detected j-peaks aligned with ground-truth ICG signals. Beat selection via removal of IBI outliers improves the ratio of reliable peaks, at rest (68%) and during exercise (63% at intervention and 52% at plateau). Other methods, such as template matching or rejecting amplitude outliers, only improve the ratio at rest. Overall, this study reveals characteristics of distortions in the head BCG signal during intervention, as a first step toward robust solutions for PTT-based BP tracking on integrated wearable devices.
Li Zhu 0004, Mehrab Bin Morshed, Jungmok Bae, Jilong Kuang
ICASSP6
2025 Earbuds Orientation Alignment Based on Markov Chain Monte Carlo Sampling
abstract
Earbuds are instrumental in health monitoring but the orientation can variate among users, which may significantly impact the health-monitoring system generalizability. To study the effect of earbuds orientation heterogeneity and align kinematics across earbuds orientations, we collected a dataset with various rotations relative to a baseline orientation. We developed the coordinate transformation by estimating Euler angles in transformation matrices with either grid search or Markov Chain Monte Carlo (MCMC) sampling. Taking ~ 17 seconds with a personal laptop, the MCMC method accurately estimated the coordinate transformation matrices to enable the transformed tri-axial linear acceleration to better match the baseline tri-axial linear acceleration with an average relative error of 1.899% (0.186 m/s2) and a maximum relative error of 2.774% averaged over all test orientations. Using the estimated transformation matrices and Samsung dataset of identification of activities of daily living (ADL), we validated the statistically significant impact of earbuds orientation heterogeneity on ADL identification (p < 0.001), which can cause 14.0% reduction in mean accuracy and 18.7% reduction in mean macro-average F1-score. To sum up, the MCMC method developed can be applied in earbuds kinematics alignment to address orientation heterogeneity and enable better earbuds-based health monitoring.
Xianghao Zhan, Ebrahim Nemati, Mohsin Y. Ahmed, Jilong Kuang
ICASSP5
2025 Know Your Heart Better: Multimodal Cardiac Output Monitoring using Earbuds
abstract
Cardiac Output (CO) is a critical indicator of health, offering insights into cardiac dysfunction, acute stress responses, and cognitive decline. Traditional CO monitoring methods, like impedance cardiography, are invasive and impractical for daily use, leading to a gap in continuous, non-invasive monitoring. Although recent advancements explored wearables on heart rate monitoring, these approaches face challenges in accurately estimating CO due to the indirect nature of the signals. To address these challenges, we introduce EarCO, a non-invasive multimodal CO monitoring system with Photoplethysmography and Ballistocardiogram signals on commodity earbuds. A novel feature fusion method is proposed to integrate raw signals and prior knowledge from both modalities, improving the system’s interpretability and accuracy. EarCO achieves an error of 1.080 L/min in the leave-one-subject-out settings with 62 subjects, making cardiovascular health monitoring accessible and practical for daily use.
Mehrab Bin Morshed, Larry Zhang, Jungmok Bae, Christina Rosa, Wendy Berry Mendes, Jilong Kuang
ICASSP10
2025 MindfulBuddy: Extracting Comprehensive Breathing Biomarkers for Breathing Exercise Biofeedback Using Earbud Motion Sensors
abstract
Slow-paced deep breathing exercises have many health benefits, including stress management, lowering blood pressure, pain management, and controlling pulmonary conditions. While biofeedback can significantly improve the efficacy of breathing exercises, existing approaches support limited biomarkers, such as breathing rate, for specific breathing exercises (e.g., equal-phase breathing) in particular conditions without considering the breath-holding phase or variation in device orientation. Therefore, there needs to be a more convenient and robust approach that can generate and deliver comprehensive digital breathing biomarkers to facilitate biofeedback for various types of breathing exercises. In this article, we present a system with lightweight algorithms to passively track mindful breathing in real-time using lower-power earbud motion sensors to extract fine-grained comprehensive breathing biomarkers for generating biofeedback on users’ breathing exercises. We utilize the earbud’s motion sensor data to detect nonbreathing head motion and develop an extensive set of breathing markers, including breathing phases, breathing depth, breathing rate, breathing symmetry, and breath-holding. Such a comprehensive set of biomarkers can enable engaging user experience and effective mindful breathing exercises toward better stress management and overall mental well-being. Moreover, we develop a physiologically informed, novel earbud orientation handling algorithm that makes our biomarkers more resilient to ear canal shape and size. Finally, we showcase potential use-cases based on the breathing biomarkers derived from our algorithms to provide biofeedback on user’s overall breathing performance.
Mehrab Bin Morshed, Sharath Chandrashekhara, Jilong Kuang
IEEE Internet Things J.5
2024 EarGo: Is Earbud a Necessary Complement to Smartwatch for Estimation of the Running Dynamics Parameters?
abstract
Human activity recognition has been an established active research area within the past few decades. While many researchers have tried to estimate some of the gait and running parameters, none was successful to provide a full suite of running dynamics parameters using commodity devices. Earbuds with their unique placement (in line with center of mass) provide an opportunity for activity recognition that never existed before with other commodity devices. Taking advantage of this opportunity, this work proposes a multi-modal approach to measure running dynamics using fusion of earbuds and smartwatch. Collecting a large dataset of 53 subjects, we developed various regression models to identify running parameters such as speed, cadence, stride length, vertical oscillation and ground contact time. These parameters were estimated in both jog and walk conditions and were evaluated in different device and context settings. Our MAPE ranges from 6.04% to 11.54% for various parameters.
Ebrahim Nemati, Mohsin Y. Ahmed, Jilong Kuang
BSN4
2024 Multimodal Breathing Rate Estimation Using Facial Motion and RPPG From RGB Camera
abstract
Camera-based respiratory monitoring is contactless, non-invasive, unobtrusive, and easily accessible compared to conventional wearable devices. This paper presents a novel multimodal approach to estimating breathing rate based on tracking the movement and color changes of the face through an RGB camera. A machine learning model determines the final breathing rate between two separately calculated ones from breathing motion and remote photoplethysmography (rPPG) to improve the measurement performance in a broader range of breathing frequencies. Our proposed pipeline is evaluated with 140 facial video recordings from 22 healthy subjects, including 6 controlled and 2 spontaneous breathing tasks ranging from 5 to 30 BPM. The estimation accuracy achieves 1.33 BPM mean absolute error and 86.53% pass rate within 2 BPM error criteria. To the best of our knowledge, our approach outperforms previous works that use a face region alone with a single RGB camera.
Migyeong Gwak, Korosh Vatanparvar, Li Zhu 0004, Mohsin Y. Ahmed, Jungmok Bae, Jilong Kuang, Jun Alex Gao
ICASSP7
2024 Ballistocardiogram-Based Heart Rate Variability Estimation for Stress Monitoring using Consumer Earbuds
abstract
Stress can potentially have detrimental effects on both physical and mental well-being, but monitoring it can be challenging, especially in free-living conditions. One approach to address this challenge is to use earbud accelerometers to capture the ballistocardiogram (BCG) response. These sensors allow for noninvasive stress monitoring by estimating physiological indicators linked to stress, such as heart rate variability (HRV). However, ear-worn devices are susceptible to motion artifacts and can exhibit significant BCG signal morphology variations. These challenges necessitate accurate algorithms to estimate HRV for everyday use. Therefore, we developed a method to measure interbeat intervals (IBI) from BCG signals collected from an earbud. To enhance IBI estimation accuracy, we employed a Bayesian method that incorporates robust apriori IBI prediction weighting and sensor fusion techniques. We have also conducted a study involving 97 participants to assess the earbuds' ability to estimate HRV metrics and classify stressful activities. Our findings demonstrate low IBI estimation error (4.16% ± 1.90%), along with lower errors in subsequent higher-order HRV metrics compared to the state-of-the-art algorithms.
David Jimmy Lin, Li Zhu 0004, Viswam Nathan, Jungmok Bae, Christina Rosa, Wendy Berry Mendes, Jilong Kuang, Jun Alex Gao
ICASSP8
2024 Core Body Temperature and its Role in Detecting Acute Stress: A Feasibility Study
abstract
Core body temperature (CBT) is one of the critical yet under-explored phenomena in the context of stress detection. Several CBT measurement methods exist, but they are often limited in continuous CBT monitoring. Furthermore, how continuous CBT can be used to model acute stress is little explored. We address these challenges by conducting an in-lab controlled study with 97 participants who participated in baseline and stress-inducing tasks while wearing prototype earbuds capable of collecting CBT. We found that accounting for changes from individual baselines in CBT results is acute stress detection with 94.88% accuracy and 94.4% F1-score, which is 29.31% and 26.07% higher in terms of accuracy and F1-score, respectively, compared to generalized features.
Mehrab Bin Morshed, Viswam Nathan, Li Zhu 0004, Jungmok Bae, Christina Rosa, Wendy Berry Mendes, Jilong Kuang, Jun Alex Gao
ICASSP8
2024 Heart Rate Variability Estimation with Dynamic Fine Filtering and Global-Local Context Outlier Removal
abstract
Consumer hearable technologies such as earbuds are increasingly embedding physiological sensors, including photoplethysmography (PPG) and inertial measurements. They create unique opportunities to passively monitor stress and deliver digital interventions such as music. However, PPG signals recorded from ear canals are often very noisy due to head movement and fit issues. This work proposes algorithms to estimate heart rate variability (HRV) features from noisy PPG signals recorded using earbuds. We have used template matching to determine the signal quality for dynamic fine filtering around the estimated heart rate. We have also improved the inter-beat interval (IBI) outlier detection and removal algorithm using the global-local context of the input PPG signal. The mean absolute error of estimating RMSSD decreased from 70.83 milliseconds (ms) to 24.88 ms, and SDNN decreased from 46.89 ms to 16.60 ms.
Ramesh Kumar Sah, Viswam Nathan, Li Zhu 0004, Jungmok Bae, Christina Rosa, Wendy Berry Mendes, Jilong Kuang, Jun Alex Gao
ICASSP8
2024 Normalization is All You Need: Robust Full-Range Contactless SpO2 Estimation Across Users
abstract
The accurate estimation of peripheral capillary oxygen saturation (SpO2) is vital for monitoring respiratory health, with applications spanning medical diagnostics and fitness tracking. Remote photoplethysmography (rPPG) offers a convenient and non-contact approach for SpO2estimation. However, existing methods predominantly rely on data within the normal SpO2range, hindering their effectiveness during hypoxemia. Moreover, cross-user variations poses significant challenges for practicality. To address these limitations, we propose a simple yet effective normalization-based SpO2estimation algorithm. By aligning individual Ratio-of-Ratios (RoR) data with a standard model at the matching SpO2level, we mitigate cross-user variation, accommodate different camera configurations, and account for lighting changes. Our experiments demonstrate that the proposed method achieves an rMSE of 2.8% with leave-one-subject-out cross-validation across the full SpO2range (70%-100%), significantly outperforming existing RoR-based and CNN-based SpO2estimation approaches. Notably, our methods excel in accurately identifying hypoxemia, a critical clinical requirement. We anticipate broader applicability of our approach in rPPG-based vital sign monitoring, underlining the potential for enhancing robustness and reliability in various domains.
Qijia Shao, Li Zhu 0004, Mohsin Y. Ahmed, Korosh Vatanparvar, Migyeong Gwak, Jungmok Bae, Jilong Kuang, Jun Alex Gao
ICASSP8
2024 Freq2Time: Weakly Supervised Learning of Camera-Based RPPG from Heart Rate
abstract
Camera-based pulse measurements from remote photoplethysmography (rPPG) have rapidly improved over recent years due to innovations in video processing and deep learning. However, modern data-driven solutions require large training datasets collected under diverse conditions. Collecting such training data is made more challenging by the need for time-synchronized video and physiological signals as ground truth. This paper presents a weakly supervised learning framework, Freq2Time, to train with heart rate (HR) labels. Our framework mitigates the need for simultaneous PPG or ECG as ground truth, since the HR changes relatively slowly and describes the target rPPG signal over a time interval. We show that 3D convolutional neural network (3DCNN) models trained with the Freq2Time framework give state-of-the-art HR performance with MAE of 2.86 bpm, when tested with challenging smartphone video data from 30 subjects. Additionally, our models still learn accurate rPPG time signals, allowing for other physiological metrics such as heart rate variability.
Jeremy Speth, Korosh Vatanparvar, Li Zhu 0004, Jilong Kuang, Jun Alex Gao
ICASSP4
2023 VTMonitor: Tidal Volume Estimation Using Earbuds
abstract
Tidal volume (VT) is defined as the volume of inhaled and exhaled air during normal breath, which is crucial for maintaining respiratory function, such as adequate air exchange in and out of the body. However, existing estimation methods either require complex setups or involve inconvenient and expensive devices, such as spirometer and chestband. Thus, leveraging the advanced artificial intelligence (AI) and wearable devices, we aim to develop a novel, accessible and convenient approach to estimate tidal volume. In this study, we propose the VTMonitor system, which utilizes consumer earbuds’ motion sensor data to estimate the tidal volume. We conducted two experiments, collecting data either in lab or at home. After analyzing the data, our VTMonitor system is effective in measuring the tidal volume.
Yincheng Jin, Tousif Ahmed, Lana Mukharesh, Jilong Kuang, Jun Alex Gao
BSN5
2023 Advancements in Face Alignment Evaluation for Contact-less Vital Sign Detection
abstract
The emergence of remote vital sign measurement techniques has provided an alternative approach for monitoring vital signs without direct physical contact. However, the performance of contactless methods such as remote photoplethysmography are dependent on the accuracy of face detection algorithms. The misalignment of the face pixels from frame to frame can introduce jitters in the generated rPPG signals, and in turn, interfere with vital sign estimations. Nonetheless, the investigation into the performance of face detectors mostly focused on quantifying the accuracy of face landmarks based on an individual image. How to assess facial alignments across video frames has largely remained unknown and understudied. To address this issue, this paper introduced three novel metrics for assessing face alignment in remote vital sign detection: (1) Circular Radius, (2) Mean Offset, and (3) Percentage of Impacted Pixels. We evaluated two face detectors using proposed metrics in static and motion scenarios, where static represents no facial movement and motion scenarios involve facial movements induced by breathing. Our experiments demonstrated that employing the superior face detector recommended by our metrics resulted in a noteworthy 12.5% reduction in second-level mean absolute error and a corresponding 3.0% improvement in 5%-accuracy for remote heart rate estimation.
Roghayeh Barmaki, Li Zhu 0004, Korosh Vatanparvar, Migyeong Gwak, Jilong Kuang, Jun Alex Gao
BSN6
2023 Activity State Tracking Under Non-Restricted Ambulatory Condition
abstract
Human Activity Recognition (HAR) is one important digital health applications to track fitness or to avoid sedentary behavior. Due to the growing popularity of consumer wearable devices, smartwatches and earbuds are being widely adopted for HAR applications. However, using just one of the devices may not be sufficient to track all activities properly. Additionally, handling motion noise becomes more challenging when a single device is used. This paper proposes a multi-modal approach to HAR by using both buds and watch. Using a large dataset of 53 subjects collected from both controlled and uncontrolled noisy environments, we demonstrate the limitations of using a single modality activity classification. We identify various noise sources imposed in uncontrolled environment and propose two novel noise handling methods to ensure the robustness of activity state tracking. We build on top of a previous activity tracking effort and demonstrate a 7.8% sensitivity improvement against current state of the art in uncontrolled noisy environment.
Ebrahim Nemati, Mohsin Y. Ahmed, Jilong Kuang, Jun Alex Gao
BSN4
2023 Remote Breathing Rate Tracking in Stationary Position Using the Motion and Acoustic Sensors of Earables
abstract
Breathing rate is critical for the user’s respiratory health and is hard to track outside the clinical context, requiring specialized devices. Earables could provide a convenient solution to track the breathing rate anywhere by leveraging the user’s breathing-related motion and sound captured through the earables’ motion sensors and microphones. However, small non-breathing head movements or background noises during the assessment affect the estimation accuracy. While noise filtering improves accuracy, it can discard valid measurements. This paper presents a multimodal approach to tracking the user’s breathing rate using a signal-processing-based algorithm on motion sensors and a lightweight machine-learning algorithm on acoustic sensors from the earables that balances the accuracy and data retention. A user study with 30 participants shows that the system can accurately calculate breathing rate (Mean Absolute Error < 2 breaths per minute) while retaining most breathing sessions (75%) performed in real-world settings. This work provides an essential direction for remote breathing rate monitoring.
Tousif Ahmed, Ebrahim Nemati, Mohsin Y. Ahmed, Jilong Kuang, Jun Alex Gao
CHI5
2023 Mouth Breathing Detection Using Audio Captured Through Earbuds
abstract
Mouth breathing has been linked to a variety of negative health outcomes, including sleep-related disorders and dental problems. Detecting mouth breathing in the daily environment could be helpful for early intervention and reversing the negative impact. However, existing research has not adequately explored methods for detecting mouth breathing in everyday settings. This study presents a machine-learning approach using audio captured by commercially available earbuds to detect mouth breathing. By leveraging the growing popularity of earbuds for health monitoring, this approach offers a more convenient and non-invasive means of detecting mouth breathing. We conducted a data collection study with 30 participants to train a convolutional neural network-based model, which achieved an accuracy of 78.4% in detecting mouth breathing. Our findings suggest that audio-based mouth breathing detection using earbuds could be a promising tool for early intervention and improved health outcomes.
Tousif Ahmed, Ebrahim Nemati, Jilong Kuang, Jun Alex Gao
ICASSP4
2023 Improving Heart Rate and Heart Rate Variability Estimation from Video Through a HR-RR-Tuned Filter
abstract
This paper presents algorithms to improve the estimation of heart rate (HR) and heart rate variability (HRV) from smartphone video. The remote photoplethysmogram (rPPG) signals are first extracted from the videos recorded. Next, we proposed an rPPG filter adaptively tuned by HR and respiratory rate (RR) to better enhance source signal that modulates HR. Additionally, we also addressed a unique smartphone artifact—occasionally seen in smartphone videos—by introducing a threshold-based algorithm. HR and HRV accuracies are assessed on 22 subjects who were instructed to breath at seven different RRs. The mean absolute errors of HR and standard deviation of the NN intervals (SDNN) are found to be 1.13 ± 0.68 bpm and 18.30 ± 10.33 ms respectively. Finally, we also conduct a few experiments to highlight the accuracy improvements made by the proposed algorithms.
Michael Chan 0006, Li Zhu 0004, Korosh Vatanparvar, Hewon Jung, Jilong Kuang, Jun Alex Gao
ICASSP5
2023 BreathIE: Estimating Breathing Inhale Exhale Ratio Using Motion Sensor Data from Consumer Earbuds
abstract
Breathing Inhale/Exhale (IE) ratio is one of the critical breathing biomarkers for pulmonary patients and healthy individuals. It can indicate the severity of lung obstruction for chronic lung patients and help detect psycho-social stress for healthy individuals. With the advancement of wearable technologies, common consumer wearables such as smartwatches offer breathing rates. However, IE ratio measurement is not available in consumer wearable devices till today. In this paper, we present a novel algorithm, BreathIE, to estimate breathing rate and IE ratio using a low-power motion sensor embedded in consumer-grade earbuds. Moreover, our algorithm is adaptive and dynamically adjusts to the user’s breathing habit by accommodating varying breathing durations at run time. We conducted a study with 30 participants, where both earbuds and a reference chestband device were used simultaneously. Experimental evaluation against the annotated reference data shows that our algorithm can estimate breathing rate with a mean absolute error (MAE) of 2.37 breaths per minute (BPM) and breathing IE ratio of 0.27 MAE while outperforming the state-of-the-art algorithms.
Tousif Ahmed, Jilong Kuang, Jun Alex Gao
ICASSP4
2022 Deep Audio Spectral Processing for Respiration Rate Estimation from Smart Commodity Earbuds
abstract
Respiration rate is an important health biomarker and a vital indicator for health and fitness. With smart earbuds gaining popularity as a commodity device, recent works have demonstrated the potential for monitoring breathing rate using such earable devices. In this work, for the first time we utilize deep image recognition techniques to infer respiration rate from earbud audio. We use image spectrograms from breathing cycle audio signals captured using Samsung earbuds as a spectral feature to train a deep convolutional neural network. Using novel earbud audio data collected from 30 subjects with both controlled breathing at a wide range (from 5 upto 45 breaths per minute), and uncontrolled natural breathing from 7-day home deployment, experimental results demonstrate that our model outperforms existing methods using earbuds for inferring respiration rates from regular intensity breathing and heavy breathing sounds with 0.77 aggregated MAE for controlled breathing and with 0.99 aggregated MAE for at-home natural breathing.
Mohsin Y. Ahmed, Tousif Ahmed, Jilong Kuang, Jun Alex Gao
BSN5
2022 Enhancement of Remote PPG and Heart Rate Estimation with Optimal Signal Quality Index
abstract
With the popularity of non-invasive vital signs detection, remote photoplethysmography (rPPG) is drawing attention in the community. Remote PPG or rPPG signals are extracted in a contactless manner that is more prone to artifacts than PPG signals collected by wearable sensors. To develop a robust and accurate pipeline to estimate heart rate (HR) from rPPG signals, we propose a novel real-time dynamic ROI tracking algorithm that applies to slight motions and light changes. Furthermore, we develop and include a signal quality index (SQI) to improve the HR estimation accuracy. Studies have explored optimal SQIs for PPG signals, but not for remote PPG signals. In this paper, we select and test six SQIs: Perfusion, Kurtosis, Skewness, Zero-crossing, Entropy, and signal-to-noise ratio (SNR) on 124 rPPG sessions from 30 participants wearing masks. Based on the mean absolute error (MAE) of HR estimation, the optimal SQI is selected and validated by Mann–Whitney U test (MWU). Lastly, we show that the HR estimation accuracy is improved by 29% after removing outliers decided by the optimal SQI, and the best result achieves the MAE of 2.308 bpm.
Jiyang Li, Korosh Vatanparvar, Li Zhu 0004, Jilong Kuang, Jun Alex Gao
BSN4
2022 Respiration Rate Estimation from Remote PPG via Camera in Presence of Non-Voluntary Artifacts
abstract
Contactless measurement of vitals has been seen as a promising alternative to contact sensors for monitoring of health condition. In this paper, we focus on respiration rate (RR) as one of the fundamental biomarkers of a person’s cardio and pulmonary activities. Remote RR estimation has gained attraction due to its various potential applications; use of RGB cameras to extract remote photoplethysmography (PPG) signal from subjects’ face has been debated as one of the enabling technologies for remote RR estimation. The technology is challenged with respect to wide range of RR and non-voluntary motion in uncontrolled settings. We propose a novel methodology to enhance the remote PPG signal and remove artifacts from the respiration signal. The method achieves 3.9bpm MAE of 90% percentile (1.3bpm decrease) for estimating RR in range of 5-25bpm. We validate the performance using smartphone video recordings of 30 subjects with uniform distribution of skin tone.
Korosh Vatanparvar, Migyeong Gwak, Li Zhu 0004, Jilong Kuang, Jun Alex Gao
BSN4
2022 Real-Time Breathing Phase Detection Using Earbuds Microphone
abstract
Tracking breathing phases (inhale and exhale) outside the hospitals can offer significant health and wellness benefits. For example, the breathing phases can provide fine-grained breathing information for breathing exercises. While previous works use smartphones and smartwatches for tracking breathing phases, in this work, we use earbuds for breathing phase detection, which can be a better form factor for breathing exercises as it requires less user attention from the user. We propose a convolutional neural network-based algorithm for detecting breathing phases using the audio captured through the earbuds during guided breathing sessions. We conducted a user study with 30 participants in both lab and home environments to develop and evaluate our algorithm. Our algorithm can detect the breathing phases with 85% accuracy by taking only a 500ms audio signal. Our work demonstrates the potential of using earbuds for tracking the breathing phases in real-time.
Tousif Ahmed, Mohsin Y. Ahmed, Ebrahim Nemati, Jilong Kuang, Jun Alex Gao
BSN6
2022 Contactless SpO2 Detection from Face Using Consumer Camera
abstract
We describe a novel computational framework for contactless oxygen saturation (SpO2) detection using videos recorded from human faces using smartphone cameras with ambient light. For contact pulse oximeter, a ratio of ratios (RoR) metric derived from selected regions of interest (ROI) combined with linear regression modeling is the standard approach. However, when used upon contactless remote PPG (rPPG), the assumptions of this standard approach do not hold automatically: 1) the rPPG signal is usually derived from the face area where the light reflection may not be uniform due to variation in skin tissue composition and/or lighting conditions (moles, hairs, beard, partial shadowing, etc.), 2) for most consumer-level cameras under ambient light, the rPPG signal is converted from light reflection associated with wide-band spectra, which creates complicated nonlinearity for SpO2mappings. We propose a computational framework to overcome these challenges by 1) determining and dynamically tracking the ROIs according to both spatial and color proximity, and calculating the RoR based on selected individual ROIs which have homogeneous skin reflections, and 2) using a nonlinear machine learning model to mapping the SpO2levels from RoRs derived from two different color combinations. We validated the framework with 30 healthy participants during various breathing tasks and achieved 1.24% Root Mean Square Error for across-subjects model and 1.06% for within-subject models, which surpassed the FDA-recognized ISO 81060-2-61:2017 standard.
Li Zhu 0004, Korosh Vatanparvar, Migyeong Gwak, Jilong Kuang, Jun Alex Gao
BSN4
2022 Ubilung: Multi-Modal Passive-Based Lung Health Assessment
abstract
Lung health assessment is traditionally done mainly through X-ray images and spirometry tests which are time-consuming, cumbersome, and costly. In this paper, we investigate the potential of passively recordable contents such as speech, cough and heart signal for such an assessment. Our regression model is the first in the literature to achieve mean absolute error (MAE) of 7.47% for estimation of forced expiratory volume in 1 sec. (FEV1) over forced vital capacity (FVC) ratio using these contents. This is comparable to the state of the art active phone-based spirometry methods. Additionally our classification models achieve a F1-score of 0.982 for healthy v.s. diseased, 0.881 for obstructive v.s. non-obstructive, 0.854 for chronic obstructive pulmonary disease (COPD) v.s. asthma, and 0.892 for severe v.s. non-severe obstruction classification.
Ebrahim Nemati, Xuhai Xu, Viswam Nathan, Korosh Vatanparvar, Tousif Ahmed, Daniel McCaffrey 0001, Jilong Kuang, Jun Alex Gao
ICASSP8
2022 Coughtrigger: Earbuds IMU Based Cough Detection Activator Using An Energy-Efficient Sensitivity-Prioritized Time Series Classifier
abstract
Persistent coughs are a major symptom of respiratory-related diseases. Increasing research attention has been paid to detecting coughs using wearables, especially during the COVID-19 pandemic. Microphone is most widely used sensor to detect coughs. However, the intense power consumption needed to process audio hinders continuous audio-based cough detection on battery-limited commercial wearables, such as earbuds. We present CoughTrigger, which utilizes a lower-power sensor, inertial measurement unit (IMU), in earbuds as a cough detection activator to trigger a higher-power sensor for audio processing and classification. It runs all-the-time as a standby service with minimal battery consumption and triggers the audio-based cough detection when a candidate cough is detected from IMU. Besides, the use of IMU brings the benefit of improved specificity of cough detection. Experiments are conducted on 45 subjects and CoughTrigger achieved 0.77 AUC score. We also validated its effectiveness on free-living data and through on-device implementation.
Ebrahim Nemati, Minh Dinh, Nathan Folkman, Tousif Ahmed, Jilong Kuang, Nabil Alshurafa, Jun Alex Gao
ICASSP7
2022 BreatheBuddy: Tracking Real-time Breathing Exercises for Automated Biofeedback Using Commodity Earbuds
abstract
Breathing exercises reduce stress and improve overall mental well-being. There are various types of breathing exercises. Performing the exercises correctly may give the best outcome and doing it in wrong ways can sometimes have adverse effect. Providing real-time biofeedback can greatly improve the user experience in doing the right exercises in the right ways. In this paper, we present methods to passively track breathing biomarkers in real-time using wireless commodity earbuds and generate feedback on users' breathing performance. We use the earbud's low-power accelerometer to generate a comprehensive set of breathing biomarkers including breathing phase, breathing rate, depth of breathing, and breathing symmetry. We have conducted studies where the subjects performed different types of guided breathing exercises while wearing the earbuds. Our algorithms detect breathing phases with 90.91% F1-score and estimate breathing rate with 95.05% accuracy. We further show that our algorithms can be used to generate biofeedback towards designing engaging smartphone's user interactions that facilitate users to accurately perform various breathing exercises.
Tousif Ahmed, Mohsin Y. Ahmed, Minh Dinh, Ebrahim Nemati, Jilong Kuang, Jun Alex Gao
Proc. ACM Hum. Comput. Interact.6
2022 Atrial Fibrillation Detection and Atrial Fibrillation Burden Estimation via Wearables
abstract
Atrial Fibrillation (AF) is an important cardiac rhythm disorder, which if left untreated can lead to serious complications such as a stroke. AF can remain asymptomatic, and it can progressively worsen over time; it is thus a disorder that would benefit from detection and continuous monitoring with a wearable sensor. We develop an AF detection algorithm, deploy it on a smartwatch, and prospectively and comprehensively validate its performance on a real-world population that included patients diagnosed with AF. The algorithm showed a sensitivity of 87.8% and a specificity of 97.4% over every 5-minute segment of PPG evaluated. Furthermore, we introduce novel algorithm blocks and system designs to increase the time of coverage and monitor for AF even during periods of motion noise and other artifacts that would be encountered in daily-living scenarios. An average of 67.8% of the entire duration the patients wore the smartwatch produced a valid decision. Finally, we present the ability of our algorithm to function throughout the day and estimate the AF burden, a first-of-this-kind measure using a wearable sensor, showing 98% correlation with the ground truth and an average error of 6.2%.
Li Zhu 0004, Viswam Nathan, Jilong Kuang, Jacob Kim, Robert Avram, Jeffrey E. Olgin, Jun Alex Gao
IEEE J. Biomed. Health Informatics3
2021 CoughBuddy: Multi-Modal Cough Event Detection Using Earbuds Platform
abstract
There has been an extensive amount of study on cough detection using acoustic features captured from smartphones and smartwatches in the past decade. However, the specificity of the algorithms has always been a concern when exposed to the unseen field data containing cough-like sounds. In this paper, we propose a novel sensor fusion algorithm that employs a hybrid of classification and template matching algorithms to tackle the problem of unseen classes. The algorithm utilizes in-ear audio signal as well as head motion captured by the inertial measurement unit (IMU). A clinical study including 45 subjects from healthy and chronic cough cohorts was conducted that contained various tasks including cough and cough-like body sounds in various conditions such as quiet/noisy and stationary/non-stationary. Our hybrid model was evaluated for sensitivity and specificity in these conditions using leave one-subject out validation (LOSOV) and achieved an average sensitivity of 83% for stationary tasks and an specificity of 91.7% for cough-like sounds reducing the false positive rate by 55%. These results indicate the feasibility and superiority of fusion in earbuds platforms for detection of cough events.
Ebrahim Nemati, Tousif Ahmed, Jilong Kuang, Jun Alex Gao
BSN5
2021 Towards Motion-Aware Passive Resting Respiratory Rate Monitoring Using Earbuds
abstract
Breathing rate is an important vital sign and an indicator of overall health and fitness. Traditionally breathing is monitored using specialized devices such as chestband or spirometers which are uncomfortable for daily use. Recent works show the feasibility of estimating breathing rate using earbuds' motion sensors. However, non-breathing head motion is one of the biggest challenges for breathing rate estimation using earbuds. In this paper, we propose algorithms to estimate breathing rate in presence of non-breathing head motion using inertial sensors embedded in commodity earbuds. Using the chestband as a reference device, we show that our algorithms can estimate breathing rate in resting positions with error rate 2.34 breaths per minute (BPM). Our algorithms can handle passive head motion and reduce the error by 27.78%. Furthermore, our algorithms can handle active head motion and help reduce the error by 45.70% when intentional non-breathing head motion is present in the data segment. It can be a big stride towards passive breathing monitoring in daily life using commodity earbuds.
Tousif Ahmed, Mohsin Y. Ahmed, Ebrahim Nemati, Minh Dinh, Nathan Folkman, Jilong Kuang, Jun Alex Gao
BSN8
2021 Real-Time 3D Arm Motion Tracking Using the 6-axis IMU Sensor of a Smartwatch
abstract
Inertial measurement unit (IMU) sensors are widely used in motion tracking for various applications, e.g., virtual physical therapy and fitness training. Traditional IMU-based motion tracking systems use 9-axis IMU sensors that include an accelerometer, gyroscope, and magnetometer. The magnetometer is essential to correct the yaw drift in orientation estimation. However, its magnetic field measurement is often disturbed by the ferromagnetic materials in the environment and requires frequent calibration. Moreover, most IMU-based systems require multiple IMU sensors to track the body motion and are not convenient for use. In this paper, we propose a novel approach that uses a single 6-axis IMU sensor of a consumer smartwatch without any magnetometer to track the user's 3D arm motion in real time. We use a recurrent neural network (RNN) model to estimate the 3D positions of both the wrist and the elbow from the noisy IMU data. Compared with the state-of-the-art approaches that use either the 9-axis IMU sensor or the combination of a 6-axis IMU and an extra device, our proposed approach significantly improves the usability and potential for pervasiveness by not requiring a magnetometer or any extra device, while achieving comparable results.
Wenchuan Wei, Keiko Kurita, Jilong Kuang, Jun Alex Gao
BSN3
2021 Better Battery Life: Towards Energy-Efficient Smartwatch-Based Atrial Fibrillation Detection in Ambulatory Free-living Environments
abstract
Atrial Fibrillation (AF) is an important medical condition that can be passively detected and tracked using a smartwatch. Diagnosis and monitoring of AF can be more effective and reliable if the smartwatch senses continuously, but this can lead to significant battery consumption by the LED in the photoplethysmography (PPG) sensor. In this paper, we explore the feasibility of leveraging downsampling to achieve energy-efficient AF detection. We collect data from participants with paroxysmal AF in real ambulatory free-living environments using a commercial smartwatch and separately study the impact of uniform downsampling and compressed sensing on AF detection. Our results reveal that downsampling enables the AF detection system to consume about 77.4% less LED power than the original sampling strategy without a significant performance drop.
Hanbin Zhang, Li Zhu 0004, Viswam Nathan, Jilong Kuang, Jacob Kim, Jun Alex Gao
BSN4
2020 Assessing Severity of Pulmonary Obstruction from Respiration Phase-Based Wheeze-Sensing Using Mobile Sensors
abstract
Obstructive pulmonary diseases cause limited airflow from the lung and severely affect patients' quality of life. Wheeze is one of the most prominent symptoms for them. High requirements imposed by traditional diagnosis methods make regular monitoring of pulmonary obstruction challenging, which hinders the opportunity of early intervention and prevention of significant exacerbation. In this work, we explore the feasibility of developing a mobile sensor-based system as a convenient means of assessing the severity of pulmonary obstruction via respiration phase-based symptomatic wheeze sensing. We conduct a 131 subjects' (91 patients and 40 healthy) study for the detection (F1: 87.96%) and characterization (F1: 79.47%) of wheeze. Subsequently, we develop novel wheeze metrics, which show a significant correlation (Pearson's correlation: -0.22, p-value: 0.024) with standard spirometry measure of pulmonary obstruction severity. This work takes a principal step towards the unobtrusive assessment of pulmonary condition from mobile sensor interactions.
Soujanya Chatterjee, Tousif Ahmed, Nazir Saleheen, Ebrahim Nemati, Viswam Nathan, Korosh Vatanparvar, Jilong Kuang
CHI8
2020 Automated Time Synchronization of Cough Events from Multimodal Sensors in Mobile Devices
abstract
Tracking the type and frequency of cough events is critical for monitoring respiratory diseases. Coughs are one of the most common symptoms of respiratory and infectious diseases like COVID-19, and a cough monitoring system could have been vital in remote monitoring during a pandemic like COVID-19. While the existing solutions for cough monitoring use unimodal (e.g., audio) approaches for detecting coughs, a fusion of multimodal sensors (e.g., audio and accelerometer) from multiple devices (e.g., phone and watch) are likely to discover additional insights and can help to track the exacerbation of the respiratory conditions. However, such multimodal and multidevice fusion requires accurate time synchronization, which could be challenging for coughs as coughs are usually concise events (0.3-0.7 seconds). In this paper, we first demonstrate the time synchronization challenges of cough synchronization based on the cough data collected from two studies. Then we highlight the performance of a cross-correlation based time synchronization algorithm on the alignment of cough events. Our algorithm can synchronize 98.9% of cough events with an average synchronization error of 0.046s from two devices.
Tousif Ahmed, Mohsin Y. Ahmed, Ebrahim Nemati, Bashima Islam, Korosh Vatanparvar, Viswam Nathan, Daniel McCaffrey 0001, Jilong Kuang, Jun Alex Gao
ICMI9
2020 BreathEasy: Assessing Respiratory Diseases Using Mobile Multimodal Sensors
abstract
Mobil respiratory assessments using commodity smartphones and smartwatches are unmet needs for patient monitoring at home. In this paper, we show the feasibility of using multimodal sensors embedded in consumer mobile devices for non-invasive, low-effort respiratory assessment. We have conducted studies with 228 chronic respiratory patients and healthy subjects, and show that our model can estimate respiratory rate with mean absolute error (MAE) 0.72$\pm$0.62 breath per minute and differentiate respiratory patients from healthy subjects with 90% recall and 76% precision when the user breathes normally by holding the device on the chest or the abdomen for a minute. Holding the device on the chest or abdomen needs significantly lower effort compared to traditional spirometry which requires a specialized device and forceful vigorous breathing. This paper shows the feasibility of developing a low-effort respiratory assessment towards making it available anywhere, anytime through users' own mobile devices.
Mohsin Y. Ahmed, Tousif Ahmed, Bashima Islam, Viswam Nathan, Korosh Vatanparvar, Ebrahim Nemati, Daniel McCaffrey 0001, Jilong Kuang, Jun Alex Gao
ICMI9
2020 Lung Function Estimation from a Monosyllabic Voice Segment Captured Using Smartphones
abstract
Chronic respiratory diseases refer to a group of lung diseases that affect the airways and cause difficulty in breathing. Respiratory diseases are one of the leading causes of death and negatively impact the patients’ quality of life. Early detection and regular monitoring of lung functions might reduce the risk of death; however, lung function assessment requires the active supervision of a medical professional in a clinical setting. To make lung function tests more accessible and ubiquitous, researchers started leveraging mobile devices, which still require active supervision and demand extraneous effort from the user. In this work, we propose a convenient mobile-based approach that uses a monosyllabic voice segment called ‘A-vowel’ sound or ‘Aaaa...’ sound to estimate lung function. We conducted two studies (a lab study and an in-clinic study) with 201 participants to develop a detection model detecting ‘A-vowel’ sound from other acoustic events and a prediction model to estimate the lung function using the detected A-vowel sound. Our study shows that A-vowel sounds can be detected with 93% accuracy, and A-vowel sounds can estimate lung functions with 7.4-11.35% mean absolute error. We also conducted a validation study with 10 participants in a noisy environment and able to detect A-vowel segments with 71% F1-Score. Our results show auspicious directions to expand the horizon of mobile-based lung assessment.
Nazir Saleheen, Tousif Ahmed, Ebrahim Nemati, Viswam Nathan, Korosh Vatanparvar, Erin Blackstock, Jilong Kuang
MobileHCI8
2020 Towards Passive Assessment of Pulmonary Function from Natural Speech Recorded Using a Mobile Phone
abstract
Chronic obstructive pulmonary disease (COPD) and asthma are the most common respiratory diseases that impact millions of people worldwide annually. With advances in mobile computing and machine learning techniques, there has been increased interest in using mobile devices to monitor pulmonary diseases. Nevertheless, the current state-of-the-art technology requires active involvement and high-effort input from the users, impeding continuous monitoring of pulmonary conditions. In this work, two algorithms are proposed for passive assessment of pulmonary condition: one for detection of obstructive pulmonary disease and the other for estimation of the pulmonary function in terms of FEV1/FVC ratio, which is an established clinical metric. The algorithms were developed and validated using the data sets from two studies: research study (healthy=40, pathological=91) and in-clinic study (healthy=10, pathological=60). From the cross-study validation where a classifier was trained on the research data set and tested on the in-clinic data set, the detection accuracy of the pathological class was obtained as 73.7% and the F1 score was 84.5% (87.2% precision and 82.0% recall). In our regression analysis, the FEV1/FVC ratio was predicted with a mean absolute error of 8.6%. Our analysis shows promising results and this work presents a meaningful milestone towards the passive assessment of pulmonary functions from spontaneous speech collected from a mobile phone.
Keum San Chun, Viswam Nathan, Korosh Vatanparvar, Ebrahim Nemati, Erin Blackstock, Jilong Kuang
PerCom7
2020 ExhaleSense: Detecting High Fidelity Forced Exhalations to Estimate Lung Obstruction on Smartphones
abstract
Spirometry is the gold standard to measure lung functions by estimating the maximum air an individual can forcefully exhale as quickly as possible. It is used not only to diagnose lung diseases such as asthma, chronic obstructive pulmonary disease (COPD) but also to assess the severity of the pulmonary condition. However, spirometry requires a specialized device called spirometer, which is mostly available in clinical facilities and cumbersome to use. Recent works have shown the feasibility of using smartphone microphone to estimate lung functions from forced exhalation effort sounds. However, maintaining the fidelity of lung function estimation on smartphones becomes challenging in unsupervised field environment in presence of other sounds such as coughs, deep inhalation, regular breathing, and speech. In this paper, we present ExhaleSense that detects forced exhalation efforts on smartphones from audio time-series data, distinguishes high fidelity efforts from poor efforts, and estimates lung obstruction. By conducting three studies with 211 pulmonary patients and healthy subjects, we show that ExhaleSense can detect forced exhalation sounds with 96.74% F1-score and estimate lung obstruction with mean absolute error as low as 7.57%. ExhaleSense shifts the gear of smartphone spirometry research from feasibility to ensuring effort quality towards high fidelity lung function estimation in unsupervised field settings.
Tousif Ahmed, Ebrahim Nemati, Viswam Nathan, Korosh Vatanparvar, Erin Blackstock, Jilong Kuang
PerCom7
2019 Extraction of Voice Parameters from Continuous Running Speech for Pulmonary Disease Monitoring
abstract
Pulmonary disease is one of the leading causes of death, and individuals with chronic diseases like asthma and chronic obstructive pulmonary disease (COPD) will have to manage their disease throughout their lifetime. Passive monitoring with mobile devices such as smartphones represents a cost-effective solution that is closely coupled with the user to detect and continuously track adverse pulmonary trends. Identifying changes in the user's voice patterns has seen recent research interest as a potentially useful biomarker that is relatively convenient and efficient to monitor. Prosodic features of the voice such as shimmer and jitter have been traditionally extracted from sustained vowel sounds in the majority of the previous work in this area. In this work, we describe a method to extract these parameters from continuous running speech, and show that they have better agreement with the corresponding parameters from sustained vowel sounds. This method was validated on real data from pulmonary patients collected under two different environments, and also compared to an existing speech processing package that is widely used in the literature to extract such features. The described method can be an important step in realizing passive monitoring of voice changes from the natural speech of pulmonary patients in their day to day lives.
Viswam Nathan, Korosh Vatanparvar, Ebrahim Nemati, Erin Blackstock, Jilong Kuang
BIBM6
2019 mLung: Privacy-Preserving Naturally Windowed Lung Activity Detection for Pulmonary Patients
abstract
mLung is a privacy preserving, naturally windowed, mobile-cloud hybrid pulmonary care service for detecting unusual lung sounds like coughing and wheezing from streaming audio and inertial sensor data from a smartphone for pulmonary patients. mLung employs a combination of: (1) natural windowing of audio data from the patient respiration cycle captured by the inertial sensors, (2) in-phone speech detection and filtering by a lightweight classifier for patient privacy, and (3) in-cloud lung and confounding sound classification by a heavyweight and expert supervised classifier. This paper describes the design and architecture of mLung and using novel lung activity data collected by smartphone from 131 patients and healthy subjects, provides empirical evidence that mLung is 15%-25% more accurate in detecting lung sounds when compared to a state-of-the-art phone based internal body sound detection system using specialized microphone hardware, with a best f-1 score of 98%.
Mohsin Y. Ahmed, Viswam Nathan, Ebrahim Nemati, Korosh Vatanparvar, Jilong Kuang
BSN6
2019 Assessment of Chronic Pulmonary Disease Patients Using Biomarkers from Natural Speech Recorded by Mobile Devices
abstract
Chronic pulmonary disease is one of the leading causes of mortality in the United States. Continuous passive monitoring of subjects using mobile sensors can help detect disease, estimate severity, track progression over time, and predict adverse exacerbation events. One of the most convenient avenues to realize this goal is through analysis of passively recorded natural speech patterns. It has been previously established that diseases such as asthma and chronic obstructive pulmonary disease (COPD) affect pause patterns and prosodic features of speech. In this study we present an exploration of the feasibility of using speech features from natural speech to detect pulmonary disease. Experiments were conducted on a cohort of 131 subjects: 91 with asthma and/or COPD, and 40 healthy controls. Patients and healthy subjects were differentiable with 68% accuracy; moreover, the subset of patients with the highest disease severity were detected with 89% accuracy.
Viswam Nathan, Korosh Vatanparvar, Ebrahim Nemati, Jilong Kuang
BSN5
2019 A Generative Model for Speech Segmentation and Obfuscation for Remote Health Monitoring
abstract
The prevalence of smart devices has enabled remote health monitoring outside of conventional clinical settings, and has reduced health care delivery cost. Passive audio recording is an essential component in remote health monitoring, however, it poses major privacy issues for subjects in uncontrolled environments like their home. There are existing voice activity detection and speech classification methodologies to identify sound events and obfuscate the human speech. However, they result in frequent false positives when distinguishing human speech from other sound events; their performance is limited to a controlled environment for a specific application; and require large amount of labeled data for training. In this paper, we present a novel speech privacy preservation methodology using generative adversarial networks to segment human speech in a recorded audio and generate human-like random speech to replace the original segment. We implemented our methodology and experimented on standard datasets of speech, environmental sounds, and cough samples generated from our internal mobile health study. Compared to current methodologies, our experimental results show much lower speech segmentation true positive rates of 17% and 14% for environmental sounds and cough datasets. Moreover, randomly generated audio samples to obfuscate the speech are shown to be likely indistinguishable from human speech (lower than 0.9% error in spectral attributes).
Korosh Vatanparvar, Viswam Nathan, Ebrahim Nemati, Jilong Kuang
BSN5
2019 DeepLung: Smartphone Convolutional Neural Network-Based Inference of Lung Anomalies for Pulmonary Patients
Mohsin Y. Ahmed, Jilong Kuang
INTERSPEECH3
2018 Configurable Pulmonary-Tuned Privacy Preservation Algorithm for Mobile Devices
Sujee Lee, Ebrahim Nemati, Jilong Kuang
BIBM3
2018 Recurrent Neural Networks Based Obesity Status Prediction Using Activity Data
abstract
Obesity, a serious public health concern worldwide, increases the risk of many diseases, including hypertension, stroke, and type 2 diabetes. To tackle this problem, researchers collect diverse types of data, which includes biomedical, behavioral and activity, and utilize machine learning techniques to mine hidden patterns for obesity status improvement prediction. While existing machine learning methods such as Recurrent Neural Networks (RNNs) provide exceptional results, it is challenging to discover hidden patterns of the sequential data due to the irregular observation time instances. Meanwhile, the lack of understanding of why those learning models are effective also limits further improvements on their architectures. Thus, we develop a RNN based time-aware architecture to handle irregular observation times and identify relevant feature extractions from longitudinal patient records for obesity status improvement pre-diction. Evaluations of real-world data involving activity data collected from wearables and electronic health records demonstrate that our proposed method can capture the underlying structures in users' time sequences with irregularities, and achieve an accuracy of 77% in predicting the obesity status improvement.
Qinghan Xue, Samuel Meehan, Jilong Kuang, Jun Alex Gao, Mooi Choo Chuah
ICMLA4
2015 Techniques for fast and scalable time series traffic generation
abstract
Many IoT applications ingest and process time series data with emphasis on 5Vs (Volume, Velocity, Variety, Value and Veracity). To design and test such systems, it is desirable to have a high-performance traffic generator specifically designed for time series data, preferably using archived data to create a truly realistic workload. However, most existing traffic generator tools either are designed for generic network applications, or only produce synthetic data based on certain time series models. In addition, few have raised their performance bar to millions-packets-per-second level with minimum time violations. In this paper, we design, implement and evaluate a highly efficient and scalable time series traffic generator for IoT applications. Our traffic generator stands out in the following four aspects: 1) it generates time-conforming packets based on high-fidelity reproduction of archived time series data; 2) it leverages an open-source Linux Exokernel middleware and a customized userspace network subsystem; 3) it includes a scalable 10G network card driver and uses "absolute" zero-copy in stack processing; and 4) it has an efficient and scalable application-level software architecture and threading model. We have conducted extensive experiments on both a quad-core Intel workstation and a 20-core Intel server equipped with Intel X540 10G network cards and Samsung's NVMe SSDs. Compared with a stock Linux baseline and a traditional mmap-based file I/O approach, we observe that our traffic generator significantly outperforms other alternatives in terms of throughput (10X), scalability (3.6X) and time violations (46.2X).
Jilong Kuang, Daniel G. Waddington, Changhui Lin
IEEE BigData1
2014 A scalable hash scheduler for decoding of multiple H.264/AVC streams on multi-core architecture
abstract
Existing scheduling schemes for decoding H.264/AVC multiple streams on multi-core are largely limited by ineffective use of multi-core architecture. Among the reasons are inefficient load balancing, in which common load metrics (e.g. tasks, frames, bytes) are unable to correctly reflect processing load at cores, unscalability of scheduling algorithms for a large scale multi-core, and bottlenecks at schedulers for multi-stream decoding. In this paper, we propose a scalable adaptive Highest Random Weight (HA-HRW) hash scheduler for distributed shared memory multi-core architecture considering the following: 1) memory access and core/cache topology of the multi-core architecture; 2) appropriate processing time load metric to enforce a true load balancing; 3) hierarchical parallel scheduling to decode multiple streams simultaneously; 4) locality characteristics of processing unit candidate to limit search within neighboring cores to enable scalable scheduling. We implement and evaluate our approach on a 32-core SGI server with realistic workload. Comparing with existing schemes, our scheme achieves higher throughput, better load balancing, better CPU utilization, and no jitter problem. Our scheme scales with multi-core and multiple streams as its time complexity is O(1).
Dung Vu, Jilong Kuang, Laxmi N. Bhuyan
ICME2
2013 Towards a Scalable Microkernel Personality for Multicore Processors
Jilong Kuang, Daniel G. Waddington, Chen Tian 0005
Euro-Par1
2012 A Scalable Physical Memory Allocation Scheme for L4 Microkernel
abstract
L4 microkernel family has become very successful on mobile devices. However, with the rapid shift from uniprocessor to multicore and manycore processor, many critical OS functions including physical memory allocator (PMA) must be re-designed in order to achieve better system throughput. While research and engineering efforts have been made for PMA in monolithic kernels such as Linux, not much work can be found for L4 microkernels. Due to the the design difference, the PMA in L4 microkernels is part of user level page fault handler (a.k.a. pager), which is executed as a stand-alone server in the least privilege mode. Memory allocation and free requests are handled through inter-process communication (IPC) rather than normal system or kernel function calls. In this work, we first study the scalability issue of the PMA implementation in L4 microkernels, and propose our solution in the context of Fiasco.OC, a state-of-the-art L4 microkernel implementation. We also discuss how to leverage the L4 microkernel design advantages to implement a PMA with more advanced features, such as load balancing, customizability and NUMA-awareness. Finally, we conduct experiments to verify the scalability result of our solution. The experiment is conducted on a 48-core AMD magny-cours server.
Chen Tian 0005, Daniel G. Waddington, Jilong Kuang
COMPSAC3
2012 Traffic-aware power optimization for network applications on multicore servers
abstract
In this paper, we design, implement, and evaluate a traffic-aware and power-efficient multicore server system by translating incoming traffic rate to appropriate system operating level, which is then translated to optimal per-core frequency configuration. According to the varying traffic rate, the system can adjust the number of active cores and per-core frequency "on-the-fly" via the use of per-core DVFS, power gating, and power migration techniques based on our new power model which considers both dynamic and static power consumption of all cores. Results on an AMD machine with two Quad-Core Opteron 2350 processors for six real network applications chosen from NetBench [19] show that our scheme reduces power consumption by an average of 41.0% compared to running with full capacity without any reduction in throughput. It also consumes less power than three other approaches, chip-wide DVFS [22], power gating [17], and chip-wide DVFS + power gating [15], by 35.2%, 24.3%, and 10.5% respectively.
Jilong Kuang, Laxmi N. Bhuyan, Raymond Klefstad
DAC1
2012 An Adaptive Dynamic Scheduling Scheme for H.264/AVC Decoding on Multicore Architecture
abstract
Parallelizing H.264/AVC decoding on multicore architectures is challenged by its inherent structural and functional dependencies at both frame and macro-block levels, as macro-blocks and certain frame types must be decoded in a sequential order. So far, dynamic scheduling scheme with recursive tail submit, as one of the best existing algorithms, provides a good throughput performance by exploiting macro-block level parallelism and mitigating global queue contention. Nevertheless, it fails to achieve an optimal performance due to 1) the use of global queue, which incurs substantial synchronization overhead when the number of cores increases and 2) the unawareness of cache locality with respect to the underlying hierarchical core/cache topology that results in unnecessary latency, communication cost and load imbalance. In this paper, we propose an adaptive dynamic scheduling scheme that employs multiple local queues to reduce lock contention, and assigns tasks in a cache locality aware and load-balancing fashion so that neighboring macro-blocks are preferably dispatched to nearby cores. We design, implement and evaluate our scheme on a 32-core cc-NUMA SGI server. Compared to existing alternatives by running real benchmark applications, we observe that our scheme produces higher throughput and lower latency with more balanced workload and less communication cost.
Dung Vu, Jilong Kuang, Laxmi N. Bhuyan
ICME2
2011 Predictive Model-Based Thermal Management for Network Applications
abstract
As processor power density has increased at an alarming rate, chip/core temperature control becomes critical in satisfying given thermal constraint and avoiding hotspots. Unlike "run-to-finish" applications whose temperature will simply rise to saturation point and then stabilize, network applications do periodic packet processing, which causes temperature to rise and fall over time. However, no existing studies have focused on characterizing the temperature variation for periodic tasks. We envision that volatile thermal behavior has to be well understood in order to optimize thermal management. In this paper, we first build a novel predictive thermal model for generic periodic tasks running on a single core. This model can dynamically derive the core temperature at any time quickly and accurately. To verify the model, we use both Hot Spot simulator and a real Linux machine to run six network applications chosen from Net Bench. Then, we propose an online model update strategy using on-chip thermal sensors, which can effectively correct incidental errors by adjusting model parameters "on-the-fly". Finally, by combining the thermal model and the online update, we design, implement and evaluate a predictive model-based thermal management scheme on an Intel Xeon E5335 core for network applications based on the Stop & Go technique. Compared with two other alternatives, our scheme achieves lower temperature, higher throughput, no thermal constraint violation, and negligible overhead cost.
Jilong Kuang, Laxmi N. Bhuyan
ANCS1
2011 E-AHRW: An Energy-Efficient Adaptive Hash Scheduler for Stream Processing on Multi-core Servers
abstract
We study a streaming network application-video transcoding to be executed on a multi-core server. It is important for the scheduler to minimize the total processing time and preserve good video quality in an energy-efficient manner. However, the performance of existing scheduling schemes is largely limited by ineffective use of the multi-core architecture characteristic and undifferentiated transcoding cost in terms of energy consumption. In this paper, we identify three key factors that collectively play important roles in affecting transcoding performance: memory access (M), core/cache topology (C) and transcoding format cost (C), or MC2for short. Based on MC2, we propose E-AHRW, an Energy-efficient Adaptive Highest Random Weight hash scheduler by extending the HRW scheduler proposed for packet scheduling on a homogeneous multiprocessor. E-AHRW achieves stream locality and load balancing at both stream and packet (frame) level by adaptively adjusting the hashing decision according to real-time weighted queue length of each processing unit (PU). Based on E-AHRW, we also design, implement and evaluate a hash-tree scheduler to further reduce the computation cost and achieve more effective load balancing on multi-core architectures. Through implementation on an Intel Xeon server and evaluations on realistic workload, we demonstrate that E-AHRW improves throughput, energy efficiency and video quality due to better load balancing, lower L2 cache miss rate and negligible scheduling overhead.
Jilong Kuang, Laxmi N. Bhuyan, Haiyong Xie 0001, Danhua Guo
ANCS1
2010 Power optimization for multimedia transcoding on multicore servers
abstract
We design, implement and evaluate a power-efficient and traffic-aware transcoding system on multicore servers that appropriately adjusts the processor operating level. The system is capable of configuring the number of active cores and core frequency "on-the-fly" according to the varying traffic rate. Results on an AMD machine show that our system saves 51.0% power consumption compared to a native system without power-saving schemes. It also outperforms three other power-aware systems, CG [2] (clock gating), C-DVFS [3] (chip-wide DVFS) and Hybrid [1] (chip-wide DVFS + power-gating), by 19.5%, 10.5% and 5.5% reduction of power consumption, respectively.
Jilong Kuang, Danhua Guo, Laxmi N. Bhuyan
ANCS1
2010 LATA: a latency and throughput-aware packet processing system
abstract
Current packet processing systems only aim at producing high throughput without considering packet latency reduction. For many real-time embedded network applications, it is essential that the processing time not exceed a given threshold. In this paper, we propose LATA, a LAtency and Throughput-Aware packet processing system for multicore architectures. Based on parallel pipeline core topology, LATA can satisfy the latency constraint and produce high throughput by exploiting fine-grained task-level parallelism. We implement LATA on an Intel machine with two Quad-Core Xeon E5335 processors and compare it with four other systems (Parallel, Greedy, Random and Bipar) for six network applications. LATA exhibits an average of 36.5% reduction of latency and a maximum of 62.2% reduction of latency for URL over Random with comparable throughput performance.
Jilong Kuang, Laxmi N. Bhuyan
DAC1
2010 Optimizing Throughput and Latency under Given Power Budget for Network Packet Processing
abstract
Current state-of-the-art task scheduling algorithms for network packet processing schedule the program into a parallel-pipeline topology on network processors to maximize the throughput. However, there has been no existing work targeting power budget for packet processing on off-the-shelf multicore architectures. As energy consumption, reliability and cooling cost for packet processing systems become increasingly important, it is necessary to integrate power-awareness into a scheduler to meet the power budget. In this paper, we propose a novel scheduling algorithm to optimize both throughput and latency given a power budget for network packet processing on multicore architectures. This algorithm addresses power-aware parallel-pipeline scheduling problem by applying per-core DVFS to optimally adjust frequency on each core. We implement our algorithm on an AMD machine with two Quad-Core Opteron 2350 processors and compare the results with existing algorithms given the same power budget. For six real packet processing applications, our algorithm improves throughput and reduces latency by an average of 64.6% and 25.2%, respectively.
Jilong Kuang, Laxmi N. Bhuyan
INFOCOM1