Jagmohan Chauhan

dblp:79/11198 · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0003-2080-3276ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 9 since 2021Computer networks · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Short Paper: The Starlink Robot: A Platform and Dataset for Mobile Satellite Communication
abstract
The integration of satellite communication into mobile devices represents a paradigm shift in connectivity, yet the performance characteristics under motion and environmental occlusion remain poorly understood. We present the Starlink Robot, the first mobile robotic platform equipped with Starlink satellite internet, comprehensive sensor suite including upward-facing camera, LiDAR, and IMU, designed to systematically study satellite communication performance during movement. Our multi-modal dataset captures synchronized communication metrics, motion dynamics, sky visibility, and 3D environmental context across diverse scenarios including steady-state motion, variable speeds, and different occlusion conditions. This platform and dataset enable researchers to develop motion-aware communication protocols, predict connectivity disruptions, and optimize satellite communication for emerging mobile applications from smartphones to autonomous vehicles. In this work, we use LEOViz for real-time data collection and visualization. The project is available at https://starlinkrobot.github.io.
Boyi Liu 0003, Qianyi Zhang, Qiang Yang 0018, Jianhao Jiao, Jagmohan Chauhan, Dimitrios Kanoulas
SenSys5
2026 PPGSpeech: A Wearable Silent Speech Interface Leveraging Neck-Worn Photoplethysmography
abstract
Silent speech interfaces (SSIs) promise private and noise-immune communication, but current solutions often sacrifice user comfort, mobility, or privacy. This paper introduces PPGSpeech, a novel SSI that overcomes these limitations by pioneering the use of photoplethysmography (PPG) acquired from a comfortable, necklace-style wearable device. Our core discovery is that subtle neck muscle movements during silent articulation induce distinct, measurable modulations in the underlying PPG signal. To harness this phenomenon, we developed a complete end-to-end system featuring (1) a custom neck-worn sensor for multi-wavelength PPG acquisition, (2) a deep learning pipeline that converts 1D PPG signals into 2D time-frequency images via Continuous Wavelet Transform (CWT) and classifies them using a lightweight CNN, and (3) a Pix2Pix GAN model to reconstruct audible speech from the captured signals. In a 16-participant study covering a vocabulary of 15 commands and four confounding actions, our user-dependent model achieved a recognition accuracy of 81.41% ± 9.74. Furthermore, our speech reconstruction achieved a Mean Opinion Score (MOS) of 3.48 and a Word Correct Rate (WCR) of 60.67%, demonstrating that the PPG signal is sufficiently rich to recover intelligible speech. By establishing the viability of neck-based PPG for silent speech, PPGSpeech offers a discreet, privacy-preserving, and continuously wearable paradigm for next-generation human-computer interaction.
Lingde Hu, Yu He 0028, Seokmin Choi, Yang Gao 0025, Jagmohan Chauhan, Zhanpeng Jin
IEEE Internet Things J.7
2026 Parallax: Runtime Parallelization for Operator Fallbacks in Heterogeneous Edge Systems
abstract
The growing demand for real-time DNN applications on edge devices necessitates faster inference of increasingly complex models. Although many devices include specialized accelerators (e.g., mobile GPUs), dynamic control-flow operators and unsupported kernels often fall back to CPU execution. Existing frameworks handle these fallbacks poorly, leaving CPU cores idle and causing high latency and memory spikes. We introduce Parallax, a framework that accelerates mobile DNN inference without model refactoring or custom operator implementations. Parallax first partitions the computation DAG to expose parallelism, then employs branch-aware memory management with dedicated arenas and buffer reuse to reduce runtime footprint. An adaptive scheduler executes branches according to device memory constraints; meanwhile, fine-grained subgraph control enables heterogeneous inference of dynamic models. By evaluating on representative DNNs across different mobile devices, Parallax achieves up to 46% latency reduction, maintains controlled memory overhead (26.5%), and delivers up to 30% energy savings compared with the state-of-the-art frameworks, offering improvements aligned with the responsiveness demands of real-time mobile inference.
Chong Tang 0006, Jagmohan Chauhan
IEEE Internet Things J.3
2026 MobileROS: A Wireless-Native Robot Operating System for Mobile Robotics
abstract
The increasing deployment of mobile robots in dynamic outdoor environments necessitates robotic systems capable of maintaining reliability amidst fluctuating wireless connectivity. While the Robot Operating System (ROS) has established itself as the de facto standard for such networked robotics, its abstraction of communication as an opaque, besteffort utility creates a critical bottleneck: it fails to leverage physical layer (PHY) information, resulting in degraded performance and unreliable execution in fluctuating networks. To address this, this paper presents MobileROS, a wireless-native robot operating system that transforms wireless communication from an external service into a core system resource. Grounded in the Symbiotic Paradigm, MobileROS establishes a bidirectional exchange where network conditions inform robotic decisions and mission requirements guide network resource allocation. Based on service mesh principles and domain-driven design, our architecture implements a Hub-Engines-Cells (HEC) model. It features a central Hub for global optimization, three specialized engines (the Radio Information Engine, the Cross Domain Engine, and the Physical Adaptive Engine) for crosslayer intelligence, and distributed Cells as functional units. A key mechanism, Application-Driven Bidirectional Dynamic Slicing, allows robots to actively reconfigure network resources based on semantic urgency, transforming the robot from a passive observer into an active network controller. We systematically evaluate MobileROS across three cities (London, Hong Kong, and Shenzhen) in five scenarios: distributed visual SLAM, cross-domain LiDAR perception, V2X autonomous driving, hybrid multi-robot collaboration against WebRTC baselines, and partition recovery validating CAP-theorem-aware failsafe mechanisms. Results demonstrate that MobileROS maintains significantly more stable performance than standard ROS in mobile wireless deployments.We provide implementation details athttps://github.com/MobileROS.
Boyi Liu 0003, Qianyi Zhang, Yongguang Lu, Jianhao Jiao, Jagmohan Chauhan, Wen Wu 0003, Jun Zhang 0004, Dimitrios Kanoulas
IEEE Trans. Robotics5
2025 Power- and Deadline-Aware Dynamic Inference on Intermittent Computing Systems
abstract
In energy-harvesting intermittent computing systems, balancing power constraints with the need for timely and accurate inference remains a critical challenge. Existing methods often sacrifice significant accuracy or fail to adapt effectively to fluctuating power conditions. This paper presents DualAdaptNet, a power- and deadline-aware neural network architecture that dynamically adapts both its width and depth to ensure reliable inference under variable power conditions. Additionally, a runtime scheduling method is introduced to select an appropriate sub-network configuration based on real-time energy-harvesting conditions and system deadlines. Experimental results on the MNIST dataset demonstrate that our approach completes up to 7.0% more inference tasks within a specified deadline, while also improving average accuracy by 15.4% compared to the state-of-the-art.
Hengrui Zhao, Lei Xun, Jagmohan Chauhan, Geoff V. Merrett
DATE3
2025 FedTMOS: Efficient One-Shot Federated Learning with Tsetlin Machine
abstract
One-Shot Federated Learning (OFL) is a promising approach that reduce communication to a single round, minimizing latency and resource consumption. However, existing OFL methods often rely on Knowledge Distillation, which introduce server-side training, increasing latency. While neuron matching and model fusion techniques bypass server-side training, they struggle with alignment when heterogeneous data is present. To address these challenges, we proposed One-Shot Federated Learning with Tsetlin Machine (FedTMOS), a novel data-free OFL framework built upon the low-complexity and class-adaptive properties of the Tsetlin Machine. FedTMOS first clusters then reassigns class-specific weights to form models using an inter-class maximization approach, efficiently generating balanced server models without requiring additional training. Our extensive experiments demonstrate that FedTMOS significantly outperforms its ensemble counterpart by an average of $6.16$%, and the leading state-of-the-art OFL baselines by $7.22$% across various OFL settings. Moreover, FedTMOS achieves at least a $2.3\times$ reduction in upload communication costs and a $75\times$ reduction in server latency compared to methods requiring server-side training. These results establish FedTMOS as a highly efficient and practical solution for OFL scenarios.
Shannon How Shi Qi, Jagmohan Chauhan, Geoff V. Merrett, Jonathon S. Hare
ICLR2
2025 Continual Generalized Category Discovery: Learning and Forgetting from a Bayesian Perspective
abstract
Continual Generalized Category Discovery (C-GCD) faces a critical challenge: incrementally learning new classes from unlabeled data streams while preserving knowledge of old classes. Existing methods struggle with catastrophic forgetting, especially when unlabeled data mixes known and novel categories. We address this by analyzing C-GCD’s forgetting dynamics through a Bayesian lens, revealing that covariance misalignment between old and new classes drives performance degradation. Building on this insight, we propose Variational Bayes C-GCD (VB-CGCD), a novel framework that integrates variational inference with covariance-aware nearest-class-mean classification. VB-CGCD adaptively aligns class distributions while suppressing pseudo-label noise via stochastic variational updates. Experiments show VB-CGCD surpasses prior art by +15.21% with the overall accuracy in the final session on standard benchmarks. We also introduce a new challenging benchmark with only 10% labeled data and extended online phases—VB-CGCD achieves a 67.86% final accuracy, significantly higher than state-of-the-art (38.55%), demonstrating its robust applicability across diverse scenarios. Code is available at: https://github.com/daihao42/VB-CGCD
Jagmohan Chauhan
ICML2
2025 PhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth Observation
abstract
Recent progress on Remote Sensing Foundation Models (RSFMs) aims toward universal representations for Earth observation imagery. However, current efforts often scale up in size significantly without addressing efficiency constraints critical for real-world applications (e.g., onboard processing, rapid disaster response) or treat multispectral (MS) data as generic imagery, overlooking valuable physical priors. We introduce PhySwin, a foundation model for MS data that integrates physical priors with computational efficiency. PhySwin combines three innovations: (i) physics-informed pretraining objectives leveraging radiometric constraints to enhance feature learning; (ii) an efficient MixMAE formulation tailored to SwinV2 for low-FLOP, scalable pretraining; and (iii) token-efficient spectral embedding to retain spectral detail without increasing token counts. Pretrained on over 1M Sentinel-2 tiles, PhySwin achieves SOTA results (+1.32\% mIoU segmentation, +0.80\% F1 change detection) while reducing inference latency by up to 14.4$\times$ and computational complexity by up to 43.6$\times$ compared to ViT-based RSFMs.
Chong Tang 0006, Joseph Powell, Dirk Koch, Robert Mullins 0001, Alex S. Weddell, Jagmohan Chauhan
NeurIPS6
2024 AdaTM: Logic Inspired Adaptive Tsetlin Machines for Efficient and Effective Continual Learning on the Edge
Chong Tang 0006, Neelam Singh, Jagmohan Chauhan
EWSN3
2024 TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce Edge
abstract
On-device training is essential for user personalisation and privacy. With the pervasiveness of IoT devices and microcontroller units (MCUs), this task becomes more challenging due to the constrained memory and compute resources, and the limited availability of labelled user data. Nonetheless, prior works neglect the data scarcity issue, require excessively long training time ($\textit{e.g.}$ a few hours), or induce substantial accuracy loss ($\geq$10%). In this paper, we propose TinyTrain, an on-device training approach that drastically reduces training time by selectively updating parts of the model and explicitly coping with data scarcity. TinyTrain introduces a task-adaptive sparse-update method that $\textit{dynamically}$ selects the layer/channel to update based on a multi-objective criterion that jointly captures user data, the memory, and the compute capabilities of the target device, leading to high accuracy on unseen tasks with reduced computation and memory footprint. TinyTrain outperforms vanilla fine-tuning of the entire network by 3.6-5.0% in accuracy, while reducing the backward-pass memory and computation cost by up to 1,098$\times$ and 7.68$\times$, respectively. Targeting broadly used real-world edge devices, TinyTrain achieves 9.5$\times$ faster and 3.5$\times$ more energy-efficient training over status-quo approaches, and 2.23$\times$ smaller memory footprint than SOTA methods, while remaining within the 1 MB memory envelope of MCU-grade platforms.
Young D. Kwon, Rui Li 0052, Stylianos I. Venieris, Jagmohan Chauhan, Nicholas D. Lane, Cecilia Mascolo
ICML4
2024 Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking
abstract
Respiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Generalizable respiratory acoustic foundation models pretrained with unlabeled data would offer appealing advantages and possibly unlock this impasse. However, given the safety-critical nature of healthcare applications, it is pivotal to also ensure openness and replicability for any proposed foundation model solution. To this end, we introduce OPERA, an OPEn Respiratory Acoustic foundation model pretraining and benchmarking system, as the first approach answering this need. We curate large-scale respiratory audio datasets ($\sim$136K samples, over 400 hours), pretrain three pioneering foundation models, and build a benchmark consisting of 19 downstream respiratory health tasks for evaluation. Our pretrained models demonstrate superior performance (against existing acoustic models pretrained with general audio on 16 out of 19 tasks) and generalizability (to unseen datasets and new respiratory audio modalities). This highlights the great promise of respiratory acoustic foundation models and encourages more studies using OPERA as an open resource to accelerate research on respiratory audio for health. The system is accessible from https://github.com/evelyn0414/OPERA.
Yuwei Zhang 0001, Tong Xia, Jing Han 0010, Yu Wu 0021, Georgios Rizos, Yang Liu 0101, Mohammed Mosuily, Jagmohan Chauhan, Cecilia Mascolo
NeurIPS8
2023 MMLung: Moving Closer to Practical Lung Health Estimation using Smartphones
abstract
Long-term respiratory illnesses like Chronic Obstructive Pulmonary Disease (COPD) and Asthma are commonly diagnosed with the gold standard spirometry, which is a lung health test that requires specialized equipment and trained healthcare experts, making it expensive and difficult to scale. Moreover, blowing into a spirometer can be quite hard for people suffering from pulmonary illnesses. To solve the aforementioned limitations, we introduce MMLung, an approach that leverages information obtained from multiple audio signals by combining multiple tasks and different modalities performed on the microphone of a smartphone to estimate lung function. Our proposed approach achieves the best mean absolute percentage error (MAPE) of 1.3% on a cohort of 40 participants. Compared to the reported performances (5%-10% MAPE) on lung health estimation using smartphones, MMLung shows that practical lung health estimation is viable by combining multiple tasks utilizing multiple modalities.
Mohammed Mosuily, Lindsay Welch, Jagmohan Chauhan
INTERSPEECH3
2023 Conditional Neural ODE Processes for Individual Disease Progression Forecasting: A Case Study on COVID-19
abstract
Time series forecasting, as one of the fundamental machine learning areas, has attracted tremendous attentions over recent years. The solutions have evolved from statistical machine learning (ML) methods to deep learning techniques. One emerging sub-field of time series forecasting is individual disease progression forecasting, e.g., predicting individuals' disease development over a few days (e.g., deteriorating trends, recovery speed) based on few past observations. Despite the promises in the existing ML techniques, a variety of unique challenges emerge for disease progression forecasting, such as irregularly-sampled time series, data sparsity, and individual heterogeneity in disease progression. To tackle these challenges, we propose novel Conditional Neural Ordinary Differential Equations Processes (CNDPs), and validate it in a COVID-19 disease progression forecasting task using audio data. CNDPs allow for irregularly-sampled time series modelling, enable accurate forecasting with sparse past observations, and achieve individual-level progression forecasting. CNDPs show strong performance with an Unweighted Average Recall (UAR) of 78.1%, outperforming a variety of commonly used Recurrent Neural Networks based models. With the proposed label-enhancing mechanism (i.e., including the initial health status as input) and the customised individual-level loss, CNDPs further boost the performance reaching a UAR of 93.6%. Additional analysis also reveals the model's capability in tracking individual-specific recovery trend, implying the potential usage of the model for remote disease progression monitoring. In general, CNDPs pave new pathways for time series forecasting, and provide considerable advantages for disease progression monitoring.
Ting Dang, Jing Han 0010, Tong Xia, Erika Bondareva, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Dimitris Spathis, Pietro Cicuta, Cecilia Mascolo
KDD6
2023 LifeLearner: Hardware-Aware Meta Continual Learning System for Embedded Computing Platforms
abstract
Continual Learning (CL) allows applications such as user personalization and household robots to learn on the fly and adapt to context. This is an important feature when context, actions, and users change. However, enabling CL on resource-constrained embedded systems is challenging due to the limited labeled data, memory, and computing capacity.
Young D. Kwon, Jagmohan Chauhan, Hong Jia, Stylianos I. Venieris, Cecilia Mascolo
SenSys2
2022 YONO: Modeling Multiple Heterogeneous Neural Networks on Microcontrollers
abstract
Internet of Things (IoT) systems provide large amounts of data on all aspects of human behavior. Machine learning techniques, especially deep neural networks (DNN), have shown promise in making sense of this data at a large scale. Also, the research community has worked to reduce the computational and resource demands of DNN to compute on low-resourced micro controllers (MCUs). However, most of the current work in embedded deep learning focuses on solving a single task efficiently, while the multi-tasking nature and applications of IoT devices demand systems that can handle a diverse range of tasks (such as activity, gesture, voice, and context recognition) with input from a variety of sensors, simultaneously. In this paper, we propose YONO, a product quantization (PQ) based approach that compresses multiple heterogeneous models and enables in-memory model execution and model switching for dissimilar multi-task learning on MCUs. We first adopt PQ to learn codebooks that store weights of different models. Also, we propose a novel network optimization and heuristics to maximize the com-pression rate and minimize the accuracy loss. Then, we develop an online component of YONO for efficient model execution and switching between multiple tasks on an MCU at run time without relying on an external storage device. YONO shows remarkable performance as it can compress multiple heterogeneous models with negligible or no loss of accuracy up to 12.37x. Furthermore, YONO's online component enables an efficient execution (latency of 16–159 ms and energy consumption of 3.8-37.9 mJ per operation) and reduces modelloading/switching la-tency and energy consumption by 93.3-94.5% and 93.9-95.0%, respectively, compared to external storage access. Interestingly, YONO can compress various architectures trained with datasets that were not shown during YONO's offline codebook learning phase showing the generalizability of our method. To summarize, YONO shows great potential and opens further doors to enable multi-task learning systems on extremely resource-constrained devices.
Young D. Kwon, Jagmohan Chauhan, Cecilia Mascolo
IPSN2
2022 PassWalk: Spatial Authentication Leveraging Lateral Shift and Gaze on Mobile Headsets
abstract
Secure and usable user authentication on mobile headsets is a challenging problem. The miniature-sized touchpad on such devices becomes a hurdle to user interactions that impact usability. However, the most common authentication methods, i.e., the standard QWERTY virtual keyboard or mid-air inputs to enter passwords are highly vulnerable to shoulder surfing attacks. In this paper, we present PassWalk, a keyboard-less authentication system leveraging multi-modal inputs on mobile headsets. PassWalk demonstrates the feasibility of user authentication driven by the user's gaze and lateral shifts (i.e., footsteps) simultaneously. The keyboard-less authentication interface in PassWalk enables users to accomplish highly mobile inputs of graphical passwords, containing digital overlays and physical objects. We conduct an evaluation with 22 recruited participants (15 legitimate users and 7 attackers). Our results show that PassWalk provides high security (only 1.1% observation attacks were successful) with a mean authentication time of 8.028s, which outperforms the commercial method of using the QWERTY virtual keyboard (21.5% successful attacks) and a research prototype LookUnLock (5.5% successful attacks). Additionally, PassWalk entails a significantly smaller workload on the user than the current commercial methods.
Abhishek Kumar 0011, Lik-Hang Lee, Jagmohan Chauhan, Xiang Su 0001, Mohammad Ashraful Hoque, Susanna Pirttikangas, Sasu Tarkoma, Pan Hui 0001
ACM Multimedia3
2021 Exploring Automatic COVID-19 Diagnosis via Voice and Symptoms from Crowdsourced Data
abstract
The development of fast and accurate screening tools, which could facilitate testing and prevent more costly clinical tests, is key to the current pandemic of COVID-19. In this context, some initial work shows promise in detecting diagnostic signals of COVID-19 from audio sounds. In this paper, we propose a voice-based framework to automatically detect individuals who have tested positive for COVID-19. We evaluate the performance of the proposed framework on a subset of data crowdsourced from our app, containing 828 samples from 343 participants. By combining voice signals and reported symptoms, an AUC of 0.79 has been attained, with a sensitivity of 0.68 and a specificity of 0.82. We hope that this study opens the door to rapid, low-cost, and convenient pre-screening tools to automatically detect the disease.
Jing Han 0010, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Cecilia Mascolo
ICASSP3
2021 Exploring System Performance of Continual Learning for Mobile and Embedded Sensing Applications
abstract
Continual learning approaches help deep neural network models adapt and learn incrementally by trying to solve catastrophic forgetting. However, whether these existing approaches, applied traditionally to image-based tasks, work with the same efficacy to the sequential time series data generated by mobile or embedded sensing systems remains an unanswered question. To address this void, we conduct the first comprehensive empirical study that quantifies the performance of three predominant continual learning schemes (i.e., regularization, replay, and replay with examples) on six datasets from three mobile and embedded sensing applications in a range of scenarios having different learning complexities. More specifically, we implement an end-to-end continual learning framework on edge devices. Then we investigate the generalizability, trade-offs between performance, storage, computational costs, and memory footprint of different continual learning methods. Our findings suggest that replay with exemplars-based schemes such as iCaRL has the best performance trade-offs, even in complex scenarios, at the expense of some storage space (few MBs) for training examples (1% to 5%). We also demonstrate for the first time that it is feasible and practical to run continual learning on-device with a limited memory budget. In particular, the latency on two types of mobile and embedded devices suggests that both incremental learning time (few seconds - 4 minutes) and training time (1 - 75 minutes) across datasets are acceptable, as training could happen on the device when the embedded device is charging thereby ensuring complete data privacy. Finally, we present some guidelines for practitioners who want to apply a continual learning paradigm for mobile sensing tasks.
Young D. Kwon, Jagmohan Chauhan, Abhishek Kumar 0011, Pan Hui 0001, Cecilia Mascolo
SEC2
2021 The Benefit of the Doubt: Uncertainty Aware Sensing for Edge Computing Platforms
Lorena Qendro, Jagmohan Chauhan, Alberto Gil C. P. Ramos, Cecilia Mascolo
SEC2
2021 FastICARL: Fast Incremental Classifier and Representation Learning with Efficient Budget Allocation in Audio Sensing Applications
abstract
Various incremental learning (IL) approaches have been proposed to help deep learning models learn new tasks/classes continuously without forgetting what was learned previously (i.e., avoid catastrophic forgetting). With the growing number of deployed audio sensing applications that need to dynamically incorporate new tasks and changing input distribution from users, the ability of IL on-device becomes essential for both efficiency and user privacy. However, prior works suffer from high computational costs and storage demands which hinders the deployment of IL on-device. In this work, to overcome these limitations, we develop an end-to-end and on-device IL framework, FastICARL, that incorporates an exemplar-based IL and quantization in the context of audio-based applications. We first employ k-nearest-neighbor to reduce the latency of IL. Then, we jointly utilize a quantization technique to decrease the storage requirements of IL. We implement FastICARL on two types of mobile devices and demonstrate that FastICARL remarkably decreases the IL time up to 78-92% and the storage requirements by 2-4 times without sacrificing its performance. FastICARL enables complete on-device IL, ensuring user privacy as the user data does not need to leave the device.
Young D. Kwon, Jagmohan Chauhan, Cecilia Mascolo
Interspeech2
2021 The INTERSPEECH 2021 Computational Paralinguistics Challenge: COVID-19 Cough, COVID-19 Speech, Escalation & Primates
abstract
The INTERSPEECH 2021 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the COVID-19 Cough and COVID-19 Speech Sub-Challenges, a binary classification on COVID-19 infection has to be made based on coughing sounds and speech; in the Escalation SubChallenge, a three-way assessment of the level of escalation in a dialogue is featured; and in the Primates Sub-Challenge, four species vs background need to be classified. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' COMPARE and BoAW features as well as deep unsupervised representation learning using the AuDeep toolkit, and deep feature extraction from pre-trained CNNs using the Deep Spectrum toolkit; in addition, we add deep end-to-end sequential modelling, and partially linguistic analysis.
Björn W. Schuller, Anton Batliner, Christian Bergler, Cecilia Mascolo, Jing Han 0010, Iulia Lefter, Heysem Kaya, Shahin Amiriparian, Alice Baird, Lukas Stappen, Sandra Ottl, Maurice Gerczuk, Panagiotis Tzirakis, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Léon J. M. Rothkrantz, Joeri A. Zwerts, Jelle Treep, Casper S. Kaandorp
Interspeech15
2020 Exploring Automatic Diagnosis of COVID-19 from Crowdsourced Respiratory Sound Data
abstract
Audio signals generated by the human body (e.g., sighs, breathing, heart, digestion, vibration sounds) have routinely been used by clinicians as indicators to diagnose disease or assess disease progression. Until recently, such signals were usually collected through manual auscultation at scheduled visits. Research has now started to use digital technology to gather bodily sounds (e.g., from digital stethoscopes) for cardiovascular or respiratory examination, which could then be used for automatic analysis. Some initial work shows promise in detecting diagnostic signals of COVID-19 from voice and coughs. In this paper we describe our data analysis over a large-scale crowdsourced dataset of respiratory sounds collected to aid diagnosis of COVID-19. We use coughs and breathing to understand how discernible COVID-19 sounds are from those in asthma or healthy controls. Our results show that even a simple binary machine learning classifier is able to classify correctly healthy and COVID-19 sounds. We also show how we distinguish a user who tested positive for COVID-19 and has a cough from a healthy user with a cough, and users who tested positive for COVID-19 and have a cough from users with asthma and a cough. Our models achieve an AUC of above 80% across all tasks. These results are preliminary and only scratch the surface of the potential of this type of data and audio-based machine learning. This work opens the door to further investigation of how automatically analysed respiratory patterns could be used as pre-screening signals to aid COVID-19 diagnosis.
Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Jing Han 0010, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Cecilia Mascolo
KDD2
2017 Are Wearables Ready for HTTPS? On the Potential of Direct Secure Communication on Wearables
abstract
The majority of available wearable computing devices require communication with Internet servers for data analysis and storage, and rely on a paired smartphone to enable secure communication. However, many wearables are equipped with WiFi network interfaces, enabling direct communication with the Internet. Secure communication protocols could then run on these wearables themselves, yet it is not clear if they can be efficiently supported.,,,,In this paper, we show that wearables are ready for direct and secure Internet communication by means of experiments with both controlled local web servers and Internet servers. We observe that the overall energy consumption and communication delay can be reduced with direct Internet connection via WiFi from wearables compared to using smartphones as relays via Bluetooth. We also show that the additional HTTPS cost caused by TLS handshake and encryption is closely related to the number of parallel connections, and has the same relative impact on wearables and smartphones.
Harini Kolamunna, Jagmohan Chauhan, Yining Hu 0001, Kanchana Thilakarathna, Diego Perino, Dwight J. Makaroff, Aruna Seneviratne
LCN2
2017 BreathPrint: Breathing Acoustics-based User Authentication
abstract
We propose BreathPrint, a new behavioural biometric signature based on audio features derived from an individual's commonplace breathing gestures. Specifically, BreathPrint uses the audio signatures associated with the three individual gestures: sniff, normal, and deep breathing, which are sufficiently different across individuals. Using these three breathing gestures, we develop the processing pipeline that identifies users via the microphone sensor on smartphones and wearable devices. In BreathPrint, a user performs breathing gestures while holding the device very close to their nose. Using off-the-shelf hardware, we experimentally evaluate the BreathPrint prototype with 10 users, observed over seven days. We show that users can be authenticated reliably with an accuracy of over 94% for all the three breathing gestures in intra-sessions and deep breathing gesture provides the best overall balance between true positives (successful authentication) and false positives (resiliency to directed impersonation and replay attacks). Moreover, we show that this breathing sound based biometric is also robust to some typical changes in both physiological and environmental context, and that it can be applied on multiple smartphone platforms. Early results suggest that breathing based biometrics show promise as either to be used as a secondary authentication modality in a multimodal biometric authentication system or as a user disambiguation technique for some daily lifestyle scenarios.
Jagmohan Chauhan, Yining Hu 0001, Suranga Seneviratne, Archan Misra, Aruna Seneviratne, Youngki Lee 0001
MobiSys1
2016 Gesture-Based Continuous Authentication for Wearable Devices: The Smart Glasses Use Case
Jagmohan Chauhan, Hassan Jameel Asghar, Anirban Mahanti, Mohamed Ali Kâafar
ACNS1
2011 VM clock synchronization measurements
abstract
A modern twist to clock synchronization is the push towards server virtualization and cloud computing. Distributed applications now often run on geographically-distributed, virtualized servers that are isolated from the underlying hardware. Despite the fact that most applications run well in an abstracted environment without direct hardware access, tight real-time synchronized clocks are still necessary. We performed a set of experiments to quantify the additional clock error when running the NTP daemon inside a Virtual Machine. We found that the delay asymmetry is relatively small when running in a non-virtualized environment, but quite significant in a virtualized environment.
Jagmohan Chauhan, Dwight J. Makaroff, Anthony J. Arkles
IPCCC1