Tauhidur Rahman

dblp:18/10649 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-1981-6395ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 5 since 2021Systems, architecture and hardware · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Computer networks · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals
abstract
Time-series foundation models excel at tasks like forecasting across diverse data types by leveraging informative waveform representations. Wearable sensing data, however, pose unique challenges due to their variability in patterns and frequency bands, especially for healthcare-related outcomes. The main obstacle lies in crafting generalizable representations that adapt efficiently across heterogeneous sensing configurations and applications. To address this, we propose NormWear, the first paradigm multi-modal and ubiquitous framework of foundation model designed to extract generalized and informative representations from wearable sensing data. Specifically, we design a channel-aware attention mechanism with a shared special liaison [CLS] token to detect signal patterns in both intra-sensor and inter-sensors. This helps the model to extract more meaningful information considering both time series themselves and the relationships between input sensors. This helps the model to be widely compatible with various sensors settings. NormWear is pretrained on a diverse set of physiological signals, including PPG, ECG, EEG, GSR, and IMU, from various public datasets. Our model shows exceptional generalizability across 11 public wearable sensing datasets, spanning 18 applications in mental health, body state inference, vital sign estimation, and disease risk evaluation. It consistently outperforms competitive baselines under zero-shot, partial-shot, and full-shot settings, indicating broad applicability in real-world health applications.
Yunfei Luo, Yuliang Chen, Asif Salekin, Tauhidur Rahman
ACM Trans. Comput. Heal.4
2025 Anti-Sensing: Defense Against Unauthorized Radar-Based Human Vital Sign Sensing with Physically Realizable Wearable Oscillators
abstract
Recent advancements in Ultra-Wideband (UWB) radar technology have enabled contactless, non-line-of-sight vital sign monitoring, making it a valuable tool for healthcare. However, UWB radar's ability to capture sensitive physiological data, even through walls, raises significant privacy concerns, particularly in human-robot interactions and autonomous systems that rely on radar for sensing human presence and physiological functions. In this paper, we present Anti-Sensing, a novel defense mechanism designed to prevent unauthorized radarbased sensing. Our approach introduces physically realizable perturbations, such as oscillatory motion from wearable devices, to disrupt radar sensing by mimicking natural cardiac motion, thereby misleading heart rate (HR) estimations. We develop a gradient-based algorithm to optimize the frequency and spatial amplitude of these oscillations for maximal disruption while ensuring physiological plausibility. Through both simulations and real-world experiments with radar data and neural networkbased HR sensing models, we demonstrate the effectiveness of Anti-Sensing in significantly degrading model accuracy, offering a practical solution for privacy preservation.
Md. Farhan Tasnim Oshim, Nigel Doering, Bashima Islam, Tsui-Wei Weng, Tauhidur Rahman
ICRA5
2025 AeroSafe: Mobile Indoor Air Purification Using Aerosol Residence Time Analysis and Robotic Cough Emulator Testbed
abstract
Indoor air quality plays an essential role in the safety and well-being of occupants, especially in the context of airborne diseases. This paper introduces AeroSafe, a novel approach aimed at enhancing the efficacy of indoor air purification systems through a robotic cough emulator testbed and a digital-twins-based aerosol residence time analysis. Current portable air filters often overlook the concentrations of respiratory aerosols generated by coughs, posing a risk, particularly in high-exposure environments like healthcare facilities and public spaces. To address this gap, we present a robotic dual-agent physical emulator comprising a maneuverable mannequin simulating cough events and a portable air purifier autonomously responding to aerosols. The generated data from this emulator trains a digital twins model, combining a physics-based compartment model with a machine learning approach, using Long Short-Term Memory (LSTM) networks and graph convolution layers. Experimental results demonstrate the model's ability to predict aerosol concentration dynamics with a mean residence time prediction error within 35 seconds. The proposed system's real-time intervention strategies outperform static air filter placement, showcasing its potential in mitigating airborne pathogen risks.
M. Tanjid Hasan Tonmoy, Rahath Malladi, Kaustubh Singh, Forsad Al Hossain, Andrés E. Tejada-Martínez, Tauhidur Rahman
ICRA7
2025 Predicting Quality of Video Gaming Experience using Global-Scale Telemetry Data and Federated Learning
abstract
Frames Per Second (FPS) significantly affects the gaming experience. Providing players with accurate FPS estimates prior to purchase benefits both players and game developers. However, we have a limited understanding of how to predict a game's technical performance on a specific device. In this paper, we first conduct a comprehensive analysis of a wide range of factors that may affect game FPS on a global-scale dataset to identify the determinants of FPS. This includes player-side and game-side characteristics, as well as country-level socio-economic statistics. Furthermore, recognizing that accurate FPS predictions require extensive user data, which raises privacy concerns, we propose a federated learning-based model to ensure user privacy. Each player and game is assigned a unique learnable knowledge kernel that gradually extracts latent features for improved accuracy. We also introduce a novel training and prediction scheme that allows these kernels to be dynamically plug-and-play, effectively addressing cold start issues. To train this model with minimal bias, we collected a large telemetry dataset from 224 countries and regions, 100,000 users, and 835 games. Our model achieved a mean Wasserstein distance of 0.469 between predicted and ground truth FPS distributions, outperforming all baseline methods.
Zhongyang Zhang, Jinhe Wen, Zixi Chen 0004, Dara Arbab, Sruti Sahani, Kent Giard, Bijan Arbab, Haojian Jin, Tauhidur Rahman
Proc. ACM Hum. Comput. Interact.9
2024 Pharmacokinetics-Informed Neural Network for Predicting Opioid Administration Moments with Wearable Sensors
abstract
Long-term and high-dose prescription opioid use places individuals at risk for opioid misuse, opioid use disorder (OUD), and overdose. Existing methods for monitoring opioid use and detecting misuse rely on self-reports, which are prone to reporting bias, and toxicology testing, which may be infeasible in outpatient settings. Although wearable technologies for monitoring day-to-day health metrics have gained significant traction in recent years due to their ease of use, flexibility, and advancements in sensor technology, their application within the opioid use space remains underexplored. In the current work, we demonstrate that oral opioid administrations can be detected using physiological signals collected from a wrist sensor. More importantly, we show that models informed by opioid pharmacokinetics increase reliability in predicting the timing of opioid administrations. Forty-two individuals who were prescribed opioids as a part of their medical treatment in-hospital and after discharge were enrolled. Participants wore a wrist sensor throughout the study, while opioid administrations were tracked using electronic medical records and self-reports. We collected 1,983 hours of sensor data containing 187 opioid administrations from the inpatient setting and 927 hours of sensor data containing 40 opioid administrations from the outpatient setting. We demonstrate that a self-supervised pre-trained model, capable of learning the canonical time series of plasma concentration of the drug derived from opioid pharmacokinetics, can reliably detect opioid administration in both settings. Our work suggests the potential of pharmacokinetic-informed, data-driven models to objectively detect opioid use in daily life.
Bhanu Teja Gullapalli, Stephanie Carreiro, Brittany P. Chapman, Eric L. Garland, Tauhidur Rahman
AAAI5
2024 Multi-stakeholder Perspectives on Mental Health Screening Tools for Children
abstract
Pediatric mental health is a growing concern around the world, affecting children’s social-emotional development and increasing the risk of poor behavioral outcomes later in life. However, obtaining a behavioral diagnosis in early childhood is challenging due to lack of access to resources, low parental mental health literacy, and children’s dependence on several stakeholders to coordinate care for them. While app-based, at-home screening tools could offer a scalable and convenient diagnostic solution for families, stakeholder perspectives on their utility and usability remain to be examined. This work reports on a survey of child mental health practitioners and interviews with parents to illustrate existing barriers to care that stakeholders encounter, the perceived benefits of app-based screening tools in meeting their needs, and the challenges in scaling up these tools. We identify where stakeholders agree or disagree, delineate key design tensions, and provide recommendations for the development of future screening technologies.
Manasa Kalanadhabhatta, Adrelys Mateo Santana, Lynnea Mayorga, Tauhidur Rahman, Deepak Ganesan, Adam S. Grabell
CHI4
2024 V2CE: Video to Continuous Events Simulator
abstract
Dynamic Vision Sensor (DVS)-based solutions have recently garnered significant interest across various computer vision tasks, offering notable benefits in terms of dynamic range, temporal resolution, and inference speed. However, as a relatively nascent vision sensor compared to Active Pixel Sensor (APS) devices such as RGB cameras, DVS suffers from a dearth of ample labeled datasets. Prior efforts to convert APS data into events often grapple with issues such as a considerable domain shift from real events, the absence of quantified validation, and layering problems within the time axis. In this paper, we present a novel method for video-to-events stream conversion from multiple perspectives, considering the specific characteristics of DVS. A series of carefully designed losses helps enhance the quality of generated event voxels significantly. We also propose a novel local dynamic-aware timestamp inference strategy to accurately recover event timestamps from event voxels in a continuous fashion and eliminate the temporal layering problem. Results from rigorous validation through quantified metrics at all stages of the pipeline establish our method unquestionably as the current state-of-the-art (SOTA). The code can be found at bit.ly/v2ce.
Zhongyang Zhang, Shuyang Cui, Kaidong Chai, Haowen Yu, Subhasis Dasgupta, Upal Mahbub, Tauhidur Rahman
ICRA7
2024 NeRF-enabled Analysis-Through-Synthesis for ISAR Imaging of Small Everyday Objects with Sparse and Noisy UWB Radar Data
abstract
Inverse Synthetic Aperture Radar (ISAR) imaging presents a formidable challenge when it comes to small everyday objects due to their limited Radar Cross-Section (RCS) and the inherent resolution constraints of radar systems. Existing ISAR reconstruction methods including backprojection (BP) often require complex setups and controlled environments, rendering them impractical for many real-world noisy scenarios. In this paper, we propose a novel Analysis-through-Synthesis (ATS) framework enabled by Neural Radiance Fields (NeRF) for high-resolution coherent ISAR imaging of small objects using sparse and noisy Ultra-Wideband (UWB) radar data with an inexpensive and portable setup. Our end-to-end framework integrates ultra-wideband radar wave propagation, reflection characteristics, and scene priors, enabling efficient 2D scene reconstruction without the need for costly anechoic chambers or complex measurement test beds. With qualitative and quantitative comparisons, we demonstrate that the proposed method outperforms traditional techniques and generates ISAR images of complex scenes with multiple targets and complex structures in Non-Line-of-Sight (NLOS) and noisy scenarios, particularly with limited number of views and sparse UWB radar scans. This work represents a significant step towards practical, cost-effective ISAR imaging of small everyday objects, with broad implications for robotics and mobile sensing applications.
Md. Farhan Tasnim Oshim, Albert W. Reed, Suren Jayasuriya, Tauhidur Rahman
IROS4
2024 Folk Models of Loot Boxes in Video Games
abstract
Regulations require video games to provide transparency regarding loot box odds to keep players informed, leading many games to disclose probabilities in various ways; yet, the extent of players' comprehension of loot box mechanics remains unclear. We performed a content analysis on 80 online posts to understand players' perceptions of loot box odds in two popular video games (Genshin Impact and Honkai: Star Rail). We then conducted semi-structured interviews with 24 players to explore the causes of these folk models across more games. Utilizing a bottom-up open coding approach, we created a taxonomy of folk models players have about loot boxes. We found that participants generally possessed inaccurate mental models of how loot boxes work, and they wanted game companies to enhance loot box transparency in three areas of probability disclosures: granularity, longitude, and scope.
Jinhe Wen, Zhongyang Zhang, Tuan M. Tran, Lianrui Mu, Tauhidur Rahman, Haojian Jin
Proc. ACM Hum. Comput. Interact.5
2023 Detecting PTSD Using Neural and Physiological Signals: Recommendations from a Pilot Study
abstract
Post-traumatic stress disorder (PTSD) is a serious condition that is characterized by negative mood and affect, hyperarousal, irritability, and reactivity, as well as deterioration of cognitive processes such as attention and memory. Timely identification and treatment of PTSD symptoms can significantly improve symptom management and recovery. However, accurate prediction of PTSD outside clinical settings is often challenging. In this work, we investigate whether deficits in cognitive performance can be used to classify individuals with and without PTSD. We further examine whether neural and physiological signals such as prefrontal cortex activity, heart rate, respiration, and electrodermal activity recorded in conjunction with cognitive task performance can be leveraged to improve PTSD classification. Our results indicate that working memory tasks can achieve an F1 score of 0.80 at classifying individuals with PTSD, which can be further improved to 0.91 by combining multimodal information from neurophysiological signals. Based on our findings, we provide recommendations for in-the-wild PTSD classification.
Manasa Kalanadhabhatta, Shaily Roy, Trevor Grant, Asif Salekin, Tauhidur Rahman, Dessa Bergen-Cico
ACII5
2023 Neuromorphic high-frequency 3D dancing pose estimation in dynamic environment
abstract
Technology-mediated dance experiences, as a medium of entertainment, are a key element in both traditional and virtual reality-based gaming platforms. These platforms predominantly depend on unobtrusive and continuous human pose estimation as a means of capturing input. Current solutions primarily employ RGB or RGB-Depth cameras for dance gaming applications; however, the former is hindered by low-light conditions due to motion blur and reduced sensitivity, while the latter exhibits excessive power consumption, diminished frame rates, and restricted operational distance. Boasting ultra-low latency, energy efficiency, and a wide dynamic range, neuromorphic cameras present a viable solution to surmount these limitations. Here, we introduce YeLan, a neuromorphic camera-driven, three-dimensional, high-frequency human pose estimation (HPE) system capable of withstanding low-light environments and dynamic backgrounds. We have compiled the first-ever neuromorphic camera dance HPE dataset and devised a fully adaptable motion-to-event, physics-conscious simulator. YeLan surpasses baseline models under strenuous conditions and exhibits resilience against varying clothing types, background motion, viewing angles, occlusions, and lighting fluctuations.
Zhongyang Zhang, Kaidong Chai, Haowen Yu, Ramzi Majaj, Francesca Walsh, Edward Jay Wang, Upal Mahbub, Hava T. Siegelmann, Donghyun Kim 0002, Tauhidur Rahman
Neurocomputing10
2022 Extracting Multimodal Embeddings via Supervised Contrastive Learning for Psychological Screening
abstract
The diagnosis of psychological disorders in early childhood is of utmost importance given their severe impact on children's academic and social skills as well as general adaptive functioning. Wearable and video-based systems have the potential to collect important diagnostic information in the form of neurophysiological and behavioral signals. However, accurate prediction of psychological disorder status from multimodal data streams necessitates their combination into meaningful features for classification models. In this work, we present a multitask supervised contrastive learning approach to learn useful multimodal embeddings from functional Near-Infrared Spectroscopy, galvanic skin response, and facial video data collected during a frustration-inducing task. The generated embeddings are able to accurately infer emotion regulation-related psychological disorders with an F1 score of 0.91, having significant implications for early-childhood mental health diagnoses.
Manasa Kalanadhabhatta, Adrelys Mateo Santana, Deepak Ganesan, Tauhidur Rahman, Adam S. Grabell
ACII4
2021 Knowledge Transfer across Imaging Modalities Via Simultaneous Learning of Adaptive Autoencoders for High-Fidelity Mobile Robot Vision
abstract
Enabling mobile robots for solving challenging and diverse shape, texture, and motion related tasks with high fidelity vision requires the integration of novel multimodal imaging sensors and advanced fusion techniques. However, it is associated with high cost, power, hardware modification, and computing requirements which limit its scalability. In this paper, we propose a novel Simultaneously Learned Auto Encoder Domain Adaptation (SAEDA)-based transfer learning technique to empower noisy sensing with advanced sensor suite capabilities. In this regard, SAEDA trains both source and target auto-encoders together on a single graph to obtain the domain invariant feature space between the source and target domains on simultaneously collected data. Then, it uses the domain invariant feature space to transfer knowledge between different signal modalities. The evaluation has been done on two collected datasets (LiDAR and Radar) and one existing dataset (LiDAR, Radar and Video) which provides a significant improvement in quadruped robot-based classification (home floor and human activity recognition) and regression (surface roughness estimation) problems. We also integrate our sensor suite and SAEDA framework on two real-time systems (vacuum cleaning and Mini-Cheetah quadruped robots) for studying the feasibility and usability.
Md Mahmudur Rahman 0005, Tauhidur Rahman, Donghyun Kim 0002, Mohammad Arif Ul Alam
IROS2
2020 MechanoBeat: Monitoring Interactions with Everyday Objects using 3D Printed Harmonic Oscillators and Ultra-Wideband Radar
abstract
In this paper we present MechanoBeat, a 3D printed mechanical tag that oscillates at a unique frequency upon user interaction. With the help of an ultra-wideband (UWB) radar array, MechanoBeat can unobtrusively monitor interactions with both stationary and mobile objects. MechanoBeat consists of small, scalable, and easy-to-install tags that do not require any batteries, silicon chips, or electronic components. Tags can be produced using commodity desktop 3D printers with cheap materials. We develop an efficient signal processing and deep learning method to locate and identify tags using only the signals reflected from the tag vibrations. MechanoBeat is capable of detecting simultaneous interactions with high accuracy, even in noisy environments. We leverage UWB radar signals' high penetration property to sense interactions behind walls in a non-line-of-sight (NLOS) scenario. A number of applications using MechanoBeat have been explored and the results have been presented in the paper.
Md. Farhan Tasnim Oshim, Julian Killingback, Dave Follette, Huaishu Peng, Tauhidur Rahman
UIST5
2019 W!NCE: eyewear solution for upper face action units monitoring
abstract
The ability to unobtrusively and continuously monitor one's facial expressions has implications for a variety of application domains ranging from affective computing to health-care and the entertainment industry The standard Facial Action Coding System (FACS) along with camera based methods have been shown to provide objective indicators of facial expressions; however, these approaches can also be fairly limited for mobile applications due to privacy concerns and awkward positioning of the camera. To bridge this gap, W!NCE re-purposes a commercially available Electrooculography-based eyeglass (J!NS MEME) for continuously and unobtrusively sensing of upper facial action units with high fidelity. W!NCE detects facial gestures using a two-stage processing pipeline involving motion artifact removal and facial action detection. We validate our system's applicability through extensive evaluation on data from 17 users under stationary and ambulatory settings.
Soha Rostaminia, Alexander Lamson, Subhransu Maji, Tauhidur Rahman, Deepak Ganesan
ETRA4
2016 Nutrilyzer: A Mobile System for Characterizing Liquid Food with Photoacoustic Effect
abstract
In this paper, we propose Nutrilyzer, a novel mobile sensing system for characterizing the nutrients and detecting adulterants in liquid food with the photoacoustic effect. By listening to the sound of the intensity modulated light or electromagnetic wave with different wavelengths, our mobile photoacoustic sensing system captures unique spectra produced by the transmitted and scattered light while passing through various liquid food. As different liquid foods with different chemical compositions yield uniquely different spectral signatures, Nutrilyzer's signal processing and machine learning algorithm learn to map the photoacoustic signature to various liquid food characteristics including nutrients and adulterants. We evaluated Nutrilyzer for milk nutrient prediction (i.e., milk protein) and milk adulterant detection. We have also explored Nutrilyzer for alcohol concentration prediction. The Nutrilyzer mobile system consists of an array of 16 LEDs in ultraviolet, visible and near-infrared region, two piezoelectric sensors and an ARM microcontroller unit, which are designed and fabricated in a printed circuit board and a 3D printed photoacoustic housing.
Tauhidur Rahman, Alexander Travis Adams, Perry Schein, Aadhar Jain, David Erickson, Tanzeem Choudhury
SenSys1
2015 DoppleSleep: a contactless unobtrusive sleep sensing system using short-range Doppler radar
abstract
In this paper, we present DoppleSleep -- a contactless sleep sensing system that continuously and unobtrusively tracks sleep quality using commercial off-the-shelf radar modules. DoppleSleep provides a single sensor solution to track sleep-related physical and physiological variables including coarse body movements and subtle and fine-grained chest, heart movements due to breathing and heartbeat. By integrating vital signals and body movement sensing, DoppleSleep achieves 89.6% recall with Sleep vs. Wake classification and 80.2% recall with REM vs. Non-REM classification compared to EEG-based sleep sensing. Lastly, it provides several objective sleep quality measurements including sleep onset latency, number of awakenings, and sleep efficiency. The contactless nature of DoppleSleep obviates the need to instrument the user's body with sensors. Lastly, DoppleSleep is implemented on an ARM microcontroller and a smartphone application that are benchmarked in terms of power and resource usage.
Tauhidur Rahman, Alexander Travis Adams, Ruth Vinisha, Mi Zhang 0002, Shwetak N. Patel, Julie A. Kientz, Tanzeem Choudhury
UbiComp1
2014 BodyBeat: a mobile system for sensing non-speech body sounds
abstract
In this paper, we propose BodyBeat, a novel mobile sensing system for capturing and recognizing a diverse range of non-speech body sounds in real-life scenarios. Non-speech body sounds, such as sounds of food intake, breath, laughter, and cough contain invaluable information about our dietary behavior, respiratory physiology, and affect. The BodyBeat mobile sensing system consists of a custom-built piezoelectric microphone and a distributed computational framework that utilizes an ARM microcontroller and an Android smartphone. The custom-built microphone is designed to capture subtle body vibrations directly from the body surface without being perturbed by external sounds. The microphone is attached to a 3D printed neckpiece with a suspension mechanism. The ARM embedded system and the Android smartphone process the acoustic signal from the microphone and identify non-speech body sounds. We have extensively evaluated the BodyBeat mobile sensing system. Our results show that BodyBeat outperforms other existing solutions in capturing and recognizing different types of important non-speech body sounds.
Tauhidur Rahman, Alexander Travis Adams, Mi Zhang 0002, Erin Cherry, Bobby Zhou, Huaishu Peng, Tanzeem Choudhury
MobiSys1
2012 A personalized emotion recognition system using an unsupervised feature adaptation scheme
abstract
A personalized emotion recognition system aims to tune the model to recognize the expressive behaviors of a targeted person. Such a system can play an important role in various domains including call center and health care applications. Adapting any general emotion recognition system for a particular individual requires speech samples and prior knowledge about their emotional content. These assumptions constrain the use of these techniques in many real scenarios in which no annotated data is available to train or adapt the models. To address this problem, this paper introduces an unsupervised feature adaptation scheme that aims to reduce the mismatch between the acoustic features used to train the system and the acoustic features extracted from the unknown targeted speaker. The adaptation scheme uses our recently proposed iterative feature normalization (IFN) framework. An emotion detection system is trained with the IEMOCAP database. For testing, a database was created by downloading videos from a video-sharing website, containing various interviews from a targeted subject (1.5 hours). The detection system is used to identify emotional speech with and without the proposed feature adaptation scheme. The experimental results indicate that the proposed approach improves the unweighted accuracy from 50.8% to 70.0%.
Tauhidur Rahman, Carlos Busso
ICASSP1
2012 Indoor robotic terrain classification via angular velocity based hierarchical classifier selection
abstract
This paper proposes a novel approach to terrain classification by wheeled mobile robots, which utilizes vibration data. In our proposed approach, a mobile robot has the ability to categorize terrain types simply by driving over them. Classification of terrain is based on measurements obtained from an inertial measurement unit strapped directly to the robot's chassis. In contrast to the previous approaches, we use acceleration and angular velocity measurements in all cardinal directions to extract over 800 features. Sequential Forward Floating Feature Selection is used to narrow down this large group of features to a set of 15 to 20 that are the most useful. The reduced set of features is used by a Linear Bayes Normal Classifier to classify terrain. Furthermore, different feature sets are generated for different velocity conditions, and the classifier switches based on the current robot velocity. Experimental results are presented that show the strong performance of the proposed system, including 90% accuracy over 20 continuous minutes of driving across different terrains.
David Tick, Tauhidur Rahman, Carlos Busso, Nicholas R. Gans
ICRA2
2012 Unveiling the Acoustic Properties that Describe the Valence Dimension
Carlos Busso, Tauhidur Rahman
INTERSPEECH2
2011 Detecting Sleepiness by Fusing Classifiers Trained with Novel Acoustic Features
abstract
Automatic sleepiness detection is a challenging task that can lead to advances in various domains including traffic safety, medicine and human-machine interaction. This paper analyzes the discriminative power of different acoustic features to detect sleepiness. The study uses the sleepy language corpus (SLC). Along with standard acoustic features, novel features are proposed including functionals across voiced segment statistics in the F0 contour, likelihoods of reference models used to contrast non-neutral speech, and a set of robust to noise spectral features. These feature sets, which have performed well in other paralinguistic tasks such as emotion recognition, are used to train classifiers that are combined at the feature and decision levels. The best unweighted accuracy (UA) is obtained by combining the classifiers at the decision level under a maximum likelihood framework (UA = 70.97%). This performance is higher than the best results reported in the corpus. Index Terms: Speaker State Recognition, Paralinguistics, Affective Computing, Sleepiness
Tauhidur Rahman, Soroosh Mariooryad, Shalini Keshavamurthy, Gang Liu 0001, John H. L. Hansen, Carlos Busso
INTERSPEECH1