Ting Dang

dblp:170/5330 · DBLP profile ↗
← Back
36ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0003-3806-1493ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 From Cheap to Chic: Enhancing Music Playback Quality of Budget Earphones via Hardware-Aware Learning
abstract
Low-end earphones are widely spread due to their affordability, but their limited speaker hardware often leads to poor music playback quality. This raises a key question: Can we compensate for hardware limitations to enhance the listening experience without modifying the device? Existing EQ-based approaches attempt this, but they rely on frequency response curves (FRCs) measured under ideal conditions, which fail to capture real-world distortions such as harmonic and intermodulation effects.
Changshuo Hu, Hung Manh Pham, Ting Dang, Jiannan Li, Rajesh Krishna Balan, Dong Ma 0001
SenSys3
2025 AER-LLM: Ambiguity-aware Emotion Recognition Leveraging Large Language Models
abstract
Recent advancements in Large Language Models (LLMs) have demonstrated great success in many Natural Language Processing (NLP) tasks. In addition to their cognitive intelligence, exploring their capabilities in emotional intelligence is also crucial, as it enables more natural and empathetic conversational AI. Recent studies have shown LLMs’ capability in recognizing emotions, but they often focus on single emotion labels and overlook the complex and ambiguous nature of human emotions. This study is the first to address this gap by exploring the potential of LLMs in recognizing ambiguous emotions, leveraging their strong generalization capabilities and in-context learning. We design zero-shot and few-shot prompting and incorporate past dialogue as context information for ambiguous emotion recognition. Experiments conducted using three datasets indicate significant potential for LLMs in recognizing ambiguous emotions, and highlight the substantial benefits of including context information. Furthermore, our findings indicate that LLMs demonstrate a high degree of effectiveness in recognizing less ambiguous emotions and exhibit potential for identifying more ambiguous emotions, paralleling human perceptual capabilities.
Yuan Gong 0001, Vidhyasaharan Sethu, Ting Dang
ICASSP4
2025 Cognitive Load Monitoring via Earable Acoustic Sensing
abstract
The rapid adoption of ear-worn devices (earables) has shown significant potential for continuous health monitoring. Despite their close proximity to the human brain and diverse sensing capabilities, the exploration of earable sensing in relation to cognitive function remains underexplored. Building on theoretical and empirical foundations regarding the interplay between cognitive load, auditory complexity, and changes in hearing characteristics influenced by brain function, this study is the first to leverage earable acoustic sensing to assess cognitive load. We specifically designed auditory tasks to elicit four levels of cognitive load and used otoacoustic emissions (OAEs) to measure cochlear response changes in response to cognitive load. By utilizing both audio content indicating auditory complexity and OAEs reflecting hearing characteristic changes, we designed machine learning pipelines to automate the assessment in a four-class cognitive detection task, achieving an accuracy of 68.88%. This research opens a new pathway for using earable acoustic sensing in monitoring cognitive function and holds great potential for future cognitive augmentation.
Jiatao Quan, Khaldoon Al-Naimi, Xijia Wei, Yang Liu 0101, Fahim Kawsar, Alessandro Montanari, Ting Dang
ICASSP7
2025 MicarVLMoE: A Modern Gated Cross-Aligned Vision-Language Mixture of Experts Model for Medical Image Captioning and Report Generation
abstract
Medical image reporting (MIR) aims to generate structured clinical descriptions from radiological images. Existing methods struggle with fine-grained feature extraction, multimodal alignment, and generalization across diverse imaging types, often relying on vanilla transformers and focusing primarily on chest X-rays. We propose MicarVLMoE, a vision-language mixture-of-experts model with gated cross-aligned fusion, designed to address these limitations. Our architecture includes: (i) a multiscale vision encoder (MSVE) for capturing anatomical details at varying resolutions, (ii) a multihead dual-branch latent attention (MDLA) module for vision-language alignment through latent bottleneck representations, and (iii) a modulated mixture-of-experts (MoE) decoder for adaptive expert specialization. We extend MIR to CT scans, retinal imaging, MRI scans, and gross pathology images, reporting state-of-the-art results on COVCTR, MMR, PGROSS, and ROCO datasets. Extensive experiments and ablations confirm improved clinical accuracy, cross-modal alignment, and model interpretability. Code is available at https://github.com/AI-14/micar-vl-moe.
Amaan Izhar, Nurul Japar, Norisma Idris, Ting Dang
IJCNN4
2025 A Study of Speech Embedding Similarities Between Australian Aboriginal and High-Resource Languages
Eliathamby Ambikairajah, Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu
INTERSPEECH3
2025 Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
Jule Valendo Halim, Hong Jia, Ting Dang
INTERSPEECH4
2025 E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models
abstract
Speech Foundation Models encounter significant performance degradation when deployed in real-world scenarios involving acoustic domain shifts, such as background noise and speaker accents. Test-time adaptation (TTA) has recently emerged as a viable strategy to address such domain shifts at inference time without requiring access to source data or labels. However, existing TTA approaches, particularly those relying on backpropagation, are memory-intensive, limiting their applicability in speech tasks and resource-constrained settings. Although backpropagation-free methods offer improved efficiency, existing ones exhibit poor accuracy. This is because they are predominantly developed for vision tasks, which fundamentally differ from speech task formulations, noise characteristics, and model architecture, posing unique transferability challenges. In this paper, we introduce E-BAT, first Efficient BAckpropagation-free TTA framework designed explicitly for speech foundation models. E-BAT achieves a balance between adaptation effectiveness and memory efficiency through three key components: (i) lightweight prompt adaptation for a forward-pass-based feature alignment, (ii) a multi-scale loss to capture both global (utterance-level) and local distribution shifts (token-level) and (iii) a test-time exponential moving average mechanism for stable adaptation across utterances. Experiments conducted on four noisy speech datasets spanning sixteen acoustic conditions demonstrate consistent improvements, with 4.1\%--13.5% accuracy gains over backpropogation-free baselines and 2.0$\times$–6.4$\times$ GPU memory savings compared to backpropogation-based methods. By enabling scalable and robust adaptation under acoustic variability, this work paves the way for developing more efficient adaptation approaches for practical speech processing systems in real-world environments.
Jiaheng Dong, Hong Jia, Soumyajit Chatterjee, Abhirup Ghosh, Ting Dang
NeurIPS6
2025 FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
abstract
Large Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications.
Yan Gao 0016, Massimo Roberto Scamarcia, Javier Fernández-Marqués, Mohammad Naseri, Chong Shen Ng, Dimitris Stripelis, Zexi Li 0001, Tao Shen 0002, Jiamu Bai, Daoyuan Chen, Zikai Zhang 0003, Rui Hu 0005, Inseo Song, Kangyoon Lee, Hong Jia, Ting Dang, Zheyuan Liu 0002, Daniel J. Beutel, Lingjuan Lyu, Nicholas D. Lane
NeurIPS16
2025 Self-Supervised rU-net With Spectrum Branch: A Novel Framework for Subject-independent Emotion Recognition based on Peripheral Physiological Signals
abstract
Frequency-domain features of peripheral physiological signals are vital for emotion recognition. However, existing end-to-end network architectures rarely extract them efficiently. To address this limitation, we propose a multimodal rU-Net model incorporating time-frequency information fusion. Specifically, the spectrum is integrated as a parallel branch alongside the time-domain branch for feature extraction. A fusion module enables direct frequency domain feature extraction and subsequent time-frequency fusion. By utilizing the rU-Net encoder with multimodal signal channels, our approach processes skin temperature (SKT), electrodermal activity (EDA), and photoplethysmography (PPG) data simultaneously, thus preventing model inflation from encoder stacking. The CASE and DEAP datasets have been validated using the leave-one-subject-out (LOSO) approach. In the 3-class classification, the best accuracy for valence (V) and arousal (A) were 69.36% and 71.34%, respectively, while in the 2-class classification V and A were 70.56% and 70.29%, respectively. This work offers valuable insights and a novel approach for future research in emotion recognition based on peripheral physiological signals collected by non-EEG wearable devices.
Lifeng You, Ting Dang, Xiao Liu 0001
SMC3
2025 Data-Efficient Psychiatric Disorder Detection via Self-Supervised Learning on Frequency-Enhanced Brain Networks
abstract
Psychiatric disorders involve complex neural activity changes, with functional magnetic resonance imaging (fMRI) data serving as key diagnostic evidence. However, data scarcity and the diverse nature of fMRI information pose significant challenges. While graph-based self-supervised learning (SSL) methods have shown promise in brain network analysis, they primarily focus on time-domain representations, often overlooking the rich information embedded in the frequency domain. To overcome these limitations, we propose F requency- E nhanced Net work (FENet), a novel SSL framework specially designed for fMRI data that integrates time-domain and frequency-domain information to improve psychiatric disorder detection in small-sample datasets. FENet constructs multi-view brain networks based on the inherent properties of fMRI data, explicitly incorporating frequency information into the learning process of representation. Additionally, it employs domain-specific encoders to capture temporal-spectral characteristics, including an efficient frequency-domain encoder that highlights disease-relevant frequency features. Finally, FENet introduces a domain consistency-guided learning objective, which balances the utilization of diverse information and generates frequency-enhanced brain graph representations. Experiments on two real-world medical datasets demonstrate that FENet outperforms state-of-the-art methods while maintaining strong performance in minimal data conditions. Furthermore, we analyze the correlation between various frequency-domain features and psychiatric disorders, emphasizing the critical role of high-frequency information in disorder detection.
Mujie Liu, Mengchu Zhu, Qichao Dong, Ting Dang, Jiangang Ma, Jing Ren 0001, Feng Xia 0001
ACM Trans. Comput. Heal.4
2025 SQUIREDL: Sparse Sequence-to-Sequence Uncertainty Estimation in Evidential Deep Learning
abstract
Machine Learning models typically assume that time series are regularly spaced; however, this is often unrealistic in healthcare, where missing data recordings are common. In this context, uncertainty estimates play a pivotal role, as they can enable confident and non-confident predictions to be distinguished. We propose SQUIREDL, a novel uncertainty-aware sequence-to-sequence prediction method for sparse healthcare time series. Specifically, we enhance the state-of-the-art evidential regression framework, widely used for uncertainty estimation, to handle missing data. Following data imputation with an Akima spline-based method, we modify the loss function of evidential regression by assigning different weights to imputed and observed data points, to offer more reliable uncertainty estimates. Additionally, we examine a variety of metrics for assessing the success of uncertainty estimations on sequence-to-sequence predictions, providing a reliable way to evaluate the models in a medical setting. Our proposal is demonstrated in two clinical applications. In continuous glucose monitoring, we use sequence-to-sequence prediction to obtain the hypoglycaemia risk from glucose sensor readings. Our approach captures the ground truth risk values 30% more accurately, bringing consistent improvements in both uncertainty-aware and accuracy-based metrics. Similarly, in COVID-19 hospital admissions data, we achieve a 22% improvement in the accuracy of uncertainty-aware predictions, enabling better resource planning.
Sotirios Vavaroutas, Ting Dang, Emma Rocheteau, Cecilia Mascolo
ACM Trans. Comput. Heal.2
2025 Speech Emotion Recognition via CNN-Transformer and multidimensional attention mechanism
Jiazheng Huang, Ting Dang, Jintao Cheng
Speech Commun.4
2024 StatioCL: Contrastive Learning for Time Series via Non-Stationary and Temporal Contrast
abstract
Contrastive learning (CL) has emerged as a promising approach for representation learning in time series data by embedding similar pairs closely while distancing dissimilar ones. However, existing CL methods often introduce false negative pairs (FNPs) by neglecting inherent characteristics and then randomly selecting distinct segments as dissimilar pairs, leading to erroneous representation learning, reduced model performance, and overall inefficiency. To address these issues, we systematically define and categorize FNPs in time series into semantic false negative pairs and temporal false negative pairs for the first time: the former arising from overlooking similarities in label categories, which correlates with similarities in non-stationarity and the latter from neglecting temporal proximity. Moreover, we introduce StatioCL, a novel CL framework that captures non-stationarity and temporal dependency to mitigate both FNPs and rectify the inaccuracies in learned representations. By interpreting and differentiating non-stationary states, which reflect the correlation between trends or temporal dynamics with underlying data patterns, StatioCL effectively captures the semantic characteristics and eliminates semantic FNPs. Simultaneously, StatioCL establishes fine-grained similarity levels based on temporal dependencies to capture varying temporal proximity between segments and to mitigate temporal FNPs. Evaluated on real-world benchmark time series classification datasets, StatioCL demonstrates a substantial improvement over state-of-the-art CL methods, achieving a 2.9% increase in Recall and a 19.2% reduction in FNPs. Most importantly, StatioCL also shows enhanced data efficiency and robustness against label scarcity.
Yu Wu 0021, Ting Dang, Dimitris Spathis, Hong Jia, Cecilia Mascolo
CIKM2
2024 Variational Connectionist Temporal Classification for Order-Preserving Sequence Modeling
abstract
Connectionist temporal classification (CTC) is commonly adopted for sequence modeling tasks like speech recognition, where it is necessary to preserve order between the input and target sequences. However, CTC is only applied to deterministic sequence models, where the latent space is discontinuous and sparse, which in turn makes them less capable of handling data variability when compared to variational models. In this paper, we integrate CTC with a variational model and derive loss functions that can be used to train more generalizable sequence models that preserve order. Specifically, we derive two versions of the novel variational CTC based on two reasonable assumptions, the first being that the variational latent variables at each time step are conditionally independent; and the second being that these latent variables are Markovian. We show that both loss functions allow direct optimization of the variational lower bound for the model log-likelihood, and present computationally tractable forms for implementing them.
Zheng Nan, Ting Dang, Vidhyasaharan Sethu, Beena Ahmed
ICASSP2
2024 Towards Enabling DPOAE Estimation on Single-Speaker Earbuds
abstract
Distortion Product OtoAcoustic Emissions (DPOAEs) represents faint cochlear responses to dual-frequency stimuli, commonly employed in hearing screening. This paper introduces an innovative approach to trigger DPOAEs using single-speaker earbuds. Due to their compact size, the speakers used in the earbuds exhibit nonlinear behavior, leading to Inter-Modulation Distortions (IMDs) that interfere with DPOAE signals. Conventional medical devices employ dual speakers to mitigate this distortion, such a solution is impractical for space-constrained earbuds. To address this challenge, we propose a method that triggers DPOAEs while circumventing IMDs by designing a stimulus signal that alternates between the two frequencies necessary for triggering DPOAEs. The performance of our system was evaluated through a preliminary user study involving 8 participants, and it demonstrated a median correlation of 0.65 when compared to a medical-grade reference device.
Irtaza Shahid, Khaldoon Al-Naimi, Ting Dang, Yang Liu 0101, Fahim Kawsar, Alessandro Montanari
ICASSP3
2024 Dual-Constrained Dynamical Neural ODEs for Ambiguity-aware Continuous Emotion Prediction
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
INTERSPEECH2
2024 Efficient and Personalized Mobile Health Event Prediction via Small Language Models
abstract
Healthcare monitoring is crucial for early detection, timely intervention, and the ongoing management of health conditions, ultimately improving individuals' quality of life. Recent research shows that Large Language Models (LLMs) have demonstrated impressive performance in supporting healthcare tasks. However, existing LLM-based healthcare solutions typically rely on cloud-based systems, which raise privacy concerns and increase the risk of personal information leakage. As a result, there is growing interest in running these models locally on devices like mobile phones and wearables to protect users' privacy. Small Language Models (SLMs) are potential candidates to solve privacy and computational issues, as they are more efficient and better suited for local deployment. However, the performance of SLMs in healthcare domains has not yet been investigated. This paper examines the capability of SLMs to accurately analyze health data, such as steps, calories, sleep minutes, and other vital statistics, to assess an individual's health status. Our results show that, TinyLlama, which has 1.1 billion parameters, utilizes 4.31 GB memory, and has 0.48s latency, showing the best performance compared other four state-of-the-art (SOTA) SLMs on various healthcare applications. Our results indicate that SLMs could potentially be deployed on wearable or mobile devices for real-time health monitoring, providing a practical solution for efficient and privacy-preserving healthcare.
Xin Wang 0215, Ting Dang, Vassilis Kostakos, Hong Jia
MobiCom2
2024 TinyTTA: Efficient Test-time Adaptation via Early-exit Ensembles on Edge Devices
abstract
The increased adoption of Internet of Things (IoT) devices has led to the generation of large data streams with applications in healthcare, sustainability, and robotics. In some cases, deep neural networks have been deployed directly on these resource-constrained units to limit communication overhead, increase efficiency and privacy, and enable real-time applications. However, a common challenge in this setting is the continuous adaptation of models necessary to accommodate changing environments, i.e., data distribution shifts. Test-time adaptation (TTA) has emerged as one potential solution, but its validity has yet to be explored in resource-constrained hardware settings, such as those involving microcontroller units (MCUs). TTA on constrained devices generally suffers from i) memory overhead due to the full backpropagation of a large pre-trained network, ii) lack of support for normalization layers on MCUs, and iii) either memory exhaustion with large batch sizes required for updating or poor performance with small batch sizes. In this paper, we propose TinyTTA, to enable, for the first time, efficient TTA on constrained devices with limited memory. To address the limited memory constraints, we introduce a novel self-ensemble and batch-agnostic early-exit strategy for TTA, which enables continuous adaptation with small batch sizes for reduced memory usage, handles distribution shifts, and improves latency efficiency. Moreover, we develop the TinyTTA Engine, a first-of-its-kind MCU library that enables on-device TTA. We validate TinyTTA on a Raspberry Pi Zero 2W and an STM32H747 MCU. Experimental results demonstrate that TinyTTA improves TTA accuracy by up to 57.6\%, reduces memory usage by up to six times, and achieves faster and more energy-efficient TTA. Notably, TinyTTA is the only framework able to run TTA on MCU STM32H747 with a 512 KB memory constraint while maintaining high performance.
Hong Jia, Young D. Kwon, Alessio Orsino, Ting Dang, Domenico Talia, Cecilia Mascolo
NeurIPS4
2024 An evaluation of heart rate monitoring with in-ear microphones under motion
abstract
With the soaring adoption of in-ear wearables, the research community has started investigating suitable in-ear heart rate detection systems. Heart rate is a key physiological marker of cardiovascular health and physical fitness. Continuous and reliable heart rate monitoring with wearable devices has therefore gained increasing attention in recent years. Existing heart rate detection systems in wearables mainly rely on photoplethysmography (PPG) sensors, however, these are notorious for poor performance in the presence of human motion. In this work, leveraging the occlusion effect that enhances low-frequency bone-conducted sounds in the ear canal, we investigate for the first time in-ear audio-based motion-resilient heart rate monitoring. We first collected heart rate-induced sounds in the ear canal using an in-ear microphone under seven stationary activities and two full-body motion activities (i.e., walking, and running). Then, we devised a novel deep learning based motion artefact (MA) mitigation framework to denoise the in-ear audio signals, followed by a heart rate estimation algorithm to extract heart rate. With data collected from 15 subjects over nine activities, we demonstrate that hEARt, our end-to-end approach, achieves a mean absolute error (MAE) of 1.88 ± 2.89 BPM, 6.83 ± 5.05 BPM, and 13.19 ± 11.37 BPM for stationary, walking, and running, respectively, opening the door to a new non-invasive and affordable heart rate monitoring with usable performance for daily activities. Not only does hEARt outperform previous in-ear heart rate monitoring work, but it outperforms reported in-ear PPG performance.
Kayla-Jade Butkow, Ting Dang, Andrea Ferlini, Dong Ma 0001, Yang Liu 0101, Cecilia Mascolo
Pervasive Mob. Comput.2
2024 Uncertainty-Aware Health Diagnostics via Class-Balanced Evidential Deep Learning
abstract
Uncertainty quantification is critical for ensuring the safety of deep learning-enabled health diagnostics, as it helps the model account for unknown factors and reduces the risk of misdiagnosis. However, existing uncertainty quantification studies often overlook the significant issue of class imbalance, which is common in medical data. In this paper, we propose a class-balanced evidential deep learning framework to achieve fair and reliable uncertainty estimates for health diagnostic models. This framework advances the state-of-the-art uncertainty quantification method of evidential deep learning with two novel mechanisms to address the challenges posed by class imbalance. Specifically, we introduce a pooling loss that enables the model to learn less biased evidence among classes and a learnable prior to regularize the posterior distribution that accounts for the quality of uncertainty estimates. Extensive experiments using benchmark data with varying degrees of imbalance and various naturally imbalanced health data demonstrate the effectiveness and superiority of our method. Our work pushes the envelope of uncertainty quantification from theoretical studies to realistic healthcare application scenarios. By enhancing uncertainty estimation for class-imbalanced data, we contribute to the development of more reliable and practical deep learning-enabled health diagnostic systems.
Tong Xia, Ting Dang, Jing Han 0010, Lorena Qendro, Cecilia Mascolo
IEEE J. Biomed. Health Informatics2
2023 Belief Mismatch Coefficient (BMC): A Novel Interpretable Measure of Prediction Accuracy for Ambiguous Emotion States
abstract
Despite of efforts made to model emotion ambiguity and develop ambiguity aware emotion prediction systems, there is a need for a quantitative and interpretable measure of the accuracy of such systems, regardless of recent advances in representing emotion ambiguity through probability distributions. In this paper, we propose a novel measure called the “Belief Mismatch Coefficient (BMC) that quantifies the differences in the belief that emotional states are perceived from certain regions within the arousal/valence space when comparing a predicted distribution to an underlying distribution inferred from ground truth ratings. The proposed metric is validated using simulated labels to demonstrate its effectiveness in quantifying various prediction errors. Furthermore, it is extended to real-case emotion prediction systems using two state-of-the-art modeling techniques on the RECOLA dataset. The experimental results confirm that the proposed metric can efficiently capture and differentiate between various prediction errors, while also offering insights into the predictions. Moreover, it demonstrates significant advantages in capturing a comprehensive view of the predicted distribution compared to traditional metrics such as Concordance Correlation Coefficients.
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
ACII2
2023 Constrained Dynamical Neural ODE for Time Series Modelling: A Case Study on Continuous Emotion Prediction
abstract
weA number of machine learning applications involve time series prediction, and in some cases additional information about dynamical constraints on the target time series may be available. For instance, it might be known that the desired quantity cannot change faster than some rate or that the rate is dependent on some known factors. However, incorporating these constraints into deep learning models, such as recurrent neural networks, is not straightforward. In this paper, we propose constrained dynamical neural ordinary differential equation (CD-NODE) models, which treat the desired time series as a dynamic process that can be described by an ODE. CD-NODEs model the rate of change of the time series as a function of both itself and the current input features, parameterised as a neural network. We explore the effect of constraining the dynamics of the model by placing explicit restrictions on the rate of change. The proposed model is evaluated on speech-based continuous emotion prediction, where such dynamical constraints are expected, using the publicly available RECOLA dataset. Results suggest that the model achieves performances comparable with the state-of-the-art despite using significantly fewer parameters. Additional analyses reveal that imposing these constraints on the model leads to faster convergence and better performance, especially with smaller training data sets.
Ting Dang, Antoni Dimitriadis, Jingyao Wu 0002, Vidhyasaharan Sethu, Eliathamby Ambikairajah
ICASSP1
2023 From Interval to Ordinal: A HMM based Approach for Emotion Label Conversion
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
INTERSPEECH2
2023 Conditional Neural ODE Processes for Individual Disease Progression Forecasting: A Case Study on COVID-19
abstract
Time series forecasting, as one of the fundamental machine learning areas, has attracted tremendous attentions over recent years. The solutions have evolved from statistical machine learning (ML) methods to deep learning techniques. One emerging sub-field of time series forecasting is individual disease progression forecasting, e.g., predicting individuals' disease development over a few days (e.g., deteriorating trends, recovery speed) based on few past observations. Despite the promises in the existing ML techniques, a variety of unique challenges emerge for disease progression forecasting, such as irregularly-sampled time series, data sparsity, and individual heterogeneity in disease progression. To tackle these challenges, we propose novel Conditional Neural Ordinary Differential Equations Processes (CNDPs), and validate it in a COVID-19 disease progression forecasting task using audio data. CNDPs allow for irregularly-sampled time series modelling, enable accurate forecasting with sparse past observations, and achieve individual-level progression forecasting. CNDPs show strong performance with an Unweighted Average Recall (UAR) of 78.1%, outperforming a variety of commonly used Recurrent Neural Networks based models. With the proposed label-enhancing mechanism (i.e., including the initial health status as input) and the customised individual-level loss, CNDPs further boost the performance reaching a UAR of 93.6%. Additional analysis also reveals the model's capability in tracking individual-specific recovery trend, implying the potential usage of the model for remote disease progression monitoring. In general, CNDPs pave new pathways for time series forecasting, and provide considerable advantages for disease progression monitoring.
Ting Dang, Jing Han 0010, Tong Xia, Erika Bondareva, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Dimitris Spathis, Pietro Cicuta, Cecilia Mascolo
KDD1
2023 hEARt: Motion-resilient Heart Rate Monitoring with In-ear Microphones
abstract
With the soaring adoption of in-ear wearables, the research community has started investigating suitable in-ear heart rate (HR) detection systems. HR is a key physiological marker of cardiovascular health and physical fitness. Continuous and reliable HR monitoring with wearable devices has therefore gained increasing attention in recent years. Existing HR detection systems in wearables mainly rely on photoplethysmography (PPG) sensors, however, these are notorious for poor performance in the presence of human motion. In this work, leveraging the occlusion effect that enhances low-frequency bone-conducted sounds in the ear canal, we investigate for the first time in-ear audio-based motion-resilient HR monitoring. We first collected HR-induced sounds in the ear canal leveraging an in-ear microphone under stationary and three different activities (i.e., walking, running, and speaking). Then, we devised a novel deep learning based motion artefact (MA) mitigation framework to denoise the in-ear audio signals, followed by an HR estimation algorithm to extract HR. With data collected from 20 subjects over four activities, we demonstrate that hEARt, our end-to-end approach, achieves a mean absolute error (MAE) of 3.02$\pm\ \boldsymbol{ 2.97}$BPM, 8.12$\pm\ \boldsymbol{6.74}$BPM, 11.23$\pm\ \boldsymbol{9.20}$BPM and 9.39$\pm\ \boldsymbol{6.97}$BPM for stationary, walking, running and speaking, respectively, opening the door to a new non-invasive and affordable HR monitoring with usable performance for daily activities. Not only does hEARt outperform previous in-ear HR monitoring work, but it outperforms reported in-ear PPG performance.
Kayla-Jade Butkow, Ting Dang, Andrea Ferlini, Dong Ma 0001, Cecilia Mascolo
PERCOM2
2023 DNN controlled adaptive front-end for replay attack detection systems
abstract
Developing robust countermeasures to protect automatic speaker verification systems against replay spoofing attacks is a well-recognized challenge. Current approaches to spoofing detection are generally based on a fixed front-end, typically a time-invariant filter bank, followed by a machine learning back-end. In this paper, we propose a novel approach whereby the front-end comprises an adaptive filter bank with a deep neural network-based controller, which is jointly trained along with a neural network back-end. Specifically, the deep neural network-based adaptive filter controller tunes the selectivity and sensitivity of the front-end filter bank at every frame to capture replay-related artefacts. We demonstrate the effectiveness of the proposed framework in spoofing attack detection on a synthesized dataset and ASVSpoof 2019 and ASVSpoof 2021 challenge datasets in terms of equal error rate and its ability to capture artefacts that differentiate replayed signals from genuine ones in comparison to conventional non-adaptive front-end.
Buddhi Wickramasinghe, Eliathamby Ambikairajah, Vidhyasaharan Sethu, Julien Epps, Haizhou Li 0001, Ting Dang
Speech Commun.6
2023 A Novel Markovian Framework for Integrating Absolute and Relative Ordinal Emotion Information
abstract
There is growing interest in affective computing for the representation and prediction of emotions along ordinal scales. However, the term ordinal emotion label has been used to refer to both absolute notions such as low or high arousal, as well as relation notions such as arousal is higher at one instance compared to another. In this paper, we introduce the terminology absolute and relative ordinal labels to make this distinction clear and investigate both with a view to integrate them and exploit their complementary nature. We propose a Markovian framework referred to as Dynamic Ordinal Markov Model (DOMM) that makes use of both absolute and relative ordinal information, to improve speech based ordinal emotion prediction. Finally, the proposed framework is validated on two speech corpora commonly used in affective computing, the RECOLA and the IEMOCAP databases, across a range of system configurations. The results consistently indicate that integrating relative ordinal information improves absolute ordinal emotion prediction.
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
IEEE Trans. Affect. Comput.2
2022 A Novel Sequential Monte Carlo Framework for Predicting Ambiguous Emotion States
abstract
When continuous emotion labelling of natural (non-acted) data is desired, it is typically collected from multiple annotators. However, most automatic emotion recognition systems trained on such data ignore disagreement between annotators and only models the average rating, despite the observation that the degree of disagreement would reflect the ambiguity and subtlety in every expression of emotions. In this paper, we propose a novel Sequential Monte Carlo framework that models the perceived emotion as time-varying distributions that allows for ambiguity to be incorporated. Additionally, we present alternative measures that consider both the similarity of prediction to the multiple labels, as well as whether the degree of ambiguity in the prediction and labels. The proposed system was validated on the publicly available RECOLA dataset.
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
ICASSP2
2022 Exploring Semi-supervised Learning for Audio-based COVID-19 Detection using FixMatch
Ting Dang, Thomas Quinnell, Cecilia Mascolo
INTERSPEECH1
2021 Uncertainty-Aware COVID-19 Detection from Imbalanced Sound Data
abstract
Recently, sound-based COVID-19 detection studies have shown great promise to achieve scalable and prompt digital prescreening.However, there are still two unsolved issues hindering the practice.First, collected datasets for model training are often imbalanced, with a considerably smaller proportion of users tested positive, making it harder to learn representative and robust features.Second, deep learning models are generally overconfident in their predictions.Clinically, false predictions aggravate healthcare costs.Estimation of the uncertainty of screening would aid this.To handle these issues, we propose an ensemble framework where multiple deep learning models for sound-based COVID-19 detection are developed from different but balanced subsets from original data.As such, data are utilized more effectively compared to traditional up-sampling and down-sampling approaches: an AUC of 0.74 with a sensitivity of 0.68 and a specificity of 0.69 is achieved.Simultaneously, we estimate uncertainty from the disagreement across multiple models.It is shown that false predictions often yield higher uncertainty, enabling us to suggest the users with certainty higher than a threshold to repeat the audio test on their phones or to take clinical tests if digital diagnosis still fails.This study paves the way for a more robust sound-based COVID-19 automated screening system.
Tong Xia, Jing Han 0010, Lorena Qendro, Ting Dang, Cecilia Mascolo
Interspeech4
2021 Compensation Techniques for Speaker Variability in Continuous Emotion Prediction
abstract
Continuous time-varying prediction of emotions based on speech in terms of attributes (i.e., arousal) has received considerable attention in the past few years. However, the variability introduced by factors not related to emotion, such as speaker and phonetic variability, which in turn may lead to less reliable models and less accurate emotion predictions, has not been fully explored yet. In particular, even though speaker variability has been shown to be a significant confounding factor in continuous emotion prediction systems, there remains a paucity of analyses about how speaker variability affects continuous emotion prediction systems and which methods can be applied to compensate for this variability. This paper first formulates speaker variability systematically in terms of probability distributions in both feature and model spaces, and quantifies the effect of speaker variability by comparing inter- and intra-speaker variability between speaker-dependent models. Second, two compensation techniques based on partial least squares dimensional reduction and feature mapping are proposed. Finally, the effectiveness of the proposed techniques is validated on three databases, across which they show consistent improvement in arousal, valence and dominance prediction. Additional quantitative analyse reveals that the two proposed techniques compensate for speaker variability in both the feature and model spaces simultaneously.
Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
IEEE Trans. Affect. Comput.1
2019 A Novel Bag-of-Optimised-Clusters Front-End for Speech based Continuous Emotion Prediction
abstract
Almost all current speech based emotion prediction systems employ front-ends that approximately represent the distribution of frame based features over a suitable window, typically via a set of statistical functionals or the use of Bag-of-Audio-Words (BoAW) features. These front-ends are designed either by manual selection of appropriate statistical functionals, via a feature selection approach or unsupervised clustering of the frame-based features and may not be optimal for the task at hand. This paper proposes a novel front-end that discriminatively learns feature clusters to generate a codebook optimised for emotion prediction, which is then used to generate a Bag-of-Optimised-Clusters (BoOC) feature set. Moreover, this front-end is implemented as a sequence of neural network layers that allow both the proposed front-end and a suitable deep learning backend to be jointly trained. The Bag-of-Optimised-Clusters frontend is tested on the RECOLA database and results show that it outperforms the well-established BoAW features.
Deboshree Bose, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Sarith Fernando
ACII2
2019 Speech Based Emotion Prediction: Can a Linear Model Work?
Anda Ouyang, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
INTERSPEECH2
2018 Dynamic Multi-Rater Gaussian Mixture Regression Incorporating Temporal Dependencies of Emotion Uncertainty Using Kalman Filters
abstract
Predicting continuous emotion in terms of affective attributes has mainly been focused on hard labels, which ignored the ambiguity of recognizing certain emotions. This ambiguity may result in high inter-rater variability and in turn causes varying prediction uncertainty with time. Based on the assumption that temporal dependencies occur in the evolution of emotion uncertainty, this paper proposes a dynamic multi-rater Gaussian Mixture Regression (GMR), aiming to obtain the emotion uncertainty prediction reflected by multi-raters by taking into account their temporal dependencies. This framework is achieved by incorporating feedforward and backward Kalman filters into GMR to estimate the time-dependent label distribution that reflects the emotion uncertainty. It also provides the benefits of relaxing the label distribution of Gaussian assumption to that of a Gaussian Mixture Model (GMM). In addition, a new measurement to estimate emotion uncertainty from GMM as the local variability is adopted. Experiments conducted on the RECOLA database reveal that incorporating temporal dependencies is critical for emotion uncertainty prediction with 17% relative improvement for arousal, and that the proposed framework for emotion uncertainty prediction shows potential in conventional emotion attribute prediction.
Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
ICASSP1
2017 An Investigation of Emotion Prediction Uncertainty Using Gaussian Mixture Regression
Ting Dang, Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah
INTERSPEECH1
2016 Factor Analysis Based Speaker Normalisation for Continuous Emotion Prediction
Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah
INTERSPEECH1