EDBT 2026 Demo / reviewers in the wild / expert
Vidhyasaharan Sethu
dblp:52/7823
· DBLP profile ↗
79ranked-venue papers
7as first author
26since 2021 · last 2025
0000-0001-8492-1787ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 41 · 4 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AER-LLM: Ambiguity-aware Emotion Recognition Leveraging Large Language ModelsabstractRecent advancements in Large Language Models (LLMs) have demonstrated great success in many Natural Language Processing (NLP) tasks. In addition to their cognitive intelligence, exploring their capabilities in emotional intelligence is also crucial, as it enables more natural and empathetic conversational AI. Recent studies have shown LLMs’ capability in recognizing emotions, but they often focus on single emotion labels and overlook the complex and ambiguous nature of human emotions. This study is the first to address this gap by exploring the potential of LLMs in recognizing ambiguous emotions, leveraging their strong generalization capabilities and in-context learning. We design zero-shot and few-shot prompting and incorporate past dialogue as context information for ambiguous emotion recognition. Experiments conducted using three datasets indicate significant potential for LLMs in recognizing ambiguous emotions, and highlight the substantial benefits of including context information. Furthermore, our findings indicate that LLMs demonstrate a high degree of effectiveness in recognizing less ambiguous emotions and exhibit potential for identifying more ambiguous emotions, paralleling human perceptual capabilities. Yuan Gong 0001, Vidhyasaharan Sethu, Ting Dang |
ICASSP | 3 |
| 2025 | Improved Out-of-domain Detection in VAE Latent Spaces with Boundary-driven RegularisationabstractIn out-of-domain (OOD) detection tasks, encoding the actual data into a suitable latent space could be beneficial since it may facilitate measurement of the spatial relationship between in-domain (IND) and OOD data. However, any such mapping of data to a latent space carries the risk that some OOD points may be mapped to in-domain regions. To address this drawback we propose a novel method that spatially separates these two domains in the latent space. It first geometrically defines the boundary of IND in the latent space and then forces some enumerated OOD data (known as surrogate OOD) to fit that boundary. This then helps the encoder to map OOD to the surrounding area of IND, consequently reducing any overlap. Following this, accurate OOD detection is achieved by a dedicated detector distinguishing IND and its boundary. We illustrate the effects of the proposed method on synthetic data and validated it over both grey-scale and RGB image datasets. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 2 |
| 2025 | Evidential Neural GPLDA: A Novel Approach to Quantify Prediction Uncertainty in Speaker Verification SystemsabstractThe uncertainty of an automatic speaker verification (ASV) system is typically estimated using its overall accuracy. However it fails to express "when" the system is uncertain in a predictive and case-by-case manner. Also, prior to interpreting each prediction made by ASV systems, there is a need to assess if the system is confident about the prediction, which remains less explored in current research. Given the sense that uncertainty of this notion should be associated with the knowledge level held by the system, we propose an Evidential Neural GPLDA back-end inspired by evidential deep learning. This approach quantifies the uncertainty in each prediction based on the density of the training data supporting that prediction. This is achieved by showing if the input representation is similar to those of the training data samples, resulting in a confidence estimate based on training data only. We find loss functions and out-of-domain samples to train the proposed model such that it parameterizes a sharp Dirichlet distribution as low uncertainty and a flat one as high uncertainty. Experiments show that the proposed novel back-end effectively quantifies uncertainty, providing an estimate that aligns with the error rate. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 2 |
| 2025 | Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance FeaturesabstractEstimating frequency-varying acoustic parameters is essential for enhancing immersive perception in realistic spatial audio creation. In this paper, we propose a unified framework that blindly estimates reverberation time (T60), direct-to-reverberant ratio (DRR), and clarity (C50) across 10 frequency bands using first-order Ambisonics (FOA) speech recordings as inputs. The proposed framework utilizes a novel feature named Spectro-Spatial Covariance Vector (SSCV), efficiently representing temporal, spectral as well as spatial information of the FOA signal. Our models significantly outperform existing single-channel methods with only spectral information, reducing estimation errors by more than half for all three acoustic parameters. Additionally, we introduce FOA-Conv3D, a novel back-end network for effectively utilising the SSCV feature with a 3D convolutional encoder. FOA-Conv3D outperforms the convolutional neural network (CNN) and recurrent convolutional neural network (CRNN) backends, achieving lower estimation errors and accounting for a higher proportion of variance (PoV) for all 3 acoustic parameters. Hanyu Meng, Jeroen Breebaart, Jeremy Stoddard, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 4 |
| 2025 | A Study of Speech Embedding Similarities Between Australian Aboriginal and High-Resource Languages
Eliathamby Ambikairajah, Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu |
INTERSPEECH | 4 |
| 2025 | Quantifying prediction uncertainties in automatic speaker verification systemsabstractFor modern automatic speaker verification (ASV) systems, explicitly quantifying the confidence for each prediction strengthens the system’s reliability by indicating in which case the system is with trust. However, current paradigms do not take this into consideration. We thus propose to express confidence in the prediction by quantifying the uncertainty in ASV predictions. This is achieved by developing a novel Bayesian framework to obtain a score distribution for each input. The mean of the distribution is used to derive the decision while the spread of the distribution represents the uncertainty arising from the plausible choices of the model parameters. To capture the plausible choices, we sample the probabilistic linear discriminant analysis (PLDA) back-end model posterior through Hamiltonian Monte-Carlo (HMC) and approximate the embedding model posterior through stochastic Langevin dynamics (SGLD) and Bayes-by-backprop. Given the resulting score distribution, a further quantification and decomposition of the prediction uncertainty are achieved by calculating the score variance, entropy, and mutual information. The quantified uncertainties include the aleatoric uncertainty and epistemic uncertainty (model uncertainty). We evaluate them by observing how they change while varying the amount of training speech, the duration, and the noise level of testing speech. The experiments indicate that the behaviour of those quantified uncertainties reflects the changes we made to the training and testing data, demonstrating the validity of the proposed method as a measure of uncertainty. • The paper emphasises the need for quantifying and separating uncertainties in ASV. • The paper proposes a novel framework incorporating various Bayesian learning methods. • The major cause of epistemic uncertainty is training data size and test data length. • The noise level in the test utterance increases the aleatoric uncertainty in ASV. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed, Kong-Aik Lee |
Comput. Speech Lang. | 2 |
| 2024 | ChatGPT in the Classroom: A Shift in Engineering Design EducationabstractArtificial intelligence tools like ChatGPT are increasingly being incorporated into our education paradigm. This paper explores how ChatGPT was used in an Electrical Engineering Design Proficiency course at the University of New South Wales in Sydney. The course is a term-long laboratory-based class that centres on independent student work in system design, implementation, and validation. Students were encouraged to consult ChatGPT for design solutions, explanations, and suggestions, with the requirement that they declare any use of AI tools. The assessment process was carefully designed to determine whether responses originated from AI tools or the students' own understanding. Notably, 70% of the fifty students in the class utilised ChatGPT to enhance their understanding of the subject. The paper will also discuss the specific design tasks given to students, the assessment process, and explore ChatGPT's potential as a supportive educational tool in other courses. Eliathamby Ambikairajah, Tharmakulasingam Sirojan, Tharmarajah Thiruvaran, Vidhyasaharan Sethu |
EDUCON | 4 |
| 2024 | A Tiered Learning Framework for Self-Guided Engineering Design EducationabstractThe Tiered Learning Framework is a multilayered system designed to aid students in assessing their progress and understanding within the learning process. Each layer corresponds to the skills students need to develop, and the framework encourages students to self-assess their current level and identify what they need to progress to higher tiers of learning. This paper explores the application of the Tiered Learning Framework in an Electrical Engineering Design Proficiency course at the University of New South Wales, Sydney. The course encompasses design tasks in critical areas such as electronic circuit design, signal processing design, and power system design. The paper focuses on the design and assessment of the signal processing design task as an example. In a class of 49 students, 96% agreed that the framework helped them achieve their learning goals, and 91% found it challenging yet encouraging for their learning. Compared to a non-tiered framework, 71% of the total students favour a tiered framework. Eliathamby Ambikairajah, Tharmarajah Thiruvaran, Vidhyasaharan Sethu, Deepak Mishra 0001, Tharmakulasingam Sirojan |
EDUCON | 3 |
| 2024 | A Probability Gradient Based Approach for Sampling Boundaries of In-Domain DataabstractIn machine learning applications, it is desirable to distinguish between in-domain and out-of-domain data. However, in most cases, only in-domain data is available and consequently identifying the ‘boundary’ between in-domain and out-of-domain is a significant challenge. In this paper we present a novel technique that can identify points on this boundary based on only in-domain data. Specifically, the proposed method operates on the hypothesis that the gradient of the probability of the data being in-domain will be highest at the boundary. It utilises an iterative approach, alternating between Monte Carlo sampling of an estimated ‘boundary distribution’ and a binary classifier trained on these points to distinguish between in-domain and out-of-domain data to improve the estimate of the boundary. This method leads to both a set of high-quality data points from the boundary and a calibrated out-of-domain detector. The operation of the proposed approach is validated on MNIST, Fashion-MNIST, and Omniglot datasets. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 2 |
| 2024 | Variational Connectionist Temporal Classification for Order-Preserving Sequence ModelingabstractConnectionist temporal classification (CTC) is commonly adopted for sequence modeling tasks like speech recognition, where it is necessary to preserve order between the input and target sequences. However, CTC is only applied to deterministic sequence models, where the latent space is discontinuous and sparse, which in turn makes them less capable of handling data variability when compared to variational models. In this paper, we integrate CTC with a variational model and derive loss functions that can be used to train more generalizable sequence models that preserve order. Specifically, we derive two versions of the novel variational CTC based on two reasonable assumptions, the first being that the variational latent variables at each time step are conditionally independent; and the second being that these latent variables are Markovian. We show that both loss functions allow direct optimization of the variational lower bound for the model log-likelihood, and present computationally tractable forms for implementing them. Zheng Nan, Ting Dang, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 3 |
| 2024 | Dual-Constrained Dynamical Neural ODEs for Ambiguity-aware Continuous Emotion Prediction
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2024 | Binaural Selective Attention Model for Target Speaker Extraction
Hanyu Meng, Qiquan Zhang, Xiangyu Zhang 0005, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2024 | Can Modelling Inter-Rater Ambiguity Lead To Noise-Robust Continuous Emotion Predictions?abstractThere has been increasing attention drawn to modelling interrater ambiguity in Continuous Emotion Recognition (CER) systems using probability distributions for arousal and valence.However, the relationship between modelling label ambiguity and robustness to noise, and more broadly, the impact of realworld noise on CER systems remains insufficiently explored.In this study, we argue that incorporating inter-rater ambiguity during training can regularize the noise response, leading to noise robustness.To this end, we propose a novel loss function that incorporates inter-rater ambiguity into model training.Experiments conducted on the RECOLA dataset demonstrate that our proposed method achieves a maximum Concordance Correlation Coefficient (CCC) improvement of 0.117 and 0.077 for mean and standard deviation predictions, respectively, across all noise conditions.We further integrate traditional noisy augmentation strategies with our proposed method and observe promising results. Ya-Tse Wu, Jingyao Wu 0002, Vidhyasaharan Sethu, Chi-Chun Lee |
INTERSPEECH | 3 |
| 2024 | Continuous Emotion Ambiguity Prediction: Modeling With Beta DistributionsabstractConventional continuous emotion prediction systems are typically trained to predict the ‘average’ of affect ratings obtained from multiple human annotators. These systems, however, ignore the ambiguity inherent in the perceived emotions, which is not captured by the ‘average rating’. This paper presents a novel ambiguity-aware continuous emotion prediction system that predicts the time-varying emotion state as a series of beta distributions. Our recent work has shown beta distributions to be an effective parametric model of a collection of affect ratings. This work develops an appropriate cost function that enables neural networks to be trained to predict beta distributions. It also investigates the choice of parameterization of the beta distribution, the choice of activation functions of the output layer, and the tractability of gradient definitions in combination with the loss function. The proposed framework is implemented using a Bag-of-Audio-Words front-end and an LSTM-based back-end and evaluated on the RECOLA dataset. In addition to comparison with baseline systems that only predict the ‘average rating’, the effectiveness with which the predictions represent ambiguity in perceived emotions is also evaluated. Experimental results reveal that the proposed approach outperforms other ambiguity-aware systems, especially when predicting valence. Deboshree Bose, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Belief Mismatch Coefficient (BMC): A Novel Interpretable Measure of Prediction Accuracy for Ambiguous Emotion StatesabstractDespite of efforts made to model emotion ambiguity and develop ambiguity aware emotion prediction systems, there is a need for a quantitative and interpretable measure of the accuracy of such systems, regardless of recent advances in representing emotion ambiguity through probability distributions. In this paper, we propose a novel measure called the “Belief Mismatch Coefficient (BMC) that quantifies the differences in the belief that emotional states are perceived from certain regions within the arousal/valence space when comparing a predicted distribution to an underlying distribution inferred from ground truth ratings. The proposed metric is validated using simulated labels to demonstrate its effectiveness in quantifying various prediction errors. Furthermore, it is extended to real-case emotion prediction systems using two state-of-the-art modeling techniques on the RECOLA dataset. The experimental results confirm that the proposed metric can efficiently capture and differentiate between various prediction errors, while also offering insights into the predictions. Moreover, it demonstrates significant advantages in capturing a comprehensive view of the predicted distribution compared to traditional metrics such as Concordance Correlation Coefficients. Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ACII | 3 |
| 2023 | Constrained Dynamical Neural ODE for Time Series Modelling: A Case Study on Continuous Emotion PredictionabstractweA number of machine learning applications involve time series prediction, and in some cases additional information about dynamical constraints on the target time series may be available. For instance, it might be known that the desired quantity cannot change faster than some rate or that the rate is dependent on some known factors. However, incorporating these constraints into deep learning models, such as recurrent neural networks, is not straightforward. In this paper, we propose constrained dynamical neural ordinary differential equation (CD-NODE) models, which treat the desired time series as a dynamic process that can be described by an ODE. CD-NODEs model the rate of change of the time series as a function of both itself and the current input features, parameterised as a neural network. We explore the effect of constraining the dynamics of the model by placing explicit restrictions on the rate of change. The proposed model is evaluated on speech-based continuous emotion prediction, where such dynamical constraints are expected, using the publicly available RECOLA dataset. Results suggest that the model achieves performances comparable with the state-of-the-art despite using significantly fewer parameters. Additional analyses reveal that imposing these constraints on the model leads to faster convergence and better performance, especially with smaller training data sets. Ting Dang, Antoni Dimitriadis, Jingyao Wu 0002, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 4 |
| 2023 | What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy ConditionsabstractThere is increasing interest in the use of the LEArnable Front-end (LEAF) in a variety of speech processing systems. However, there is a dearth of analyses of what is actually learnt and the relative importance of training the different components of the front-end. In this paper, we investigate this question on keyword spotting, speech-based emotion recognition and language identification tasks and find that the filters for spectral decomposition and the low pass filter used to estimate spectral energy variations exhibit no learning and the per-channel energy normalisation (PCEN) is the key component that is learnt. Following this, we explore the potential of adapting only the PCEN layer with a small amount of noisy data to enable it to learn appropriate dynamic range compression that better suits the noise conditions. This in turn enables a system trained on clean speech to work more accurately on noisy test data as demonstrated by the experimental results reported in this paper. Hanyu Meng, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2023 | Improving wav2vec2-based Spoken Language Identification by Learning Phonological Features
Mostafa Shahin, Zheng Nan, Vidhyasaharan Sethu, Beena Ahmed |
INTERSPEECH | 3 |
| 2023 | From Interval to Ordinal: A HMM based Approach for Emotion Label Conversion
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2023 | DNN controlled adaptive front-end for replay attack detection systemsabstractDeveloping robust countermeasures to protect automatic speaker verification systems against replay spoofing attacks is a well-recognized challenge. Current approaches to spoofing detection are generally based on a fixed front-end, typically a time-invariant filter bank, followed by a machine learning back-end. In this paper, we propose a novel approach whereby the front-end comprises an adaptive filter bank with a deep neural network-based controller, which is jointly trained along with a neural network back-end. Specifically, the deep neural network-based adaptive filter controller tunes the selectivity and sensitivity of the front-end filter bank at every frame to capture replay-related artefacts. We demonstrate the effectiveness of the proposed framework in spoofing attack detection on a synthesized dataset and ASVSpoof 2019 and ASVSpoof 2021 challenge datasets in terms of equal error rate and its ability to capture artefacts that differentiate replayed signals from genuine ones in comparison to conventional non-adaptive front-end. Buddhi Wickramasinghe, Eliathamby Ambikairajah, Vidhyasaharan Sethu, Julien Epps, Haizhou Li 0001, Ting Dang |
Speech Commun. | 3 |
| 2023 | A Novel Markovian Framework for Integrating Absolute and Relative Ordinal Emotion InformationabstractThere is growing interest in affective computing for the representation and prediction of emotions along ordinal scales. However, the term ordinal emotion label has been used to refer to both absolute notions such as low or high arousal, as well as relation notions such as arousal is higher at one instance compared to another. In this paper, we introduce the terminology absolute and relative ordinal labels to make this distinction clear and investigate both with a view to integrate them and exploit their complementary nature. We propose a Markovian framework referred to as Dynamic Ordinal Markov Model (DOMM) that makes use of both absolute and relative ordinal information, to improve speech based ordinal emotion prediction. Finally, the proposed framework is validated on two speech corpora commonly used in affective computing, the RECOLA and the IEMOCAP databases, across a range of system configurations. The results consistently indicate that integrating relative ordinal information improves absolute ordinal emotion prediction. Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | A Novel Sequential Monte Carlo Framework for Predicting Ambiguous Emotion StatesabstractWhen continuous emotion labelling of natural (non-acted) data is desired, it is typically collected from multiple annotators. However, most automatic emotion recognition systems trained on such data ignore disagreement between annotators and only models the average rating, despite the observation that the degree of disagreement would reflect the ambiguity and subtlety in every expression of emotions. In this paper, we propose a novel Sequential Monte Carlo framework that models the perceived emotion as time-varying distributions that allows for ambiguity to be incorporated. Additionally, we present alternative measures that consider both the similarity of prediction to the multiple labels, as well as whether the degree of ambiguity in the prediction and labels. The proposed system was validated on the publicly available RECOLA dataset. Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2021 | AusKidTalk: An Auditory-Visual Corpus of 3- to 12-Year-Old Australian Children's SpeechabstractHere we present AusKidTalk [1], an audio-visual (AV) corpus of Australian children’s speech collected to facilitate the development of speech based technological solutions for children. It builds upon the technology and expertise developed through the collection of an earlier corpus of Australian adult speech, AusTalk [2,3]. This multi-site initiative was established to remedy the dire shortage of children’s speech corpora in Australia and around the world that are sufficiently sized to train accurate automated speech processing tools for children. We are collecting ~600 hours of speech from children aged 3–12 years that includes single word and sentence productions as well as narrative and emotional speech. In this paper, we discuss the key requirements for AusKidTalk and how we designed the recording setup and protocol to meet them. We also discuss key findings from our feasibility study of the recording protocol, recording tools, and user interface. Beena Ahmed, Kirrie J. Ballard, Denis Burnham, Tharmakulasingam Sirojan, Hadi Mehmood, Dominique Estival, Elise Baker, Felicity Cox, Joanne Arciuli, Titia Benders, Katherine Demuth, Barbara Kelly, Chloé Diskin-Holdaway, Mostafa Shahin, Vidhyasaharan Sethu, Julien Epps, Chwee Beng Lee, Eliathamby Ambikairajah |
Interspeech | 15 |
| 2021 | Parametric Distributions to Model Numerical Emotion Labels
Deboshree Bose, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
Interspeech | 2 |
| 2021 | An adaptive transmission line cochlear model based front-end for replay attack detection
Tharshini Gunendradasan, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001 |
Speech Commun. | 4 |
| 2021 | Compensation Techniques for Speaker Variability in Continuous Emotion PredictionabstractContinuous time-varying prediction of emotions based on speech in terms of attributes (i.e., arousal) has received considerable attention in the past few years. However, the variability introduced by factors not related to emotion, such as speaker and phonetic variability, which in turn may lead to less reliable models and less accurate emotion predictions, has not been fully explored yet. In particular, even though speaker variability has been shown to be a significant confounding factor in continuous emotion prediction systems, there remains a paucity of analyses about how speaker variability affects continuous emotion prediction systems and which methods can be applied to compensate for this variability. This paper first formulates speaker variability systematically in terms of probability distributions in both feature and model spaces, and quantifies the effect of speaker variability by comparing inter- and intra-speaker variability between speaker-dependent models. Second, two compensation techniques based on partial least squares dimensional reduction and feature mapping are proposed. Finally, the effectiveness of the proposed techniques is validated on three databases, across which they show consistent improvement in arousal, valence and dominance prediction. Additional quantitative analyse reveals that the two proposed techniques compensate for speaker variability in both the feature and model spaces simultaneously. Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Cochlear Signal Processing: A Platform for Learning the Fundamentals of Digital Signal ProcessingabstractThe first digital signal processing course in most electrical engineering programmes around the world tends to be a significant jump in abstraction for most students. This is a consequence of them being introduced to a large number of mathematical concepts with insufficient time to consolidate the ideas with practical examples. In this paper we propose that cochlear signal processing is an excellent platform that brings together almost all the fundamental DSP concepts taught in an introductory course and also lends itself as a suitable candidate for project-based learning. The paper provides details on how such a project was setup and run in the introductory 3rdyear DSP course at UNSW Sydney. In addition, we also provide measure of student interest and engagement with the project as well as their feedback on its impact on their learning patterns and effectiveness at reinforcing theoretical concepts. Eliathamby Ambikairajah, Vidhyasaharan Sethu |
ICASSP | 2 |
| 2020 | Adversarial Multi-Task Learning for Speaker Normalization in Replay DetectionabstractSpoofing detection algorithms in voice biometrics are adversely affected by differences in the speech characteristics of the various target users. In this paper, we propose a novel speaker normalisation technique that employs adversarial multi-task learning to compensate for this speaker variability. The proposed system is designed to learn a feature space that discriminates between genuine and replayed speech while simultaneously reduces the discrimination between different speakers. We initially characterise the impact of speaker variability and quantify the effect of the proposed speaker normalisation technique directly on the feature distributions. Following this, we validate the technique on spoofing detection experiments carried out on two different corpora, ASVSpoof 2017 v2.0 and BTAS 2016 replay, and demonstrate its effectiveness. We obtain EER of 7.11% and 0.83% on the two corpora respectively, lower than that of all relevant baselines. Gajan Suthokumar, Vidhyasaharan Sethu, Kaavya Sriskandaraja, Eliathamby Ambikairajah |
ICASSP | 2 |
| 2020 | Generalized Two-Stage Rank Regression Framework for Depression Score Prediction from SpeechabstractThis paper introduces a novel speech-based depression score prediction paradigm, the 2-stage ranking prediction framework, and highlights the benefits it brings to depression prediction. Conventional regression approaches aim to discern a single functional relationship between speech features and depression scores, making an implicit assumption about the existence of a single fixed relationship between the features and scores. However, as the relationship between severity of depression and the clinical score may vary over the range of the assessment scale, this style of analysis may not be suited to depression prediction. The proposed framework on the other hand, imposes a series of partitions on the feature space, with each partition corresponding to a distinct predefined range of depression scores, and predicts the score based on measures of membership to each partition. This approach provides additional flexibility by allowing different rankings to be learnt for different depression scores, and relaxes assumptions made by conventional regression approaches. Results demonstrate the framework's suitability for depression score prediction: different 2-stage implementations, based on heterogeneous feature extraction and modelling approaches, produce state-of-the-art results on the AVEC-2013 dataset. It is also demonstrated that, unlike fusion of conventional regression systems, the fusion of two-stage systems consistently improves prediction performance. Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, James R. Williamson, Thomas F. Quatieri, Jarek Krajewski |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | Using Gaussian Processes with LSTM Neural Networks to Predict Continuous-Time, Dimensional Emotion in Ambiguous SpeechabstractIn continuous emotion recognition (CER) applications, commonly used models of the emotional content of speech cannot represent some aspects of the behavior of the dimensional emotion values that are the targets of prediction, such as ambiguity, as these models include only a single value for each point in time for each emotional dimension. In this paper, we first propose a model for the emotional content of speech that uses a Gaussian process (GP) to define a distribution that incorporates the inherent ambiguity of emotional speech. Then, we propose a predictive CER system which combines this model alongside LSTM neural network techniques that have that have previously been shown to perform well on this task. When tested on a practical CER task using the RECOLA dataset, this combined LSTM-GP approach is shown to achieve similar or higher performance to a LSTM neural network on its own when measured on mean-only terms. When the variance of the distribution is also incorporated into the performance measure, the LSTM-GP system, which predicts a distribution over the continuous-time outputs rather than just a time series of predicted values, is shown to be able to more realistically model the underlying ambiguity of the emotional content of the speech recordings. Mia Atcheson, Vidhyasaharan Sethu, Julien Epps |
ACII | 2 |
| 2019 | A Novel Bag-of-Optimised-Clusters Front-End for Speech based Continuous Emotion PredictionabstractAlmost all current speech based emotion prediction systems employ front-ends that approximately represent the distribution of frame based features over a suitable window, typically via a set of statistical functionals or the use of Bag-of-Audio-Words (BoAW) features. These front-ends are designed either by manual selection of appropriate statistical functionals, via a feature selection approach or unsupervised clustering of the frame-based features and may not be optimal for the task at hand. This paper proposes a novel front-end that discriminatively learns feature clusters to generate a codebook optimised for emotion prediction, which is then used to generate a Bag-of-Optimised-Clusters (BoOC) feature set. Moreover, this front-end is implemented as a sequence of neural network layers that allow both the proposed front-end and a suitable deep learning backend to be jointly trained. The Bag-of-Optimised-Clusters frontend is tested on the RECOLA database and results show that it outperforms the well-established BoAW features. Deboshree Bose, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Sarith Fernando |
ACII | 3 |
| 2019 | Phoneme Specific Modelling and Scoring Techniques for Anti Spoofing SystemabstractReplay attack refers to the use of recorded speech in an attempt to spoof an automatic speaker verification system and the development of countermeasures that can detect these attacks is an active area of research. This paper investigates the effect of phoneme specific information on replay attack detection. It then develops a replay detection system that employs phoneme specific genuine and spoof models and compares novel scoring methods that take into account phonetic information obtained from a suitable phoneme recogniser. Experiment result on the ASVSpoof 2017 V2.0 corpus indicated that replayed speech may be easier to detect from speech corresponding to some phonemes compared to others and consequently judicious use of phoneme specific models can improve replay detection systems. Gajan Suthokumar, Kaavya Sriskandaraja, Vidhyasaharan Sethu, Chamith Wijenayake, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2019 | Auditory Inspired Spatial Differentiation for Replay Spoofing Attack DetectionabstractThe security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. This paper explores the effect of `spatial differentiation' used in auditory system modelling to improve frequency selectivity and hence provide a more selective front-end for replay attack detection. Experiments were done using a parallel filter bank consisting of simple 2ndorder IIR bandpass filters following which, processing analogous to spatial differentiation was employed to obtain higher order stable IIR filters, in turn leading to highly selective filter banks. Two novel features based on spatially differentiated higher order filter bank have been proposed. Together they yield a relative improvement of 29.9% in replay speech detection over a constant Q transform based baseline system, when evaluated on the ASVspoof 2017 Version 2.0 database. Buddhi Wickramasinghe, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001 |
ICASSP | 4 |
| 2019 | Speech Based Emotion Prediction: Can a Linear Model Work?
Anda Ouyang, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2019 | Estimating cognitive load from speech gathered in a complex real-life training exercise
Maria Vukovic, Vidhyasaharan Sethu, Jessica Parker, Lawrence Cavedon, Margaret Lech, John Thangarajah |
Int. J. Hum. Comput. Stud. | 2 |
| 2018 | Dynamic Multi-Rater Gaussian Mixture Regression Incorporating Temporal Dependencies of Emotion Uncertainty Using Kalman FiltersabstractPredicting continuous emotion in terms of affective attributes has mainly been focused on hard labels, which ignored the ambiguity of recognizing certain emotions. This ambiguity may result in high inter-rater variability and in turn causes varying prediction uncertainty with time. Based on the assumption that temporal dependencies occur in the evolution of emotion uncertainty, this paper proposes a dynamic multi-rater Gaussian Mixture Regression (GMR), aiming to obtain the emotion uncertainty prediction reflected by multi-raters by taking into account their temporal dependencies. This framework is achieved by incorporating feedforward and backward Kalman filters into GMR to estimate the time-dependent label distribution that reflects the emotion uncertainty. It also provides the benefits of relaxing the label distribution of Gaussian assumption to that of a Gaussian Mixture Model (GMM). In addition, a new measurement to estimate emotion uncertainty from GMM as the local variability is adopted. Experiments conducted on the RECOLA database reveal that incorporating temporal dependencies is critical for emotion uncertainty prediction with 17% relative improvement for arousal, and that the proposed framework for emotion uncertainty prediction shows potential in conventional emotion attribute prediction. Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 2 |
| 2018 | Factorized Hidden Variability Learning for Adaptation of Short Duration Language Identification ModelsabstractBidirectional long short term memory (BLSTM) recurrent neural networks (RNNs) have recently outperformed other state-of-the-art approaches, such as i-vector and deep neural networks (DNNs) in automatic language identification (LID), particularly when testing with very short utterances (`3s). Mismatches conditions between training and test data, e.g. speaker, channel, duration and environmental noise, are a major source of performance degradation for LID. A factorized hidden variability subspace (FHVS) learning technique is proposed for the adaptation of BLSTM RNNs to compensate for these types of mismatches in recording conditions. In the proposed approach, condition dependent parameters are estimated to adapt the hidden layer weights of the BLSTM in the FHVS. We evaluate FHVS on the AP17-OLR data set. Experimental results show that the FHVS method outperforms the standard BLSTM approach, achieving 27% relative improvements with utterance-level adaptation over the standard BLSTM for 1 s duration utterances. Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 2 |
| 2018 | End-to-End Hierarchical Language Identification SystemabstractRecently, hierarchical language identification systems have shown significant improvement over single level systems in both closed and open set language identification tasks. However, developing such a system requires the features and classifier selection at each node in the hierarchical structure to be hand crafted. Motivated by the superior ability of end-to-end deep neural network architecture to jointly optimize the feature extraction and classification process, we propose a novel approach developing an end-to-end hierarchical language identification system. The proposed approach also demonstrates the in -built ability of the end-to-end hierarchical structure training that enables an out-of-set language model, without using any additional out-of-set language training data. Experiments are conducted on the NIST LRE 2015 data set. The overall results show relative improvements of 18.6% and 27.3% in terms of Cavgin closed and open set tasks over the corresponding baseline systems. Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
ICASSP | 2 |
| 2018 | Speaker-Phonetic Vector Estimation for Short Duration Speaker VerificationabstractPhonetic variability is one of the primary challenges in short duration speaker verification. This paper proposes a novel method that modifies the standard normal distribution prior in the total variability model to use a mixture of Gaussians as the prior distribution. The proposed speaker-phonetic vectors are then estimated from the posterior probability of latent variables, and each vector has a phonetic meaning. Unlike the standard total variability model, the proposed method can incorporate a phoneme classifier to perform soft content matching, which has the potential to solve the phonetic variability problem. Parameter estimation and scoring formulae for speaker-phonetic vectors method are presented. Experimental results obtained using NIST 2010 data show that the proposed technique leads to relative improvements of more than 30% when fused with total variability model and tested on 3 second duration test files. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
ICASSP | 2 |
| 2018 | Demonstrating and Modelling Systematic Time-varying Annotator Disagreement in Continuous Emotion Annotation
Mia Atcheson, Vidhyasaharan Sethu, Julien Epps |
INTERSPEECH | 2 |
| 2018 | Sub-band Envelope Features Using Frequency Domain Linear Prediction for Short Duration Language Identification
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2018 | Deep Siamese Architecture Based Replay Detection for Secure Voice Biometric
Kaavya Sriskandaraja, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2018 | Modulation Dynamic Features for the Detection of Replay AttacksabstractThe development of automatic systems that can detect replayed speech has emerged as a significant research challenge for securing voice biometric systems and is the focus of this paper. Specifically, this paper proposes two novel features to capture the static and dynamic characteristics of the signal from the modulation spectrum, which complement short term spectral features for use in replay detection. The modulation spectral centroid frequency feature is proposed as a vector representation of the first order spectral moments of the modulation spectrum. In conjunction to this, the long term spectral average serves to capture the static characteristics of the modulation spectrum. The proposed system, employing a GMM back-end, was evaluated on the ASVSpoof 2017 dataset and found to yield an EER of 6.54%. Gajan Suthokumar, Vidhyasaharan Sethu, Chamith Wijenayake, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2018 | Using language cluster models in hierarchical language identification
Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
Speech Commun. | 2 |
| 2018 | Generalized Variability Model for Speaker VerificationabstractIn this letter, we propose a generalized variability model as an extension to the total variability model. While the total variability model employs a standard normal prior distribution in its typical setup, the proposed generalized variability model relaxes this assumption and allows the latent variable distribution to be a mixture of Gaussians. The conventional total variability model can then be viewed as a special case of this generalized version where the number of mixture components is constrained to one. This proposed model is validated in the context of speaker verification tasks on both the standard and extended NIST SRE 2010 datasets. Experimental results show that modeling the distribution of the latent variables as a mixture of Gaussians leads to a better performance under all conditions and a greater gain can be expected for speaker verification using short utterances. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
IEEE Signal Process. Lett. | 2 |
| 2017 | Modeling variable length phoneme sequences - A step towards linguistic information for speech emotion recognition in wider worldabstractVocal gestures play an important role in emotion expression and can be used by speech based emotion recognition systems. This paper proposes the use of BLSTM neural networks to model salient variable length phoneme sequences, which in turn can represent relevant vocal gestures. Unlike existing techniques, the proposed approach is not restricted to modelling phoneme sequences of a fixed length and both salience and optimal modelling length of phoneme sequences are learnt from the training data. Three possible phoneme representations that can be modelled by BLSTMs are compared and experimental results suggest that sequences of Phone Log Likelihood Ratios are more representative of emotions when compared to sequences of phoneme labels represented as one — hot vectors. On the IEMOCAP database, the proposed approach achieves an Unweighted Average Recall (UAR) of 56.4%, an improvement of 6.5% in absolute terms over the previous approach of modelling fixed length phoneme sequences on a 4-class classification problem. The proposed linguistic system is complementary to acoustic features with a fused system leading to an absolute improvement of 5% to the UAR. Kalani Wataraka Gamage, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ACII | 2 |
| 2017 | Salience based lexical features for emotion recognitionabstractIn this paper we focus on the usefulness of verbal events for speech based emotion recognition. In particular, the use of phoneme sequences to encode verbal cues related to the expression of emotions is proposed and lexical features based on these phoneme sequences are introduced for use in automatic emotion recognition systems where manual transcripts are not available. Secondly, a novel estimate of emotional salience of verbal cues, applicable to both phoneme sequences and words, is presented. Experimental results on the IEMOCAP database show that the proposed automatic phoneme sequence based features can achieve an Unweighted Average Recall (UAR) of 49% with proposed salience measure. Further, the proposed salience measure can lead to an UAR of 64% when using manual word transcriptions. Both of these are the highest UARs reported on the IEMOCAP database for systems using lexical features extracted from automatic and manual transcripts respectively. Kalani Wataraka Gamage, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 2 |
| 2017 | An Investigation of Emotion Prediction Uncertainty Using Gaussian Mixture Regression
Ting Dang, Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2017 | Bidirectional Modelling for Short Duration Language Identification
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2017 | Investigating Scalability in Hierarchical Language Identification System
Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 53 |
| 2017 | Incorporating Local Acoustic Variability Information into Short Duration Speaker Verification
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
INTERSPEECH | 2 |
| 2017 | Independent Modelling of High and Low Energy Speech Frames for Spoofing DetectionabstractSpoofing detection systems for automatic speaker verification have moved from only modelling voiced frames to modelling all speech frames. Unvoiced speech has been shown to carry information about spoofing attacks and anti-spoofing systems may further benefit by treating voiced and unvoiced speech differently. In this paper, we separate speech into low and high energy frames and independently model the distributions of both to form two spoofing detection systems that are then fused at the score level. Experiments conducted on the ASVspoof 2015, BTAS 2016 and Spoofing and Anti-Spoofing (SAS) corpora demonstrate that the proposed approach of fusing two independent high and low energy spoofing detection systems consistently outperforms the standard approach that does not distinguish between high and low energy frames. Gajan Suthokumar, Kaavya Sriskandaraja, Vidhyasaharan Sethu, Chamith Wijenayake, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2016 | A hierarchical framework for language identificationabstractMost current language recognition systems model different levels of information such as acoustic, prosodic, phonotactic, etc. independently and combine the model likelihoods in order to make a decision. However, these are single level systems that treat all languages identically and hence incapable of exploiting any similarities that may exist within groups of languages. In this paper, a hierarchical language identification (HLID) framework is proposed that involves a series of classification decisions at multiple levels involving language clusters of decreasing sizes with individual languages identified only at the final level. The performance of proposed hierarchical framework is compared with a state-of-the-art LID system on the NIST 2007 database and the results indicate that the proposed approach outperforms state-of-the-art systems. Saad Irtza, Vidhyasaharan Sethu, Haris Bavattichalil, Eliathamby Ambikairajah, Haizhou Li 0001 |
ICASSP | 2 |
| 2016 | Factor Analysis Based Speaker Normalisation for Continuous Emotion Prediction
Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2016 | A Feature Normalisation Technique for PLLR Based Language Identification Systems
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2016 | Out of Set Language Modelling in Hierarchical Language Identification
Saad Irtza, Vidhyasaharan Sethu, Sarith Fernando, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2016 | Parallel Speaker and Content Modelling for Text-Dependent Speaker Verification
Saad Irtza, Kaavya Sriskandaraja, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2016 | Twin Model G-PLDA for Duration Mismatch Compensation in Text-Independent Speaker Verification
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
INTERSPEECH | 2 |
| 2016 | Investigation of Sub-Band Discriminative Information Between Spoofed and Genuine Speech
Kaavya Sriskandaraja, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2015 | Weighted pairwise Gaussian likelihood regression for depression score predictionabstractThis paper presents a technique in which feature vectors are mapped onto ordinal ranges of clinical depression scores using weighted pairwise Gaussians. The position of a test vector with respect to these partitions is used to perform depression score prediction. Results found on a set of spectral and formant based speech characteristics indicate the potential of this technique for performing depression score prediction. Key results on the AVEC 2013 development set indicate that the inclusion of weights and Bayesian adaptation improves system performance by 16.5% - 18.5% when compared to using an unweighted non-adapted system. Fusing results from Bayesian adapted models corresponding to different feature spaces offers up to 8% further improvement. Further, fusion consistently improves performance on both the AVEC 2013 development and test set, in contrast to conventional regressor fusion. Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Jarek Krajewski |
ICASSP | 3 |
| 2015 | Relevance vector machine for depression prediction
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Jarek Krajewski |
INTERSPEECH | 2 |
| 2015 | Phonemes frequency based PLLR dimensionality reduction for language recognition
Saad Irtza, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2015 | A model based voice activity detector for noisy environments
Kaavya Sriskandaraja, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2015 | Analysis of acoustic space variability in speech affected by depression
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Sebastian Schnieder, Jarek Krajewski |
Speech Commun. | 2 |
| 2014 | Variability compensation in small data: Oversampled extraction of i-vectors for the classification of depressed speechabstractVariations in the acoustic space due to changes in speaker mental state are potentially overshadowed by variability due to speaker identity and phonetic content. Using the Audio/Visual Emotion Challenge and Workshop 2013 Depression Dataset we explore the suitability of i-vectors for reducing these latter sources of variability for distinguishing between low or high levels of speaker depression. In addition we investigate whether supervised variability compensation methods such as Linear Discriminant Analysis (LDA), and Within Class Covariance Normalisation (WCCN), applied in the i-vector domain, could be used to compensate for speaker and phonetic variability. Classification results show that i-vectors formed using an over-sampling methodology outperform a baseline set by KL-means supervectors. However the effect of these two compensation methods does not appear to improve system accuracy. Visualisations afforded by the t-Distributed Stochastic Neighbour Embedding (t-SNE) technique suggest that despite the application of these techniques, speaker variability is still a strong confounding effect. Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Jarek Krajewski |
ICASSP | 3 |
| 2014 | Probabilistic acoustic volume analysis for speech affected by depressionabstractAlterations in speech motor control in depressed individuals have been found to manifest as a reduction in spectral variability. In this paper we present a novel method for measuring acoustic volume a model-based measure that is reflective of this decrease in spectral variability and assess the ability of features resulting from this measure for indexing a speaker’s level of depression. A Monte Carlo approximation that enables the computation of this measure is also outlined in this paper. Results found using the AVEC 2013 Challenge Dataset indicate there is a statistically significant reduction in acoustic variation with increasing levels of speaker depression, and using features designed to capture this change it is possible to outperform a range of conventional spectral measures when predicting a speaker’s level of depression. Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Jarek Krajewski |
INTERSPEECH | 2 |
| 2014 | The UNSW submission to INTERSPEECH 2014 compare cognitive load challengeabstractSpeech based cognitive load estimation is a new field of research. Due to this relative ‘lack of maturity’, a single best approach to building cognitive load estimation systems has not been established yet. The primary aim of this submission is to report the performance of various basic utterance level classification frameworks developed using important elements of state-of-the-art speaker recognition systems. This may lead to a suitable basis for future cognitive load estimation systems. As a consequence of being a part of a challenge, it is expected that these frameworks will be compared to a much larger number of alternative approaches than what would otherwise be possible. In keeping with this focused aim, the GMM supervector approaches along with some variants are utilised. The systems outlined in this paper include a frame-level MFCC-GMM system along with utterance level GMMsupervector-SVM, GMM-ivector-SVM and GMM-JFA-SVM systems. The best combined system has an accuracy (UAR) of 66.6% as evaluated on the challenge development set and 63.7% as evaluated on the test set. Jia Min Karen Kua, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2013 | Speaker variability in speech based emotion models - Analysis and normalisationabstractAll features commonly utilised in speech based emotion classification systems capture both emotion-specific information and speaker-specific information. This paper proposes a novel method to gauge the effect of speaker-specific information on emotion modelling based on two measures: a Monte Carlo approximation to KL divergence and an estimate of feature variability based on diagonal covariance matrices. In addition, a novel speaker normalisation technique based on joint factor analysis is also proposed. This method is analogous to channel compensation in speaker verification systems, with one significant extension. The model domain compensation is mapped back to frame-level features, allowing for use in a wider range of emotion classification frameworks and in conjuncture with other normalisation techniques. Preliminary evaluations on the IEMOCAP database suggests that the proposed technique improves the performance of GMM based classification systems based on widely employed features such as pitch, MFCCs and deltas. Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah |
ICASSP | 1 |
| 2013 | Modeling spectral variability for the classification of depressed speechabstractQuantifying how the spectral content of speech relates to changes in mental state may be crucial in building an objective speech-based depression classification system with clinical utility. This paper investigates the hypothesis that important depression based information can be captured within the covariance structure of a Gaussian Mixture Model (GMM) of recorded speech. Significant negative correlations found between a speaker’s average weighted variance- a GMM-based indicator of speaker variability- and their level of depression support this hypothesis. Further evidence is provided by the comparison of classification accuracies from seven different GMM-UBM systems, each formed by varying different parameter combinations during MAP adaption. This analysis shows that variance-only adaptation either outperforms or matches the de facto standard mean-only adaptation when classifying both the presence and severity of depression. This result is perhaps the first of its kind seen in GMM-UBM speech classification. Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Michael Breakspear, Roland Göcke |
INTERSPEECH | 3 |
| 2013 | GMM based speaker variability compensated system for interspeech 2013 compare emotion challengeabstractThis paper describes the University of New South Wales system for the Interspeech 2013 ComParE emotion subchallenge. The primary aim of the submission is to explore the performance of model based variability compensation techniques applied to emotion classification and as a consequence of being a part of a challenge, to enable a comparison of these methods to alternative approaches. In keeping with this focused aim, a simple frame based front-end of MFCC and ΔMFCC is utilised. The systems outlined in this paper consists of a joint factor analysis based system and one based on a library of speaker-specific emotion models along with a basic GMM based system. The best combined system has an accuracy (UAR) of 47.8% as evaluated on the challenge development set and 35.7% as evaluated on the test set. Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2012 | Speaker variability in emotion recognition - an adaptation based approachabstractNone of the features commonly utilised in automatic emotion classification systems completely disassociate emotion-specific information from speaker-specific information. Consequently, this speaker-specific variability adversely affects the performance of the emotion classification system and in existing systems is frequently mitigated by some form of speaker normalisation. Speaker adaptation offers an alternative to normalisation and this paper proposes a novel bootstrapping technique which involves selecting appropriate initial models from a large training pool, prior to speaker adaptation of emotion models in the context of GMM based emotion classification as an alternative to speaker normalisation. Evaluations on the LDC Emotional Prosody and the FAU Aibo corpora reveal that an emotion classification system based on the proposed bootstrapping method outperforms systems based on speaker normalisation as long as a small amount of labelled adaptation data is available. It also outperforms speaker adaption from common initial models estimated from all training speakers. Ni Ding, Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah |
ICASSP | 2 |
| 2011 | Investigation of spectral centroid features for cognitive load classification
Phu Ngoc Le, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Eric H. C. Choi |
Speech Commun. | 4 |
| 2009 | Speaker dependency of spectral features and speech production cues for automatic emotion classificationabstractSpectral and excitation features, commonly used in automatic emotion classification systems, parameterise different aspects of the speech signal. This paper groups these features as speech production cues, broad spectral measures and detailed spectral measures and looks at how they differ in their performance in both speaker dependent and speaker independent systems. The extent of speaker normalisation on these features is also considered. Combinations of different features are then compared in terms of classification accuracies. Evaluations were conducted on the LDC emotional speech corpus for a five-class problem. Results indicate that MFCCs are very discriminative but suffer from speaker variability. Further, results suggest that the best front end for a speaker independent system is a combination of pitch, energy and formant information. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
ICASSP | 1 |
| 2009 | Pitch contour parameterisation based on linear stylisation for emotion recognitionabstractThe pitch contour contains information that characterises the emotion being expressed by speech, and consequently features extracted from pitch form an integral part of many automatic emotion recognition systems. While pitch contours may have many small variations and hence are difficult to represent compactly, it may be possible to parameterise them by approximating the contour for each voiced segment by a straight line. This paper looks at such a parameterisation method in the context of emotion recognition. Listening tests were performed to subjectively determine if the linearly stylised contours were able to sufficiently capture information pertaining to emotions expressed in speech. Furthermore these parameters were used as features for an automatic 5-class emotion classification system. The use of the proposed parameters rather than pitch statistics resulted in a relative increase in accuracy of about 20%. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 1 |
| 2008 | Empirical mode decomposition based weighted frequency feature for speech-based emotion classificationabstractThis paper focuses on speech based emotion classification utilizing acoustic data. The most commonly used acoustic features are pitch and energy, along with prosodic information like rate of speech. We propose the use of a novel feature based on instantaneous frequency obtained from the speech, in addition to the aforementioned features, in order to take into account the vocal tract parameters as well as vocal chord excitation. The proposed features employ the recently emerged empirical mode decomposition to decompose speech into AM-FM signals that are symmetric about zero and suitable for Hilbert transformation to extract the instantaneous frequency. The proposed features provide a relative increase in classification accuracy of approximately 9% when appended to established acoustic features. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
ICASSP | 1 |
| 2008 | Phonetic and speaker variations in automatic emotion classificationabstractThe speech signal contains information that characterises the speaker and the phonetic content, together with the emotion being expressed. This paper looks at the effect of this speakerand phoneme-specific information on speech-based automatic emotion classification. The performances of a classification system using established acoustic and prosodic features for different phonemes are compared, in both speaker-dependent and speaker-independent modes, using the LDC Emotional Prosody speech corpus. Results from these evaluations indicate that speaker variability is more significant than phonetic variations. They also suggest that some phonemes are easier to classify than others. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 1 |
| 2007 | Group delay features for emotion detectionabstractThis paper focuses on speech based emotion classification utilizing acoustic data. The most commonly used acoustic features are pitch and energy, along with prosodic information like the rate of speech. We propose the use of a novel feature based on the phase response of an all-pole model of the vocal tract obtained from linear predictive coefficients (LPC), in addition to the aforementioned features. We compare this feature to other commonly used acoustic features based on classification accuracy. The back-end of our system employs a probabilistic neural network based classifier. Evaluations conducted on the LDC Emotional Prosody speech corpus indicate the proposed features are well suited to the task of emotion classification. The proposed features are able to provide a relative increase in classification accuracy of about 14% over established features when combined with them to form a larger feature vector. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 1 |
| 2007 | A Novel Technique for Noise Reduction in InSAR ImagesabstractThis letter proposes a new technique for noise reduction applied to synthetic aperture radar interferometry. This technique involves a nonlinear filter that separates the interferogram into two components: one containing the smooth (low frequency) part and the other containing the detail (high frequency) part. The smooth part is obtained using a combination of a median filter and a smoothing filter. The detail component is obtained by subtracting the smooth component from the original signal. This detail component is filtered to remove noise and then added to the smooth component to generate the final output. Both simulated and real data are used to evaluate the performance of the proposed technique under different conditions. The experimental results show that the proposed technique outperforms most commonly used interferometric phase filters Vidhyasaharan Sethu, Eliathamby Ambikairajah, Linlin Ge |
IEEE Geosci. Remote. Sens. Lett. | 2 |