Julien Epps

dblp:92/6424 · DBLP profile ↗
← Back
138ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0001-6624-5551ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 98 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 71 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 15 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3
YearPublicationVenuePosition
2026 Automatic speech-based alcohol intoxication detection for automotive safety applications
abstract
• This research is the first to report automatic alcohol detection classification using omni-microphone speaker recordings from the Alcohol Language Corpus (ALC), which less invasive and more natural than ALC headset speaker microphone recordings typically reported on in previous publications. • Five different balanced-class BAC ranges (e.g., <0.000, 0.001-0.049, 0.050-0.095, 0.100-0.149, >0.150) were evaluated to ascertain the intoxication detection robustness various speech-based feature and speech task types per BAC range. • A binary class classification ('non-intoxicated' versus 'intoxicated') approach using task-dependent recordings demonstrated accuracy improvements over the traditional task-agnostic approach; and further, cross-fold leave-one-out classification results show that task-dependent majority-vote method provided added classification performance gains. • The task-dependent majority-vote method using a relatively small number of features (320) outperformed previous published ALC automatic alcohol intoxication baselines that utilized higher-quality headset mic recordings, a higher BAC 'intoxication' class threshold (>0.05), and relied on a larger number of features (i.e., up to 4k). There is a responsibility to advance automatic alcohol intoxication screening capabilities in modern automobiles to reduce the high rate of alcohol-related accidents and fatalities worldwide. Automatic speech-based alcohol intoxication screening offers a tremendous safety opportunity in the automotive industry due to its non-invasive convenience, comparatively inexpensive cost, and rapid result processing. Using the Alcohol Language Corpus (ALC), this study examines automatic alcohol intoxication classification based on participants’ non-intoxicated/intoxicated omni-microphone speech recordings. Experimentation of many different speech features (e.g., glottal, landmarks, linguistic, prosodic, spectral, syllabic, vocal tract coordination) across different blood alcohol concentration (BAC) ranges and specific verbal tasks show significant changes as participants' BAC increases. Intoxicated participants produce lower average fundamental frequency (F0) with an increase in F0 frequency modulation, breathiness and creakiness voice qualities in intoxicated recordings when compared to their non-intoxicated recordings. For the picture description and tongue twister tasks, manual irregularity disfluency and pause linguistic features significantly increase in intoxicated recordings. Further, for all verbal tasks, automatically extracted syllabic pause features show a significant increase in intoxicated recordings. Implementation of task-dependent support vector machine classifier model with a ≥0.001 BAC 'intoxication' sensitivity threshold increases alcohol classification by up to 8% absolute gain over a task-agnostic approach. Moreover, intoxication classification results demonstrate that task-dependent modeling with majority vote decision improves classification accuracy with up to 20% absolute gain depending on task when compared to file-by-file task-agnostic method results reported previously in ALC baseline studies that used higher quality headset microphone recordings.
Brian Stasak, Julien Epps
Comput. Speech Lang.2
2026 The eyes show it: Exploring eye behavior and impact on human mental state during task interruption
abstract
• Both gaze and paragaze can indicate task transitions due to interruption or not • Both gaze and paragaze can reflect task change due to interruption • Only strong emotional interruption affects eye behaviors afterwards • Task type significantly affects the time to attend an interruption • Unproportionally increased completion time for low and high cognitive load tasks One technology adoption outcome is ever more frequent task interruptions, which makes interruption management a central human computer interface design problem. Studies on interruption often focused on the link between interruptions and genesis of error and eye gaze to validate associated task cues engagement. How eye behaviors reflect load type and level changes between primary and interrupting tasks and how emotional interruptions impact mental state in primary tasks – factors that contribute to cognitive processing – remain unanswered. In this study, we analyzed performance metrics and eye behaviors from 18 participants while they completed an uninterrupted and interrupted roleplay task. Our results reveal that although interruption notification could affect the time to attend to an interruption, high cognitive load resulted in a detrimental impact on completion time and low cognitive load led to time reset for primary tasks. No existing evidence suggests that affect type influences the time metrics, however, high arousal and valence induced by interruptions altered eye behavior upon returning to the primary task, indicating less visual information taken and under a high load level, which showed evidence of the impact of affective interruption for the first time. Noteworthily, the eye behaved differently at interruption beginning compared with the end and uninterrupted task transitions, indicating the feasibility of including eye behavior for interruption estimation. Overall, these results are the first experimental evidence for theories that interruption can cause high mental load, posit extra mental effort, forgetting primary tasks and behaviors being directed by the current most active goal.
Siyuan Chen 0002, Julien Epps
Int. J. Hum. Comput. Stud.2
2025 Rethinking Mamba in Speech Processing by Self-Supervised Models
abstract
The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model’s performance varies across different tasks. For instance, in tasks such as speech enhancement and spectrum reconstruction, the Mamba model performs well when used independently. However, for tasks like speech recognition, additional modules are required to surpass the performance of attention-based models. We propose the hypothesis that the Mamba-based model excels in "reconstruction" tasks within speech processing. However, for "classification tasks" such as Speech Recognition, additional modules are necessary to accomplish the "reconstruction" step. To validate our hypothesis, we analyze the previous Mamba-based Speech Models from an information theory perspective. Furthermore, we leveraged the properties of HuBERT in our study. We trained a Mamba-based HuBERT model, and the mutual information patterns, along with the model’s performance metrics, confirmed our assumptions.
Xiangyu Zhang 0005, Mostafa Shahin, Beena Ahmed, Julien Epps
ICASSP5
2025 Multimodal Task Analysis in Wearable Contexts
Julien Epps
ICMI1
2025 Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
Xiangyu Zhang 0005, Daijiao Liu, Tianyi Xiao, Cihan Xiao, Tünde Szalay, Mostafa Shahin, Beena Ahmed, Julien Epps
INTERSPEECH8
2025 Phonological level wav2vec2-based Mispronunciation Detection and Diagnosis method
abstract
The automatic identification and analysis of pronunciation errors, known as Mispronunciation Detection and Diagnosis (MDD) plays a crucial role in Computer Aided Pronunciation Learning (CAPL) tools such as Second-Language (L2) learning or speech therapy applications. Existing MDD methods relying on analysing phonemes can only detect categorical errors of phonemes that have an adequate amount of training data to be modelled. Due to the unpredictable nature of pronunciation errors made by non-native or disordered speakers and the scarcity of training datasets, it is unfeasible to model all types of mispronunciations. Moreover, phoneme-level MDD approaches can provide only limited diagnostic information about the error made. To address this, in this paper, we propose a low-level MDD approach based on the detection of phonological features. Phonological features break down phoneme production into elementary components that are directly related to the articulatory system leading to more formative feedback for the learner. We further propose a multi-label variant of the Connectionist Temporal Classification (CTC) approach to jointly model the non-mutually exclusive phonological features using a single model. The pre-trained wav2vec2 model was employed as a core model for the phonological feature detector. The proposed method was applied to L2 speech corpora collected from English learners from different native languages. The proposed phonological level MDD method was further compared to the traditional phoneme-level MDD and achieved a significantly lower False Acceptance Rate (FAR), False Rejection Rate (FRR), and Diagnostic Error Rate (DER) over all phonological features compared to the phoneme-level equivalent.
Mostafa Shahin, Julien Epps, Beena Ahmed
Speech Commun.2
2025 Eye Action Units as Combinations of Discrete Eye Behaviors for Wearable Mental State Analysis
abstract
Mental state induced by different task contexts and load levels is of interest for human health and wellness, and eye activity extracted from infrared eye images is well-suited to estimate it. As a useful tool for emotion recognition, facial action units (FAUs) extracted from facial images are well-established, however these are insufficiently detailed for the eye. In this paper, we extract discrete eye behaviors from eyelid, iris and pupil boundaries and propose eye action units (EAUs) based on the eye appearance (behavior), providing a detailed and interpretable representation that shares the advantages of FAUs and describes the wide range of eye states and shapes. Eight volunteers annotated 11 EAUs for 120 eye images, represented by a series of discrete eye behaviors. Analysis shows that the EAUs can be viably characterized by fundamental eye behaviors with moderate to substantial agreement. When evaluating discrete eye behaviors and EAUs for recognition of four mental states and two load levels, the former achieved significantly higher accuracy than conventional features of pupil size change and blink rate, especially using behavior duration features, and EAUs outperformed combinations of discrete eye behaviors in general, implying their utility as an action unit.
Siyuan Chen 0002, Julien Epps
IEEE Trans. Affect. Comput.2
2024 When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
abstract
Depression is a critical concern in global mental health, prompting extensive research into AIbased detection methods.Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications.However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities.Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped.In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection.We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks.By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts.This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals.Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines.In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals.
Xiangyu Zhang 0005, Hexin Liu, Kaishuai Xu, Qiquan Zhang, Daijiao Liu, Beena Ahmed, Julien Epps
EMNLP7
2023 Recognizing Conversational State from the Eye Using Wearable Eyewear
abstract
During a naturalistic conversation, we can easily perceive whether a person is speaking or listening and when they are about to start or finish speaking, which we refer to as conversational states in this paper. Building this ability into a machine can be helpful in a wide variety of human-machine collaboration contexts for effective communication. A wealth of evidence from psychology and neuroscience suggests that the eye is a reliable window into human internal states since it senses the perceived changes in the outside world. In this study, we investigate the relationship between eye behavior and four conversational states for the first time and examine the viability of automatically recognizing the conversational states. The results demonstrate that eye center relative to the head is not a good indicator to distinguish conversational states, but the distribution, frequency and duration of eye states are strongly correlated with conversational states while eyelid is associated with conversational state in certain extent. The accuracy of recognition of the four conversational states using the proposed eye behaviors is well above chance level, and outperforms baselines using pupillary response, pupil center position, and blink rate. This finding suggests that eye state and eyelid shape contains valuable information about conversational states but have been overlooked in previous studies. It is promising to include eye behavior in wearable interfaces for dialogue systems and for social signal processing, with the further advantages of being ‘always on’, less sensitive to luminance variability, and better privacy preserving compared with using facial image and speech data.
Siyuan Chen 0002, Julien Epps
ACII2
2023 DNN controlled adaptive front-end for replay attack detection systems
abstract
Developing robust countermeasures to protect automatic speaker verification systems against replay spoofing attacks is a well-recognized challenge. Current approaches to spoofing detection are generally based on a fixed front-end, typically a time-invariant filter bank, followed by a machine learning back-end. In this paper, we propose a novel approach whereby the front-end comprises an adaptive filter bank with a deep neural network-based controller, which is jointly trained along with a neural network back-end. Specifically, the deep neural network-based adaptive filter controller tunes the selectivity and sensitivity of the front-end filter bank at every frame to capture replay-related artefacts. We demonstrate the effectiveness of the proposed framework in spoofing attack detection on a synthesized dataset and ASVSpoof 2019 and ASVSpoof 2021 challenge datasets in terms of equal error rate and its ability to capture artefacts that differentiate replayed signals from genuine ones in comparison to conventional non-adaptive front-end.
Buddhi Wickramasinghe, Eliathamby Ambikairajah, Vidhyasaharan Sethu, Julien Epps, Haizhou Li 0001, Ting Dang
Speech Commun.4
2023 A High-Quality Landmarked Infrared Eye Video Dataset (IREye4Task): Eye Behaviors, Insights and Benchmarks for Wearable Mental State Analysis
abstract
Sensing the mental state induced by different task contexts, where cognition is a focus, is as important as sensing the affective state where emotion is induced in the foreground of consciousness, because completing tasks is part of every waking moment of life. However, few datasets are publicly available to advance mental state analysis, especially those using the eye as the sensing modality with detailed ground truth for eye behaviors. In this study, we contribute a high-quality publicly accessible eye video dataset, IREye4Task, where the eyelid, pupil and iris boundary are annotated for each frame to obtain eye behaviors as responses to four different task contexts and two load levels of tasks, over more than a million frames. Meanwhile, we propose a series of eye behavior representations to provide insights into how the eye behaves during different mental states. Finally, we benchmark three mental-state recognition tasks for this dataset to demonstrate the effectiveness of the eye behavior representations. This is the first public wearable eye video dataset for mental state analysis with high quality eye landmarks and a variety of mental states, and is the first study analyzing comprehensive eye behaviors far beyond using pupil size and blink in previous studies.
Siyuan Chen 0002, Julien Epps
IEEE Trans. Affect. Comput.2
2023 Ordinal Logistic Regression With Partial Proportional Odds for Depression Prediction
abstract
Like many psychological scales, depression scales are ordinal in nature. Depression prediction from behavioral signals has so far been posed either as classification or regression problems. However, these naive approaches have fundamental issues because they are not focused on ranking, unlike ordinal regression, which is the most appropriate approach. Ordinal regression to date has comparatively few methods when compared with other branches in machine learning, and its usage is limited to specific research domains. Ordinal logistic regression (OLR) is one such method, which is an extension for ordinal data of the well-known logistic regression, but is not familiar in speech processing, affective computing or depression prediction. The primary aim of this article is to investigate proportionality structures and model selection for the design of ordinal regression systems within the logistic regression framework. A new greedy-based algorithm for partial proportional odds model selection (GREP) is proposed that allows the parsimonious design of effective ordinal logistic regression models, which avoids an exhaustive search and outperforms model selection using the Brant test. Evaluations on the DAIC-WOZ and AViD depression corpora show that OLR models exploiting GREP can outperform two competitive baseline systems (GSR and CNN), in terms of both RMSE and Spearman correlation.
Sadari Jayawardena, Julien Epps, Eliathamby Ambikairajah
IEEE Trans. Affect. Comput.2
2022 Investigation of Speech Landmark Patterns for Depression Detection
abstract
The massive and growing burden imposed on modern society by depression has motivated investigations into early detection through automated, scalable and non-invasive methods, including those based on speech. However, speech-based methods that capture articulatory information effectively across different recording devices and in naturalistic environments are still needed. This article proposes two feature sets associated with speech articulation events based on counts and durations of sequential landmark groups orn-grams. Statistical analysis of the duration-based features reveals that durations from several consecutive landmark bigrams and onset-offset landmark pairs are significant in discriminating depressed from non-depressed speakers. In addition to investigating different normalization approaches and values ofnfor landmarkn-gram features, experiments across different elicitation tasks suggest that the features can be tailored to capture different articulatory aspects of depressed voices. Evaluations of both landmark duration features and landmarkn-gram features on the DAIC-WOZ and SH2 datasets show that they are highly effective, either alone or fused, relative to existing approaches.
Zhaocheng Huang, Julien Epps, Dale Joachim
IEEE Trans. Affect. Comput.2
2022 Multi-Task Semi-Supervised Adversarial Autoencoding for Speech Emotion Recognition
abstract
Inspite the emerging importance of Speech Emotion Recognition (SER), the state-of-the-art accuracy is quite low and needs improvement to make commercial applications of SER viable. A key underlying reason for the low accuracy is the scarcity of emotion datasets, which is a challenge for developing any robust machine learning model in general. In this article, we propose a solution to this problem: a multi-task learning framework that uses auxiliary tasks for which data is abundantly available. We show that utilisation of this additional data can improve the primary task of SER for which only limited labelled data is available. In particular, we use gender identifications and speaker recognition as auxiliary tasks, which allow the use of very large datasets, e. g., speaker classification datasets. To maximise the benefit of multi-task learning, we further use an adversarial autoencoder (AAE) within our framework, which has a strong capability to learn powerful and discriminative features. Furthermore, the unsupervised AAE in combination with the supervised classification networks enables semi-supervised learning which incorporates a discriminative component in the AAE unsupervised training pipeline. This semi-supervised learning essentially helps to improve generalisation of our framework and thus leads to improvements in SER performance. The proposed model is rigorously evaluated for categorical and dimensional emotion, and cross-corpus scenarios. Experimental results demonstrate that the proposed model achieves state-of-the-art performance on two publicly available datasets.
Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Julien Epps, Björn W. Schuller
IEEE Trans. Affect. Comput.5
2021 Automatic Elicitation Compliance for Short-Duration Speech Based Depression Detection
abstract
Detecting depression from the voice in naturalistic environments is challenging, particularly for short-duration audio recordings. This enhances the need to interpret and make optimal use of elicited speech. The rapid consonant-vowel syllable combination ‘pataka’ has frequently been selected as a clinical motor-speech task. However, there is significant variability in elicited recordings, which remains to be investigated. In this multi-corpus study of over 25,000 ‘pataka’ utterances, it was discovered that speech landmark- based features were sensitive to the number of ‘pataka’ utterances per recording. This landmark feature sensitivity was newly exploited to automatically estimate ‘pataka’ count and rate, achieving root mean square errors nearly three times lower than chance-level. Leveraging count-rate knowledge of the elicited speech for depression detection, results show that the estimated ‘pataka’ number and rate are important for normalizing evaluative ‘pataka’ speech data. Count and/or rate normalized ‘pataka’ models produced relative reductions in depression classification error of up to 26% compared with non-normalized models.
Brian Stasak, Zhaocheng Huang, Dale Joachim, Julien Epps
ICASSP4
2021 AusKidTalk: An Auditory-Visual Corpus of 3- to 12-Year-Old Australian Children's Speech
abstract
Here we present AusKidTalk [1], an audio-visual (AV) corpus of Australian children’s speech collected to facilitate the development of speech based technological solutions for children. It builds upon the technology and expertise developed through the collection of an earlier corpus of Australian adult speech, AusTalk [2,3]. This multi-site initiative was established to remedy the dire shortage of children’s speech corpora in Australia and around the world that are sufficiently sized to train accurate automated speech processing tools for children. We are collecting ~600 hours of speech from children aged 3–12 years that includes single word and sentence productions as well as narrative and emotional speech. In this paper, we discuss the key requirements for AusKidTalk and how we designed the recording setup and protocol to meet them. We also discuss key findings from our feasibility study of the recording protocol, recording tools, and user interface.
Beena Ahmed, Kirrie J. Ballard, Denis Burnham, Tharmakulasingam Sirojan, Hadi Mehmood, Dominique Estival, Elise Baker, Felicity Cox, Joanne Arciuli, Titia Benders, Katherine Demuth, Barbara Kelly, Chloé Diskin-Holdaway, Mostafa Shahin, Vidhyasaharan Sethu, Julien Epps, Chwee Beng Lee, Eliathamby Ambikairajah
Interspeech16
2021 What Does The Eye Best Tell? An Investigation of Eye Activity for Emotion, Cognitive load and Task Transition Recognition
abstract
Eye activity has previously been found to be relevant to emotion, cognitive load and task transition, but studies have mainly focused on one of them and used a single type of task. This motivates an investigation of eye activity for emotion, cognitive load and task transition recognition in a single study with the goal of understanding the capability of eye activity. We recorded 15 participants’ eye data while they completed a sequence of free-viewing emotional image tasks and a sequence of arithmetic tasks. Eye activity features within fixed analysis windows were extracted to classify (i) between four levels of arousal, (ii) three levels of valence, (iii) between three levels of cognitive load, (iv) between affective and cognitive tasks, and (v) between task transition and non-transition. The results suggest that eye activity can best be used for task transition recognition, with an accuracy of 77%, followed by cognitive load level recognition, then affective and cognitive task recognition. Implications of this finding include automatically segmenting long and continuous signals into task units with eye activity before analyzing human behavior.
Siyuan Chen 0002, Julien Epps
SMC2
2021 Wearable Fatigue Detection Based on Blink-Saccade Synchronisation
abstract
Automatic detection of fatigue based on remote cameras and computer vision has been investigated for many years, and is a viable solution in many contexts, but not for mobile contexts. This paper investigates eye activity extraction methods for wearable fatigue detection using low-cost hardware that is becoming ubiquitous in glasses form-factors, and proposes a new measure for fatigue based on the synchronization between blink and saccade. Evaluation on a novel dataset shows that the novel blink-saccade synchronization measure achieves statistically significant separation of control and fatigued participants, and provides automatic fatigue detection accuracy improvements of 10% relative to existing saccade measures based on velocity and amplitude.
Colin Lam, Julien Epps, Siyuan Chen 0002
SMC2
2021 An Investigation of Automatic Saccade and Fixation Detection from Wearable Infrared Cameras
abstract
Eye movement plays an important role in cognition and perception, and the detection of saccade and fixation has been studied for human-computer applications, however often under conditions where head movement is constrained, and often using calibration-dependent gaze information rather than the raw pupil position. In order to investigate the performance of saccade and fixation detection using gaze and pupil data, three representative saccade detection algorithms are applied to both pupil data and gaze data collected with and without head movement, and their performance is evaluated against a stimulus-induced ground truth under different measures. Results indicate that saccade/fixation detection using pupil data generally provides better performance than using gaze data with an 8.6% improvement in Cohen’s Kappa (averaged across the three algorithms), even when moderate head movement is involved. Hence, pupil data can be used as an alternative to gaze data for saccade and fixation detection in wearable contexts with less effort in calibration and higher accuracy.
Zishan Wang, Julien Epps, Siyuan Chen 0002
SMC2
2021 An adaptive transmission line cochlear model based front-end for replay attack detection
Tharshini Gunendradasan, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001
Speech Commun.3
2021 Read speech voice quality and disfluency in individuals with recent suicidal ideation or suicide attempt
Brian Stasak, Julien Epps, Heather T. Schatten, Ivan W. Miller, Emily Mower Provost, Michael F. Armey
Speech Commun.2
2021 Task Load Estimation from Multimodal Head-Worn Sensors Using Event Sequence Features
abstract
For longitudinal behavior analysis, task type is an inevitable and important variable. In this article, we propose an event-based behavior modeling approach and employ non-invasive wearable sensing modalities (eye activity, speech and head movement) to recognize task load level under four different task load types. The novelty lies in converting physiological and behavioral signals into meaningful events and utilizing their sequence across multiple modalities to distinguish load levels and types. We evaluated this approach on head-worn sensor data from 24 participants completing four different tasks for recognizing (i) low and high load level for a given task load type, (ii) low and high load level regardless of load type, and (iii) both load level and load type. Findings show that the recognition rate is reasonable in (i), close to chance level in (ii), and well above chance level in (iii) for 8 classes using participant-dependent and -independent schemes. Further, a fusion of the proposed event-based features and conventional continuous features achieved the best or similar performance in most cases. These results suggest that task type needs to be considered when using continuous features and that the proposed event-based modeling paradigm is promising for longitudinal behavior analysis.
Siyuan Chen 0002, Julien Epps
IEEE Trans. Affect. Comput.2
2020 Exploiting Vocal Tract Coordination Using Dilated CNNS For Depression Detection In Naturalistic Environments
abstract
Depression detection from speech continues to attract significant research attention but remains a major challenge, particularly when the speech is acquired from diverse smartphones in natural environments. Analysis methods based on vocal tract coordination have shown great promise in depression and cognitive impairment detection for quantifying relationships between features over time through eigenvalues of multi-scale cross-correlations. Motivated by the success of these methods, this paper proposes a novel way to extract full vocal tract coordination (FVTC) features by use of convolutional neural networks (CNNs), overcoming earlier shortcomings. Evaluations of the proposed FVTC-CNN structure on depressed speech data show improvements in mean F1 scores of at least 16.4% under clean conditions and comparable results under noisy conditions relative to existing VTC baseline systems.
Zhaocheng Huang, Julien Epps, Dale Joachim
ICASSP2
2020 Multimodal Event-based Task Load Estimation from Wearables
abstract
Humans always engage multiple modalities when performing tasks, such as eye activity, speech and head movement, which contain rich information indicative of task load that can help understand and predict human psychological state and behavior. In recent research into multimodal signal processing, the ideas of sequence- and coordination-based event features have been proposed, which explicitly utilize the interaction information among different modalities. In this paper, we propose event intensity and event duration-based features, which capture the extent and duration of onset events that denote major changes in behavior signal. These features are combined with sequence- and coordination-based event features to achieve state-of-the-art performance in assessing task load levels and load types. In experimental work, we collected eye activity, speech and head movement data from 24 participants during cognitive, perceptual, physical and communication tasks. Results suggest that by fusing these four compact, interpretable event-based features, strong accuracy can be achieved: 84% for two load level classification, 89% for four load type classification and 76% for 8-class classification, outperforming conventional statistical features and deep neural network self-learned features by up to 9% and 25% respectively. These features do not need to be selected during training and can generalize well for different participants and different task types.
Siyuan Chen 0002, Julien Epps
IJCNN2
2020 Domain Adaptation for Enhancing Speech-Based Depression Detection in Natural Environmental Conditions Using Dilated CNNs
abstract
Depression disorders are a major growing concern worldwide, especially given the unmet need for widely deployable depression screening for use in real-world environments. Speech-based depression screening technologies have shown promising results, but primarily in systems that are trained using laboratory-based recorded speech. They do not generalize well on data from more naturalistic settings. This paper addresses the generalizability issue by proposing multiple adaptation strategies that update pre-trained models based on a dilated convolutional neural network (CNN) framework, which improve depression detection performance in both clean and naturalistic environments. Experimental results on two depression corpora show that feature representations in CNN layers need to be adapted to accommodate environmental changes, and that increases in data quantity and quality are helpful for pre-training models for adaptation. The cross-corpus adapted systems produce relative improvements of 29.4% and 17.2% in unweighted average recall over non-adapted systems for both clean and naturalistic corpora, respectively.
Zhaocheng Huang, Julien Epps, Dale Joachim, Brian Stasak, James R. Williamson, Thomas F. Quatieri
INTERSPEECH2
2020 How Ordinal Are Your Data?
Sadari Jayawardena, Julien Epps, Zhaocheng Huang
INTERSPEECH2
2020 Augmenting Turn-Taking Prediction with Wearable Eye Activity During Conversation
Siyuan Chen 0002, Julien Epps
INTERSPEECH3
2020 Investigating Light-ResNet Architecture for Spoofing Detection Under Mismatched Conditions
Prasanth Parasu, Julien Epps, Kaavya Sriskandaraja, Gajan Suthokumar
INTERSPEECH2
2020 UNSW System Description for the Shared Task on Automatic Speech Recognition for Non-Native Children's Speech
Mostafa Shahin, Renée Lu, Julien Epps, Beena Ahmed
INTERSPEECH3
2020 Think before you speak: An investigation of eye activity patterns during conversations using eyewear
Julien Epps, Siyuan Chen 0002
Int. J. Hum. Comput. Stud.2
2020 Generalized Two-Stage Rank Regression Framework for Depression Score Prediction from Speech
abstract
This paper introduces a novel speech-based depression score prediction paradigm, the 2-stage ranking prediction framework, and highlights the benefits it brings to depression prediction. Conventional regression approaches aim to discern a single functional relationship between speech features and depression scores, making an implicit assumption about the existence of a single fixed relationship between the features and scores. However, as the relationship between severity of depression and the clinical score may vary over the range of the assessment scale, this style of analysis may not be suited to depression prediction. The proposed framework on the other hand, imposes a series of partitions on the feature space, with each partition corresponding to a distinct predefined range of depression scores, and predicts the score based on measures of membership to each partition. This approach provides additional flexibility by allowing different rankings to be learnt for different depression scores, and relaxes assumptions made by conventional regression approaches. Results demonstrate the framework's suitability for depression score prediction: different 2-stage implementations, based on heterogeneous feature extraction and modelling approaches, produce state-of-the-art results on the AVEC-2013 dataset. It is also demonstrated that, unlike fusion of conventional regression systems, the fusion of two-stage systems consistently improves prediction performance.
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, James R. Williamson, Thomas F. Quatieri, Jarek Krajewski
IEEE Trans. Affect. Comput.3
2020 An Investigation of Partition-Based and Phonetically-Aware Acoustic Features for Continuous Emotion Prediction from Speech
abstract
Phonetic variability has long been considered a confounding factor for emotional speech processing, so phonetic features have been rarely explored. However, surprisingly some features with purely phonetic information have shown state-of-the-art performance for continuous prediction of emotions (e.g., arousal and valence), for which the underlying causes are unknown to date. In this article, we present in-depth investigations into phonetic features on three widely used corpora - RECOLA, SEMAINE and USC CreativeIT - to explore this from two perspectives: acoustic space partitioning information and phonetic content. First, comparisons of multiple different partitioning methods confirm the significance of partitioning information in speech, and reveal the new understanding that varying the number of partitions has a greater effect on valence than arousal prediction: a detailed representation of the acoustic space is needed for valence, whilst a general one is adequate for arousal. Second, phoneme-specific examination of phonetic features suggests that phonetic content is less emotionally informative than partitioning information, and is more important for arousal than for valence. Furthermore, we propose a novel set of phonetically-aware acoustic features, attaining significant improvements for valence (in particular) and arousal prediction across RECOLA, SEMAINE and CreativeIT respectively, compared with conventional reference acoustic features.
Zhaocheng Huang, Julien Epps
IEEE Trans. Affect. Comput.2
2020 Multimodal Coordination Measures to Understand Users and Tasks
abstract
Physiological and behavioral measures allow computing devices to augment user interaction experience by understanding their mental load. Current techniques often utilize complementary information between different modalities to index load level typically within a specific task. In this study, we propose a new approach utilizing the timing between physiology/behavior change events to index low and high load level of four task types. Findings from a user study where eye, speech, and head movement data were collected from 24 participants demonstrate that the proposed measures are significantly different between low and high load levels with high effect size. It was also found that voluntary actions are more likely to be coordinated during tasks. Implications for the design of multimodal-multisensor interfaces include (i) utilizing event change and interaction in multiple modalities is feasible to distinguish task load levels and load types and (ii) voluntary actions should be allowed for effective task completion.
Siyuan Chen 0002, Julien Epps
ACM Trans. Comput. Hum. Interact.2
2019 Using Gaussian Processes with LSTM Neural Networks to Predict Continuous-Time, Dimensional Emotion in Ambiguous Speech
abstract
In continuous emotion recognition (CER) applications, commonly used models of the emotional content of speech cannot represent some aspects of the behavior of the dimensional emotion values that are the targets of prediction, such as ambiguity, as these models include only a single value for each point in time for each emotional dimension. In this paper, we first propose a model for the emotional content of speech that uses a Gaussian process (GP) to define a distribution that incorporates the inherent ambiguity of emotional speech. Then, we propose a predictive CER system which combines this model alongside LSTM neural network techniques that have that have previously been shown to perform well on this task. When tested on a practical CER task using the RECOLA dataset, this combined LSTM-GP approach is shown to achieve similar or higher performance to a LSTM neural network on its own when measured on mean-only terms. When the variance of the distribution is also incorporated into the performance measure, the LSTM-GP system, which predicts a distribution over the continuous-time outputs rather than just a time series of predicted values, is shown to be able to more realistically model the underlying ambiguity of the emotional content of the speech recordings.
Mia Atcheson, Vidhyasaharan Sethu, Julien Epps
ACII3
2019 Transmission Line Cochlear Model Based AM-FM Features for Replay Attack Detection
abstract
This paper focuses on providing a countermeasure to replay attack which is the simplest and more accessible form of attack used to spoof automatic speaker verification systems. Specifically, it proposes the use of the transmission line cochlear model, which resembles the human cochlea more accurately than parallel filter bank models, in the front-end of replay detection systems. Here the basilar membrane is modeled as a cascade of digital filters with decreasing resonant frequencies. In this context, we propose two features - transmission line cochlea-amplitude modulation (TLC-AM) and frequency modulation (TLC-FM) - to extract the modulation features of the speech from the simulated membrane displacements. TLC-AM is analogous to the output of the inner hair cell bending movement, which accurately captures the amplitude modulation component of the speech. TLC-FM is extracted by deriving the in-phase and out of phase signals of basilar membrane displacement. Results show that individual TLC-AM and TLC-FM features perform better than the best parallel filter bank baseline system. Experiments suggest that higher frequency selectivity is beneficial for replay detection, especially for AM, and the proposed TLC model is better able to achieve this property than parallel filter bank models.
Tharshini Gunendradasan, Saad Irtza, Eliathamby Ambikairajah, Julien Epps
ICASSP4
2019 Speech Landmark Bigrams for Depression Detection from Naturalistic Smartphone Speech
abstract
Detection of depression from speech has attracted significant research attention in recent years but remains a challenge, particularly for speech from diverse smartphones in natural environments. This paper proposes two sets of novel features based on speech landmark bigrams associated with abrupt speech articulatory events for depression detection from smartphone audio recordings. Combined with techniques adapted from natural language text processing, the proposed features further exploit landmark bigrams by discovering latent articulatory events. Experimental results on a large, naturalistic corpus containing various spoken tasks recorded from diverse smartphones suggest that speech landmark bigram features provide a 30.1% relative improvement in F1 (depressed) relative to an acoustic feature baseline system. As might be expected, a key finding was the importance of tailoring the choice of landmark bigrams to each elicitation task, revealing that different aspects of speech articulation are elicited by different tasks, which can be effectively captured by the landmark approaches.
Zhaocheng Huang, Julien Epps, Dale Joachim
ICASSP2
2019 Evaluation Measures for Depression Prediction and Affective Computing
abstract
A variety of evaluation measures are being used to validate systems in depression prediction and affective computing. Among them, the most common measures focus on the error between the ground truth and predictions. However, when the ground truth is ordinal such as in psychiatric scores, ranking information is more important than the actual error. Therefore, this study systematically analyses the properties of classification, error-based and ranking measures particularly using classification accuracy, root mean square error (RMSE) and Spearman rank correlation coefficient, with the aim of identifying suitable measures for evaluating depression prediction and affective computing. For the purpose of analysis, we employed both synthetic data and real depression prediction systems evaluated with the AVEC2017 depression corpus. Outcomes of the experiments suggest that RMSE and classification accuracy, which are frequently used, are not sensitive to ordering and that rank correlation measures are more appropriate for depression prediction, which is an ordinal problem.
Sadari Jayawardena, Julien Epps, Eliathamby Ambikairajah
ICASSP2
2019 Auditory Inspired Spatial Differentiation for Replay Spoofing Attack Detection
abstract
The security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. This paper explores the effect of `spatial differentiation' used in auditory system modelling to improve frequency selectivity and hence provide a more selective front-end for replay attack detection. Experiments were done using a parallel filter bank consisting of simple 2ndorder IIR bandpass filters following which, processing analogous to spatial differentiation was employed to obtain higher order stable IIR filters, in turn leading to highly selective filter banks. Two novel features based on spatially differentiated higher order filter bank have been proposed. Together they yield a relative improvement of 29.9% in replay speech detection over a constant Q transform based baseline system, when evaluated on the ASVspoof 2017 Version 2.0 database.
Buddhi Wickramasinghe, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001
ICASSP3
2019 An Adaptive-Q Cochlear Model for Replay Spoofing Detection
Tharshini Gunendradasan, Eliathamby Ambikairajah, Julien Epps, Haizhou Li 0001
INTERSPEECH3
2019 Direct Modelling of Speech Emotion from Raw Speech
abstract
Speech emotion recognition is a challenging task and heavily depends on hand-engineered acoustic features, which are typically crafted to echo human perception of speech signals. However, a filter bank that is designed from perceptual evidence is not always guaranteed to be the best in a statistical modelling framework where the end goal is for example emotion classification. This has fuelled the emerging trend of learning representations from raw speech especially using deep learning neural networks. In particular, a combination of Convolution Neural Networks (CNNs) and Long Short Term Memory (LSTM) have gained great traction for the intrinsic property of LSTM in learning contextual information crucial for emotion recognition; and CNNs been used for its ability to overcome the scalability problem of regular neural networks. In this paper, we show that there are still opportunities to improve the performance of emotion recognition from the raw speech by exploiting the properties of CNN in modelling contextual information. We propose the use of parallel convolutional layers to harness multiple temporal resolutions in the feature extraction block that is jointly trained with the LSTM based classification network for the emotion recognition task. Our results suggest that the proposed model can reach the performance of CNN trained with hand-engineered features from both IEMOCAP and MSP-IMPROV datasets.
Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Julien Epps
INTERSPEECH5
2019 Biologically Inspired Adaptive-Q Filterbanks for Replay Spoofing Attack Detection
Buddhi Wickramasinghe, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH3
2019 An investigation of linguistic stress and articulatory vowel characteristics for automatic depression classification
Brian Stasak, Julien Epps, Roland Göcke
Comput. Speech Lang.2
2019 Automatic depression classification based on affective read sentences: Opportunities for text-dependent analysis
Brian Stasak, Julien Epps, Roland Göcke
Speech Commun.2
2019 Atomic Head Movement Analysis for Wearable Four-Dimensional Task Load Recognition
abstract
Physical activity recognition using wearable sensors has achieved good performance in discriminating heterogeneous activities for health monitoring, but there has been less investigation of sedentary activities, e.g., desk work, which is often physically homogenous, to improve health in office environments. In this study, we explored head movement as a new sensing modality for physical and mental activity analysis. A new algorithm which segments gyroscope signals into atomic head movement events is proposed. Instead of recognizing activities in terms of predefined categories, we recognized four dimensions of task load: cognitive, perceptual, communicative, and physical, analogous to current manual workload assessment methods like NASA-TLX. We collected head movement data from 24 participants who wore a tri-axial inertial sensor at head while performing multiple tasks with varying load levels at office. An average of 70% accuracy was achieved for recognizing cognitive load levels, and more than 80% for the other three load types. The proposed event features outperformed a set of 181 features from previous physical activity recognition studies. We also demonstrated that these atomic event features are diagnostic of different load types in cross-load type classification, showing the promise of physical and mental load monitoring for health.
Siyuan Chen 0002, Julien Epps
IEEE J. Biomed. Health Informatics2
2018 Demonstrating and Modelling Systematic Time-varying Annotator Disagreement in Continuous Emotion Annotation
Mia Atcheson, Vidhyasaharan Sethu, Julien Epps
INTERSPEECH3
2018 Detection of Replay-Spoofing Attacks Using Frequency Modulation Features
Tharshini Gunendradasan, Buddhi Wickramasinghe, Phu Ngoc Le, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH5
2018 Depression Detection from Short Utterances via Diverse Smartphones in Natural Environmental Conditions
abstract
Depression is a leading cause of disease burden worldwide, however there is an unmet need for screening and diagnostic measures that can be widely deployed in real-world environments. Voice-based diagnostic methods are convenient, non-invasive to elicit, and can be collected and processed in near real-time using modern smartphones, smart speakers, and other devices. Studies in voice-based depression detection to date have primarily focused on laboratory-collected voice samples, which are not representative of typical user environments or devices. This paper conducts the first investigation of voice-based depression assessment techniques on real-world data from 887 speakers, recorded using a variety of different smartphones. Evaluations on 16 hours of speech show that conservative segment selection strategies using highly thresholded voice activity detection, coupled with tailored normalization approaches are effective for mitigating smartphone channel variability and background environmental noise. Together, these strategies can achieve F1 scores comparable with or better than those from a combination of clean recordings, a single recording environment and long utterances. The scalability of speech elicitation via smartphone allows detailed models dependent on gender, smartphone manufacturer and/or elicitation task. Interestingly, results herein suggest that normalization based on these criteria may be more effective than tailored models for detecting depressed speech.
Zhaocheng Huang, Julien Epps, Dale Joachim
INTERSPEECH2
2018 Variational Autoencoders for Learning Latent Representations of Speech Emotion: A Preliminary Study
abstract
Learning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, generating effective features is crucial. Currently, handcrafted features are mostly used for speech emotion recognition, however, features learned automatically using deep learning have shown strong success in many problems, especially in image processing. In particular, deep generative models such as Variational Autoencoders (VAEs) have gained enormous success for generating features for natural images. Inspired by this, we propose VAEs for deriving the latent representation of speech signals and use this representation to classify emotions. To the best of our knowledge, we are the first to propose VAEs for speech emotion classification. Evaluations on the IEMOCAP dataset demonstrate that features learned by VAEs can produce state-of-the-art results for speech emotion classification.
Siddique Latif, Rajib Rana, Junaid Qadir 0001, Julien Epps
INTERSPEECH4
2018 Transfer Learning for Improving Speech Emotion Classification Accuracy
abstract
The majority of existing speech emotion recognition research focuses on automatic emotion detection using training and testing data from same corpus collected under the same conditions.The performance of such systems has been shown to drop significantly in cross-corpus and cross-language scenarios.To address the problem, this paper exploits a transfer learning technique to improve the performance of speech emotion recognition systems that is novel in cross-language and cross-corpus scenarios.Evaluations on five different corpora in three different languages show that Deep Belief Networks (DBNs) offer better accuracy than previous approaches on cross-corpus emotion recognition, relative to a Sparse Autoencoder and SVM baseline system.Results also suggest that using a large number of languages for training and using a small fraction of the target data in training can significantly boost accuracy compared with baseline also for the corpus with limited training examples.
Siddique Latif, Rajib Rana, Shahzad Younis, Junaid Qadir 0001, Julien Epps
INTERSPEECH5
2018 Frequency Domain Linear Prediction Features for Replay Spoofing Attack Detection
Buddhi Wickramasinghe, Saad Irtza, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH4
2018 A Novel Framework for Distress Detection through an Automated Speech Processing System
abstract
Based on our ongoing work, this work in progress project aims to develop an automated system to detect distress in people to enable early referral for interventions to target anxiety and depression, to mitigate suicidal ideation and to improve adherence to treatment. The project will utilize either use existing voice data to assess people into various scales of distress, or will collect voice data as per existing standards of distress measurement, to develop basic computing algorithms required to detect various attributes associated with distress, detected through a person's voice in a telephone call to a helpline. This will be then matched with the already available psychological assessment instruments such as the Distress Thermometer for these persons. In order to trigger interventions, organizational contexts are essential as interventions rely on the type of distress. Therefore, the model will be tested on various organizational settings such as the Police, Emergency and Health along with the Distress detection instruments normally used in a psychological assessment for accuracy and validation. The outcome of the project will culminate in a fully automated integrated system, and will save significant resources to organizations. The translation of the project will be realized in step-change improvements to quality of life within the gamut of public policy.
Rajib Rana, Raj Gururajan, Geraldine Mackenzie, Jeff Dunn, Anthony Gray, Xujuan Zhou, Prabal Datta Barua, Julien Epps, Gerald Humphris
WI8
2018 Multimodal Depression Detection: Fusion Analysis of Paralinguistic, Head Pose and Eye Gaze Behaviors
abstract
An estimated 350 million people worldwide are affected by depression. Using affective sensing technology, our long-term goal is to develop an objective multimodal system that augments clinical opinion during the diagnosis and monitoring of clinical depression. This paper steps towards developing a classification system-oriented approach, where feature selection, classification and fusion-based experiments are conducted to infer which types of behaviour (verbal and nonverbal) and behaviour combinations can best discriminate between depression and non-depression. Using statistical features extracted from speaking behaviour, eye activity, and head pose, we characterise the behaviour associated with major depression and examine the performance of the classification of individual modalities and when fused. Using a real-world, clinically validated dataset of 30 severely depressed patients and 30 healthy control subjects, a Support Vector Machine is used for classification with several feature selection techniques. Given the statistical nature of the extracted features, feature selection based on T-tests performed better than other methods. Individual modality classification results were considerably higher than chance level (83 percent for speech, 73 percent for eye, and 63 percent for head). Fusing all modalities shows a remarkable improvement compared to unimodal systems, which demonstrates the complementary nature of the modalities. Among the different fusion approaches used here, feature fusion performed best with up to 88 percent average accuracy. We believe that is due to the compatible nature of the extracted statistical features.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Matthew Hyett, Gordon Parker, Michael Breakspear
IEEE Trans. Affect. Comput.4
2017 A PLLR and multi-stage Staircase Regression framework for speech-based emotion prediction
abstract
Continuous prediction of dimensional emotions (e.g. arousal and valence) has attracted increasing research interest recently. When processing emotional speech signals, phonetic features have been rarely used due to the assumption that phonetic variability is a confounding factor that degrades emotion recognition/prediction performance. In this paper, instead of eliminating phonetic variability, we investigated whether Phone Log-Likelihood Ratio (PLLR) features could be used to index arousal and valence in a pairwise low/high framework. A multi-stage staircase regression (SR) framework which enables fusion at three different stages is also investigated. Results on the RECOLA database show that PLLR outperforms EGEMAPS features for arousal and valence. Interestingly, long-term averaged PLLR proved to be more robust and emotionally informative than local frame-level PLLR, which contains more phoneme-specific information. Within the multistage SR framework, PLLR yielded an 8.2% and 11.6% relative improvement in CCC for arousal and valence respectively, showing great promise for including phonetic features in emotion prediction systems.
Zhaocheng Huang, Julien Epps
ICASSP2
2017 An Investigation of Crowd Speech for Room Occupancy Estimation
Siyuan Chen 0002, Julien Epps, Eliathamby Ambikairajah, Phu Ngoc Le
INTERSPEECH2
2017 An Investigation of Emotion Prediction Uncertainty Using Gaussian Mixture Regression
Ting Dang, Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah
INTERSPEECH3
2017 Bidirectional Modelling for Short Duration Language Identification
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH4
2017 An Investigation of Emotion Dynamics and Kalman Filtering for Speech-Based Emotion Prediction
Zhaocheng Huang, Julien Epps
INTERSPEECH2
2017 Elicitation Design for Acoustic Depression Classification: An Investigation of Articulation Effort, Linguistic Complexity, and Word Affect
Brian Stasak, Julien Epps, Roland Göcke
INTERSPEECH2
2016 Detecting the instant of emotion change from speech using a martingale framework
abstract
Towards a better understanding of emotion in speech, it is important to understand how emotion changes and when it changes. Recognizing emotions using pre-segmented speech utterances results in a loss in continuity of emotions and does not provide insights into emotion changes. In this paper, we propose an investigation into emotion change detection from the perspective of exchangeability of data points observed sequentially using a martingale framework. Within the framework, a per-frame GMM likelihood based approach is proposed as a measure of strangeness from a particular emotion class. Experimental results on the IEMOCAP database demonstrate that the proposed martingale framework offers significant improvements over the baseline GLR method for detecting emotion changes not only between neutral and emotional speech, but also between positive and negative classes along the arousal and valence emotion dimensions.
Zhaocheng Huang, Julien Epps
ICASSP2
2016 Cross-Cultural Depression Recognition from Vocal Biomarkers
abstract
No studies have investigated cross-cultural and cross-language characteristics of depressed speech. We investigated the generalisability of a vocal biomarker-based approach to depression detection in clinical interviews recorded in three countries (Australia, the USA and Germany), two languages (German and English) and different accents (Australian and American). Several approaches to training and testing within and between datasets were evaluated. Using the same experimental protocol separately within each dataset, (cross-classification) accuracy was high.combining datasets, high accuracy was high again and consistent across language, recording environment, and culture. Training and testing between datasets, however, attenuated accuracy. These finding emphasize the importance of heterogeneous training sets for robust depression detection.
Sharifa Alghowinem, Roland Göcke, Julien Epps, Michael Wagner 0004, Jeffrey F. Cohn
INTERSPEECH3
2016 Automatic Classification of Lexical Stress in English and Arabic Languages Using Deep Learning
Mostafa Shahin, Julien Epps, Beena Ahmed
INTERSPEECH2
2016 An Investigation of Emotional Speech in Depression Classification
Brian Stasak, Julien Epps, Nicholas Cummins, Roland Göcke
INTERSPEECH2
2016 The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing
abstract
Work on voice sciences over recent decades has led to a proliferation of acoustic parameters that are used quite selectively and are not always extracted in a similar fashion. With many independent teams working in different research areas, shared standards become an essential safeguard to ensure compliance with state-of-the-art methods allowing appropriate comparison of results across studies and potential integration and combination of extraction and recognition systems. In this paper we propose a basic standard acoustic parameter set for various areas of automatic voice analysis, such as paralinguistic or clinical speech analysis. In contrast to a large brute-force parameter set, we present a minimalistic set of voice parameters here. These were selected based on a) their potential to index affective physiological changes in voice production, b) their proven value in former studies as well as their automatic extractability, and c) their theoretical significance. The set is intended to provide a common baseline for evaluation of future research and eliminate differences caused by varying parameter sets or even different implementations of the same parameters. Our implementation is publicly available with the openSMILE toolkit. Comparative evaluations of the proposed feature set and large baseline feature sets of INTERSPEECH challenges show a high performance of the proposed set in relation to its size.
Florian Eyben, Klaus R. Scherer, Björn W. Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shri Narayanan, Khiet P. Truong
IEEE Trans. Affect. Comput.8
2015 Weighted pairwise Gaussian likelihood regression for depression score prediction
abstract
This paper presents a technique in which feature vectors are mapped onto ordinal ranges of clinical depression scores using weighted pairwise Gaussians. The position of a test vector with respect to these partitions is used to perform depression score prediction. Results found on a set of spectral and formant based speech characteristics indicate the potential of this technique for performing depression score prediction. Key results on the AVEC 2013 development set indicate that the inclusion of weights and Bayesian adaptation improves system performance by 16.5% - 18.5% when compared to using an unweighted non-adapted system. Fusing results from Bayesian adapted models corresponding to different feature spaces offers up to 8% further improvement. Further, fusion consistently improves performance on both the AVEC 2013 development and test set, in contrast to conventional regressor fusion.
Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Jarek Krajewski
ICASSP2
2015 Relevance vector machine for depression prediction
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Jarek Krajewski
INTERSPEECH3
2015 An investigation of emotion change detection from speech
abstract
Emotion recognition based on speech plays an important role in Human Computer Interaction (HCI), which has motivated extensive recent investigation into this area. However, current research on emotion recognition is focused on recognizing emotion on a per-file basis and mostly does not provide insight into emotion changes. In this paper, we report on an initial investigation into detecting the instant of emotion change using Gaussian Mixture Models (GMM) based methods, either without or with prior knowledge of emotions: the Generalized Likelihood Ratio and Emotion Pair Likelihood Ratios, together with a novel normalization scheme to improve emotion change detection accuracy. Experimental results based on the IEMOCAP corpus are presented that demonstrate a promising baseline. Despite the challenging nature of the problem, this work provides a path towards systems that detect and understand emotion changes, and also presents very interesting questions for further investigation.
Zhaocheng Huang, Julien Epps, Eliathamby Ambikairajah
INTERSPEECH2
2015 Analysis of acoustic space variability in speech affected by depression
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Sebastian Schnieder, Jarek Krajewski
Speech Commun.3
2015 A review of depression and suicide risk assessment using speech analysis
Nicholas Cummins, Stefan Scherer, Jarek Krajewski, Sebastian Schnieder, Julien Epps, Thomas F. Quatieri
Speech Commun.5
2015 Voice source under cognitive load: Effects and classification
Tet Fei Yap, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi
Speech Commun.2
2014 Variability compensation in small data: Oversampled extraction of i-vectors for the classification of depressed speech
abstract
Variations in the acoustic space due to changes in speaker mental state are potentially overshadowed by variability due to speaker identity and phonetic content. Using the Audio/Visual Emotion Challenge and Workshop 2013 Depression Dataset we explore the suitability of i-vectors for reducing these latter sources of variability for distinguishing between low or high levels of speaker depression. In addition we investigate whether supervised variability compensation methods such as Linear Discriminant Analysis (LDA), and Within Class Covariance Normalisation (WCCN), applied in the i-vector domain, could be used to compensate for speaker and phonetic variability. Classification results show that i-vectors formed using an over-sampling methodology outperform a baseline set by KL-means supervectors. However the effect of these two compensation methods does not appear to improve system accuracy. Visualisations afforded by the t-Distributed Stochastic Neighbour Embedding (t-SNE) technique suggest that despite the application of these techniques, speaker variability is still a strong confounding effect.
Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Jarek Krajewski
ICASSP2
2014 Probabilistic acoustic volume analysis for speech affected by depression
abstract
Alterations in speech motor control in depressed individuals have been found to manifest as a reduction in spectral variability. In this paper we present a novel method for measuring acoustic volume a model-based measure that is reflective of this decrease in spectral variability and assess the ability of features resulting from this measure for indexing a speaker’s level of depression. A Monte Carlo approximation that enables the computation of this measure is also outlined in this paper. Results found using the AVEC 2013 Challenge Dataset indicate there is a statistically significant reduction in acoustic variation with increasing levels of speaker depression, and using features designed to capture this change it is possible to outperform a range of conventional spectral measures when predicting a speaker’s level of depression.
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Jarek Krajewski
INTERSPEECH3
2014 The INTERSPEECH 2014 computational paralinguistics challenge: cognitive & physical load
abstract
The INTERSPEECH 2014 Computational Paralinguistics Challenge provides for the first time a unified test-bed for the automatic recognition of speakers’ cognitive and physical load in speech. In this paper, we describe these two Sub-Challenges, their conditions, baseline results and experimental procedures, as well as the COMPARE baseline features generated with the openSMILE toolkit and provided to the participants in the Challenge.
Björn W. Schuller, Stefan Steidl, Anton Batliner, Julien Epps, Florian Eyben, Fabien Ringeval, Erik Marchi, Yue Zhang 0014
INTERSPEECH4
2014 Using Task-Induced Pupil Diameter and Blink Rate to Infer Cognitive Load
abstract
Minimizing user cognitive load is suggested as an integral part of human-centered design, where a more intuitive, easy to learn, and adaptive interface is desired. In this context, it is difficult to develop optimal strategies to improve the design without first knowing how user cognitive load fluctuates during interaction. In this study, we investigate how cognitive load measurement is affected by different task types from the perspective of the load theory of attention, using pupil diameter and blink measures. We induced five levels of cognitive load during low and high perceptual load tasks and found that although pupil diameter showed significant effects on cognitive load when the perceptual load was low, neither blink rate nor pupil diameter showed significant effects on cognitive load when the perceptual load was high. The results indicate that pupil diameter can index cognitive load only in the situation of low perceptual load and are the first to provide empirical support for the cognitive control aspect of the load theory of attention, in the context of cognitive load measurement. Meanwhile, blink is a better indicator of perceptual load than cognitive load. This study also implies that perceptual load should be considered in cognitive load measurement using pupil diameter and blink measures. Automatic detection of the type and level of load in this manner helps pave the way for better reasoning about user internal processes for human-centered interface design.
Siyuan Chen 0002, Julien Epps
Hum. Comput. Interact.2
2014 Efficient and Robust Pupil Size and Blink Estimation From Near-Field Video Sequences for Human-Machine Interaction
abstract
Monitoring pupil and blink dynamics has applications in cognitive load measurement during human-machine interaction. However, accurate, efficient, and robust pupil size and blink estimation pose significant challenges to the efficacy of real-time applications due to the variability of eye images, hence to date, require manual intervention for fine tuning of parameters. In this paper, a novel self-tuning threshold method, which is applicable to any infrared-illuminated eye images without a tuning parameter, is proposed for segmenting the pupil from the background images recorded by a low cost webcam placed near the eye. A convex hull and a dual-ellipse fitting method are also proposed to select pupil boundary points and to detect the eyelid occlusion state. Experimental results on a realistic video dataset show that the measurement accuracy using the proposed methods is higher than that of widely used manually tuned parameter methods or fixed parameter methods. Importantly, it demonstrates convenience and robustness for an accurate and fast estimate of eye activity in the presence of variations due to different users, task types, load, and environments. Cognitive load measurement in human-machine interaction can benefit from this computationally efficient implementation without requiring a threshold calibration beforehand. Thus, one can envisage a mini IR camera embedded in a lightweight glasses frame, like Google Glass, for convenient applications of real-time adaptive aiding and task management in the future.
Siyuan Chen 0002, Julien Epps
IEEE Trans. Cybern.2
2013 Detecting depression: A comparison between spontaneous and read speech
abstract
Major depressive disorders are mental disorders of high prevalence, leading to a high impact on individuals, their families, society and the economy. In order to assist clinicians to better diagnose depression, we investigate an objective diagnostic aid using affective sensing technology with a focus on acoustic features. In this paper, we hypothesise that (1) classifying the general characteristics of clinical depression using spontaneous speech will give better results than using read speech, (2) that there are some acoustic features that are robust and would give good classification results in both spontaneous and read, and (3) that a `thin-slicing' approach using smaller parts of the speech data will perform similarly if not better than using the whole speech data. By examining and comparing recognition results for acoustic features on a real-world clinical dataset of 30 depressed and 30 control subjects using SVM for classification and a leave-one-out cross-validation scheme, we found that spontaneous speech has more variability, which increases the recognition rate of depression. We also found that jitter, shimmer, energy and loudness feature groups are robust in characterising both read and spontaneous depressive speech. Remarkably, thin-slicing the read speech, using either the beginning of each sentence or the first few sentences performs better than using all reading task data.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Michael Breakspear, Gordon Parker
ICASSP4
2013 A comparative study of different classifiers for detecting depression from spontaneous speech
abstract
Accurate detection of depression from spontaneous speech could lead to an objective diagnostic aid to assist clinicians to better diagnose depression. Little thought has been given so far to which classifier performs best for this task. In this study, using a 60-subject real-world clinically validated dataset, we compare three popular classifiers from the affective computing literature - Gaussian Mixture Models (GMM), Support Vector Machines (SVM) and Multilayer Perceptron neural networks (MLP) - as well as the recently proposed Hierarchical Fuzzy Signature (HFS) classifier. Among these, a hybrid classifier using GMM models and SVM gave the best overall classification results. Comparing feature, score, and decision fusion, score fusion performed better for GMM, HFS and MLP, while decision fusion worked best for SVM (both for raw data and GMM models). Feature fusion performed worse than other fusion methods in this study. We found that loudness, root mean square, and intensity were the voice features that performed best to detect depression in this dataset.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Tom Gedeon, Michael Breakspear, Gordon Parker
ICASSP4
2013 Spectro-temporal analysis of speech affected by depression and psychomotor retardation
abstract
To enhance current diagnostic methods used when assessing a depressed individual, an objective screening mechanism, ideally based on non-intrusive behavioral signals, is needed. Given the clinical description of depression speech as `dull, monotonous and flat' and promising previous results from spectral features, we hypothesize that the effects of depression on speech are embedded in spectro-temporal events. To test this hypothesis we explore different methodologies, based on the modulation spectrum, for extracting long-term spectro-temporal information from speech and assess their suitability as a clinical marker of depression. Results indicate that: depressive speech information is captured in the modulation spectrum, long-term spectro-temporal information is important in depressed speech identification and there are potential differences in the effects that depression and psychomotor retardation have on speech production mechanisms.
Nicholas Cummins, Julien Epps, Eliathamby Ambikairajah
ICASSP2
2013 Speaker variability in speech based emotion models - Analysis and normalisation
abstract
All features commonly utilised in speech based emotion classification systems capture both emotion-specific information and speaker-specific information. This paper proposes a novel method to gauge the effect of speaker-specific information on emotion modelling based on two measures: a Monte Carlo approximation to KL divergence and an estimate of feature variability based on diagonal covariance matrices. In addition, a novel speaker normalisation technique based on joint factor analysis is also proposed. This method is analogous to channel compensation in speaker verification systems, with one significant extension. The model domain compensation is mapped back to frame-level features, allowing for use in a wider range of emotion classification frameworks and in conjuncture with other normalisation techniques. Preliminary evaluations on the IEMOCAP database suggests that the proposed technique improves the performance of GMM based classification systems based on widely employed features such as pitch, MFCCs and deltas.
Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah
ICASSP2
2013 Mental Workload Classification via Online Writing Features
abstract
Mental workload is an important factor during writing, which may affect the writing efficiency and user experience. This paper aims at a method to classify the mental workload levels during writing process, via examination of online writing features in a two-stage algorithm structure. At the first stage, a curvature tracking method is applied to the handwriting script, to examine the curvature for individual writing points. Then a selection process allocates writing points into subsets, each corresponding to one curvature span. The second stage extracts velocity features, used to characterize mental workload, from points in each curvature span. A Parzen-window classifier is applied on velocity features from each curvature span. The classification decisions from individual classifiers are fused with a selective voting scheme for the overall mental workload classification decision. This paper finally discusses the classification accuracy for three mental workload levels and compares it with previous work.
Kun Yu 0001, Julien Epps, Fang Chen 0001
ICDAR2
2013 Characterising depressed speech for classification
abstract
Depression is a serious psychiatric disorder that affects mood, thoughts, and the ability to function in everyday life. This pa-per investigates the characteristics of depressed speech for the purpose of automatic classification by analysing the effect of different speech features on the classification results. We anal-ysed voiced, unvoiced and mixed speech in order to gain a better understanding of depressed speech and to bridge the gap be-tween physiological and affective computing studies. This un-derstanding may ultimately lead to an objective affective sens-ing system that supports clinicians in their diagnosis and mon-itoring of clinical depression. The characteristics of depressed speech were statistically analysed using ANOVA and linked to their classification results using GMM and SVM. Features were extracted and classified over speech utterances of 30 clinically depressed patients against 30 controls (both gender-matched) in a speaker-independent manner. Most feature classification re-sults were consistent with their statistical characteristics, pro-viding a link between physiological and affective computing studies. The classification results from low-level features were slightly better than the statistical functional features, which in-dicates a loss of information in the latter. We found that both mixed and unvoiced speech were as useful in detecting depres-sion as voiced speech, if not better. Index Terms: depression, speech characteristics, mood classi-fication
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Gordon Parker, Michael Breakspear
INTERSPEECH4
2013 Modeling spectral variability for the classification of depressed speech
abstract
Quantifying how the spectral content of speech relates to changes in mental state may be crucial in building an objective speech-based depression classification system with clinical utility. This paper investigates the hypothesis that important depression based information can be captured within the covariance structure of a Gaussian Mixture Model (GMM) of recorded speech. Significant negative correlations found between a speaker’s average weighted variance- a GMM-based indicator of speaker variability- and their level of depression support this hypothesis. Further evidence is provided by the comparison of classification accuracies from seven different GMM-UBM systems, each formed by varying different parameter combinations during MAP adaption. This analysis shows that variance-only adaptation either outperforms or matches the de facto standard mean-only adaptation when classifying both the presence and severity of depression. This result is perhaps the first of its kind seen in GMM-UBM speech classification.
Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Michael Breakspear, Roland Göcke
INTERSPEECH2
2013 GMM based speaker variability compensated system for interspeech 2013 compare emotion challenge
abstract
This paper describes the University of New South Wales system for the Interspeech 2013 ComParE emotion subchallenge. The primary aim of the submission is to explore the performance of model based variability compensation techniques applied to emotion classification and as a consequence of being a part of a challenge, to enable a comparison of these methods to alternative approaches. In keeping with this focused aim, a simple frame based front-end of MFCC and ΔMFCC is utilised. The systems outlined in this paper consists of a joint factor analysis based system and one based on a library of speaker-specific emotion models along with a basic GMM based system. The best combined system has an accuracy (UAR) of 47.8% as evaluated on the challenge development set and 35.7% as evaluated on the test set.
Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah, Haizhou Li 0001
INTERSPEECH2
2013 I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verification
abstract
I4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort.
Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah
INTERSPEECH25
2013 Automatic and continuous user task analysis via eye activity
abstract
A day in the life of a user can be segmented into a series of tasks: a user begins a task, becomes loaded perceptually and cognitively to some extent by the objects and mental challenge that comprise that task, then at some point switches or is distracted to a new task, and so on. Understanding the contextual task characteristics and user behavior in interaction can benefit the development of intelligent systems to aid user task management. Applications that aid the user in one way or another have proliferated as computing devices become more and more of a constant companion. However, direct and continuous observations of individual tasks in a naturalistic context and subsequent task analysis, for example the diary method, have traditionally been a manual process. We propose a method for automatic task analysis system, which monitors the user's current task and analyzes it in terms of the task transition, and perceptual and cognitive load imposed by the task. An experiment was conducted in which participants were required to work continuously on groups of three sequential tasks of different types. Three classes of eye activity, namely pupillary response, blink and eye movement, were analyzed to detect the task transition and non-transition states, and to estimate three levels of perceptual load and three levels of cognitive load every second to infer task characteristics. This paper reports statistically significant classification accuracies in all cases and demonstrates the feasibility of this approach for task monitoring and analysis.
Siyuan Chen 0002, Julien Epps, Fang Chen 0001
IUI2
2013 i-Vector with sparse representation classification for speaker verification
Jia Min Karen Kua, Julien Epps, Eliathamby Ambikairajah
Speech Commun.2
2012 Speaker variability in emotion recognition - an adaptation based approach
abstract
None of the features commonly utilised in automatic emotion classification systems completely disassociate emotion-specific information from speaker-specific information. Consequently, this speaker-specific variability adversely affects the performance of the emotion classification system and in existing systems is frequently mitigated by some form of speaker normalisation. Speaker adaptation offers an alternative to normalisation and this paper proposes a novel bootstrapping technique which involves selecting appropriate initial models from a large training pool, prior to speaker adaptation of emotion models in the context of GMM based emotion classification as an alternative to speaker normalisation. Evaluations on the LDC Emotional Prosody and the FAU Aibo corpora reveal that an emotion classification system based on the proposed bootstrapping method outperforms systems based on speaker normalisation as long as a small amount of labelled adaptation data is available. It also outperforms speaker adaption from common initial models estimated from all training speakers.
Ni Ding, Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah
ICASSP3
2012 A study of automatic phonetic segmentation for forensic voice comparison
abstract
Forensic voice comparison (FVC) systems have often involved manual annotation of usable phonetic units, requiring substantial human labor. Recent research has shown the efficacy of automatic methods in FVC, and this paper investigates automatic phonetic segmentation in FVC systems. Nasals and vowels were found to contribute the most in terms of improvements in both the validity and reliability of the system. Results show that as a function of the duration of the recognized tokens there is a trade-off in which an improvement in validity corresponds to a degradation in reliability and vice versa. An implication is that minimizing the error of automatically estimated monophone boundaries may not necessarily result in the best system validity or reliability. A substantial improvement in log-likelihood-ratio cost (validity) of 17.02% and in 95% credible interval (reliability) of 5.97% over the baseline system was possible by fusing baseline scores with those from nasal and vowel segments.
Chee Cheun Huang, Julien Epps
ICASSP2
2012 Classification of Working Memory Load Using Wavelet Complexity Features of EEG Signals
Pega Zarjam, Julien Epps, Fang Chen 0001, Nigel H. Lovell
ICONIP (2)2
2012 A Comparison of Classification Paradigms for Speaker Likeability Determination
abstract
In this paper we investigate the performance of different classification paradigms, testing each with a range of acoustic features, to find a system that is well suited to speaker likeability classification. We introduce a Sparse Representation Classifier for paralinguistic classification and explore the role of training data selection for a GMM classifier. Results demonstrate that (1) Single dimensional features of pitch direction, shimmer and spectral roll-off were the most suitable features found when testing on the development set but we were unable to reproduce their performance in the final classification task, (2) Using UBM training data selection increased accuracy of MFCC's and (3) Sparse Representation showed promise as a paralinguistic classifier with results comparable to that of SVM.
Nicholas Cummins, Julien Epps, Jia Min Karen Kua
INTERSPEECH2
2012 Speaker Clustering in Emotion Recognition
abstract
Speaker variability is a known challenge for emotion recognition, however little work has been done on speaker similarity in terms of its contribution to the performance in the emotion classification task. In this paper, we investigate this topic, and find a clear link between speaker proximity and the recognition accuracy. Motivated by this result, emotion based speaker clustering is proposed as a new strategy for speaker adaptation. It involves using speaker proximity to cluster individual speakers’ emotion models in the training set on a per-emotion basis, and adapting the test speaker’s emotion from the closest cluster. A series of tests were conducted to explore how system performance varies with clustering method, the number of clusters and the amount of adapting data. Results on the LDC Emotion Prosody and FAU Aibo Corpora show that this method outperforms speaker bootstrap, both in terms of relieving computation load and producing higher accuracy.
Ni Ding, Julien Epps
INTERSPEECH2
2012 Multimodal behavior and interaction as indicators of cognitive load
abstract
High cognitive load arises from complex time and safety-critical tasks, for example, mapping out flight paths, monitoring traffic, or even managing nuclear reactors, causing stress, errors, and lowered performance. Over the last five years, our research has focused on using the multimodal interaction paradigm to detect fluctuations in cognitive load in user behavior during system interaction. Cognitive load variations have been found to impact interactive behavior: by monitoring variations in specific modal input features executed in tasks of varying complexity, we gain an understanding of the communicative changes that occur when cognitive load is high. So far, we have identified specific changes in: speech, namely acoustic, prosodic, and linguistic changes; interactive gesture; and digital pen input, both interactive and freeform. As ground-truth measurements, galvanic skin response, subjective, and performance ratings have been used to verify task complexity. The data suggest that it is feasible to use features extracted from behavioral changes in multiple modal inputs as indices of cognitive load. The speech-based indicators of load, based on data collected from user studies in a variety of domains, have shown considerable promise. Scenarios include single-user and team-based tasks; think-aloud and interactive speech; and single-word, reading, and conversational speech, among others. Pen-based cognitive load indices have also been tested with some success, specifically with pen-gesture, handwriting, and freeform pen input, including diagraming. After examining some of the properties of these measurements, we present a multimodal fusion model, which is illustrated with quantitative examples from a case study. The feasibility of employing user input and behavior patterns as indices of cognitive load is supported by experimental evidence. Moreover, symptomatic cues of cognitive load derived from user behavior such as acoustic speech signals, transcribed text, digital pen trajectories of handwriting, and shapes pen, can be supported by well-established theoretical frameworks, including O'Donnell and Eggemeier's workload measurement [1986] Sweller's Cognitive Load Theory [Chandler and Sweller 1991], and Baddeley's model of modal working memory [1992] as well as McKinstry et al.'s [2008] and Rosenbaum's [2005] action dynamics work. The benefit of using this approach to determine the user's cognitive load in real time is that the data can be collected implicitly that is, during day-to-day use of intelligent interactive systems, thus overcomes problems of intrusiveness and increases applicability in real-world environments, while adapting information selection and presentation in a dynamic computer interface with reference to load.
Fang Chen 0001, Natalie Ruiz, Eric H. C. Choi, Julien Epps, M. Asif Khawaja, Ronnie Taib, Bo Yin 0002, Yang Wang 0002
ACM Trans. Interact. Intell. Syst.4
2011 Speaker verification using sparse representation classification
abstract
Sparse representations of signals have received a great deal of attention in recent years, and the sparse representation classifier has very lately appeared in a speaker recognition system. This approach represents the (sparse) GMM mean supervector of an unknown speaker as a linear combination of an over-complete dictionary of GMM supervectors of many speaker models, and ℓ1-norm minimization results in a non-zero coefficient corresponding to the unknown speaker class index. Here this approach is tested on large databases, introducing channel-/session-variability compensation, and fused with a GMM-SVM system. Evaluations on the NIST 2001 SRE and NIST 2006 SRE database show that when the outputs of the MFCC UBM-GMM based classifier (for NIST 2001 SRE) or MFCC GMM-SVM based classifier (for NIST 2006 SRE) are fused with the MFCC GMM Sparse Representation Classifier (GMM-SRC) based classifier, an absolute gain of 1.27% and 0.25% in EER can be achieved respectively.
Jia Min Karen Kua, Eliathamby Ambikairajah, Julien Epps, Roberto Togneri
ICASSP3
2011 Using clustering comparison measures for speaker recognition
abstract
Recent results seem to cast some doubt over the assumption that improvements in fused recognition accuracy for speaker recognition systems based on different acoustic features are due mainly to the different origins of the features (e.g. magnitude, phase, modulation information). In this study, we utilize clustering comparison measures to investigate acoustic and speaker modelling aspects of the speaker recognition task separately and demonstrate that front-end diversity can be achieved purely through different 'partitioning' of the acoustic space. Further, features that exhibit good 'stability' with respect to repeated clustering are shown to also give good EER performance in speaker recognition. This has implications for feature choice, fusion of systems employing different features, and for UBM data selection. A method for the latter problem is presented that gives up to an 11% relative reduction in EER using only 20-30% of the usual UBM training data set.
Jia Min Karen Kua, Julien Epps, Mohaddeseh Nosratighods, Eliathamby Ambikairajah, Eric H. C. Choi
ICASSP2
2011 Voice source features for cognitive load classification
abstract
Previous work in speech-based cognitive load classification has shown that the glottal source contains important information for cognitive load discrimination. However, the reliability of glottal flow features depends on the accuracy of the glottal flow estimation, which is a non-trivial process. In this paper, we propose the use of acoustic voice source features extracted directly from the speech spectrum (or cepstrum) for cognitive load classification. We also propose pre and post-processing techniques to improve the estimation of the cepstral peak prominence (CPP). 3-class classification results on two databases showed CPP as a promising cognitive load classification feature that outperforms glottal flow features. Score-level fusion of the CPP-based classification system with a formant frequency-based system yielded a final improved accuracy of 62.7%, suggesting that CPP contains useful voice source information that complements the information captured by vocal tract features.
Tet Fei Yap, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi
ICASSP2
2011 Measuring Cognitive Workload with Low-Cost Electroencephalograph
Avi Knoll, Yang Wang 0002, Fang Chen 0001, Jie Xu 0008, Natalie Ruiz, Julien Epps, Pega Zarjam
INTERACT (4)6
2011 Building an Audio-Visual Corpus of Australian English: Large Corpus Collection with an Economical Portable and Replicable Black Box
abstract
The Big Australian Speech Corpus project incorporates the strategic goals of 30 Chief Investigators from various speech science areas. Speech from 1000 geographically and socially diverse speakers is being recorded using a uniform and automated protocol plus standardized hardware and software to produce a widely applicable and extensible database – AusTalk. Here we describe the project’s major components and organization; share the lessons learnt from difficulties and challenges; and present the results achieved so far.
Denis Burnham, Dominique Estival, Steven Fazio, Jette Viethen, Felicity Cox, Robert Dale, Steve Cassidy, Julien Epps, Roberto Togneri, Michael Wagner 0004, Yuko Kinoshita, Roland Göcke, Joanne Arciuli, Mark Onslow, Trent W. Lewis, Andrew Butcher, John Hajek
INTERSPEECH8
2011 An Investigation of Depressed Speech Detection: Features and Normalization
abstract
In recent years, the problem of automatic detection of mental illness from the speech signal has gained some initial interest, however questions remaining include how speech segments should be selected, what features provide good discrimination, and what benefits feature normalization might bring given the speaker-specific nature of mental disorders. In this paper, these questions are addressed empirically using classifier configurations employed in emotion recognition from speech, evaluated on a 47-speaker depressed/neutral read sentence speech database. Results demonstrate that (1) detailed spectral features are well suited to the task, (2) speaker normalization provides benefits mainly for less detailed features, and (3) dynamic information appears to provide little benefit. Classification accuracy using a combination of MFCC and formant based features approached 80 % for this database. Index Terms: mental state recognition, depressed speech, feature comparison, MFCC, Gaussian mixture models
Nicholas Cummins, Julien Epps, Michael Breakspear, Roland Göcke
INTERSPEECH2
2011 Eye activity as a measure of human mental effort in HCI
abstract
The measurement of a user's mental effort is a problem whose solutions may have important applications to adaptive interfaces and interface evaluation. Previous studies have empirically shown links between eye activity and mental effort; however these have usually investigated only one class of eye activity on tasks atypical of HCI. This paper reports on research into eight eye activity based features, spanning eye blink, pupillary response and eye movement information, for real time mental effort measurement. Results from an experiment conducted using a computer-based training system show that the three classes of eye features are capable of discriminating different cognitive load levels. Correlation analysis between various pairs of features suggests that significant improvements in discriminating different effort levels can be made by combining multiple features. This shows an initial step towards a real-time cognitive load measurement system in human-computer interaction.
Siyuan Chen 0002, Julien Epps, Natalie Ruiz, Fang Chen 0001
IUI2
2011 Cognitive load evaluation of handwriting using stroke-level features
abstract
This paper examines several writing features for the evaluation of cognitive load. Our analysis is focused on writing features within and between written strokes, including writing pressure, writing velocity, stroke length and inter-stroke movements. Based on a study of 20 subjects performing a sentence composition task, the reported findings reveal that writing pressure and writing velocity information are very good indicators of cognitive load. A stroke selection threshold was investigated for constraining the feature extraction to long strokes, which resulted in a small further improvement. Differing from most previous research investigating cognitive load during writing based on task performance criteria, this work proposes a new approach to cognitive load measurement using writing dynamics, with the potential to allow new or improve existing handwriting interfaces.
Kun Yu 0001, Julien Epps, Fang Chen 0001
IUI2
2011 Investigation of spectral centroid features for cognitive load classification
Phu Ngoc Le, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Eric H. C. Choi
Speech Commun.3
2010 Glottal features for speech-based cognitive load classification
abstract
Cognitive load measurement is important when designing adaptive interfaces that optimize the performance of users working on high mental load tasks. Recent research on automatic speech-based measurement system indicates that cognitive load information is more prominent in the frequency region below 1 kHz. This study investigates the effects of cognitive load on glottal parameters (open quotient, normalized amplitude quotient and speed quotient), and proposes a system employing these parameters as features for cognitive load classification. Analysis of the glottal parameter distributions suggests that an increase in cognitive load can be related to a more creaky voice quality. Additionally, three-class classification results show that score-level fusion of systems based on the glottal features and baseline features (MFCCs, pitch, intensity and shifted delta cepstra) improves the baseline accuracy from 79% to 84%.
Tet Fei Yap, Julien Epps, Eric H. C. Choi, Eliathamby Ambikairajah
ICASSP2
2010 minCEntropy: A Novel Information Theoretic Approach for the Generation of Alternative Clusterings
abstract
Traditional clustering has focused on creating a single good clustering solution, while modern, high dimensional data can often be interpreted, and hence clustered, in different ways. Alternative clustering aims at creating multiple clustering solutions that are both of high quality and distinctive from each other. Methods for alternative clustering can be divided into objective-function-oriented and data-transformation-oriented approaches. This paper presents a novel information theoretic-based, objective-function-oriented approach to generate alternative clusterings, in either an unsupervised or semi-supervised manner. We employ the conditional entropy measure for quantifying both clustering quality and distinctiveness, resulting in an analytically consistent combined criterion. Our approach employs a computationally efficient nonparametric entropy estimator, which does not impose any assumption on the probability distributions. We propose a partitional clustering algorithm, named minCEntropy, to concurrently optimize both clustering quality and distinctiveness. minCEntropy requires setting only some rather intuitive parameters, and performs competitively with existing methods for alternative clustering.
Xuan Vinh Nguyen, Julien Epps
ICDM2
2010 A Study of Voice Source and Vocal Tract Filter Based Features in Cognitive Load Classification
abstract
Speech has been recognized as an attractive method for the measurement of cognitive load. Previous approaches have used mel frequency cepstral coefficients (MFCCs) as discriminative features to classify cognitive load. The MFCCs contain information from both the voice source and the vocal tract, so that the individual contributions of each to cognitive load variation are unclear. This paper aims to extract speech features related to either the voice source or the vocal tract and use them to discriminate between cognitive load levels in order to identify the individual contribution of each for cognitive load measurement. Voice source-related features are then used to improve the performance of current cognitive load classification systems, using adapted Gaussian mixture models. Our experimental result shows that the use of voice source feature could yield around 12% reduction in relative error rate compared with the baseline system based on MFCCs, intensity, and pitch contour.
Phu Ngoc Le, Julien Epps, Eric H. C. Choi, Eliathamby Ambikairajah
ICPR2
2010 An investigation of formant frequencies for cognitive load classification
Tet Fei Yap, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi
INTERSPEECH2
2010 Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance
Xuan Vinh Nguyen, Julien Epps, James Bailey 0001
J. Mach. Learn. Res.2
2010 Perceptual speech enhancement exploiting temporal masking properties of human auditory system
Teddy Surya Gunawan, Eliathamby Ambikairajah, Julien Epps
Speech Commun.3
2010 A segment selection technique for speaker verification
Mohaddeseh Nosratighods, Eliathamby Ambikairajah, Julien Epps, Michael J. Carey 0002
Speech Commun.3
2009 A Novel Approach for Automatic Number of Clusters Detection in Microarray Data Based on Consensus Clustering
abstract
Estimating the true number of clusters in a data set is one of the major challenges in cluster analysis. Yet in certain domains,knowing the true number of clusters is of high importance. For example, in medical research, detecting the true number of groups and sub-groups of cancer would be of utmost importance for their effective treatment. In this paper we propose a novel method to estimate the number of clusters in a micro array data set based on the consensus clustering approach. Although the main objective of consensus clustering is to discover a robust and high quality cluster structure in a data set, closer inspection of the set of clusterings obtained can often give valuable information about the appropriate number of clusters present. More specifically, the set off clusterings obtained when the specified number of clusters coincides with the true number of clusters tends to be less diverse.To quantify this diversity we develop a novel index, namely the Consensus Index (CI), which is built upon a suitable clustering similarity measure such as the well known Adjusted Rand Index (ARI)or our recently developed, information theoretic based index, namely the Adjusted Mutual Information (AMI). Our experiments on both synthetic and real microarray data sets indicate that the CI is a useful indicator for determining the appropriate number of clusters.
Xuan Vinh Nguyen, Julien Epps
BIBE2
2009 The I4U system in NIST 2008 speaker recognition evaluation
abstract
This paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU).
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin
ICASSP17
2009 Evaluation of a fused FM and cepstral-based speaker recognition system on the NIST 2008 SRE
abstract
In this paper, the fusion of two speaker recognition subsystems, one based on Frequency Modulation (FM) and another on MFCC features, is reported. The motivation for their fusion was to improve the recognition accuracy across different types of channel variations, since the two features are believed to contain complementary information. It was found that the MFCC-based subsystem outperformed the FM-based subsystem on telephone conversations from NIST SRE-06 dataset, while the opposite was true for NIST SRE-08 telephone data. As a result, the FM-based subsystem performed as well as the MFCC-based subsystem and their fusion gave up to 23% relative improvement in terms of EER over the MFCC subsystem alone, when evaluated on the NIST 2008 core condition.
Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Bin Ma 0001, Haizhou Li 0001
ICASSP3
2009 Speaker dependency of spectral features and speech production cues for automatic emotion classification
abstract
Spectral and excitation features, commonly used in automatic emotion classification systems, parameterise different aspects of the speech signal. This paper groups these features as speech production cues, broad spectral measures and detailed spectral measures and looks at how they differ in their performance in both speaker dependent and speaker independent systems. The extent of speaker normalisation on these features is also considered. Combinations of different features are then compared in terms of classification accuracies. Evaluations were conducted on the LDC emotional speech corpus for a five-class problem. Results indicate that MFCCs are very discriminative but suffer from speaker variability. Further, results suggest that the best front end for a speaker independent system is a combination of pitch, energy and formant information.
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps
ICASSP3
2009 Information theoretic measures for clusterings comparison: is a correction for chance necessary?
abstract
Information theoretic based measures form a fundamental class of similarity measures for comparing clusterings, beside the class of pair-counting based and set-matching based measures. In this paper, we discuss the necessity of correction for chance for information theoretic based measures for clusterings comparison. We observe that the baseline for such measures, i.e. average value between random partitions of a data set, does not take on a constant value, and tends to have larger variation when the ratio between the number of data points and the number of clusters is small. This effect is similar in some other non-information theoretic based measures such as the well-known Rand Index. Assuming a hypergeometric model of randomness, we derive the analytical formula for the expected mutual information value between a pair of clusterings, and then propose the adjusted version for several popular information theoretic based measures. Some examples are given to demonstrate the need and usefulness of the adjusted measures.
Xuan Vinh Nguyen, Julien Epps, James Bailey 0001
ICML2
2009 LS regularization of group delay features for speaker recognition
Jia Min Karen Kua, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi
INTERSPEECH2
2009 Pitch contour parameterisation based on linear stylisation for emotion recognition
abstract
The pitch contour contains information that characterises the emotion being expressed by speech, and consequently features extracted from pitch form an integral part of many automatic emotion recognition systems. While pitch contours may have many small variations and hence are difficult to represent compactly, it may be possible to parameterise them by approximating the contour for each voiced segment by a straight line. This paper looks at such a parameterisation method in the context of emotion recognition. Listening tests were performed to subjectively determine if the linearly stylised contours were able to sufficiently capture information pertaining to emotions expressed in speech. Furthermore these parameters were used as features for an automatic 5-class emotion classification system. The use of the proposed parameters rather than pitch statistics resulted in a relative increase in accuracy of about 20%.
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH3
2009 Analysis of band structures for speaker-specific information in FM feature extraction
Tharmarajah Thiruvaran, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH3
2008 Optimizing period-3 methods for eukaryotic gene prediction
abstract
In this paper, we firstly investigate the effect of window lengths on selected signal processing-based gene and exon prediction methods. We then optimize these methods to improve their prediction accuracy by employing the best DNA representation, a suitable window length, and boosting the output signals to enhance protein coding and suppress the non-coding regions. It is shown herein that the proposed method outperforms major existing time-domain, frequency- domain, and combined time-frequency approaches. By comparison with the existing DFT-based methods, the proposed method not only requires 50% less processing but also exhibits relative improvements of 53.3%, 46.7%, and 24.2% respectively over spectral content, spectral rotation, and paired and weighted spectral rotation measures in terms of prediction accuracy of exonic nucleotides at a 5% false positive rate using the GENSCAN test set.
Mahmood Akhtar, Eliathamby Ambikairajah, Julien Epps
ICASSP3
2008 A self-directed learning approach to signal processing education
abstract
This paper describes a self-directed, project-based learning scheme implemented in an introductory Signal Processing course at the University of New South Wales. The course was structured around a major laboratory project in which students were required to research course material, understand the relevant theory, and apply this in order to arrive at a solution. Lectures were delivered via prerecorded DVDs, allowing students to self-pace their absorption of new content and allowing teaching staff to concentrate on specific student issues during face-to-face classes. Evaluation of the course structure by the lecturer and instructors suggested that students gained a better conceptual understanding of signal processing theory than in previous years. Students were generally positive towards the process, but found it difficult to adjust to.
Eliathamby Ambikairajah, Julien Epps, Samuel J. Freney, Ming Sheng
ICASSP2
2008 Empirical mode decomposition based weighted frequency feature for speech-based emotion classification
abstract
This paper focuses on speech based emotion classification utilizing acoustic data. The most commonly used acoustic features are pitch and energy, along with prosodic information like rate of speech. We propose the use of a novel feature based on instantaneous frequency obtained from the speech, in addition to the aforementioned features, in order to take into account the vocal tract parameters as well as vocal chord excitation. The proposed features employ the recently emerged empirical mode decomposition to decompose speech into AM-FM signals that are symmetric about zero and suitable for Hilbert transformation to extract the instantaneous frequency. The proposed features provide a relative increase in classification accuracy of approximately 9% when appended to established acoustic features.
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps
ICASSP3
2008 Digital Signal Processing Techniques for Gene Finding in Eukaryotes
Mahmood Akhtar, Eliathamby Ambikairajah, Julien Epps
ICISP3
2008 Phonetic and speaker variations in automatic emotion classification
abstract
The speech signal contains information that characterises the speaker and the phonetic content, together with the emotion being expressed. This paper looks at the effect of this speakerand phoneme-specific information on speech-based automatic emotion classification. The performances of a classification system using established acoustic and prosodic features for different phonemes are compared, in both speaker-dependent and speaker-independent modes, using the LDC Emotional Prosody speech corpus. Results from these evaluations indicate that speaker variability is more significant than phonetic variations. They also suggest that some phonemes are easier to classify than others.
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH3
2008 FM features for automatic forensic speaker recognition
abstract
Frequency modulation (FM) information from the speech signal is herein proposed to complement the conventional amplitude based features for automatic forensic speaker recognition systems. In addition to presenting the AM-FM model of speech used to generate the proposed frequency modulation features, the significance of frequency modulation for speaker recognition is discussed. Evaluation results from an automatic forensic speaker recognition system combining FM and MFCC features are shown to out-perform those of a system employing MFCC features alone, in terms of all typical metrics, such as detection error trade-off curves, Tippett curves and applied probability of error curves. Index Terms: frequency modulation, automatic forensic speaker recognition.
Tharmarajah Thiruvaran, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH3
2007 Time and Frequency Domain Methods for Gene and Exon Prediction in Eukaryotes
abstract
The detection of period-3 components in exons of eukaryotic gene sequences enables signal processing based time-domain and frequency-domain methods to predict these regions. In this paper, we improve the prediction accuracy of frequency-domain methods by proposing a new algorithm known as the paired and weighted spectral rotation (PWSR) measure, which exploits both period-3 behaviour and another useful statistical property of genomic sequences. By comparison with existing frequency-domain approaches, the proposed PWSR method reveals relative improvements of 15.2% and 10.7% respectively over spectral content and spectral rotation measures in terms of prediction accuracy of exonic nucleotides at a 10% false positive rate using the GENSCAN test set. Finally, we combine the proposed PWSR with an existing time-domain method to demonstrate further signal processing-based improvements in gene and exon prediction accuracy.
Mahmood Akhtar, Julien Epps, Eliathamby Ambikairajah
ICASSP (2)2
2007 P-Value Segment Selection Technique for Speaker Verification
abstract
This paper presents a segment selection technique for discarding portions of speech that result in poor discrimination ability in speaker verification tasks. Theory supporting the significance of a frame selection procedure for test segments, prior to making decisions, is also developed. This approach has the ability to reduce the effect of the acoustic regions of speech that are not accurately represented due to a lack of training data. Compared with a baseline system using both CMS and variance normalization, the proposed segment selection technique brings 24% relative reduction in error rate over the entire testing data of the 2002 NIST Dataset in terms of minimum DCF. For short test segments, i.e. less than 15 seconds, the novel frame dropping technique produces a significant relative error rate reduction of 23% in terms of minimum DCF.
Mohaddeseh Nosratighods, Eliathamby Ambikairajah, Julien Epps, Michael J. Carey 0002
ICASSP (4)3
2007 Group delay features for emotion detection
abstract
This paper focuses on speech based emotion classification utilizing acoustic data. The most commonly used acoustic features are pitch and energy, along with prosodic information like the rate of speech. We propose the use of a novel feature based on the phase response of an all-pole model of the vocal tract obtained from linear predictive coefficients (LPC), in addition to the aforementioned features. We compare this feature to other commonly used acoustic features based on classification accuracy. The back-end of our system employs a probabilistic neural network based classifier. Evaluations conducted on the LDC Emotional Prosody speech corpus indicate the proposed features are well suited to the task of emotion classification. The proposed features are able to provide a relative increase in classification accuracy of about 14% over established features when combined with them to form a larger feature vector.
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH3
2006 Warped Magnitude and Phase-Based Features for Language Identification
abstract
To date, systems for the identification of spoken languages have normally used magnitude-based parameterization methods such as the MFCC and PLP. This paper investigates the use of the recently proposed modified group delay function (MODGDF) coefficients in combination with traditional magnitude-based features in a Gaussian mixture model (GMM) based system. We also examine the application of feature warping to magnitude-based features and the MODGDF and find that it can offer a significant cumulative improvement. We find that the addition of a modified regression-based shifted delta cepstrum (SDC) further improves system performance beyond that obtained by a more standard SDC configuration. The combination of PLP, feature warping and the proposed regression-based SDC achieved an accuracy of 88.4% in tests on 10 languages in the OGI TS Corpus, which compares very favourably with alternative language identification systems reported in the literature
Felicity Allen, Eliathamby Ambikairajah, Julien Epps
ICASSP (1)3
2006 Improving Separability of Eeg Signals During Motor Imagery With An Efficient Circular Laplacian
abstract
This paper reports on a new EEG re-referencing scheme, known as the circular Laplacian, for processing multichannel EEG signals. The new reference signals are derived from the average potentials on the circles around the electrodes. The radii of the circles can be adjusted to achieve spatial filtering of EEG at different frequencies. Evaluation with motor imagery recordings suggests that the circular Laplacian leads to a maximum of 5% improvement over the traditional discrete Laplacian in a motor imagery classification task.
Julien Epps
ICASSP (2)2
2006 Classifying EEG for brain-computer interfaces: learning optimal filters for dynamical system features
abstract
Classification of multichannel EEG recordings during motor imagination has been exploited successfully for brain-computer interfaces (BCI). In this paper, we consider EEG signals as the outputs of a networked dynamical system (the cortex), and exploit novel features from the collective dynamics of the system for classification. Herein, we also propose a new framework for learning optimal filters automatically from the data, by employing a Fisher ratio criterion. Experimental evaluations comparing the proposed dynamical system features with the CSP and the AR features reveal their competitive performance during classification. Results also show the benefits of employing the spatial and the temporal filters optimized using the proposed learning approach.
Julien Epps
ICML2
2005 Experiences with an electronic whiteboard teaching laboratory and tablet PC based lecture presentations [DSP courses]
abstract
This paper presents our experience in constructing an electronic whiteboard-based computer laboratory for teaching digital signal processing (DSP) courses in Australian undergraduate and postgraduate programs. Student interaction with the electronic whiteboard-based tutorial class environment is also reported. Away from the laboratory, DSP lectures were presented using a tablet PC as a digital whiteboard. This supported high quality handwriting annotation of lecture slides, and overcame the limited flexibility present in the existing PowerPoint mode of lecture delivery. For selected self-paced tutorial questions, solutions were provided in electronic format comprising the lecturer's handwritten explanation on a blank slide, input using the tablet PC, combined with audio commentary. An evaluation of student opinions towards this multi-mode delivery of DSP education was illuminating, and the overall experience with these technological aids was that signal processing could be effectively and naturally taught with high student attention span.
Eliathamby Ambikairajah, Julien Epps, Ming Sheng, Branko G. Celler, Peter Chen
ICASSP (5)2
2005 A study of manual gesture-based selection for the PEMMI multimodal transport management interface
abstract
Operators of traffic control rooms are often required to quickly respond to critical incidents using a complex array of multiple keyboards, mice, very large screen monitors and other peripheral equipment. To support the aim of finding more natural interfaces for this challenging application, this paper presents PEMMI (Perceptually Effective Multimodal Interface), a transport management system control prototype taking video-based manual gesture and speech recognition as inputs. A specific theme within this research is determining the optimum strategy for gesture input in terms of both single-point input selection and suitable multimodal feedback for selection. It has been found that users tend to prefer larger selection areas for targets in gesture interfaces, and tend to select within 44% of this selection radius. The minimum effective size for targets when using 'device-free' gesture interfaces was found to be 80 pixels (on a 1280x1024 screen). This paper also shows that feedback on gesture input via large screens is enhanced by the use of both audio and visual cues to guide the user's multimodal input. Audio feedback in particular was found to improve user response time by an average of 20% over existing gesture selection strategies for multimodal tasks.
Fang Chen 0001, Eric H. C. Choi, Julien Epps, Serge Lichman, Natalie Ruiz, Yu (David) Shi, Ronnie Taib, Mike Wu
ICMI3
2005 An energy search approach to variable frame rate front-end processing for robust ASR
abstract
Extensive research has been devoted to robustness in the presence of various types and degrees of environmental noise over the past several years, however this remains one of the main problems facing automatic speech recognition systems. This paper describes a new variable frame rate analysis technique, based upon searching a predefined lookahead interval for the next frame position that maximizes the firstorder difference of the log energy (ΔE) between the consecutive frames. The application of this novel technique to noise-robust ASR front-end processing is also reported. In comparison with existing variable frame rate methods in the literature, the proposed energy search approach is simpler and achieves similar recognition accuracy improvements at lower complexity. Experimental work on the Aurora II connected digits database reveals that the proposed front-end, together with cumulative distribution mapping, achieves average digit recognition accuracies of 78.32% for a model set trained from clean data and 89.95% for a model set trained from data with multiple noise conditions, representing 6.1% and 2.3% reductions in word error rates respectively over a cumulative distribution mapping baseline.
Julien Epps, Eric H. C. Choi
INTERSPEECH1
2005 Language Identification using Warping and the Shifted Delta Cepstrum
abstract
This paper proposes the novel use of feature warping for automatic language identification, in combination with the shifted delta cepstrum (SDC) and perceptual linear predictive coefficients in a Gaussian mixture model (GMM) based system. Experimental results on various configurations of front-end techniques reported herein demonstrate that, besides providing robustness against channel mismatch and noise as found in existing literature, feature warping is useful more generally as a technique for pre-mapping data for improved compatibility with a GMM back-end. The configuration reported in this paper provides a language identification performance of 76.4% using the OGI/NIST database, a 46.5% relative reduction in error rate when compared with a benchmark system employing Mel frequency cepstral coefficients and the SDC
Felicity Allen, Eliathamby Ambikairajah, Julien Epps
MMSP3
2004 Perceptual wavelet packet audio coder
abstract
Traditional wavelet packet audio compression algorithms do not utilize the temporal masking properties of the human auditory system, relying instead on simultaneous masking models. This paper presents the design and implementation of a perceptual wavelet audio coder by incorporating temporal and simultaneous masking models. The efficiency of the encoder was assessed based upon the number of bits required to code wavelet packet coefficients in each critical band, while retaining perceptual transparency. Subjective listening tests conforming to ITU-R BS.1116 revealed the bit rate is reduced by more than 17% compared to using a coder that only employs a simultaneous masking model.
Teddy Surya Gunawan, Eliathamby Ambikairajah, Julien Epps
INTERSPEECH3
2003 Evaluation of a virtual teaching laboratory for signal processing education
abstract
The paper presents our experience in teaching a digital signal processing (DSP) course in an Australian postgraduate program entirely using virtual tele-lectures. A virtual teaching laboratory was designed for this purpose, allowing students to receive fully interactive, real time lectures delivered from a remote international location. We present the methodology and technology used to develop a complete set of tele-lectures and online tools for a course entitled 'Signal processing and applications'. An evaluation of student opinions towards the virtual teaching laboratory revealed that 90% of students rapidly became comfortable with the use of this new educational facility, among other results. The overall experience with the VTL was that signal processing can be effectively and naturally taught in this mode and that there are great potential benefits in connecting the signal processing research and educational community.
Eliathamby Ambikairajah, Julien Epps, Ming Sheng, Branko G. Celler
ICASSP (3)2
2003 Temporal structure constrained transformation for speaker adaptation
abstract
We suggest that rather than modeling speaker mismatch as an affine transform of the entire feature vector, it can be modeled by an affine transform of the static coefficients with additional constraints imposed by the temporal relationships of the streams of coefficients. This results in the different streams sharing the same rotation matrix, and thus reduces the complexity and memory requirements for speaker adaptation, as well as minimizes the adaptation data requirements. We present the solution for the case where temporal structure constrained transforms (TSCT) are optimized using the maximum likelihood criterion. The experiments presented in the paper show that with the proposed approach, the same accuracy after adaptation for the Wall Street Journal (WSJ) task can be achieved by using only 60% of the total number of transformation parameters that it would require if conventional block-diagonal transformation is used. In addition, TSCT provides better recognition accuracy when there is only a very limited amount of adaptation data.
Eric H. C. Choi, Trym Holter, Julien Epps, Arun Gopalakrishnan
ICASSP (1)3
2003 Decomposition of speech into voiced and unvoiced components based on a state-space signal model
abstract
We present a novel method for decomposing speech into voiced and unvoiced components. After demodulating the variations in the spectral envelope, energy and pitch, the method involves applying a bank of Kalman filters to separate the harmonic and non-harmonic components of the signal. This approach relies on a state-space representation of the composite signal, and provides a way to estimate accurately the harmonic component without the large delay required by a linear phase comb filter. However it also requires prior knowledge of the variance of the unvoiced component and the state transition parameters. We present a novel method to determine these parameters accurately based on a variant of the expectation-maximization algorithm. Modifications for dealing with unvoiced segments and voicing onset are also described.
Mark Thomson, Simon Boland, Mike Wu, Julien Epps, Michael Smithers
ICASSP (1)4
2001 Wideband speech and audio coding using gammatone filter banks
abstract
Considerable research attention has been directed towards speech and audio coding algorithms capable of producing high quality coded speech and audio, however few of these use signal representations which account for temporal as well as spectral detail. This paper presents a new technique for 16 kHz wideband speech and audio coding, whereby analysis and synthesis are performed using a linear phase gammatone filter bank. The outputs of these critical band filters are processed to obtain a series of pulse trains that represent neural firing. Auditory masking is then applied to reduce the number of pulses, producing a more compact time-frequency parameterization. The critical band gains and pulse amplitudes and positions are then coded using a combination of non-uniform quantization, arithmetic coding and vector quantization. This coding paradigm produces high quality coded speech and audio, is based upon well-known models of the auditory system, is highly scalable, and has moderate complexity.
Eliathamby Ambikairajah, Julien Epps, Lee Lin
ICASSP2
1998 Speech enhancement using STC-based bandwidth extension
abstract
Telephone speech is typically bandlimited to 4 kHz, resulting in a 'muffled' quality.Coding speech with bandwidth greater than 4 kHz reduces this distortion, but requires a higher bit rate to avoid other types of distortion.An alternative to coding wider bandwidth speech is to exploit correlation between the 0-4 kHz and 4-8 kHz speech bands to resynthesize wideband speech from narrowband speech.This paper presents a method for re-synthesizing narrowband coded speech using sinusoidal transform coding (STC), modified codebook mapping and a novel method for the synthesis of highband unvoiced components.Informal listening test results indicate that this method produces a significant quality improvement in speech which has been coded using narrowband standards.
Julien Epps, W. Harvey Holmes
ICSLP1
1997 Real time measurements of the vocal tract resonances during speech
abstract
The formants of speech sounds are usually attributed to resonances of the vocal tract. Formant frequencies are usually estimated by inspection of spectrograms or by automated techniques such as linear prediction. In this paper we measure the frequencies of the first two resonances of the vocal tract directly, in real time, using acoustic impedance spectrometry. The vocal tract is excited by a carefully calibrated, broad band, acoustic current signal applied outside the lips while the subject is speaking. The sound pressure response is analysed to give the resonant frequencies. We compare this new method (Real-time Acoustic Vocal tract Excitation or RAVE) with linear prediction and we report the vocal tract resonances for eleven vowels of Australian English. We also report preliminary results of using feedback from vocal tract excitation as a speech trainer, and its effect on improving the pronunciation of foreign vowel sounds by monolingual anglophones.
Julien Epps, Annette Dowd, Joe Wolfe
EUROSPEECH1