EDBT 2026 Demo / reviewers in the wild / expert
Eliathamby Ambikairajah
dblp:24/3343
· DBLP profile ↗
143ranked-venue papers
13as first author
24since 2021 · last 2025
0000-0003-4673-6534ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 127 · 11 first-author · 14 since 2021Artificial intelligence and machine learning · 68 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance FeaturesabstractEstimating frequency-varying acoustic parameters is essential for enhancing immersive perception in realistic spatial audio creation. In this paper, we propose a unified framework that blindly estimates reverberation time (T60), direct-to-reverberant ratio (DRR), and clarity (C50) across 10 frequency bands using first-order Ambisonics (FOA) speech recordings as inputs. The proposed framework utilizes a novel feature named Spectro-Spatial Covariance Vector (SSCV), efficiently representing temporal, spectral as well as spatial information of the FOA signal. Our models significantly outperform existing single-channel methods with only spectral information, reducing estimation errors by more than half for all three acoustic parameters. Additionally, we introduce FOA-Conv3D, a novel back-end network for effectively utilising the SSCV feature with a 3D convolutional encoder. FOA-Conv3D outperforms the convolutional neural network (CNN) and recurrent convolutional neural network (CRNN) backends, achieving lower estimation errors and accounting for a higher proportion of variance (PoV) for all 3 acoustic parameters. Hanyu Meng, Jeroen Breebaart, Jeremy Stoddard, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 5 |
| 2025 | Experimental Demonstration of Integrated WiFi Communication and Occupancy MonitoringabstractAccurate occupancy counting is essential for effective space management, safety protocols, and resource optimisation in various environments. In recent years, the rapid development of the Internet of Things (IoT) has advanced traditional WiFi sensing into a new domain of research known as Integrated Sensing and Communication (ISAC). Such technology enables simultaneous communication and sensing, allowing for efficient data transmission while gathering environmental information. In this paper, we conduct a novel experimental demonstration and robustness validation of ISAC for occupancy counting. Compared to traditional WiFi sensing technology, we first demonstrated that human presence can be detected by sensing applications using communication packets. Specifically, using our Machine Learning (ML) algorithm, we got a sensing accuracy of more than 96 %. However, we found that this accuracy decreases with increased communication rates and higher occupancy levels, as they introduce noise and interference into the WiFi channel. We also experimentally investigated edge communication, which reveals a tradeoff between occupancy monitoring and communication rate regarding the number of Transmission Control Protocol (TCP) retransmission requests for different distances. We observed that improved accuracy results in an increase in TCP retransmission requests and a reduced communication quantity. Both sensing and communication performance lie in an optimal deployment of the transmit and receive with devices. Further tests at various distances revealed novel insight that sensing performance is better at both short and long distances. Specifically, we present that sensing benefits from diverse reflections, whereas communication relies on a balance best achieved at intermediate distances. Xihao Liang, Deepak Mishra 0001, Aruna Seneviratne, Eliathamby Ambikairajah |
ICC | 5 |
| 2025 | A Study of Speech Embedding Similarities Between Australian Aboriginal and High-Resource Languages
Eliathamby Ambikairajah, Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu |
INTERSPEECH | 1 |
| 2024 | ChatGPT in the Classroom: A Shift in Engineering Design EducationabstractArtificial intelligence tools like ChatGPT are increasingly being incorporated into our education paradigm. This paper explores how ChatGPT was used in an Electrical Engineering Design Proficiency course at the University of New South Wales in Sydney. The course is a term-long laboratory-based class that centres on independent student work in system design, implementation, and validation. Students were encouraged to consult ChatGPT for design solutions, explanations, and suggestions, with the requirement that they declare any use of AI tools. The assessment process was carefully designed to determine whether responses originated from AI tools or the students' own understanding. Notably, 70% of the fifty students in the class utilised ChatGPT to enhance their understanding of the subject. The paper will also discuss the specific design tasks given to students, the assessment process, and explore ChatGPT's potential as a supportive educational tool in other courses. Eliathamby Ambikairajah, Tharmakulasingam Sirojan, Tharmarajah Thiruvaran, Vidhyasaharan Sethu |
EDUCON | 1 |
| 2024 | A Tiered Learning Framework for Self-Guided Engineering Design EducationabstractThe Tiered Learning Framework is a multilayered system designed to aid students in assessing their progress and understanding within the learning process. Each layer corresponds to the skills students need to develop, and the framework encourages students to self-assess their current level and identify what they need to progress to higher tiers of learning. This paper explores the application of the Tiered Learning Framework in an Electrical Engineering Design Proficiency course at the University of New South Wales, Sydney. The course encompasses design tasks in critical areas such as electronic circuit design, signal processing design, and power system design. The paper focuses on the design and assessment of the signal processing design task as an example. In a class of 49 students, 96% agreed that the framework helped them achieve their learning goals, and 91% found it challenging yet encouraging for their learning. Compared to a non-tiered framework, 71% of the total students favour a tiered framework. Eliathamby Ambikairajah, Tharmarajah Thiruvaran, Vidhyasaharan Sethu, Deepak Mishra 0001, Tharmakulasingam Sirojan |
EDUCON | 1 |
| 2024 | An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech EnhancementabstractTransformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encoding exactly impacts speech enhancement based on Transformer architectures. In this paper, we perform a comprehensive empirical study evaluating five positional encoding methods, i.e., Sinusoidal and learned absolute position embedding (APE), T5-RPE, KERPLE, as well as the Transformer without positional encoding (No-Pos), across both causal and noncausal configurations. We conduct extensive speech enhancement experiments, involving spectral mapping and masking methods. Our findings establish that positional encoding is not quite helpful for the models in a causal configuration, which indicates that causal attention may implicitly incorporate position information. In a noncausal configuration, the models significantly benefit from the use of positional encoding. In addition, we find that among the four position embeddings, relative position embeddings outperform APEs. Qiquan Zhang, Meng Ge, Hongxu Zhu, Eliathamby Ambikairajah, Zhaoheng Ni, Haizhou Li 0001 |
ICASSP | 4 |
| 2024 | Dual-Constrained Dynamical Neural ODEs for Ambiguity-aware Continuous Emotion Prediction
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2024 | Binaural Selective Attention Model for Target Speaker Extraction
Hanyu Meng, Qiquan Zhang, Xiangyu Zhang 0005, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 5 |
| 2024 | An Exploration of Length Generalization in Transformer-Based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2024 | Continuous Emotion Ambiguity Prediction: Modeling With Beta DistributionsabstractConventional continuous emotion prediction systems are typically trained to predict the ‘average’ of affect ratings obtained from multiple human annotators. These systems, however, ignore the ambiguity inherent in the perceived emotions, which is not captured by the ‘average rating’. This paper presents a novel ambiguity-aware continuous emotion prediction system that predicts the time-varying emotion state as a series of beta distributions. Our recent work has shown beta distributions to be an effective parametric model of a collection of affect ratings. This work develops an appropriate cost function that enables neural networks to be trained to predict beta distributions. It also investigates the choice of parameterization of the beta distribution, the choice of activation functions of the output layer, and the tractability of gradient definitions in combination with the loss function. The proposed framework is implemented using a Bag-of-Audio-Words front-end and an LSTM-based back-end and evaluated on the RECOLA dataset. In addition to comparison with baseline systems that only predict the ‘average rating’, the effectiveness with which the predictions represent ambiguity in perceived emotions is also evaluated. Experimental results reveal that the proposed approach outperforms other ambiguity-aware systems, especially when predicting valence. Deboshree Bose, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Belief Mismatch Coefficient (BMC): A Novel Interpretable Measure of Prediction Accuracy for Ambiguous Emotion StatesabstractDespite of efforts made to model emotion ambiguity and develop ambiguity aware emotion prediction systems, there is a need for a quantitative and interpretable measure of the accuracy of such systems, regardless of recent advances in representing emotion ambiguity through probability distributions. In this paper, we propose a novel measure called the “Belief Mismatch Coefficient (BMC) that quantifies the differences in the belief that emotional states are perceived from certain regions within the arousal/valence space when comparing a predicted distribution to an underlying distribution inferred from ground truth ratings. The proposed metric is validated using simulated labels to demonstrate its effectiveness in quantifying various prediction errors. Furthermore, it is extended to real-case emotion prediction systems using two state-of-the-art modeling techniques on the RECOLA dataset. The experimental results confirm that the proposed metric can efficiently capture and differentiate between various prediction errors, while also offering insights into the predictions. Moreover, it demonstrates significant advantages in capturing a comprehensive view of the predicted distribution compared to traditional metrics such as Concordance Correlation Coefficients. Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ACII | 4 |
| 2023 | Constrained Dynamical Neural ODE for Time Series Modelling: A Case Study on Continuous Emotion PredictionabstractweA number of machine learning applications involve time series prediction, and in some cases additional information about dynamical constraints on the target time series may be available. For instance, it might be known that the desired quantity cannot change faster than some rate or that the rate is dependent on some known factors. However, incorporating these constraints into deep learning models, such as recurrent neural networks, is not straightforward. In this paper, we propose constrained dynamical neural ordinary differential equation (CD-NODE) models, which treat the desired time series as a dynamic process that can be described by an ODE. CD-NODEs model the rate of change of the time series as a function of both itself and the current input features, parameterised as a neural network. We explore the effect of constraining the dynamics of the model by placing explicit restrictions on the rate of change. The proposed model is evaluated on speech-based continuous emotion prediction, where such dynamical constraints are expected, using the publicly available RECOLA dataset. Results suggest that the model achieves performances comparable with the state-of-the-art despite using significantly fewer parameters. Additional analyses reveal that imposing these constraints on the model leads to faster convergence and better performance, especially with smaller training data sets. Ting Dang, Antoni Dimitriadis, Jingyao Wu 0002, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 5 |
| 2023 | What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy ConditionsabstractThere is increasing interest in the use of the LEArnable Front-end (LEAF) in a variety of speech processing systems. However, there is a dearth of analyses of what is actually learnt and the relative importance of training the different components of the front-end. In this paper, we investigate this question on keyword spotting, speech-based emotion recognition and language identification tasks and find that the filters for spectral decomposition and the low pass filter used to estimate spectral energy variations exhibit no learning and the per-channel energy normalisation (PCEN) is the key component that is learnt. Following this, we explore the potential of adapting only the PCEN layer with a small amount of noisy data to enable it to learn appropriate dynamic range compression that better suits the noise conditions. This in turn enables a system trained on clean speech to work more accurately on noisy test data as demonstrated by the experimental results reported in this paper. Hanyu Meng, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2023 | From Interval to Ordinal: A HMM based Approach for Emotion Label Conversion
Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2023 | DNN controlled adaptive front-end for replay attack detection systemsabstractDeveloping robust countermeasures to protect automatic speaker verification systems against replay spoofing attacks is a well-recognized challenge. Current approaches to spoofing detection are generally based on a fixed front-end, typically a time-invariant filter bank, followed by a machine learning back-end. In this paper, we propose a novel approach whereby the front-end comprises an adaptive filter bank with a deep neural network-based controller, which is jointly trained along with a neural network back-end. Specifically, the deep neural network-based adaptive filter controller tunes the selectivity and sensitivity of the front-end filter bank at every frame to capture replay-related artefacts. We demonstrate the effectiveness of the proposed framework in spoofing attack detection on a synthesized dataset and ASVSpoof 2019 and ASVSpoof 2021 challenge datasets in terms of equal error rate and its ability to capture artefacts that differentiate replayed signals from genuine ones in comparison to conventional non-adaptive front-end. Buddhi Wickramasinghe, Eliathamby Ambikairajah, Vidhyasaharan Sethu, Julien Epps, Haizhou Li 0001, Ting Dang |
Speech Commun. | 2 |
| 2023 | Ordinal Logistic Regression With Partial Proportional Odds for Depression PredictionabstractLike many psychological scales, depression scales are ordinal in nature. Depression prediction from behavioral signals has so far been posed either as classification or regression problems. However, these naive approaches have fundamental issues because they are not focused on ranking, unlike ordinal regression, which is the most appropriate approach. Ordinal regression to date has comparatively few methods when compared with other branches in machine learning, and its usage is limited to specific research domains. Ordinal logistic regression (OLR) is one such method, which is an extension for ordinal data of the well-known logistic regression, but is not familiar in speech processing, affective computing or depression prediction. The primary aim of this article is to investigate proportionality structures and model selection for the design of ordinal regression systems within the logistic regression framework. A new greedy-based algorithm for partial proportional odds model selection (GREP) is proposed that allows the parsimonious design of effective ordinal logistic regression models, which avoids an exhaustive search and outperforms model selection using the Brant test. Evaluations on the DAIC-WOZ and AViD depression corpora show that OLR models exploiting GREP can outperform two competitive baseline systems (GSR and CNN), in terms of both RMSE and Spearman correlation. Sadari Jayawardena, Julien Epps, Eliathamby Ambikairajah |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | A Novel Markovian Framework for Integrating Absolute and Relative Ordinal Emotion InformationabstractThere is growing interest in affective computing for the representation and prediction of emotions along ordinal scales. However, the term ordinal emotion label has been used to refer to both absolute notions such as low or high arousal, as well as relation notions such as arousal is higher at one instance compared to another. In this paper, we introduce the terminology absolute and relative ordinal labels to make this distinction clear and investigate both with a view to integrate them and exploit their complementary nature. We propose a Markovian framework referred to as Dynamic Ordinal Markov Model (DOMM) that makes use of both absolute and relative ordinal information, to improve speech based ordinal emotion prediction. Finally, the proposed framework is validated on two speech corpora commonly used in affective computing, the RECOLA and the IEMOCAP databases, across a range of system configurations. The results consistently indicate that integrating relative ordinal information improves absolute ordinal emotion prediction. Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | A Time-Frequency Attention Module for Neural Speech EnhancementabstractSpeech enhancement plays an essential role in a wide range of speech processing applications. Recent studies on speech enhancement tend to investigate how to effectively capture the long-term contextual dependencies of speech signals to boost performance. However, these studies generally neglect the time-frequency (T-F) distribution information of speech spectral components, which is equally important for speech enhancement. In this paper, we propose a simple yet very effective network module, which we term the T-F attention (TFA) module, that uses two parallel attention branches, i.e., time-frame attention and frequency-channel attention, to explicitly exploit position information to generate a 2-D attention map to characterise the salient T-F speech distribution. We validate our TFA module as part of two widely used backbone networks (residual temporal convolution network and Transformer) and conduct speech enhancement with four most popular training objectives. Our extensive experiments demonstrate that our proposed TFA module consistently leads to substantial enhancement performance improvements in terms of the five most widely used objective metrics, with negligible parameter overheads. In addition, we further evaluate the efficacy of speech enhancement as a front-end for a downstream speech recognition task. Our evaluation results show that the TFA module significantly improves the robustness of the system to noisy conditions. Qiquan Zhang, Xinyuan Qian 0001, Zhaoheng Ni, Aaron Nicolson, Eliathamby Ambikairajah, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | A Novel Sequential Monte Carlo Framework for Predicting Ambiguous Emotion StatesabstractWhen continuous emotion labelling of natural (non-acted) data is desired, it is typically collected from multiple annotators. However, most automatic emotion recognition systems trained on such data ignore disagreement between annotators and only models the average rating, despite the observation that the degree of disagreement would reflect the ambiguity and subtlety in every expression of emotions. In this paper, we propose a novel Sequential Monte Carlo framework that models the perceived emotion as time-varying distributions that allows for ambiguity to be incorporated. Additionally, we present alternative measures that consider both the similarity of prediction to the multiple labels, as well as whether the degree of ambiguity in the prediction and labels. The proposed system was validated on the publicly available RECOLA dataset. Jingyao Wu 0002, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 4 |
| 2022 | Sustainable Deep Learning at Grid Edge for Real-Time High Impedance Fault DetectionabstractHigh impedance faults (HIFs) on overhead power lines are known to cause fires. They are difficult to detect using conventional protection relays because the fault current is insufficient to cause tripping. The delay in detecting HIFs can result in severe bushfires and energy losses; hence a high throughput, low latency detection scheme needs to be developed for HIF detection. Moreover, the complexities associated with HIF detection demands signal processing techniques combined with artificial intelligence to achieve higher detection accuracy. This paper proposes a sustainable deep learning-based approach in an edge device, that can be mounted on top of a power pole to detect HIFs in real-time. Data acquisition, feature extraction, and deep learning based fault identification are performed in an embedded edge node to achieve higher throughput, reduced latency as well as offload the network traffic. Furthermore, optimization techniques such as hardware parallelism and pipelining are adapted to achieve real-time fault identification on edge devices while ensuring the efficient usage of its limited resources. Real-time implementation of the proposed system is validated through laboratory experiments and the results demonstrate the suitability of edge computing to detect HIFs in terms of reduced detection latency (115.2 ms) and higher detection accuracy (98.67 percent). Tharmakulasingam Sirojan, Shibo Lu, Bao Toan Phung, Daming Zhang 0001, Eliathamby Ambikairajah |
IEEE Trans. Sustain. Comput. | 5 |
| 2021 | AusKidTalk: An Auditory-Visual Corpus of 3- to 12-Year-Old Australian Children's SpeechabstractHere we present AusKidTalk [1], an audio-visual (AV) corpus of Australian children’s speech collected to facilitate the development of speech based technological solutions for children. It builds upon the technology and expertise developed through the collection of an earlier corpus of Australian adult speech, AusTalk [2,3]. This multi-site initiative was established to remedy the dire shortage of children’s speech corpora in Australia and around the world that are sufficiently sized to train accurate automated speech processing tools for children. We are collecting ~600 hours of speech from children aged 3–12 years that includes single word and sentence productions as well as narrative and emotional speech. In this paper, we discuss the key requirements for AusKidTalk and how we designed the recording setup and protocol to meet them. We also discuss key findings from our feasibility study of the recording protocol, recording tools, and user interface. Beena Ahmed, Kirrie J. Ballard, Denis Burnham, Tharmakulasingam Sirojan, Hadi Mehmood, Dominique Estival, Elise Baker, Felicity Cox, Joanne Arciuli, Titia Benders, Katherine Demuth, Barbara Kelly, Chloé Diskin-Holdaway, Mostafa Shahin, Vidhyasaharan Sethu, Julien Epps, Chwee Beng Lee, Eliathamby Ambikairajah |
Interspeech | 18 |
| 2021 | Parametric Distributions to Model Numerical Emotion Labels
Deboshree Bose, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
Interspeech | 3 |
| 2021 | An adaptive transmission line cochlear model based front-end for replay attack detection
Tharshini Gunendradasan, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001 |
Speech Commun. | 2 |
| 2021 | Compensation Techniques for Speaker Variability in Continuous Emotion PredictionabstractContinuous time-varying prediction of emotions based on speech in terms of attributes (i.e., arousal) has received considerable attention in the past few years. However, the variability introduced by factors not related to emotion, such as speaker and phonetic variability, which in turn may lead to less reliable models and less accurate emotion predictions, has not been fully explored yet. In particular, even though speaker variability has been shown to be a significant confounding factor in continuous emotion prediction systems, there remains a paucity of analyses about how speaker variability affects continuous emotion prediction systems and which methods can be applied to compensate for this variability. This paper first formulates speaker variability systematically in terms of probability distributions in both feature and model spaces, and quantifies the effect of speaker variability by comparing inter- and intra-speaker variability between speaker-dependent models. Second, two compensation techniques based on partial least squares dimensional reduction and feature mapping are proposed. Finally, the effectiveness of the proposed techniques is validated on three databases, across which they show consistent improvement in arousal, valence and dominance prediction. Additional quantitative analyse reveals that the two proposed techniques compensate for speaker variability in both the feature and model spaces simultaneously. Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | Cochlear Signal Processing: A Platform for Learning the Fundamentals of Digital Signal ProcessingabstractThe first digital signal processing course in most electrical engineering programmes around the world tends to be a significant jump in abstraction for most students. This is a consequence of them being introduced to a large number of mathematical concepts with insufficient time to consolidate the ideas with practical examples. In this paper we propose that cochlear signal processing is an excellent platform that brings together almost all the fundamental DSP concepts taught in an introductory course and also lends itself as a suitable candidate for project-based learning. The paper provides details on how such a project was setup and run in the introductory 3rdyear DSP course at UNSW Sydney. In addition, we also provide measure of student interest and engagement with the project as well as their feedback on its impact on their learning patterns and effectiveness at reinforcing theoretical concepts. Eliathamby Ambikairajah, Vidhyasaharan Sethu |
ICASSP | 1 |
| 2020 | Adversarial Multi-Task Learning for Speaker Normalization in Replay DetectionabstractSpoofing detection algorithms in voice biometrics are adversely affected by differences in the speech characteristics of the various target users. In this paper, we propose a novel speaker normalisation technique that employs adversarial multi-task learning to compensate for this speaker variability. The proposed system is designed to learn a feature space that discriminates between genuine and replayed speech while simultaneously reduces the discrimination between different speakers. We initially characterise the impact of speaker variability and quantify the effect of the proposed speaker normalisation technique directly on the feature distributions. Following this, we validate the technique on spoofing detection experiments carried out on two different corpora, ASVSpoof 2017 v2.0 and BTAS 2016 replay, and demonstrate its effectiveness. We obtain EER of 7.11% and 0.83% on the two corpora respectively, lower than that of all relevant baselines. Gajan Suthokumar, Vidhyasaharan Sethu, Kaavya Sriskandaraja, Eliathamby Ambikairajah |
ICASSP | 4 |
| 2020 | Low-Complexity Real-Time Light Field Compression using 4-D Approximate DCTabstractA low-complexity codec and a hardware architecture are proposed for achieving real-time compression of four-dimensional (4-D) light field (LF) signals captured from camera/lenslet arrays. The proposed system employs the 4-D extension of the two-dimensional (2-D) 8×8 approximate discrete cosine transform (ADCT) that has recently appeared in the literature. Motivated by the partial separability of the multidimensional spectrum of LFs, the proposed 4-D ADCT is obtained by cascading 2-D inter-view and 2-D intra-view transform stages. Software simulations are provided to confirm the performance of the 4-D ADCT based compression and comparisons are made with respect to 2-D inter-view only and 2-D intra-view only ADCT-based compression. Proposed digital architectures are validated using stepped hardware co-simulation on a Xilinx Virtex-7 VC-707 FPGA platform verifying 597 MHz maximum possible clock frequency, implying an ideal throughput of 18×103LFs/sec for performing 4-D ADCT on (8× 8×432×624×3) size LFs. When 10% of the ADCT coefficients per each (8×8×8×8) hypercube are retained sub aperture images show 38 dB average PSNR and 0.95 average SSIM. Namalka Liyanage, Chamith Wijenayake, Chamira U. S. Edussooriya, Arjuna Madanayake, Renato J. Cintra, Eliathamby Ambikairajah |
ISCAS | 6 |
| 2020 | Multi-depth filtering and occlusion suppression in 4-D light fields: Algorithms and architectures
Namalka Liyanage, Chamith Wijenayake, Chamira U. S. Edussooriya, Arjuna Madanayake, Panajotis Agathoklis, Leonard T. Bruton, Eliathamby Ambikairajah |
Signal Process. | 7 |
| 2019 | A Novel Bag-of-Optimised-Clusters Front-End for Speech based Continuous Emotion PredictionabstractAlmost all current speech based emotion prediction systems employ front-ends that approximately represent the distribution of frame based features over a suitable window, typically via a set of statistical functionals or the use of Bag-of-Audio-Words (BoAW) features. These front-ends are designed either by manual selection of appropriate statistical functionals, via a feature selection approach or unsupervised clustering of the frame-based features and may not be optimal for the task at hand. This paper proposes a novel front-end that discriminatively learns feature clusters to generate a codebook optimised for emotion prediction, which is then used to generate a Bag-of-Optimised-Clusters (BoOC) feature set. Moreover, this front-end is implemented as a sequence of neural network layers that allow both the proposed front-end and a suitable deep learning backend to be jointly trained. The Bag-of-Optimised-Clusters frontend is tested on the RECOLA database and results show that it outperforms the well-established BoAW features. Deboshree Bose, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Sarith Fernando |
ACII | 4 |
| 2019 | Transmission Line Cochlear Model Based AM-FM Features for Replay Attack DetectionabstractThis paper focuses on providing a countermeasure to replay attack which is the simplest and more accessible form of attack used to spoof automatic speaker verification systems. Specifically, it proposes the use of the transmission line cochlear model, which resembles the human cochlea more accurately than parallel filter bank models, in the front-end of replay detection systems. Here the basilar membrane is modeled as a cascade of digital filters with decreasing resonant frequencies. In this context, we propose two features - transmission line cochlea-amplitude modulation (TLC-AM) and frequency modulation (TLC-FM) - to extract the modulation features of the speech from the simulated membrane displacements. TLC-AM is analogous to the output of the inner hair cell bending movement, which accurately captures the amplitude modulation component of the speech. TLC-FM is extracted by deriving the in-phase and out of phase signals of basilar membrane displacement. Results show that individual TLC-AM and TLC-FM features perform better than the best parallel filter bank baseline system. Experiments suggest that higher frequency selectivity is beneficial for replay detection, especially for AM, and the proposed TLC model is better able to achieve this property than parallel filter bank models. Tharshini Gunendradasan, Saad Irtza, Eliathamby Ambikairajah, Julien Epps |
ICASSP | 3 |
| 2019 | Evaluation Measures for Depression Prediction and Affective ComputingabstractA variety of evaluation measures are being used to validate systems in depression prediction and affective computing. Among them, the most common measures focus on the error between the ground truth and predictions. However, when the ground truth is ordinal such as in psychiatric scores, ranking information is more important than the actual error. Therefore, this study systematically analyses the properties of classification, error-based and ranking measures particularly using classification accuracy, root mean square error (RMSE) and Spearman rank correlation coefficient, with the aim of identifying suitable measures for evaluating depression prediction and affective computing. For the purpose of analysis, we employed both synthetic data and real depression prediction systems evaluated with the AVEC2017 depression corpus. Outcomes of the experiments suggest that RMSE and classification accuracy, which are frequently used, are not sensitive to ordering and that rank correlation measures are more appropriate for depression prediction, which is an ordinal problem. Sadari Jayawardena, Julien Epps, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2019 | Phoneme Specific Modelling and Scoring Techniques for Anti Spoofing SystemabstractReplay attack refers to the use of recorded speech in an attempt to spoof an automatic speaker verification system and the development of countermeasures that can detect these attacks is an active area of research. This paper investigates the effect of phoneme specific information on replay attack detection. It then develops a replay detection system that employs phoneme specific genuine and spoof models and compares novel scoring methods that take into account phonetic information obtained from a suitable phoneme recogniser. Experiment result on the ASVSpoof 2017 V2.0 corpus indicated that replayed speech may be easier to detect from speech corresponding to some phonemes compared to others and consequently judicious use of phoneme specific models can improve replay detection systems. Gajan Suthokumar, Kaavya Sriskandaraja, Vidhyasaharan Sethu, Chamith Wijenayake, Eliathamby Ambikairajah |
ICASSP | 5 |
| 2019 | Auditory Inspired Spatial Differentiation for Replay Spoofing Attack DetectionabstractThe security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. This paper explores the effect of `spatial differentiation' used in auditory system modelling to improve frequency selectivity and hence provide a more selective front-end for replay attack detection. Experiments were done using a parallel filter bank consisting of simple 2ndorder IIR bandpass filters following which, processing analogous to spatial differentiation was employed to obtain higher order stable IIR filters, in turn leading to highly selective filter banks. Two novel features based on spatially differentiated higher order filter bank have been proposed. Together they yield a relative improvement of 29.9% in replay speech detection over a constant Q transform based baseline system, when evaluated on the ASVspoof 2017 Version 2.0 database. Buddhi Wickramasinghe, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001 |
ICASSP | 2 |
| 2019 | An Adaptive-Q Cochlear Model for Replay Spoofing Detection
Tharshini Gunendradasan, Eliathamby Ambikairajah, Julien Epps, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2019 | Speech Based Emotion Prediction: Can a Linear Model Work?
Anda Ouyang, Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2019 | Biologically Inspired Adaptive-Q Filterbanks for Replay Spoofing Attack Detection
Buddhi Wickramasinghe, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2019 | Reduced-Complexity Depth Filtering and Occlusion Suppression using Modulated-Sparse Light FieldsabstractApplication of depth-selective filtering to modulated-sparse light fields towards reducing the DSP and memory complexities in real-time light field processing is investigated. A modulated-sparse light field is obtained by spatially windowing an original 4-D light-field signal, which is subsequently processed by 4-D depth-selective filters to achieve planar focus and occlusion suppression. Multidimensional spectral properties of modulated-sparse light fields are explored and demonstrative examples with real light fields are provided to show that almost similar depth filtering and occlusion suppression performance can be achieved by processing such modulated-sparse light fields, leading to significant reduction in DSP hardware and memory complexities. Namalka Liyanage, Chamith Wijenayake, Chamira U. S. Edussooriya, Eliathamby Ambikairajah |
ISCAS | 4 |
| 2018 | Dynamic Multi-Rater Gaussian Mixture Regression Incorporating Temporal Dependencies of Emotion Uncertainty Using Kalman FiltersabstractPredicting continuous emotion in terms of affective attributes has mainly been focused on hard labels, which ignored the ambiguity of recognizing certain emotions. This ambiguity may result in high inter-rater variability and in turn causes varying prediction uncertainty with time. Based on the assumption that temporal dependencies occur in the evolution of emotion uncertainty, this paper proposes a dynamic multi-rater Gaussian Mixture Regression (GMR), aiming to obtain the emotion uncertainty prediction reflected by multi-raters by taking into account their temporal dependencies. This framework is achieved by incorporating feedforward and backward Kalman filters into GMR to estimate the time-dependent label distribution that reflects the emotion uncertainty. It also provides the benefits of relaxing the label distribution of Gaussian assumption to that of a Gaussian Mixture Model (GMM). In addition, a new measurement to estimate emotion uncertainty from GMM as the local variability is adopted. Experiments conducted on the RECOLA database reveal that incorporating temporal dependencies is critical for emotion uncertainty prediction with 17% relative improvement for arousal, and that the proposed framework for emotion uncertainty prediction shows potential in conventional emotion attribute prediction. Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2018 | Factorized Hidden Variability Learning for Adaptation of Short Duration Language Identification ModelsabstractBidirectional long short term memory (BLSTM) recurrent neural networks (RNNs) have recently outperformed other state-of-the-art approaches, such as i-vector and deep neural networks (DNNs) in automatic language identification (LID), particularly when testing with very short utterances (`3s). Mismatches conditions between training and test data, e.g. speaker, channel, duration and environmental noise, are a major source of performance degradation for LID. A factorized hidden variability subspace (FHVS) learning technique is proposed for the adaptation of BLSTM RNNs to compensate for these types of mismatches in recording conditions. In the proposed approach, condition dependent parameters are estimated to adapt the hidden layer weights of the BLSTM in the FHVS. We evaluate FHVS on the AP17-OLR data set. Experimental results show that the FHVS method outperforms the standard BLSTM approach, achieving 27% relative improvements with utterance-level adaptation over the standard BLSTM for 1 s duration utterances. Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2018 | End-to-End Hierarchical Language Identification SystemabstractRecently, hierarchical language identification systems have shown significant improvement over single level systems in both closed and open set language identification tasks. However, developing such a system requires the features and classifier selection at each node in the hierarchical structure to be hand crafted. Motivated by the superior ability of end-to-end deep neural network architecture to jointly optimize the feature extraction and classification process, we propose a novel approach developing an end-to-end hierarchical language identification system. The proposed approach also demonstrates the in -built ability of the end-to-end hierarchical structure training that enables an out-of-set language model, without using any additional out-of-set language training data. Experiments are conducted on the NIST LRE 2015 data set. The overall results show relative improvements of 18.6% and 27.3% in terms of Cavgin closed and open set tasks over the corresponding baseline systems. Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
ICASSP | 3 |
| 2018 | Speaker-Phonetic Vector Estimation for Short Duration Speaker VerificationabstractPhonetic variability is one of the primary challenges in short duration speaker verification. This paper proposes a novel method that modifies the standard normal distribution prior in the total variability model to use a mixture of Gaussians as the prior distribution. The proposed speaker-phonetic vectors are then estimated from the posterior probability of latent variables, and each vector has a phonetic meaning. Unlike the standard total variability model, the proposed method can incorporate a phoneme classifier to perform soft content matching, which has the potential to solve the phonetic variability problem. Parameter estimation and scoring formulae for speaker-phonetic vectors method are presented. Experimental results obtained using NIST 2010 data show that the proposed technique leads to relative improvements of more than 30% when fused with total variability model and tested on 3 second duration test files. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
ICASSP | 3 |
| 2018 | Sub-band Envelope Features Using Frequency Domain Linear Prediction for Short Duration Language Identification
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2018 | Detection of Replay-Spoofing Attacks Using Frequency Modulation Features
Tharshini Gunendradasan, Buddhi Wickramasinghe, Phu Ngoc Le, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 4 |
| 2018 | Deep Siamese Architecture Based Replay Detection for Secure Voice Biometric
Kaavya Sriskandaraja, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2018 | Modulation Dynamic Features for the Detection of Replay AttacksabstractThe development of automatic systems that can detect replayed speech has emerged as a significant research challenge for securing voice biometric systems and is the focus of this paper. Specifically, this paper proposes two novel features to capture the static and dynamic characteristics of the signal from the modulation spectrum, which complement short term spectral features for use in replay detection. The modulation spectral centroid frequency feature is proposed as a vector representation of the first order spectral moments of the modulation spectrum. In conjunction to this, the long term spectral average serves to capture the static characteristics of the modulation spectrum. The proposed system, employing a GMM back-end, was evaluated on the ASVSpoof 2017 dataset and found to yield an EER of 6.54%. Gajan Suthokumar, Vidhyasaharan Sethu, Chamith Wijenayake, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2018 | Frequency Domain Linear Prediction Features for Replay Spoofing Attack Detection
Buddhi Wickramasinghe, Saad Irtza, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 3 |
| 2018 | Low-Complexity 4-D IIR Filters for Multi-Depth Filtering and Occlusion Suppression in Light FieldsabstractLight field signal processing allows manipulation of a rich set of information captured from a scene to achieve real-time depth filtering and occlusion suppression. Low-complexity four-dimensional infinite impulse response digital filters for simultaneous depth filtering and occlusion suppression over multiple depths in light fields are proposed. A low-complexity two-dimensional separable approach is employed to design the proposed filters having multiple frequency-planar pass-bands/stopbands in the four-dimensional spatial frequency domain, that can be electronically tuned to enhance/reject planar objects at multiple depths in a light field. Filter synthesis details are provided with specific design examples corresponding to 2-passband and 1-stopband cases. Numerically generated and Lytro camera captured light fields are used to verify the effectiveness of the proposed multi-depth-pass and multi-depth-reject filters. For synthetic light field inputs these filters confirm an average denoising, depth filtering and occlusion suppression performance of 20 dB, 20 dB, and 30 dB, respectively. Namalka Liyanage, Chamith Wijenayake, Chamira U. S. Edussooriya, Arjuna Madanayake, Panajotis Agathoklis, Eliathamby Ambikairajah, Leonard T. Bruton |
ISCAS | 6 |
| 2018 | Using language cluster models in hierarchical language identification
Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
Speech Commun. | 3 |
| 2018 | Generalized Variability Model for Speaker VerificationabstractIn this letter, we propose a generalized variability model as an extension to the total variability model. While the total variability model employs a standard normal prior distribution in its typical setup, the proposed generalized variability model relaxes this assumption and allows the latent variable distribution to be a mixture of Gaussians. The conventional total variability model can then be viewed as a special case of this generalized version where the number of mixture components is constrained to one. This proposed model is validated in the context of speaker verification tasks on both the standard and extended NIST SRE 2010 datasets. Experimental results show that modeling the distribution of the latent variables as a mixture of Gaussians leads to a better performance under all conditions and a greater gain can be expected for speaker verification using short utterances. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
IEEE Signal Process. Lett. | 3 |
| 2017 | Modeling variable length phoneme sequences - A step towards linguistic information for speech emotion recognition in wider worldabstractVocal gestures play an important role in emotion expression and can be used by speech based emotion recognition systems. This paper proposes the use of BLSTM neural networks to model salient variable length phoneme sequences, which in turn can represent relevant vocal gestures. Unlike existing techniques, the proposed approach is not restricted to modelling phoneme sequences of a fixed length and both salience and optimal modelling length of phoneme sequences are learnt from the training data. Three possible phoneme representations that can be modelled by BLSTMs are compared and experimental results suggest that sequences of Phone Log Likelihood Ratios are more representative of emotions when compared to sequences of phoneme labels represented as one — hot vectors. On the IEMOCAP database, the proposed approach achieves an Unweighted Average Recall (UAR) of 56.4%, an improvement of 6.5% in absolute terms over the previous approach of modelling fixed length phoneme sequences on a 4-class classification problem. The proposed linguistic system is complementary to acoustic features with a fused system leading to an absolute improvement of 5% to the UAR. Kalani Wataraka Gamage, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ACII | 3 |
| 2017 | Salience based lexical features for emotion recognitionabstractIn this paper we focus on the usefulness of verbal events for speech based emotion recognition. In particular, the use of phoneme sequences to encode verbal cues related to the expression of emotions is proposed and lexical features based on these phoneme sequences are introduced for use in automatic emotion recognition systems where manual transcripts are not available. Secondly, a novel estimate of emotional salience of verbal cues, applicable to both phoneme sequences and words, is presented. Experimental results on the IEMOCAP database show that the proposed automatic phoneme sequence based features can achieve an Unweighted Average Recall (UAR) of 49% with proposed salience measure. Further, the proposed salience measure can lead to an UAR of 64% when using manual word transcriptions. Both of these are the highest UARs reported on the IEMOCAP database for systems using lexical features extracted from automatic and manual transcripts respectively. Kalani Wataraka Gamage, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2017 | An Investigation of Crowd Speech for Room Occupancy Estimation
Siyuan Chen 0002, Julien Epps, Eliathamby Ambikairajah, Phu Ngoc Le |
INTERSPEECH | 3 |
| 2017 | An Investigation of Emotion Prediction Uncertainty Using Gaussian Mixture Regression
Ting Dang, Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2017 | Bidirectional Modelling for Short Duration Language Identification
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 3 |
| 2017 | Investigating Scalability in Hierarchical Language Identification System
Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 62 |
| 2017 | Incorporating Local Acoustic Variability Information into Short Duration Speaker Verification
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
INTERSPEECH | 3 |
| 2017 | Independent Modelling of High and Low Energy Speech Frames for Spoofing DetectionabstractSpoofing detection systems for automatic speaker verification have moved from only modelling voiced frames to modelling all speech frames. Unvoiced speech has been shown to carry information about spoofing attacks and anti-spoofing systems may further benefit by treating voiced and unvoiced speech differently. In this paper, we separate speech into low and high energy frames and independently model the distributions of both to form two spoofing detection systems that are then fused at the score level. Experiments conducted on the ASVspoof 2015, BTAS 2016 and Spoofing and Anti-Spoofing (SAS) corpora demonstrate that the proposed approach of fusing two independent high and low energy spoofing detection systems consistently outperforms the standard approach that does not distinguish between high and low energy frames. Gajan Suthokumar, Kaavya Sriskandaraja, Vidhyasaharan Sethu, Chamith Wijenayake, Eliathamby Ambikairajah |
INTERSPEECH | 5 |
| 2016 | A hierarchical framework for language identificationabstractMost current language recognition systems model different levels of information such as acoustic, prosodic, phonotactic, etc. independently and combine the model likelihoods in order to make a decision. However, these are single level systems that treat all languages identically and hence incapable of exploiting any similarities that may exist within groups of languages. In this paper, a hierarchical language identification (HLID) framework is proposed that involves a series of classification decisions at multiple levels involving language clusters of decreasing sizes with individual languages identified only at the final level. The performance of proposed hierarchical framework is compared with a state-of-the-art LID system on the NIST 2007 database and the results indicate that the proposed approach outperforms state-of-the-art systems. Saad Irtza, Vidhyasaharan Sethu, Haris Bavattichalil, Eliathamby Ambikairajah, Haizhou Li 0001 |
ICASSP | 4 |
| 2016 | Factor Analysis Based Speaker Normalisation for Continuous Emotion Prediction
Ting Dang, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2016 | A Feature Normalisation Technique for PLLR Based Language Identification Systems
Sarith Fernando, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2016 | Out of Set Language Modelling in Hierarchical Language Identification
Saad Irtza, Vidhyasaharan Sethu, Sarith Fernando, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2016 | Parallel Speaker and Content Modelling for Text-Dependent Speaker Verification
Saad Irtza, Kaavya Sriskandaraja, Vidhyasaharan Sethu, Eliathamby Ambikairajah |
INTERSPEECH | 5 |
| 2016 | Twin Model G-PLDA for Duration Mismatch Compensation in Text-Independent Speaker Verification
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
INTERSPEECH | 3 |
| 2016 | Investigation of Sub-Band Discriminative Information Between Spoofed and Genuine Speech
Kaavya Sriskandaraja, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2015 | An investigation of emotion change detection from speechabstractEmotion recognition based on speech plays an important role in Human Computer Interaction (HCI), which has motivated extensive recent investigation into this area. However, current research on emotion recognition is focused on recognizing emotion on a per-file basis and mostly does not provide insight into emotion changes. In this paper, we report on an initial investigation into detecting the instant of emotion change using Gaussian Mixture Models (GMM) based methods, either without or with prior knowledge of emotions: the Generalized Likelihood Ratio and Emotion Pair Likelihood Ratios, together with a novel normalization scheme to improve emotion change detection accuracy. Experimental results based on the IEMOCAP corpus are presented that demonstrate a promising baseline. Despite the challenging nature of the problem, this work provides a path towards systems that detect and understand emotion changes, and also presents very interesting questions for further investigation. Zhaocheng Huang, Julien Epps, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2015 | Phonemes frequency based PLLR dimensionality reduction for language recognition
Saad Irtza, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2015 | A model based voice activity detector for noisy environments
Kaavya Sriskandaraja, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2015 | Voice source under cognitive load: Effects and classification
Tet Fei Yap, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi |
Speech Commun. | 3 |
| 2014 | The UNSW submission to INTERSPEECH 2014 compare cognitive load challengeabstractSpeech based cognitive load estimation is a new field of research. Due to this relative ‘lack of maturity’, a single best approach to building cognitive load estimation systems has not been established yet. The primary aim of this submission is to report the performance of various basic utterance level classification frameworks developed using important elements of state-of-the-art speaker recognition systems. This may lead to a suitable basis for future cognitive load estimation systems. As a consequence of being a part of a challenge, it is expected that these frameworks will be compared to a much larger number of alternative approaches than what would otherwise be possible. In keeping with this focused aim, the GMM supervector approaches along with some variants are utilised. The systems outlined in this paper include a frame-level MFCC-GMM system along with utterance level GMMsupervector-SVM, GMM-ivector-SVM and GMM-JFA-SVM systems. The best combined system has an accuracy (UAR) of 66.6% as evaluated on the challenge development set and 63.7% as evaluated on the test set. Jia Min Karen Kua, Vidhyasaharan Sethu, Phu Ngoc Le, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2013 | Spectro-temporal analysis of speech affected by depression and psychomotor retardationabstractTo enhance current diagnostic methods used when assessing a depressed individual, an objective screening mechanism, ideally based on non-intrusive behavioral signals, is needed. Given the clinical description of depression speech as `dull, monotonous and flat' and promising previous results from spectral features, we hypothesize that the effects of depression on speech are embedded in spectro-temporal events. To test this hypothesis we explore different methodologies, based on the modulation spectrum, for extracting long-term spectro-temporal information from speech and assess their suitability as a clinical marker of depression. Results indicate that: depressive speech information is captured in the modulation spectrum, long-term spectro-temporal information is important in depressed speech identification and there are potential differences in the effects that depression and psychomotor retardation have on speech production mechanisms. Nicholas Cummins, Julien Epps, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2013 | Predicting the effect of AGC on speech intelligibility of cochlear implant recipients in noiseabstractThe aim of this study was to predict the effects of automatic gain control (AGC) on the speech intelligibility of cochlear implant recipients in noise. Two simple signal metrics were calculated: the proportion of clipping, and the output signal-to-noise ratio (SNR). Psychometric functions were fitted to the percent correct scores averaged over five cochlear implant recipients for three different AGC conditions, at two different input SNRs, for a range of presentation levels. The output SNR was a good predictor of the recipients' mean scores. Phyu P. Khing, Eliathamby Ambikairajah, Brett A. Swanson |
ICASSP | 2 |
| 2013 | Speaker variability in speech based emotion models - Analysis and normalisationabstractAll features commonly utilised in speech based emotion classification systems capture both emotion-specific information and speaker-specific information. This paper proposes a novel method to gauge the effect of speaker-specific information on emotion modelling based on two measures: a Monte Carlo approximation to KL divergence and an estimate of feature variability based on diagonal covariance matrices. In addition, a novel speaker normalisation technique based on joint factor analysis is also proposed. This method is analogous to channel compensation in speaker verification systems, with one significant extension. The model domain compensation is mapped back to frame-level features, allowing for use in a wider range of emotion classification frameworks and in conjuncture with other normalisation techniques. Preliminary evaluations on the IEMOCAP database suggests that the proposed technique improves the performance of GMM based classification systems based on widely employed features such as pitch, MFCCs and deltas. Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah |
ICASSP | 3 |
| 2013 | GMM based speaker variability compensated system for interspeech 2013 compare emotion challengeabstractThis paper describes the University of New South Wales system for the Interspeech 2013 ComParE emotion subchallenge. The primary aim of the submission is to explore the performance of model based variability compensation techniques applied to emotion classification and as a consequence of being a part of a challenge, to enable a comparison of these methods to alternative approaches. In keeping with this focused aim, a simple frame based front-end of MFCC and ΔMFCC is utilised. The systems outlined in this paper consists of a joint factor analysis based system and one based on a library of speaker-specific emotion models along with a basic GMM based system. The best combined system has an accuracy (UAR) of 47.8% as evaluated on the challenge development set and 35.7% as evaluated on the test set. Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2013 | I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verificationabstractI4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort. Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah |
INTERSPEECH | 34 |
| 2013 | i-Vector with sparse representation classification for speaker verification
Jia Min Karen Kua, Julien Epps, Eliathamby Ambikairajah |
Speech Commun. | 3 |
| 2012 | Speaker variability in emotion recognition - an adaptation based approachabstractNone of the features commonly utilised in automatic emotion classification systems completely disassociate emotion-specific information from speaker-specific information. Consequently, this speaker-specific variability adversely affects the performance of the emotion classification system and in existing systems is frequently mitigated by some form of speaker normalisation. Speaker adaptation offers an alternative to normalisation and this paper proposes a novel bootstrapping technique which involves selecting appropriate initial models from a large training pool, prior to speaker adaptation of emotion models in the context of GMM based emotion classification as an alternative to speaker normalisation. Evaluations on the LDC Emotional Prosody and the FAU Aibo corpora reveal that an emotion classification system based on the proposed bootstrapping method outperforms systems based on speaker normalisation as long as a small amount of labelled adaptation data is available. It also outperforms speaker adaption from common initial models estimated from all training speakers. Ni Ding, Vidhyasaharan Sethu, Julien Epps, Eliathamby Ambikairajah |
ICASSP | 4 |
| 2012 | A Non-Uniform Filterbank for Speaker Recognition
Jia Min Karen Kua, Tharmarajah Thiruvaran, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2011 | Effect of fast AGC on cochlear implant speech intelligibilityabstractThis study investigated the effect of fast-acting Automatic Gain Control (AGC) on the speech intelligibility of cochlear implant users as a function of presentation level. Both low and high signal-to-noise ratio conditions were investigated. The AGC substantially reduced the amount of clipping, but did not give consistent improvements in intelligibility. With no AGC, and high signal-to-noise ratio, speech scores were not significantly degraded until more than 25% of stimulation pulses were affected by clipping. Phyu P. Khing, Eliathamby Ambikairajah, Brett A. Swanson |
ICASSP | 2 |
| 2011 | Speaker verification using sparse representation classificationabstractSparse representations of signals have received a great deal of attention in recent years, and the sparse representation classifier has very lately appeared in a speaker recognition system. This approach represents the (sparse) GMM mean supervector of an unknown speaker as a linear combination of an over-complete dictionary of GMM supervectors of many speaker models, and ℓ1-norm minimization results in a non-zero coefficient corresponding to the unknown speaker class index. Here this approach is tested on large databases, introducing channel-/session-variability compensation, and fused with a GMM-SVM system. Evaluations on the NIST 2001 SRE and NIST 2006 SRE database show that when the outputs of the MFCC UBM-GMM based classifier (for NIST 2001 SRE) or MFCC GMM-SVM based classifier (for NIST 2006 SRE) are fused with the MFCC GMM Sparse Representation Classifier (GMM-SRC) based classifier, an absolute gain of 1.27% and 0.25% in EER can be achieved respectively. Jia Min Karen Kua, Eliathamby Ambikairajah, Julien Epps, Roberto Togneri |
ICASSP | 2 |
| 2011 | Using clustering comparison measures for speaker recognitionabstractRecent results seem to cast some doubt over the assumption that improvements in fused recognition accuracy for speaker recognition systems based on different acoustic features are due mainly to the different origins of the features (e.g. magnitude, phase, modulation information). In this study, we utilize clustering comparison measures to investigate acoustic and speaker modelling aspects of the speaker recognition task separately and demonstrate that front-end diversity can be achieved purely through different 'partitioning' of the acoustic space. Further, features that exhibit good 'stability' with respect to repeated clustering are shown to also give good EER performance in speaker recognition. This has implications for feature choice, fusion of systems employing different features, and for UBM data selection. A method for the latter problem is presented that gives up to an 11% relative reduction in EER using only 20-30% of the usual UBM training data set. Jia Min Karen Kua, Julien Epps, Mohaddeseh Nosratighods, Eliathamby Ambikairajah, Eric H. C. Choi |
ICASSP | 4 |
| 2011 | Voice source features for cognitive load classificationabstractPrevious work in speech-based cognitive load classification has shown that the glottal source contains important information for cognitive load discrimination. However, the reliability of glottal flow features depends on the accuracy of the glottal flow estimation, which is a non-trivial process. In this paper, we propose the use of acoustic voice source features extracted directly from the speech spectrum (or cepstrum) for cognitive load classification. We also propose pre and post-processing techniques to improve the estimation of the cepstral peak prominence (CPP). 3-class classification results on two databases showed CPP as a promising cognitive load classification feature that outperforms glottal flow features. Score-level fusion of the CPP-based classification system with a formant frequency-based system yielded a final improved accuracy of 62.7%, suggesting that CPP contains useful voice source information that complements the information captured by vocal tract features. Tet Fei Yap, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi |
ICASSP | 3 |
| 2011 | Investigation of spectral centroid features for cognitive load classification
Phu Ngoc Le, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Eric H. C. Choi |
Speech Commun. | 2 |
| 2010 | Glottal features for speech-based cognitive load classificationabstractCognitive load measurement is important when designing adaptive interfaces that optimize the performance of users working on high mental load tasks. Recent research on automatic speech-based measurement system indicates that cognitive load information is more prominent in the frequency region below 1 kHz. This study investigates the effects of cognitive load on glottal parameters (open quotient, normalized amplitude quotient and speed quotient), and proposes a system employing these parameters as features for cognitive load classification. Analysis of the glottal parameter distributions suggests that an increase in cognitive load can be related to a more creaky voice quality. Additionally, three-class classification results show that score-level fusion of systems based on the glottal features and baseline features (MFCCs, pitch, intensity and shifted delta cepstra) improves the baseline accuracy from 79% to 84%. Tet Fei Yap, Julien Epps, Eric H. C. Choi, Eliathamby Ambikairajah |
ICASSP | 4 |
| 2010 | A Study of Voice Source and Vocal Tract Filter Based Features in Cognitive Load ClassificationabstractSpeech has been recognized as an attractive method for the measurement of cognitive load. Previous approaches have used mel frequency cepstral coefficients (MFCCs) as discriminative features to classify cognitive load. The MFCCs contain information from both the voice source and the vocal tract, so that the individual contributions of each to cognitive load variation are unclear. This paper aims to extract speech features related to either the voice source or the vocal tract and use them to discriminate between cognitive load levels in order to identify the individual contribution of each for cognitive load measurement. Voice source-related features are then used to improve the performance of current cognitive load classification systems, using adapted Gaussian mixture models. Our experimental result shows that the use of voice source feature could yield around 12% reduction in relative error rate compared with the baseline system based on MFCCs, intensity, and pitch contour. Phu Ngoc Le, Julien Epps, Eric H. C. Choi, Eliathamby Ambikairajah |
ICPR | 4 |
| 2010 | An investigation of formant frequencies for cognitive load classification
Tet Fei Yap, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi |
INTERSPEECH | 3 |
| 2010 | Perceptual speech enhancement exploiting temporal masking properties of human auditory system
Teddy Surya Gunawan, Eliathamby Ambikairajah, Julien Epps |
Speech Commun. | 2 |
| 2010 | A segment selection technique for speaker verification
Mohaddeseh Nosratighods, Eliathamby Ambikairajah, Julien Epps, Michael J. Carey 0002 |
Speech Commun. | 2 |
| 2009 | Linear predictive modelling of gait patternsabstractThe use of a wearable triaxial accelerometer for unsupervised monitoring of human movement has become a major research focus in recent years. In this paper, the relationship between accelerometry signals and human gait is analysed using a linear prediction (LP) model. We explore the use of the LP model for analysing five gait patterns and show that the LP cepstrum can be used for gait pattern classification with high accuracy. This is then compared to a filterbank based approach to estimate the cepstral coefficients. Fifty subjects participated in collection of gait pattern data involving walking on level surfaces, and walking up and down stairs and ramps. The results show that an overall accuracy of 93% can be achieved using features derived from the cepstral coefficients for the five different walking patterns. Ronny K. Ibrahim, Eliathamby Ambikairajah, Branko G. Celler, Nigel H. Lovell |
ICASSP | 2 |
| 2009 | The I4U system in NIST 2008 speaker recognition evaluationabstractThis paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU). Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin |
ICASSP | 18 |
| 2009 | Evaluation of a fused FM and cepstral-based speaker recognition system on the NIST 2008 SREabstractIn this paper, the fusion of two speaker recognition subsystems, one based on Frequency Modulation (FM) and another on MFCC features, is reported. The motivation for their fusion was to improve the recognition accuracy across different types of channel variations, since the two features are believed to contain complementary information. It was found that the MFCC-based subsystem outperformed the FM-based subsystem on telephone conversations from NIST SRE-06 dataset, while the opposite was true for NIST SRE-08 telephone data. As a result, the FM-based subsystem performed as well as the MFCC-based subsystem and their fusion gave up to 23% relative improvement in terms of EER over the MFCC subsystem alone, when evaluated on the NIST 2008 core condition. Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 4 |
| 2009 | Speaker dependency of spectral features and speech production cues for automatic emotion classificationabstractSpectral and excitation features, commonly used in automatic emotion classification systems, parameterise different aspects of the speech signal. This paper groups these features as speech production cues, broad spectral measures and detailed spectral measures and looks at how they differ in their performance in both speaker dependent and speaker independent systems. The extent of speaker normalisation on these features is also considered. Combinations of different features are then compared in terms of classification accuracies. Evaluations were conducted on the LDC emotional speech corpus for a five-class problem. Results indicate that MFCCs are very discriminative but suffer from speaker variability. Further, results suggest that the best front end for a speaker independent system is a combination of pitch, energy and formant information. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
ICASSP | 2 |
| 2009 | Phase based features for cognitive load measurement systemabstractThe current automatic cognitive load measurement system based on MFCC and prosodic features does not take into account phase based speech information. This paper aims to improve the performance of the baseline system by introducing phase based features into the system. The additional features proposed are group delay features, all-pole model based FM features and zero crossing count based FM features. Decrease in performance is observed when phase based features are considered individually or when concatenated with baseline features. However, significant performance improvement is observed when group delay features are fused with baseline features using linear combination score level fusion. Tet Fei Yap, Eliathamby Ambikairajah, Eric H. C. Choi, Fang Chen 0001 |
ICASSP | 2 |
| 2009 | Voiced/unvoiced pattern-based duration modeling for language identificationabstractMost existing duration modeling approaches facilitates phone recognizer and require manually annotated corpus to train the segmentation models, which is usually cost- and time-consuming. In this paper, a novel duration modeling approach is proposed, which does not require phone recognizer/annotated training data, and facilitates fast computation of language identification. In this approach, the segmentation is implemented by using articulatory features like voicing status. A pair of connected unvoiced and voiced segments is considered as the unit, and the duration of each segment is normalized for each utterance and then quantized into 20 discrete ranges. The ranges of units are later considered as symbol sequences and are modeled by n-gram models, to capture the temporal pattern, which is hypothesized to vary in different languages. The experiments based on the NIST LRE 2005 tasks show a relative 19.7% EER improvement by introducing the proposed duration modeling-based system into a fusion system containing two GMM-UBM based acoustic systems using MFCC and pitch+intensity features. Bo Yin 0002, Eliathamby Ambikairajah, Fang Chen 0001 |
ICASSP | 2 |
| 2009 | Robust language identification based on fused phonotactic information with MLKSFM pre-classifierabstractIn this paper we propose a novel language identification system which utilizes fused phonotactic information. The phase spectrum of speech signals is used with the magnitude spectrum in order to obtain a more robust feature representation. Parallel Broad Phoneclass Recognition followed by Language Model (PBPRLM) is used in order to remove the bias of the likelihood scores introduced by the size inequality of phone inventories in traditional PPRLM systems. The likelihood scores from the MFCC-based and group-delay-based PPRLM and PBPRLM systems are fused together by using a Gaussian Mixture Model. Furthermore, a pre-classification based on Kohonen's map is used in order to maintain the system robustness while handling a large number of target languages. Using this proposed novel system we achieve an EER of 6.7% on the 2005 NIST LRE, and a LID recognition rate of 83.9% on a 22-language task. Liang Wang 0003, Eliathamby Ambikairajah, Eric H. C. Choi |
ICME | 2 |
| 2009 | LS regularization of group delay features for speaker recognition
Jia Min Karen Kua, Julien Epps, Eliathamby Ambikairajah, Eric H. C. Choi |
INTERSPEECH | 3 |
| 2009 | Pitch contour parameterisation based on linear stylisation for emotion recognitionabstractThe pitch contour contains information that characterises the emotion being expressed by speech, and consequently features extracted from pitch form an integral part of many automatic emotion recognition systems. While pitch contours may have many small variations and hence are difficult to represent compactly, it may be possible to parameterise them by approximating the contour for each voiced segment by a straight line. This paper looks at such a parameterisation method in the context of emotion recognition. Listening tests were performed to subjectively determine if the linearly stylised contours were able to sufficiently capture information pertaining to emotions expressed in speech. Furthermore these parameters were used as features for an automatic 5-class emotion classification system. The use of the proposed parameters rather than pitch statistics resulted in a relative increase in accuracy of about 20%. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2009 | Analysis of band structures for speaker-specific information in FM feature extraction
Tharmarajah Thiruvaran, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2008 | Optimizing period-3 methods for eukaryotic gene predictionabstractIn this paper, we firstly investigate the effect of window lengths on selected signal processing-based gene and exon prediction methods. We then optimize these methods to improve their prediction accuracy by employing the best DNA representation, a suitable window length, and boosting the output signals to enhance protein coding and suppress the non-coding regions. It is shown herein that the proposed method outperforms major existing time-domain, frequency- domain, and combined time-frequency approaches. By comparison with the existing DFT-based methods, the proposed method not only requires 50% less processing but also exhibits relative improvements of 53.3%, 46.7%, and 24.2% respectively over spectral content, spectral rotation, and paired and weighted spectral rotation measures in terms of prediction accuracy of exonic nucleotides at a 5% false positive rate using the GENSCAN test set. Mahmood Akhtar, Eliathamby Ambikairajah, Julien Epps |
ICASSP | 2 |
| 2008 | A self-directed learning approach to signal processing educationabstractThis paper describes a self-directed, project-based learning scheme implemented in an introductory Signal Processing course at the University of New South Wales. The course was structured around a major laboratory project in which students were required to research course material, understand the relevant theory, and apply this in order to arrive at a solution. Lectures were delivered via prerecorded DVDs, allowing students to self-pace their absorption of new content and allowing teaching staff to concentrate on specific student issues during face-to-face classes. Evaluation of the course structure by the lecturer and instructors suggested that students gained a better conceptual understanding of signal processing theory than in previous years. Students were generally positive towards the process, but found it difficult to adjust to. Eliathamby Ambikairajah, Julien Epps, Samuel J. Freney, Ming Sheng |
ICASSP | 1 |
| 2008 | Empirical mode decomposition based weighted frequency feature for speech-based emotion classificationabstractThis paper focuses on speech based emotion classification utilizing acoustic data. The most commonly used acoustic features are pitch and energy, along with prosodic information like rate of speech. We propose the use of a novel feature based on instantaneous frequency obtained from the speech, in addition to the aforementioned features, in order to take into account the vocal tract parameters as well as vocal chord excitation. The proposed features employ the recently emerged empirical mode decomposition to decompose speech into AM-FM signals that are symmetric about zero and suitable for Hilbert transformation to extract the instantaneous frequency. The proposed features provide a relative increase in classification accuracy of approximately 9% when appended to established acoustic features. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
ICASSP | 2 |
| 2008 | Language identification using MLKSFM for pre-classification with novel front-end featuresabstractThis paper presents two novel contributions to automatic language identification. The first one is the use of the modified multi-layer Kohonen self-organizing feature map (MLKSFM) as a pre-classification for language identification (LID). Secondly, we discuss the novel application of empirical mode decomposition (EMD) to generate features for the LID pre-classification task. The use of instantaneous frequency (IF) and instantaneous amplitude (IA) of a speech signal as features for the pre-classifier is investigated. The experiment results on a 16-language speech database indicates that, the EMD by itself cannot perform well in the LID task, however it helps to improve the pre-classification rate when concatenated with other cepstral features. The overall LID performance is also increased when pre-classification is applied. We achieve LID rates of 85.2% and 62.3% for 45-sec and 10-sec test utterances, respectively. Liang Wang 0003, Eliathamby Ambikairajah, Eric H. C. Choi |
ICASSP | 2 |
| 2008 | Accelerometry based classification of gait patterns using empirical mode decompositionabstractThis paper describes accelerometry based classification of walking patterns. A feature extraction technique based on empirical mode decomposition (EMD) is proposed for the classification of unsupervised walking activities from accelerometry data. The front-end 20 dimensional features representing the gait patterns were obtained from the first three modes of decomposition of the acceleration data in anterior-posterior, medio-lateral, and vertical direction. The back-end of the system was a 64-mixture Gaussian Mixture Model (GMM) classifier. Overall classification accuracy of 96.02% was achieved for the five different human gait patterns including walking on flat surfaces, walking up and down paved ramps and walking up and down stairways. Eliathamby Ambikairajah, Branko G. Celler, Nigel H. Lovell |
ICASSP | 2 |
| 2008 | Improvements on hierarchical language identification based on automatic language clusteringabstractHierarchical language identification (HLID) is a novel framework for combining multiple features or primary systems in language identification. In this paper, several key components of HLID are investigated and developed. Crossing likelihood ratio and Kullback-Leibler distance measures are introduced for faster and more accurate clustering. A novel feature selection scheme based on fusion is proposed to incorporate multiple features at each classification level. Further, a phone recognizer followed by language model (PRLM) system is introduced in addition to the other three acoustic systems to provide phonetic information. These proposed techniques improve the performance of HLID system to an EER of 6.3% on the NIST LRE 2003 30s task. Bo Yin 0002, Eliathamby Ambikairajah, Fang Chen 0001 |
ICASSP | 2 |
| 2008 | Speech-based cognitive load monitoring systemabstractMonitoring cognitive load is important for the prevention of faulty errors in task-critical operations, and the development of adaptive user interfaces, to maintain productivity and efficiency in work performance. Speech, as an objective and non-intrusive measure, is a suitable method for monitoring cognitive load. Existing approaches for cognitive load monitoring are limited in speaker-dependent recognition and need manually labeled data. We propose a novel automatic, speaker-independent classification approach to monitor, in real-time, the person's cognitive load level by using speech features. In this approach, a Gaussian mixture model (GMM) based classifier is created with unsupervised training. Channel and speaker normalization are deployed for improving robustness. Different delta techniques are investigated for capturing temporal information. And a background model is introduced to reduce the impact of insufficient training data. The final system achieves 71.1% and 77.5% accuracy on two different tasks, each of which has three discrete cognitive load levels. This performance shows a great potential in real-world applications. Bo Yin 0002, Fang Chen 0001, Natalie Ruiz, Eliathamby Ambikairajah |
ICASSP | 4 |
| 2008 | Digital Signal Processing Techniques for Gene Finding in Eukaryotes
Mahmood Akhtar, Eliathamby Ambikairajah, Julien Epps |
ICISP | 2 |
| 2008 | Performance improvement of text-independent speaker verification systems based on histogram enhancement in noisy environments
C. H. Kwon, J. K. Choi, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2008 | Phonetic and speaker variations in automatic emotion classificationabstractThe speech signal contains information that characterises the speaker and the phonetic content, together with the emotion being expressed. This paper looks at the effect of this speakerand phoneme-specific information on speech-based automatic emotion classification. The performances of a classification system using established acoustic and prosodic features for different phonemes are compared, in both speaker-dependent and speaker-independent modes, using the LDC Emotional Prosody speech corpus. Results from these evaluations indicate that speaker variability is more significant than phonetic variations. They also suggest that some phonemes are easier to classify than others. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2008 | FM features for automatic forensic speaker recognitionabstractFrequency modulation (FM) information from the speech signal is herein proposed to complement the conventional amplitude based features for automatic forensic speaker recognition systems. In addition to presenting the AM-FM model of speech used to generate the proposed frequency modulation features, the significance of frequency modulation for speaker recognition is discussed. Evaluation results from an automatic forensic speaker recognition system combining FM and MFCC features are shown to out-perform those of a system employing MFCC features alone, in terms of all typical metrics, such as detection error trade-off curves, Tippett curves and applied probability of error curves. Index Terms: frequency modulation, automatic forensic speaker recognition. Tharmarajah Thiruvaran, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2008 | Exploring classification techniques in speech based cognitive load monitoring
Bo Yin 0002, Natalie Ruiz, Fang Chen 0001, Eliathamby Ambikairajah |
INTERSPEECH | 4 |
| 2008 | Introducing a FM based feature to hierarchical language identificationabstractAlthough relatively neglected in auditory analysis, phase information plays an important role in human auditory intelligibility. This paper investigates a Frequency Modulation (FM) based feature and its contribution to a Language Identification (LID) system, using a Hierarchical LID framework. FM components represent the phase information of a given signal in an AM-FM model. In this paper, we extract a FM-based feature using a technique which produces consistent and continuous FM components, and build a LID system on this feature with GMM based modeling. The performance is improved by combining this system with existing MFCC, Prosody based systems and a PRLM system. When compared to the baseline system without integrating a FM-based system, the proposed Hierarchical LID system shows improvements. Additionally, the proposed system outperforms the GMM fusion-based system integrating the same four primary systems, showing that the Hierarchical LID framework is more effective in integrating additional features. Index Terms: frequency modulation, language identification, hierarchical classification, fusion Bo Yin 0002, Tharmarajah Thiruvaran, Eliathamby Ambikairajah, Fang Chen 0001 |
INTERSPEECH | 3 |
| 2008 | Investigating speech features and automatic measurement of cognitive loadabstractThe ability to measure cognitive load level in real time is extremely useful for improving the efficiency of interfaces and contents delivering, especially when interfaces and contents get complex in a multimedia environment. Speech is highly suitable for measuring cognitive load due to its non-intrusive nature and ease of collection. In this paper, we investigated the patterns of prosodic features and confirmed it is relevant to cognitive load. We also explored varied classification techniques to capture those relevant patterns of speech features. Gaussian Mixture Model (GMM), Support Vector Machine (SVM), and a hybrid SVM-GMM based classifiers were investigated with MFCC and pitch features. Individual systems and a fusion based system were evaluated on two different task scenarios - reading comprehension and Stroop test. The SVM-GMM based system achieved the highest performance on both tasks and improved the accuracy of three levels classification to 75.6% and 82.2%, respectively. Bo Yin 0002, Natalie Ruiz, Fang Chen 0001, Eliathamby Ambikairajah |
MMSP | 4 |
| 2007 | A comparisonal study of the multi-layer Kohonen self-organizing feature maps for spoken language identificationabstractOur previous research indicates that the multi-layer Kohonen self-organizing feature map (MLKSFM) gives a promising performance for spoken language identification (LID). In this paper, we enhance this approach in two distinct ways. Firstly, by considering the phase information, we propose a new type of feature vector which combines the modified group delay function (MODGDF) and the traditional MFCC. Secondly, we propose a hierarchical structure of the MLKSFM, in which the pre-classification is performed at the lower level MLKSFM and the final language identification is performed at the top level MLKSFM. For the OGI-TS speech corpus, the best LID rate is achieved at 87.3% for the 45-sec test speech utterances by using the hierarchical MLKSFM with 4 classes pre-classified at the lower level MLKSFM. For the 10-sec test speech utterances, the best LID rated is achieved at 60.0% by using the non-hierarchical MLKSFM LID system. Liang Wang 0003, Eliathamby Ambikairajah, Eric H. C. Choi |
ASRU | 2 |
| 2007 | A novel weighting technique for fusing Language Identification systems based on pair-wise performancesabstractOne of the key research issues in modern language identification (LID) research is how best to combine multiple approaches with different features. Existing statistical fusion techniques are popular but have serious limitations when development data is insufficient, since the data is used for training the statistical fuser. In this paper we compare existing fusion techniques for LID systems and propose an alternative to reduce this problem. By deriving the language-specific weighting directly from pair-wise LID performance, a novel weighting approach is introduced and implemented. Experiments on the NIST LRE 2003 task (CallFriend database) and OGI-TS databases demonstrate that the proposed weighting technique outperforms other recent fusion techniques when the available development data is limited. Bo Yin 0002, Eliathamby Ambikairajah, Fang Chen 0001 |
ASRU | 2 |
| 2007 | Time and Frequency Domain Methods for Gene and Exon Prediction in EukaryotesabstractThe detection of period-3 components in exons of eukaryotic gene sequences enables signal processing based time-domain and frequency-domain methods to predict these regions. In this paper, we improve the prediction accuracy of frequency-domain methods by proposing a new algorithm known as the paired and weighted spectral rotation (PWSR) measure, which exploits both period-3 behaviour and another useful statistical property of genomic sequences. By comparison with existing frequency-domain approaches, the proposed PWSR method reveals relative improvements of 15.2% and 10.7% respectively over spectral content and spectral rotation measures in terms of prediction accuracy of exonic nucleotides at a 10% false positive rate using the GENSCAN test set. Finally, we combine the proposed PWSR with an existing time-domain method to demonstrate further signal processing-based improvements in gene and exon prediction accuracy. Mahmood Akhtar, Julien Epps, Eliathamby Ambikairajah |
ICASSP (2) | 3 |
| 2007 | P-Value Segment Selection Technique for Speaker VerificationabstractThis paper presents a segment selection technique for discarding portions of speech that result in poor discrimination ability in speaker verification tasks. Theory supporting the significance of a frame selection procedure for test segments, prior to making decisions, is also developed. This approach has the ability to reduce the effect of the acoustic regions of speech that are not accurately represented due to a lack of training data. Compared with a baseline system using both CMS and variance normalization, the proposed segment selection technique brings 24% relative reduction in error rate over the entire testing data of the 2002 NIST Dataset in terms of minimum DCF. For short test segments, i.e. less than 15 seconds, the novel frame dropping technique produces a significant relative error rate reduction of 23% in terms of minimum DCF. Mohaddeseh Nosratighods, Eliathamby Ambikairajah, Julien Epps, Michael J. Carey 0002 |
ICASSP (4) | 2 |
| 2007 | A Novel Method for Automatic Tonal and Non-Tonal Language ClassificationabstractThis paper describes a novel method for tonal and non-tonal language classification using prosodic information. Normalized feature parameters that measure the speed and level of pitch change are used to perform the classification task. To demonstrate the effectiveness of the proposed method, the classification rates of different system configurations are compared. Evaluating a 16-language classification task using a GMM classifier, the novel system can achieve a classification rate of 87.1% for 45-sec speech segments and 81.9% for 10-sec speech segments. Possible applications of the new method to perform pre-classification in language identification are also discussed. Liang Wang 0003, Eliathamby Ambikairajah, Eric H. C. Choi |
ICME | 2 |
| 2007 | Group delay features for emotion detectionabstractThis paper focuses on speech based emotion classification utilizing acoustic data. The most commonly used acoustic features are pitch and energy, along with prosodic information like the rate of speech. We propose the use of a novel feature based on the phase response of an all-pole model of the vocal tract obtained from linear predictive coefficients (LPC), in addition to the aforementioned features. We compare this feature to other commonly used acoustic features based on classification accuracy. The back-end of our system employs a probabilistic neural network based classifier. Evaluations conducted on the LDC Emotional Prosody speech corpus indicate the proposed features are well suited to the task of emotion classification. The proposed features are able to provide a relative increase in classification accuracy of about 14% over established features when combined with them to form a larger feature vector. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2007 | Multi-layer kohonen self-organizing feature map for language identification
Liang Wang 0003, Eliathamby Ambikairajah, Eric H. C. Choi |
INTERSPEECH | 2 |
| 2007 | Hierarchical language identification based on automatic language clusteringabstractDue to the limitation of single-level classification, existing fusion techniques experience difficulty in improving the performance of language identification when the number of languages and features are further increased. Given that the similarity of feature distribution between different languages may vary, we propose a novel hierarchical language identification framework with multi-level classification. In this approach, target languages are hierarchically clustered into groups according to the distance between them, models are trained both for individual languages and language groups, and classification is hierarchically done in multi-levels. This framework is implemented and evaluated in this paper, the results showing an relative 15.1% error-rate improvement in 30s case on OGI 10-language database compared to modern GMM fusion system. Bo Yin 0002, Eliathamby Ambikairajah, Fang Chen 0001 |
INTERSPEECH | 2 |
| 2007 | A Novel Technique for Noise Reduction in InSAR ImagesabstractThis letter proposes a new technique for noise reduction applied to synthetic aperture radar interferometry. This technique involves a nonlinear filter that separates the interferogram into two components: one containing the smooth (low frequency) part and the other containing the detail (high frequency) part. The smooth part is obtained using a combination of a median filter and a smoothing filter. The detail component is obtained by subtracting the smooth component from the original signal. This detail component is filtered to remove noise and then added to the smooth component to generate the final output. Both simulated and real data are used to evaluate the performance of the proposed technique under different conditions. The experimental results show that the proposed technique outperforms most commonly used interferometric phase filters Vidhyasaharan Sethu, Eliathamby Ambikairajah, Linlin Ge |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2006 | Warped Magnitude and Phase-Based Features for Language IdentificationabstractTo date, systems for the identification of spoken languages have normally used magnitude-based parameterization methods such as the MFCC and PLP. This paper investigates the use of the recently proposed modified group delay function (MODGDF) coefficients in combination with traditional magnitude-based features in a Gaussian mixture model (GMM) based system. We also examine the application of feature warping to magnitude-based features and the MODGDF and find that it can offer a significant cumulative improvement. We find that the addition of a modified regression-based shifted delta cepstrum (SDC) further improves system performance beyond that obtained by a more standard SDC configuration. The combination of PLP, feature warping and the proposed regression-based SDC achieved an accuracy of 88.4% in tests on 10 languages in the OGI TS Corpus, which compares very favourably with alternative language identification systems reported in the literature Felicity Allen, Eliathamby Ambikairajah, Julien Epps |
ICASSP (1) | 2 |
| 2006 | A New Forward Masking Model and its Application to Speech EnhancementabstractThis paper presents a new forward masking model, which is applied to speech enhancement. The model develops a novel expression for forward masking, where the parameters are related to the masker level, the delay and the frequency obtained by curve-fitting the psychoacoustic data. This model is then incorporated, in a novel way, into a speech enhancement scheme. Objective measures using PESQ demonstrates that our enhancement scheme, provides significant improvements over four existing speech enhancement methods, when tested with speech signals corrupted by various noises at very low signal to noise ratios. Hence, the new forward masking model provides a greater and more accurate masking threshold calculation that leads to better PESQ scores Teddy Surya Gunawan, Eliathamby Ambikairajah |
ICASSP (1) | 2 |
| 2005 | Experiences with an electronic whiteboard teaching laboratory and tablet PC based lecture presentations [DSP courses]abstractThis paper presents our experience in constructing an electronic whiteboard-based computer laboratory for teaching digital signal processing (DSP) courses in Australian undergraduate and postgraduate programs. Student interaction with the electronic whiteboard-based tutorial class environment is also reported. Away from the laboratory, DSP lectures were presented using a tablet PC as a digital whiteboard. This supported high quality handwriting annotation of lecture slides, and overcame the limited flexibility present in the existing PowerPoint mode of lecture delivery. For selected self-paced tutorial questions, solutions were provided in electronic format comprising the lecturer's handwritten explanation on a blank slide, input using the tablet PC, combined with audio commentary. An evaluation of student opinions towards this multi-mode delivery of DSP education was illuminating, and the overall experience with these technological aids was that signal processing could be effectively and naturally taught with high student attention span. Eliathamby Ambikairajah, Julien Epps, Ming Sheng, Branko G. Celler, Peter Chen |
ICASSP (5) | 1 |
| 2005 | Language Identification using Warping and the Shifted Delta CepstrumabstractThis paper proposes the novel use of feature warping for automatic language identification, in combination with the shifted delta cepstrum (SDC) and perceptual linear predictive coefficients in a Gaussian mixture model (GMM) based system. Experimental results on various configurations of front-end techniques reported herein demonstrate that, besides providing robustness against channel mismatch and noise as found in existing literature, feature warping is useful more generally as a technique for pre-mapping data for improved compatibility with a GMM back-end. The configuration reported in this paper provides a language identification performance of 76.4% using the OGI/NIST database, a 46.5% relative reduction in error rate when compared with a benchmark system employing Mel frequency cepstral coefficients and the SDC Felicity Allen, Eliathamby Ambikairajah, Julien Epps |
MMSP | 2 |
| 2004 | Perceptual wavelet packet audio coderabstractTraditional wavelet packet audio compression algorithms do not utilize the temporal masking properties of the human auditory system, relying instead on simultaneous masking models. This paper presents the design and implementation of a perceptual wavelet audio coder by incorporating temporal and simultaneous masking models. The efficiency of the encoder was assessed based upon the number of bits required to code wavelet packet coefficients in each critical band, while retaining perceptual transparency. Subjective listening tests conforming to ITU-R BS.1116 revealed the bit rate is reduced by more than 17% compared to using a coder that only employs a simultaneous masking model. Teddy Surya Gunawan, Eliathamby Ambikairajah, Julien Epps |
INTERSPEECH | 2 |
| 2004 | Audio indexing using feature warping and fusion techniquesabstractThis paper reports on the improvement of speech and music indexation performance under various noisy conditions for radio broadcast using warped features fused with traditional features at the output stage. The system employs a bank of four parallel front ends followed by a classification in speech and music by Gaussian mixture models, where each front end employs a different feature extraction technique. Then an automatic gathering in macro classes is made. Indexing was performed on 8 hours of manually labelled radio broadcast from multilingual Radio France International recordings containing diverse speech and music content with different speaking styles, speakers, noise conditions and channels. For speech signal classification under the noisiest conditions, the warped features fused with traditional features produced an error rate three times smaller than that of either the warped features or the traditional features alone. Significant improvements were also found for speech classification under less noisy conditions. Christine Sénac, Eliathamby Ambikairajah |
MMSP | 2 |
| 2003 | Evaluation of a virtual teaching laboratory for signal processing educationabstractThe paper presents our experience in teaching a digital signal processing (DSP) course in an Australian postgraduate program entirely using virtual tele-lectures. A virtual teaching laboratory was designed for this purpose, allowing students to receive fully interactive, real time lectures delivered from a remote international location. We present the methodology and technology used to develop a complete set of tele-lectures and online tools for a course entitled 'Signal processing and applications'. An evaluation of student opinions towards the virtual teaching laboratory revealed that 90% of students rapidly became comfortable with the use of this new educational facility, among other results. The overall experience with the VTL was that signal processing can be effectively and naturally taught in this mode and that there are great potential benefits in connecting the signal processing research and educational community. Eliathamby Ambikairajah, Julien Epps, Ming Sheng, Branko G. Celler |
ICASSP (3) | 1 |
| 2003 | Subband noise estimation for speech enhancement using a perceptual Wiener filterabstractThe paper proposes a fast noise estimation algorithm for speech enhancement using a perceptual Wiener filter. The noisy speech is decomposed using a critical-band-rate filterbank so that a perceptual modification of Wiener filtering can be applied in speech denoising. The subband noise estimate is updated by adaptively smoothing the noisy signal power. The smoothing parameter is chosen as a function of the estimated signal-to-noise ratio. This noise estimation technique gives accurate results even at very low signal-to-noise ratios, and works continuously, even in the presence of speech. It is effective for both non-stationary and coloured noise. Enhanced speech of good quality is obtained by the perceptual Wiener filter. Lee Lin, W. Harvey Holmes, Eliathamby Ambikairajah |
ICASSP (1) | 3 |
| 2002 | Speech enhancement based on a perceptual modification of wiener filtering
Lee Lin, W. Harvey Holmes, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2001 | Wideband speech and audio coding using gammatone filter banksabstractConsiderable research attention has been directed towards speech and audio coding algorithms capable of producing high quality coded speech and audio, however few of these use signal representations which account for temporal as well as spectral detail. This paper presents a new technique for 16 kHz wideband speech and audio coding, whereby analysis and synthesis are performed using a linear phase gammatone filter bank. The outputs of these critical band filters are processed to obtain a series of pulse trains that represent neural firing. Auditory masking is then applied to reduce the number of pulses, producing a more compact time-frequency parameterization. The critical band gains and pulse amplitudes and positions are then coded using a combination of non-uniform quantization, arithmetic coding and vector quantization. This coding paradigm produces high quality coded speech and audio, is based upon well-known models of the auditory system, is highly scalable, and has moderate complexity. Eliathamby Ambikairajah, Julien Epps, Lee Lin |
ICASSP | 1 |
| 2001 | Log-magnitude modelling of auditory tuning curvesabstractWe propose the novel application of a technique for filter design that can accurately fit measured tuning curves for the auditory fibres in the log-magnitude domain. This method provides pole-zero filters with guaranteed stability, and its log-magnitude domain criterion allows tuning curves with very steep slopes to be accurately modelled with an 8/sup th/ to 10/sup th/ order pole-zero filter. Thus, this technique can also be used to design a new set of critical band filters with superior frequency domain characteristics compared with the well-known gammatone filter bank. The filter bank designed using this technique has applications in auditory-based speech and audio analysis. Lee Lin, Eliathamby Ambikairajah, W. Harvey Holmes |
ICASSP | 2 |
| 2001 | Auditory filter bank design using masking curvesabstractIt is very difficult and costly to experimentally observe the motion of the basilar membrane in a fully functional cochlea with the view to obtaining amplitude response at points along the membrane. This paper presents an inexpensive method of generating psychoacoustic tuning curves from the well-known masking curves in critical band rate. We present a method for designing critical band auditory filters from the tuning curves. It is also known that the auditory filter frequency response becomes broader with increasing input signal levels and becomes narrower with decreasing signal levels. We also propose a method for designing level dependent auditory filters. The proposed filter bank is applicable to various types of signal processing required to model human auditory filtering. Lee Lin, Eliathamby Ambikairajah, W. Harvey Holmes |
INTERSPEECH | 2 |
| 1998 | Wavelet transform-based speech enhancementabstractThis paper describes a speech enhancement system using a novel combination of a Fast Wavelet Transform structure, together with “Wiener filtering” in the wavelet domain. The specific application of interest is the enhancement of speech when a cellular phone is used within a moving vehicle. Subjective tests carried out using speech with additive vehicle noise at a signal-to-noise ratio of 10 dB indicate that the Wavelet transform-based Wiener filtering approach works well. In particular, the technique was compared to several other common enhancement methods such as thresholding applied in the wavelet domain, FFT-based Wiener filtering, and spectral subtraction, and was found to outperform these other techniques. Eliathamby Ambikairajah, Graham Tattersall |
ICSLP | 1 |
| 1997 | Comparison of auditory masking models for speech codingabstractIn this paper various auditory masking models recently developed for audio coding are compared and evaluated for telephone bandwidth speech coding applications. Four such models are outlined and their performance evaluated using a Wavelet Packet Transform based subband coder. The models are compared on the basis of the resulting perceptual speech quality and bit rate requirements. Results show that masking models 3 and 4 outlined in this paper provide near transparent quality at the lowest bit rates. M. Lynch, Eliathamby Ambikairajah |
EUROSPEECH | 2 |
| 1995 | Pitch extraction of telephone bandwidth speech using a place-temporal approach
Edward Jones, Eliathamby Ambikairajah |
EUROSPEECH | 2 |
| 1994 | A speech recognition system using both auditory and afferent pathway signal processing
Eliathamby Ambikairajah, Owen Friel, William Millar |
ICSLP | 1 |
| 1993 | The application of the wavelet transform for speech processing
Eliathamby Ambikairajah, M. Keane, Liam Kilmartin, Graham Tattersall |
EUROSPEECH | 1 |
| 1993 | Comparison of various adaptation mechanisms in an auditory model for the purpose of speech processing
Edward Jones, Eliathamby Ambikairajah |
EUROSPEECH | 2 |
| 1993 | Predictive models for speaker verification
Eliathamby Ambikairajah, M. Keane, A. Kelly, Liam Kilmartin, Graham Tattersall |
Speech Commun. | 1 |
| 1992 | Transputer implementation of front-end processors for speech recognition systems
S. Lennon, Eliathamby Ambikairajah |
ICSLP | 2 |
| 1991 | An adaptive cochlear model for speech recognition
Eliathamby Ambikairajah, Liam Kilmartin |
EUROSPEECH | 1 |
| 1991 | A perceptually-based pitch extractor for band-limited speech
Edward Jones, Eliathamby Ambikairajah |
EUROSPEECH | 2 |