Roberto Togneri

dblp:86/3744 · also Roberto B. Togneri · DBLP profile ↗
← Back
109ranked-venue papers
7as first author
13since 2021 · last 2024
0000-0002-3778-4633ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 68 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 58 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-authorSecurity and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 A Unified Loss Function to Tackle Inter-Class and Intra-Class Data Imbalance in Sound Event Detection
abstract
Data imbalance is an important issue in data-driven deep-learning methodologies. In sound event detection (SED), there are two types of data imbalance issues caused by the diverse time duration of sound events: the data imbalance between sound event classes (inter-class imbalance) and the active/inactive imbalance within the class (intra-class imbalance). In this paper, we propose a unified loss function (ULF), which adeptly addresses both the inter-class imbalance and intra-class imbalance simultaneously. Evaluation experiments substantiate that the ULF consistently yields superior and more stable performance compared to existing loss functions that singularly address either type of imbalance. Furthermore, the ULF loss also enhances the model's capacity to detect hard-to-detect sound events.
Roberto Togneri, Defeng Huang
ICASSP2
2024 Spatio-Temporal Graph Representation Learning for Fraudster Group Detection
abstract
Motivated by potential financial gain, companies may hire fraudster groups to write fake reviews to either demote competitors or promote their own businesses. Such groups are considerably more successful in misleading customers, as people are more likely to be influenced by the opinion of a large group. To detect such groups, a common model is to represent fraudster groups' static networks, consequently overlooking the longitudinal behavior of a reviewer, thus, the dynamics of coreview relations among reviewers in a group. Hence, these approaches are incapable of excluding outlier reviewers, which are fraudsters intentionally camouflaging themselves in a group and genuine reviewers happen to coreview in fraudster groups. To address this issue, we propose "FGDT," a framework for "fraudster group detection through temporal relations." FGDT first capitalizes on the effectiveness of the HIN-recurrent neural network (RNN) in both reviewers' representation learning while capturing the collaboration between reviewers. The HIN-RNN models the coreview relations of reviewers in a group in a fixed time window of 28 days. We refer to this as spatial relation learning representation to signify the generalizability of this work to other networked scenarios. Then, we use an RNN on the spatial relations to predict the spatio-temporal relations of reviewers in the group. In the third step, a graph convolution network (GCN) refines the reviewers' vector representations using these predicted relations. These refined representations are then used to remove outlier reviewers. The average of the remaining reviewers' representation is then fed to a simple fully connected layer to predict if the group is a fraudster group or not. Exhaustive experiments of FGDT showed a 5% (4%), 12% (5%), and 12% (5%) improvement over three of the most recent approaches on precision, recall, and F1-value over the Yelp (Amazon) dataset, respectively.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Neural Networks Learn. Syst.2
2023 A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech Enhancement
abstract
Beamforming weights prediction via deep neural networks has been one of the main methods in multi-channel speech enhancement tasks. The spectral-spatial cues are crucial in beamforming weights estimation, however, many existing works fail to optimally predict the beamforming weights with an absence of adequate spectral-spatial information learning. To tackle this challenge, we propose a Fourier convolutional attention encoder (FCAE) to provide a global receptive field over the frequency axis and boost the learning of spectral contexts and cross-channel features. Besides, a new convolutional recurrent encoder-decoder (CRED) structure is proposed in this work, within which FCAEs, attention blocks with skip connections and a deep feedback sequential memory network (DFSMN) serving as recurrent module are involved. The proposed CRED structure is exploited to capture the spectral-spatial joint information to obtain accurate estimation of beamforming weights. Experimental results demonstrate the superiority of the proposed approach with only 0.74M parameters and a PESQ improvement from 2.225 to 2.359 on the ConferencingSpeech2021 challenge development test set.
Xianjun Xia, Yijian Xiao, Piao Ding, Shenyi Song, Roberto Togneri
ICASSP9
2023 A New Approach to Extract Fetal Electrocardiogram Using Affine Combination of Adaptive Filters
abstract
The detection of abnormal fetal heartbeats during pregnancy is important for monitoring the health conditions of the fetus. While adult ECG has made several advances in modern medicine, noninvasive fetal electrocardiography (FECG) remains a great challenge. In this paper, we introduce a new method based on affine combinations of adaptive filters to extract FECG signals. The affine combination of multiple filters is able to precisely fit the reference signal, and thus obtain more accurate FECGs. We proposed a method to combine the Least Mean Square (LMS) and Recursive Least Squares (RLS) filters. Our approach found that the Combined Recursive Least Squares (CRLS) filter achieves the best performance among all proposed combinations. In addition, we found that CRLS is more advantageous in extracting FECG from abdominal electrocardiograms (AECG) with a small signal-to-noise ratio (SNR). Compared with the state-of-the-art Multiple Sub-Filter Adaptive Noise Canceller (MSF-ANC) method, CRLS shows improved performance. The sensitivity, accuracy and F1 score are improved by 3.58%, 2.39% and 1.36%, respectively.
Yu Xuan, Xiangyu Zhang 0005, Shuyue Stella Li, Zihan Shen, L. Paola García-Perera, Roberto Togneri
ICASSP7
2023 A Two-stage Progressive Neural Network for Acoustic Echo Cancellation
abstract
Recent studies in deep learning based acoustic echo cancellation proves the benefits of introducing a linear echo cancellation module. However, the convergence problem and potential target speech distortion impose an additional learning burden for the neural network. In this paper, we propose a two-stage progressive neural network consisting of a coarse-stage and a fine-stage module. For the coarse-stage, a light-weighted network module is designed to suppress partial echo and potential noise, where a voice activity detection path is used to enhance the learned features. For the fine-stage, a larger network is employed to deal with the more complex echo path and restore the near-end speech. We have conducted extensive experiments to verify the proposed method, and the results show that the proposed two-stage method provides a superior performance to other state-of-the-art methods.
Zhuangqi Chen, Xianjun Xia, Xianke Wang, Yanhong Leng, Roberto Togneri, Yijian Xiao, Piao Ding, Shenyi Song, Pingjian Zhang
INTERSPEECH7
2023 A novel quantum calculus-based complex least mean square algorithm (q-CLMS)
Alishba Sadiq, Imran Naseem, Shujaat Khan, Muhammad Moinuddin, Roberto Togneri, Mohammed Bennamoun
Appl. Intell.5
2023 Multi-Kernel Fusion for RBF Neural Networks
abstract
Abstract A simple yet effective architectural design of radial basis function neural networks (RBFNN) makes them amongst the most popular conventional neural networks. The current generation of radial basis function neural network is equipped with multiple kernels which provide significant performance benefits compared to the previous generation using only a single kernel. In existing multi-kernel RBF algorithms, multi-kernel is formed by the convex combination of the base/primary kernels. In this paper, we propose a novel multi-kernel RBFNN in which every base kernel has its own (local) weight. This novel flexibility in the network provides better performance such as faster convergence rate, better local minima and resilience against stucking in poor local minima. These performance gains are achieved at a competitive computational complexity compared to the contemporary multi-kernel RBF algorithms. The proposed algorithm is thoroughly analysed for performance gain using mathematical and graphical illustrations and also evaluated on three different types of problems namely: (i) pattern classification, (ii) system identification and (iii) function approximation. Empirical results clearly show the superiority of the proposed algorithm compared to the existing state-of-the-art multi-kernel approaches.
Atif Muhammad Syed, Shujaat Khan, Imran Naseem, Roberto Togneri, Mohammed Bennamoun
Neural Process. Lett.4
2023 Cross domain 2D-3D descriptor matching for unconstrained 6-DOF pose estimation
abstract
This paper presents a novel approach for cross-domain descriptor matching between 2D and 3D modalities. The 2D-3D matching is applied to localize 2D images in 3D point clouds. Direct cross-domain matching allows our technique to localize images in any type of 3D point cloud without any constraints on the nature or mechanism by which it is obtained. We propose a learning based framework, called Desc-Matcher, to directly match features between the two modalities. A dataset of 2D and 3D features with corresponding locations in images and point clouds is generated to train the Desc-Matcher. To estimate the pose of an image in any 3D cloud, keypoints and feature descriptors are extracted from the query image and the point cloud. The trained Desc-Matcher is then used to match the features from the image and the point cloud. A robust pose estimator is used to predict the location and orientation of the query image from the corresponding positions of the matched 2D and 3D features. We carried out an extensive evaluation of the proposed method for indoor and outdoor scenarios and with different types of point clouds to verify the feasibility of our approach. Experimental results show that the proposed approach can reliably estimate the 6-DOF poses of query cameras in any type of 3D point cloud with high precision. We achieved average median errors of 1.09cm/0.27∘ and 19cm/0.39∘ on the Stanford and Cambridge datasets, respectively.
Uzair Nadeem, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel, Aref Miri Rekavandi, Farid Boussaïd
Pattern Recognit.3
2023 Acoustic characterization and machine prediction of perceived masculinity and femininity in adults
Fuling Chen, Roberto Togneri, Murray Maybery, Diana Weiting Tan
Speech Commun.2
2023 HIN-RNN: A Graph Representation Learning Neural Network for Fraudster Group Detection With No Handcrafted Features
abstract
Social reviews are indispensable resources for modern consumers' decision making. For financial gain, companies pay fraudsters preferably in groups to demote or promote products and services since consumers are more likely to be misled by a large number of similar reviews from groups. Recent approaches on fraudster group detection employed handcrafted features of group behaviors without considering the semantic relation between reviews from the reviewers in a group. In this paper, we propose the first neural approach, HIN-RNN, a Heterogeneous Information Network (HIN) Compatible RNN for fraudster group detection that requires no handcrafted features. HIN-RNN provides a unifying architecture for representation learning of each reviewer, with the initial vector as the sum of word embeddings of all review text written by the same reviewer, concatenated by the ratio of negative reviews. Given a co-review network representing reviewers who have reviewed the same items with the same ratings and the reviewers' vector representation, a collaboration matrix is acquired through HIN-RNN training. The proposed approach is confirmed to be effective with marked improvement over state-of-the-art approaches on both the Yelp (22% and 12% in terms of recall and F1-value, respectively) and Amazon (4% and 2% in terms of recall and F1-value, respectively) datasets.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Neural Networks Learn. Syst.2
2022 ScoreGAN: A Fraud Review Detector Based on Regulated GAN With Data Augmentation
abstract
The promising performance of Deep Neural Networks (DNNs) in text classification has attracted researchers to use them for fraud review detection. However, the lack of trusted labeled data has limited the performance of the current solutions in detecting fraud reviews. The Generative Adversarial Network (GAN) as a semi-supervised method has been demonstrated to be effective for data augmentation purposes. The state-of-the-art solutions utilize GANs to overcome the data scarcity problem. However, they fail to incorporate the behavioral clues in fraud generation. Additionally, state-of-the-art approaches overlook the possible bot-generated reviews in the dataset. Finally, they also suffer from a common limitation in the generalization and stability of the GAN, slowing down the training procedure. In this work, we propose ScoreGAN for fraud review detection that makes use of both review text and review rating scores in the generation and detection process. Scores are incorporated through Information Gain Maximization (IGM) into the loss function for three reasons. One is to generate score-correlated reviews based on the scores given to the generator. Second, the generated reviews are employed to train the discriminator, allowing the discriminator to correctly label the possible bot-generated reviews through joint representations learned from the concatenation of GLobal Vector for Word representation (GLoVe) extracted from the text and the score. Finally, it can be used to improve the stability and generalization of the GAN. Results show that the proposed framework outperformed the existing state-of-the-art FakeGAN framework, in terms of AP by 7%, and 5% on the Yelp and TripAdvisor datasets, respectively.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Inf. Forensics Secur.2
2021 Real time surveillance for low resolution and limited data scenarios: An image set classification approach
Uzair Nadeem, Syed Afaq Ali Shah, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel
Inf. Sci.4
2021 DFraud³: Multi-Component Fraud Detection Free of Cold-Start
abstract
Fraud review detection is a hot research topic in recent years. The Cold-start is a particularly new but significant problem referring to the failure of a detection system to recognize the authenticity of a new user. State-of-the-art solutions employ a translational knowledge graph embedding approach (TransE) to model the interaction of the components of a review system. However, these approaches suffer from the limitation of TransE in handling N-1 relations and the narrow scope of a single classification task, i.e., detecting fraudsters only. In this paper, we model a review system as a Heterogeneous Information Network (HIN) which enables a unique representation to every component and performs graph inductive learning on the review data through aggregating features of nearby nodes. HIN with graph induction helps to address the camouflage issue (fraudsters with genuine reviews) which has shown to be more severe when it is coupled with cold-start, i.e., new fraudsters with genuine first reviews. In this research, instead of focusing only on one component, detecting either fraud reviews or fraud users (fraudsters), vector representations are learned for each component, enabling multi-component classification. In other words, we can detect fraud reviews, fraudsters, and fraud-targeted items, thus the name of our approach DFraud3. DFraud3demonstrates a significant accuracy increase of 13% over the state of the art on Yelp.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Inf. Forensics Secur.2
2020 A Differential Approach for Rain Field Tomographic Reconstruction Using Microwave Signals from Leo Satellites
abstract
A differential approach is proposed for tomographic rain field reconstruction using the estimated signal-to-noise ratio of microwave signals from low earth orbit satellites at the ground receivers, with the unknown baseline values eliminated before using least squares to reconstruct the attenuation field. Simulations are done when the baseline is modelled by an autoregressive process and when the baseline is assumed fixed. Comparisons between the reconstruction results for the differential and non-differential approaches suggest that the differential approach performs better in both scenarios. For high correlation coefficient and low model noise in the autoregressive process, the differential approach surpasses the non-differential approach significantly.
Xi Shen 0002, Defeng Huang, Claire Vincent, Wenxiao Wang 0005, Roberto Togneri
ICASSP5
2020 Performance Analysis for Path Attenuation Estimation of Microwave Signals Due to Rainfall and Beyond
abstract
The attenuation of microwave signals can be used for meteorological observations. For example, the received signal level (RSL) of backhaul links of cellular systems, which usually has a quantization error of 0.1 dB or more for commercial systems, has been used to measure rainfall. In this work, through the mean square error (MSE) analysis of an ideal RSL estimator, it is found that the estimation error can be lower than 0.01 dB for high signal-to-noise ratio (SNR), thereby making it feasible to measure other meteorological variables such as water vapor and clouds. However, the RSL-based estimator has poor performance in low SNR. To improve the performance, we propose a new path attenuation measurement method based on SNR estimation. Although the performance of the SNR-based estimator is better than the RSL based one for low SNR, it becomes worse in high SNR when the path attenuation is small. To solve the problem, another method is proposed based on estimating the signal power (SP) only. Both MSE analysis and simulation results show that the SP-based method is superior to both RSL and SNR based estimators for most scenarios.
Boming Song, Defeng Huang, Xi Shen 0002, Roberto Togneri
ICASSP4
2020 An Objective Voice Gender Scoring System and Identification of the Salient Acoustic Measures
abstract
Human voices vary in their perceived masculinity or femininity, and subjective gender scores provided by human raters have long been used in psychological studies to understand the complex psychosocial relationships between people. However, there has been limited research on developing objective gender scoring of voices and examining the correlation between objective gender scores (including the weighting of each acoustic factor) and subjective gender scores (i.e., perceived masculinity/ femininity). In this work we propose a gender scoring model based on Linear Discriminant Analysis (LDA) and using weakly labelled data to objectively rate speakers' masculinity and femininity. For 434 speakers, we investigated 29 acoustic measures of voice characteristics and their relationships to both the objective scores and subjective masculinity/femininity ratings. The results revealed close correspondence between objective scores and subjective ratings of masculinity for males and femininity for females (correlations of 0.667 and 0.505 respectively). Among the 29 measures, F0 was found to be the most important vocal characteristic influencing both objective and subjective ratings for both sexes. For female voices, local absolute jitter and Harmonic-to-Noise Ratio (HNR) were moderately associated with objective scores. For male voices, F0 variance influenced objective gender scores more than the subjective ratings provided by human listeners. Copyright © 2020 ISCA
Fuling Chen, Roberto Togneri, Murray Maybery, Diana Weiting Tan
INTERSPEECH2
2020 Replay anti-spoofing countermeasure based on data augmentation with post selection
Roberto Togneri, Victor Sreeram
Comput. Speech Lang.2
2020 Sound Event Detection Using Multiple Optimized Kernels
abstract
Sound event detection (SED) has been widely applied in real world applications. Convolutional recurrent neural network based SED approaches have achieved state-of-the-art performance. However, the convolution process is typically performed by using a fixed sized kernel, which adversely affects the detection accuracy especially when the acoustic features of different event classes are characterized by high variations. To deal with this, this article proposes a sound event detection technique using a convolutional recurrent neural network framework with multiple convolutional kernels of different sizes. The top performing kernels are selected from a kernel pool based on the unsupervised clustering errors and the accuracies of the temporarily trained models. Afterwards, the selected kernels are fed to multiple convolution layers to deal with the acoustic feature variations. Experimental results on different subsets of AudioSet, namely the DCASE Challenge 2017 Task 4 and DCASE Challenge 2018 Task 4, demonstrate the performance of the proposed approach compared to state-of-the-art systems.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Multi-Task Learning for Acoustic Event Detection Using Event and Frame Position Information
abstract
Acoustic event detection deals with the acoustic signals to determine the sound type and to estimate the audio event boundaries. Multi-label classification based approaches are commonly used to detect the frame wise event types with a median filter applied to determine the happening acoustic events. However, the multi-label classifiers are trained only on the acoustic event types ignoring the frame position within the audio events. To deal with this, this paper proposes to construct a joint learning based multi-task system. The first task performs the acoustic event type detection and the second task is to predict the frame position information. By sharing representations between the two tasks, we can enable the acoustic models to generalize better than the original classifier by averaging respective noise patterns to be implicitly regularized. Experimental results on the monophonic UPC-TALP and the polyphonic TUT Sound Event datasets demonstrate the superior performance of the joint learning method by achieving lower error rate and higher F-score compared to the baseline AED system.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
IEEE Trans. Multim.2
2019 Signals and Systems: Casting It as an Action-adventure Rather than a Horror Genre
abstract
Simple but effective strategies for an undergraduate introductory course in signals and systems are described in this paper. These include peer facilitated tutorials, optional class tests, in-class only lab assessment and use of interactive animations. Peer facilitated tutorials were designed to support students to help other students. The optional class tests removed the stress and anxiety students face. With in-class only lab assessment the time students spent writing lab reports was replaced with time devoted to preparing and doing the lab together as a group. The use of interactive animations enabled students to visualise key concepts. Survey results and feedback confirm the positive response and improved student satisfaction.
Roberto Togneri, Sally A. Male
ICASSP1
2019 Learning-Based Confidence Estimation for Multi-modal Classifier Fusion
Uzair Nadeem, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri
ICONIP (2)4
2019 Direct Image to Point Cloud Descriptors Matching for 6-DOF Camera Localization in Dense 3D Point Clouds
Uzair Nadeem, Mohammad A. A. K. Jalwana, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel
ICONIP (2)4
2019 3-D Tomographic Reconstruction of Rain Field Using Microwave Signals From LEO Satellites: Principle and Simulation Results
abstract
In this paper, we propose a novel approach for 3-D rain field reconstruction using satellite signals. It uses the estimated signal-to-noise ratio (SNR) at the ground receivers for low-earth orbit (LEO) satellites and, thus, indirectly estimates the path-integrated rain attenuation of the microwave communication links. A least-squares algorithm is employed to perform the 3-D tomographic reconstruction of the rain field. The proposed system model consists of an LEO satellite with a realistic overpass trajectory and multiple ground receivers with SNR estimators. Two synthetic rain events near the Great Barrier Reef in Australia are used to test the reconstruction outcome. Simulation results suggest that the reconstructed rain field has close agreement with the synthetic rain field.
Xi Shen 0002, Defeng Huang, Boming Song, Claire Vincent, Roberto Togneri
IEEE Trans. Geosci. Remote. Sens.5
2019 Auxiliary Classifier Generative Adversarial Network With Soft Labels in Imbalanced Acoustic Event Detection
abstract
In acoustic event detection, the training data size of some acoustic events is often small and imbalanced. To deal with this, this paper proposes generating the virtual training data categorically using the auxiliary classifier generative adversarial networks. Soft labels of acoustic events are first calculated to represent the acoustic event localization information. The closer the current frame is to the middle of the manually labeled acoustic event, the higher the soft label will be, which makes the soft labels positively correlated with the acoustic event localization. Then, the acoustic event class and the quantized soft labels are used as the input condition to the auxiliary classifier generative adversarial networks to generate an arbitrary number of training samples. Experimental results on the TUT Sound Event 2016 under the home environment and TUT Sound Event 2017 under the street environment demonstrate the improved performance of the proposed technique compared to existing acoustic event detection systems.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
IEEE Trans. Multim.2
2018 Acoustic Scene Classification Using Joint Time-Frequency Image-Based Feature Representations
abstract
The classification of acoustic scenes is important in emerging applications such as automatic audio surveillance, machine listening and multimedia content analysis. In this paper, we present an approach for acoustic scene classification by using joint time-frequency image-based feature representations. In acoustic scene classification, joint time-frequency representation (TFR) is shown to better represent important information across a wide range of low and middle frequencies in the audio signal. The audio signal is converted to Constant-Q Transform (CQT) and Mel-spectrum TFRs and local binary patterns (LBP) are used to extract the features from these TFRs. To ensure localized spectral information is not lost, the TFRs are divided into a number of zones. Then, we perform score level fusion to further improve the classification performance accuracy. Our technique achieves a competitive performance with a classification accuracy of 83.4% on the DCASE 2016 development dataset compared to the existing current state of the art.
Shamsiah Abidin, Roberto Togneri, Ferdous Sohel
AVSS2
2018 Finding Word Sense Embeddings of Known Meaning
Lyndon White, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
CICLing (2)2
2018 Confidence Based Acoustic Event Detection
abstract
Acoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt the multi-label classification technique to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the manually labeled boundaries are error-prone and cannot always be accurate, especially when the frame length is too short to be accurately labeled by human annotators. To deal with this, a confidence is assigned to each frame and acoustic event detection is performed using a multi-variable regression approach in this paper. Experimental results on the latest TUT sound event 2017 database of polyphonic events demonstrate the superior performance of the proposed approach compared to the multi-label classification based AED method.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
ICASSP2
2018 Local Binary Pattern with Random Forest for Acoustic Scene Classification
abstract
This paper presents an approach for acoustic scene classification using the local binary pattern (LBP) and random forest (RF). The audio signal is converted to a Constant-Q transform (CQT) representation and LBP is used to extract the features from this time-frequency representation. The CQT representations are divided into a number of sub-bands to obtain more localized features relevant to the spectral information. We then use random forest to select the most important features for each band of extracted LBP features. For further performance enhancement, we use feature level fusion of LBP and HOG features. The proposed system has achieved an accuracy of 85% on the DCASE 2016 dataset.
Shamsiah Abidin, Xianjun Xia, Roberto Togneri, Ferdous Sohel
ICME3
2018 Spoofing Detection Using Adaptive Weighting Framework and Clustering Analysis
abstract
Security of Automatic Speaker Verification (ASV) systems against imposters are now focusing on anti-spoofing countermeasures. Under the severe threat of various speech spoofing techniques, ASV systems can easily be'fooled' by spoofed speech which sounds as real as human-beings. As two effective solutions, the Constant Q Cepstral Coefficients (CQCC) and the Scattering Cepstral Coefficients (SCC) perform well on the detection of artificial speech signals, especially for attacks from speech synthesis (SS) and voice conversion (VC). However, for spoofing subsets generated by different approaches, a low Equal Error Rate (EER) cannot be maintained. In this paper, an adaptive weighting based standalone detector is proposed to address the selective detection degradation. The clustering property of the genuine and the spoofed subsets are analysed for the selection of suitable weighting factors. With a Gaussian Mixture Model (GMM) classifier as the back-end, the proposed detector is evaluated on the ASVspoof 2015 database. The EERs of 0.01% and 0.20% are obtained on the known and the unknown attacks, respectively. This presents an essential complementation between the CQCC and the SCC and also promotes the future research on generalized countermeasures.
Roberto Togneri, Victor Sreeram
INTERSPEECH2
2018 A new proof of a contrast function for bounded component analysis and further analysis
Wei Gao 0009, Shen Fan, Roberto Togneri, Victor Sreeram
Comput. Speech Lang.3
2018 Random forest classification based acoustic event detection utilizing contextual-information and bottleneck features
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
Pattern Recognit.2
2018 Spectrotemporal Analysis Using Local Binary Pattern Variants for Acoustic Scene Classification
abstract
In this paper, we present an approach for acoustic scene classification, which aggregates spectral and temporal features. We do this by proposing the first use of the variable-Q transform (VQT) to generate the time-frequency representation for acoustic scene classification. The VQT provides finer control over the resolution compared to the constant-Q transform (CQT) or short time fourier transform and can be tuned to better capture acoustic scene information. We then adopt a variant of the local binary pattern (LBP), the adjacent evaluation completed LBP (AECLBP), which is better suited to extracting features from acoustic time-frequency images. Our results yield a 5.2% improvement on the DCASE 2016 dataset compared to the application of standard CQT with LBP. Fusing our proposed AECLBP with HOG features, we achieve a classification accuracy of 85.5%, which outperforms one of the top performing systems.
Shamsiah Abidin, Roberto Togneri, Ferdous Sohel
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 RAFP-Pred: Robust Prediction of Antifreeze Proteins Using Localized Analysis of n-Peptide Compositions
abstract
In extreme cold weather, living organisms produce Antifreeze Proteins (AFPs) to counter the otherwise lethal intracellular formation of ice. Structures and sequences of various AFPs exhibit a high degree of heterogeneity, consequently the prediction of the AFPs is considered to be a challenging task. In this research, we propose to handle this arduous manifold learning task using the notion of localized processing. In particular, an AFP sequence is segmented into two sub-segments each of which is analyzed for amino acid and di-peptide compositions. We propose to use only the most significant features using the concept of information gain (IG) followed by a random forest classification approach. The proposed RAFP-Pred achieved an excellent performance on a number of standard datasets. We report a high Youden's index (sensitivity+specificity-1) value of 0.75 on the standard independent test data set outperforming the AFP-PseAAC, AFP_PSSM, AFP-Pred, and iAFP by a margin of 0.05, 0.06, 0.14, and 0.68, respectively. The verification rate on the UniProKB dataset is found to be 83.19 percent which is substantially superior to the 57.18 percent reported for the iAFP method.
Shujaat Khan, Imran Naseem, Roberto Togneri, Mohammed Bennamoun
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Cost-Sensitive Learning of Deep Feature Representations From Imbalanced Data
abstract
Class imbalance is a common problem in the case of real-world object detection and classification tasks. Data of some classes are abundant, making them an overrepresented majority, and data of other classes are scarce, making them an underrepresented minority. This imbalance makes it challenging for a classifier to appropriately learn the discriminating boundaries of the majority and minority classes. In this paper, we propose a cost-sensitive (CoSen) deep neural network, which can automatically learn robust feature representations for both the majority and minority classes. During training, our learning procedure jointly optimizes the class-dependent costs and the neural network parameters. The proposed approach is applicable to both binary and multiclass problems without any modification. Moreover, as opposed to data-level approaches, we do not alter the original data distribution, which results in a lower computational cost during the training process. We report the results of our experiments on six major image classification data sets and show that the proposed approach significantly outperforms the baseline algorithms. Comparisons with popular data sampling techniques and CoSen classifiers demonstrate the superior performance of our proposed method.
Salman Khan 0001, Munawar Hayat, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri
IEEE Trans. Neural Networks Learn. Syst.5
2017 Enhanced LBP texture features from time frequency representations for acoustic scene classification
abstract
This paper introduces the use of local binary patterns (LBP) extracted from a time-frequency representation (TFR) for acoustic scene classification. As LBP provides a description of the global TFR texture we propose a novel zoning mechanism that provides a simple solution to extract spectrally relevant local features which better characterize the audio TFRs. To further improve the classification performance, we perform feature and score level fusion of the proposed LBP (with zoning) with histogram of gradients (HOG) of the TFR images. Our technique demonstrates an improved performance by achieving a classification accuracy of 95.2% using a fusion of time-frequency derived features.
Shamsiah Abidin, Roberto Togneri, Ferdous Sohel
ICASSP2
2017 Random forest regression based acoustic event detection with bottleneck features
abstract
This paper deals with random forest regression based acoustic event detection (AED) by combining acoustic features with bottleneck features (BN). The bottleneck features have a good reputation of being inherently discriminative in acoustic signal processing. To deal with the unstructured and complex real-world acoustic events, an acoustic event detection system is constructed using bottleneck features combined with acoustic features. Evaluations were carried out on the UPC-TALP and ITC-Irst databases which consist of highly variable acoustic events. Experimental results demonstrate the usefulness of the low-dimensional and discriminative bottleneck features with relative 5.33% and 5.51% decreases in error rates respectively.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
ICME2
2017 Random forest classification based acoustic event detection
abstract
This paper deals with the acoustic event detection (AED) to improve the detection accuracy of acoustic events. Acoustic event detection task is performed by a regression via classification (RvC) based approach along with the random forest technique. A discretization process is used to convert the continuous frame positions within acoustic events into event duration class labels. Outputs of the category-specific random forest classifiers are then reversed back to the event boundary information. Evaluations on the UPC-TALP database which consists of highly variable acoustic events demonstrate the efficiency of the proposed approaches with improvements in detection error rate compared to the best baseline system.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
ICME2
2017 Learning deep structured network for weakly supervised change detection
abstract
Conventional change detection methods require a large number of images to learn background models or depend on tedious pixel-level labeling by humans. In this paper, we present a weakly supervised approach that needs only image-level labels to simultaneously detect and localize changes in a pair of images. To this end, we employ a deep neural network with DAG topology to learn patterns of change from image-level labeled training data. On top of the initial CNN activations, we define a CRF model to incorporate the local differences and context with the dense connections between individual pixels. We apply a constrained mean-field algorithm to estimate the pixel-level labels, and use the estimated labels to update the parameters of the CNN in an iterative EM framework. This enables imposing global constraints on the observed foreground probability mass function. Our evaluations on four benchmark datasets demonstrate superior detection and localization performance.
Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri
IJCAI6
2017 A Contrast Function and Algorithm for Blind Separation of Audio Signals
Wei Gao 0009, Roberto Togneri, Victor Sreeram
INTERSPEECH2
2017 Frame-Wise Dynamic Threshold Based Polyphonic Acoustic Event Detection
abstract
Acoustic event detection, the determination of the acoustic event type and the localisation of the event, has been widely applied in many real-world applications. Many works adopt multi-label classification techniques to perform the polyphonic acoustic event detection with a global threshold to detect the active acoustic events. However, the global threshold has to be set manually and is highly dependent on the database being tested. To deal with this, we replaced the fixed threshold method with a frame-wise dynamic threshold approach in this paper. Two novel approaches, namely contour and regressor based dynamic threshold approaches are proposed in this work. Experimental results on the popular TUT Acoustic Scenes 2016 database of polyphonic events demonstrated the superior performance of the proposed approaches.
Xianjun Xia, Roberto Togneri, Ferdous Sohel, Defeng Huang
INTERSPEECH2
2017 Deep feature learning for dummies: A simple auto-encoder training method using Particle Swarm Optimisation
Chao Sui, Mohammed Bennamoun, Roberto Togneri
Pattern Recognit. Lett.3
2017 A cascade gray-stereo visual feature extraction method for visual and audio-visual speech recognition
Chao Sui, Roberto Togneri, Mohammed Bennamoun
Speech Commun.2
2017 A Joint Deep Boltzmann Machine (jDBM) Model for Person Identification Using Mobile Phone Data
abstract
We propose an audio-visual person identification approach based on a joint deep Boltzmann machine (jDBM) model. The proposed jDBM model is trained in three steps: 1) learning the unimodal DBM models corresponding to the speech and facial image modalities, 2) learning the shared layer parameters using a joint restricted Boltzmann machine (jRBM) model, and 3) the fine-tuning of the jDBM model after the initialization with the parameters of the unimodal DBMs and the shared layer. The activation probabilities of the units of the shared layer are used as the joint features and a logistic regression classifier is used for the combined speech and facial image recognition. We show that by learning the shared layer parameters using a jRBM, a higher accuracy can be achieved compared to the greedy layer-wise initialization. The performance of our proposed model is also compared with a state-of-the art support vector machine (SVM), deep belief network (DBN), and the deep auto-encoder (DAE) models. In addition, our experimental results show that the joint representations obtained from the proposed jDBM model are robust to noise and missing information. Experiments were carried out on the challenging MOBIO database, which includes audio-visual data captured using mobile phones.
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel
IEEE Trans. Multim.3
2016 Generating Bags of Words from the Sums of Their Word Embeddings
Lyndon White, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
CICLing (1)2
2016 Audio-visual biometric recognition via joint sparse representations
abstract
In this paper we present a novel audio-visual (AV) person identification system based on joint sparse representation. Video features used were vectorized raw pixel values, while i-vectors were used as the audio features. Classification is performed by solving the joint sparsity optimization problem, and fusion is carried out by using the quality (confidence) assigned to each matcher. Our experimental results on the challenging MOBIO database using 100 subjects show that the system based on joint sparse representation outperforms the system based on separate sparse representations for each modality. Furthermore, we show that our newly introduced quality measure improves the system's performance, when compared to conventionally used quality measures for sparse representation - based systems.
Rudi Primorac, Roberto Togneri, Mohammed Bennamoun, Ferdous Sohel
ICPR2
2016 Integrating Geometrical Context for Semantic Labeling of Indoor Scenes using RGBD Images
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri, Imran Naseem
Int. J. Comput. Vis.4
2016 Automatic Shadow Detection and Removal from a Single Image
abstract
We present a framework to automatically detect and remove shadows in real world scenes from a single image. Previous works on shadow detection put a lot of effort in designing shadow variant and invariant hand-crafted features. In contrast, our framework automatically learns the most relevant features in a supervised manner using multiple convolutional deep neural networks (ConvNets). The features are learned at the super-pixel level and along the dominant boundaries in the image. The predicted posteriors based on the learned features are fed to a conditional random field model to generate smooth shadow masks. Using the detected shadow masks, we propose a Bayesian formulation to accurately extract shadow matte and subsequently remove shadows. The Bayesian formulation is based on a novel model which accurately models the shadow generation process in the umbra and penumbra regions. The model parameters are efficiently estimated using an iterative optimization procedure. Our proposed framework consistently performed better than the state-of-the-art on all major shadow databases collected under a variety of conditions.
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 A Discriminative Representation of Convolutional Features for Indoor Scene Recognition
abstract
Indoor scene recognition is a multi-faceted and challenging problem due to the diverse intra-class variations and the confusing inter-class similarities. This paper presents a novel approach which exploits rich mid-level convolutional features to categorize indoor scenes. Traditionally used convolutional features preserve the global spatial structure, which is a desirable property for general object recognition. However, we argue that this structuredness is not much helpful when we have large variations in scene layouts, e.g., in indoor scenes. We propose to transform the structured convolutional activations to another highly discriminative feature space. The representation in the transformed space not only incorporates the discriminative aspects of the target dataset, but it also encodes the features in terms of the general object categories that are present in indoor scenes. To this end, we introduce a new large-scale dataset of 1300 object categories which are commonly present in indoor scenes. Our proposed approach achieves a significant performance boost over previous state of the art approaches on five major scene classification datasets.
Salman Khan 0001, Munawar Hayat, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel
IEEE Trans. Image Process.4
2015 An Investigation of Neural Embeddings for Coreference Resolution
Varun Godbole, Wei Liu 0006, Roberto Togneri
CICLing (1)3
2015 Separating objects and clutter in indoor scenes
abstract
Objects' spatial layout estimation and clutter identification are two important tasks to understand indoor scenes. We propose to solve both of these problems in a joint framework using RGBD images of indoor scenes. In contrast to recent approaches which focus on either one of these two problems, we perform ‘fine grained structure categorization’ by predicting all the major objects and simultaneously labeling the cluttered regions. A conditional random field model is proposed to incorporate a rich set of local appearance, geometric features and interactions between the scene elements. We take a structural learning approach with a loss of 3D localisation to estimate the model parameters from a large annotated RGBD dataset, and a mixed integer linear programming formulation for inference. We demonstrate that our approach is able to detect cuboids and estimate cluttered regions across many different object and scene categories in the presence of occlusion, illumination and appearance variations.
Salman Khan 0001, Xuming He 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri
CVPR5
2015 Extracting deep bottleneck features for visual speech recognition
abstract
Motivated by the recent progresses in the use of deep learning techniques for acoustic speech recognition, we present in this paper a visual deep bottleneck feature (DBNF) learning scheme using a stacked auto-encoder combined with other techniques. Experimental results show that our proposed deep feature learning scheme yields approximately 24% relative improvement for visual speech accuracy. To the best of our knowledge, this is the first study which uses deep bottleneck feature on visual speech recognition. Our work firstly shows that the deep bottleneck visual feature is able to achieve a significant accuracy improvement on visual speech recognition.
Chao Sui, Roberto Togneri, Mohammed Bennamoun
ICASSP2
2015 Listening with Your Eyes: Towards a Practical Visual Speech Recognition System Using Deep Boltzmann Machines
abstract
This paper presents a novel feature learning method for visual speech recognition using Deep Boltzmann Machines (DBM). Unlike all existing visual feature extraction techniques which solely extracts features from video sequences, our method is able to explore both acoustic information and visual information to learn a better visual feature representation in the training stage. During the test stage, instead of using both audio and visual signals, only the videos are used for generating the missing audio feature, and both the given visual and given audio features are used to obtain a joint representation. We carried out our experiments on a large scale audio-visual data corpus, and experimental results show that our proposed techniques outperforms the performance of the hadncrafted features and features learned by other commonly used deep learning techniques.
Chao Sui, Mohammed Bennamoun, Roberto Togneri
ICCV3
2015 Deep Boltzmann Machines for i-Vector Based Audio-Visual Person Identification
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel
PSIVT3
2015 A confidence-based late fusion framework for audio-visual biometric identification
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel
Pattern Recognit. Lett.3
2014 Automatic Feature Learning for Robust Shadow Detection
abstract
We present a practical framework to automatically detect shadows in real world scenes from a single photograph. Previous works on shadow detection put a lot of effort in designing shadow variant and invariant hand-crafted features. In contrast, our framework automatically learns the most relevant features in a supervised manner using multiple convolutional deep neural networks (ConvNets). The 7-layer network architecture of each ConvNet consists of alternating convolution and sub-sampling layers. The proposed framework learns features at the super-pixel level and along the object boundaries. In both cases, features are extracted using a context aware window centered at interest points. The predicted posteriors based on the learned features are fed to a conditional random field model to generate smooth shadow contours. Our proposed framework consistently performed better than the state-of-the-art on all major shadow databases collected under a variety of conditions.
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri
CVPR4
2014 Geometry Driven Semantic Labeling of Indoor Scenes
Salman Khan 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri
ECCV (1)4
2014 On the use of contextual time-frequency information for full-band clustering-based convolutive blind source separation
abstract
In this paper we propose to incorporate contextual time-frequency information for clustering-based blind source separation. Previous clustering-based approaches have successfully used clustering techniques to estimate time-frequency separation masks; however, these approaches generally do not consider the contextual information of each time-frequency slot. Motivated by the homogenous behavior of speech signals, we modify the fuzzy c-means clustering to bias the results in favor of cluster membership homogeneity within localized neighborhoods in the time-frequency space. Experimental evaluations in both simulated and real-world underdetermined environments demonstrate improvement in source separation performance over previous clustering approaches.
Matt Atcheson, Ingrid Jafari, Roberto Togneri, Sven Nordholm
ICASSP3
2014 Source number estimation in reverberant conditions via full-band weighted, adaptive fuzzy c-means clustering
abstract
We introduce a novel approach for source number estimation through an adaptive fuzzy c-means clustering. Spatial feature vectors are extracted from microphone observations, weighted for reliability and then clustered in a full-band manner using an adaptive variation on the fuzzy c-means. A number of quality measures are combined to produce a weighted sum which is used to find the optimal number of clusters at each iteration of the clustering algorithm. Experimental evaluations using real-world recordings from a reverberant room (RT60= 390 ms) demonstrated encouraging performance in both even- and under-determined conditions.
Joshua Hollick, Ingrid Jafari, Roberto Togneri, Sven Nordholm
ICASSP3
2014 Confidence-based Rank-level Fusion for Audio-visual Person Identification System
abstract
A multibiometric identification system establishes the identity of a person based on the input biometric data presented to its sub-systems. Each sub-system compares the features extracted from the input against the templates of all identities stored in gallery. The best matched identity is ranked highest in the ranked list. In rank-level fusion, the ranked lists from different sub-systems are combined to reach a final decision. However, the state-of-the-art rank-level fusion methods consider that all sub-systems are equally reliable in terms of classifying the probe data. In practice, the probe data may be affected by different sources of degradation (e.g., illumination and pose variation on the face image, environmental noise) and thus affecting the overall recognition accuracy. In this paper, robust rank-level fusion methods (e.g., confidence based highest rank and Borda count) are proposed by using confidence measures for each sub-system in the decision making process. Experimental results show that the proposed confidence based rank-level fusion achieved higher recognition rates than state-of-the-art rank-level fusion methods.
Mohammad Rafiqul Alam, Mohammed Bennamoun, Roberto Togneri, Ferdous Sohel
ICPRAM3
2014 On the use of the Watson mixture model for clustering-based under-determined blind source separation
abstract
Copyright © 2014 ISCA. In this paper, we investigate the application of a generative clustering technique for the estimation of time-frequency source separation masks. Recent advances in time-frequency clustering-based approaches to blind source separation have touched upon the Watson mixture model (WMM) as a tool for source separation. However, most methods have been frequency bin-wise and have thus required the additional permutation alignment stage, and previous full-band methods which employ the WMM have yet to be applied to the under-determined setting. We propose to evaluate the clustering ability of the WMM within the clustering-based source separation framework. Evaluations confirm the superiority of the WMM against other previously used clustering techniques such as the fuzzy c-means.
Ingrid Jafari, Roberto Togneri, Sven Nordholm
INTERSPEECH2
2014 Quantifying target spotting performances with complex geoscientific imagery using ERP P300 responses
Yathunanthan Sivarajah, Eun-Jung Holden, Roberto Togneri, Greg Price, Tele Tan
Int. J. Hum. Comput. Stud.3
2014 Hybrid Metaheuristic Approaches to the Expectation Maximization for Estimation of the Hidden Markov Model for Signal Modeling
abstract
The expectation maximization (EM) is the standard training algorithm for hidden Markov model (HMM). However, EM faces a local convergence problem in HMM estimation. This paper attempts to overcome this problem of EM and proposes hybrid metaheuristic approaches to EM for HMM. In our earlier research, a hybrid of a constraint-based evolutionary learning approach to EM (CEL-EM) improved HMM estimation. In this paper, we propose a hybrid simulated annealing stochastic version of EM (SASEM) that combines simulated annealing (SA) with EM. The novelty of our approach is that we develop a mathematical reformulation of HMM estimation by introducing a stochastic step between the EM steps and combine SA with EM to provide better control over the acceptance of stochastic and EM steps for better HMM estimation. We also extend our earlier work and propose a second hybrid which is a combination of an EA and the proposed SASEM, (EA-SASEM). The proposed EA-SASEM uses the best constraint-based EA strategies from CEL-EM and stochastic reformulation of HMM. The complementary properties of EA and SA and stochastic reformulation of HMM of SASEM provide EA-SASEM with sufficient potential to find better estimation for HMM. To the best of our knowledge, this type of hybridization and mathematical reformulation have not been explored in the context of EM and HMM training. The proposed approaches have been evaluated through comprehensive experiments to justify their effectiveness in signal modeling using the speech corpus: TIMIT. Experimental results show that proposed approaches obtain higher recognition accuracies than the EM algorithm and CEL-EM as well.
Md. Shamsul Huda, John Yearwood, Roberto Togneri
IEEE Trans. Cybern.3
2013 A lip extraction algorithm using region-based ACM with automatic contour initialization
abstract
In a lipreading system, lip extraction is a fundamental method that directly affects the final speech recognition results. However, most existing systems need to detect some facial features as prior-knowledge to construct the initial contour, and any erroneous feature detection will lead to an incorrect lip extraction. In order to solve this problem, this paper presents a new framework which integrates both global region-based Active Contour Model (ACM) and localized region-based ACM. With the utilization of the proposed framework, the initial contour does not need to be specified according to the speaker facial features before extracting the lip, so that any erroneous extraction introduced by an incorrect initial contour is effectively eliminated. Experimental results show the efficiency of the proposed method in comparison with the existing methods.
Chao Sui, Mohammed Bennamoun, Roberto Togneri, Serajul Haque
WACV3
2013 Speech enhancement strategy for speech recognition microcontroller under noisy environments
Kit Yan Chan, Sven Nordholm, Ka Fai Cedric Yiu, Roberto Togneri
Neurocomputing4
2012 A robust approach to reverberant blind source separation in the presence of noise for arbitrarily arranged sensors
abstract
Considerable attention has been devoted to the reverberant blind source separation problem: in particular, the concept of time-frequency masking. However, realistic acoustic scenarios often comprise not only reverberation, but also additive noise due to factors such as non-ideal channels. This paper presents robust evaluations of a time-frequency masking approach for separation in such realistic conditions. The fuzzy c-means clustering algorithm is used to cluster spatial feature cues into a time-frequency mask. Experimental results demonstrated superiority in separation, with notable improvements in the SNR additionally observed. Not only does this establish the proposed scheme viable for reverberant blind source separation, but also as a credible means of speech enhancement in the presence of additive broadband noise.
Ingrid Jafari, Roberto Togneri, Sven Nordholm
ICASSP2
2012 Iterative Group Selection-Based Enhancement of Time-Frequency masks for Missing Data Recognition
abstract
Missing data approaches have recently been applied to speech recognition tasks to increase noise robustness. The drawback of missing data techniques is the vulnerability of the recognizer to errors in the reliability mask. This work proposes a novel group selection algorithm to perform top-down refinement of initial bottom-up reliability mask estimates with the goal of removing these errors. A novel probabilistic decision process based on normalized likelihood distances is proposed and used to evaluate the quality of a reliability mask without any a priori noise knowledge. Experimental results on a speaker identification task illustrate the ability of the combined bottom-up top-down system to significantly outperform traditional bottom-up only missing data techniques for various types of mask corruption.
Daniel Pullella, Roberto Togneri
Int. J. Pattern Recognit. Artif. Intell.2
2012 Robust regression for face recognition
Imran Naseem, Roberto Togneri, Mohammed Bennamoun
Pattern Recognit.2
2011 Speaker verification using sparse representation classification
abstract
Sparse representations of signals have received a great deal of attention in recent years, and the sparse representation classifier has very lately appeared in a speaker recognition system. This approach represents the (sparse) GMM mean supervector of an unknown speaker as a linear combination of an over-complete dictionary of GMM supervectors of many speaker models, and ℓ1-norm minimization results in a non-zero coefficient corresponding to the unknown speaker class index. Here this approach is tested on large databases, introducing channel-/session-variability compensation, and fused with a GMM-SVM system. Evaluations on the NIST 2001 SRE and NIST 2006 SRE database show that when the outputs of the MFCC UBM-GMM based classifier (for NIST 2001 SRE) or MFCC GMM-SVM based classifier (for NIST 2006 SRE) are fused with the MFCC GMM Sparse Representation Classifier (GMM-SRC) based classifier, an absolute gain of 1.27% and 0.25% in EER can be achieved respectively.
Jia Min Karen Kua, Eliathamby Ambikairajah, Julien Epps, Roberto Togneri
ICASSP4
2011 Building an Audio-Visual Corpus of Australian English: Large Corpus Collection with an Economical Portable and Replicable Black Box
abstract
The Big Australian Speech Corpus project incorporates the strategic goals of 30 Chief Investigators from various speech science areas. Speech from 1000 geographically and socially diverse speakers is being recorded using a uniform and automated protocol plus standardized hardware and software to produce a widely applicable and extensible database – AusTalk. Here we describe the project’s major components and organization; share the lessons learnt from difficulties and challenges; and present the results achieved so far.
Denis Burnham, Dominique Estival, Steven Fazio, Jette Viethen, Felicity Cox, Robert Dale, Steve Cassidy, Julien Epps, Roberto Togneri, Michael Wagner 0004, Yuko Kinoshita, Roland Göcke, Joanne Arciuli, Mark Onslow, Trent W. Lewis, Andrew Butcher, John Hajek
INTERSPEECH9
2011 Underdetermined Blind Source Separation with Fuzzy Clustering for Arbitrarily Arranged Sensors
abstract
Recently, the concept of time-frequency masking has developed as an important approach to the blind source separation problem, particularly when in the presence of reverberation. However, previous research has been limited by factors such as the sensor arrangement and/or the mask estimation technique implemented. This paper presents a novel integration of two established approaches to BSS in an effort to overcome such limitations. A multidimensional feature vector is extracted from a non-linear sensor arrangement, and the fuzzy c-means algorithm is then applied to cluster the feature vectors into representations of the source speakers. Fuzzy time-frequency masks are estimated and applied to the observations for source recovery. The evaluations on the proposed study demonstrated improved separation quality over all test conditions. This establishes the potential of multidimensional fuzzy c-means clustering for mask estimation in the context of blind source separation
Ingrid Jafari, Serajul Haque, Roberto Togneri, Sven Nordholm
INTERSPEECH3
2011 An Auditory Motivated Asymmetric Compression Technique for Speech Recognition
abstract
The Mel-frequency cepstral coefficient (MFCC) parameterization for automatic speech recognition (ASR) utilizes several perceptual features of the human auditory system, one of which is the static compression. Motivated by the human auditory system, the conventional static logarithmic compression applied in the MFCC is analyzed using psychophysical loudness perception curves. Following the property of the auditory system that the dynamic range compression is higher in the basal regions than the apical regions of the basilar membrane, we propose a method of unequal (asymmetric) compression, i.e., higher compression applied in the higher frequency regions than the lower frequency regions. The methods is applied and tested in the MFCC and the PLP parameterizations in the spectral domain, and the ZCPA auditory model used as an ASR front-end in the temporal domain. The extent of the asymmetric compression is applied as a multiplicative gain to the existing static compression, and is determined from the gradient of the piece-wise linear segment of the perceptual compression curve. The proposed method has the advantage of adjusting compression parametrically for improved ASR performance and audibility in noise conditions by low-frequency spectral enhancement, particularly of vowels with lowerF1 andF2 formants. Continuous-density HMM recognition using the Aurora 2 corpus and the TIdigits show performance improvements in additive noise conditions.
Serajul Haque, Roberto Togneri, Anthony Zaknich
IEEE Trans. Speech Audio Process.2
2011 A New Evidence Model for Missing Data Speech Recognition With Applications in Reverberant Multi-Source Environments
abstract
Conventional hidden Markov model (HMM) decoders often experience severe performance degradations in practice due to their inability to cope with uncertain data in time-varying environments. In order to address this issue, we propose the bounded-Gauss-Uniform mixture probability density function (pdf) as a new class of evidence model for missing data speech recognition. Exemplary for a hands-free speech recognition scenario, we illustrate how the parameters of the new mixture pdf can be estimated with the help of a multi-channel source separation front-end. In comparison with other models the new evidence pdf retains a fuller description of the available data and provides a more effective link between source separation and recognition. The superiority of the bounded-Gauss-Uniform mixture pdf over conventional approaches is demonstrated for a connected digits recognition task under varying test conditions.
Marco Kühne, Roberto Togneri, Sven Nordholm
IEEE Trans. Speech Audio Process.2
2010 A psychoacoustic spectral subtraction method for noise suppression in automatic speech recognition
abstract
A time-frequency spectral subtraction method based on the knowledge of several psychoacoustic properties of human perception is presented. These effects are the critical band filtering, synaptic adaptation which also introduces temporal forward masking, equal loudness preemphasis, power law of hearing, and simultaneous masking effect. The perceptual speech and noise is estimated separately by a detailed psychoacoustic non-linear transformation undergoing in the human auditory system. The spectral subtraction using a over-subtraction factor and a spectral floor is measured by a speech recognition front-end using a continuous density HMM recognizer. The method shows reduced residual noise and improved word recognition performance in broadband Gaussian noise conditions compared to conventional spectral subtraction method.
Serajul Haque, Roberto Togneri
ICASSP2
2010 Robust Regression for Face Recognition
abstract
In this paper we address the problem of illumination invariant face recognition. Using a fundamental concept that in general, patterns from a single object class lie on a linear subspace, we develop a linear model representing a probe image as a linear combination of class-specific galleries. In the presence of noise, the well-conditioned inverse problem is solved using the robust Huber estimation and the decision is ruled in favor of the class with the minimum reconstruction error. The proposed Robust Linear Regression Classification (RLRC) algorithm is extensively evaluated for two standard databases and has shown good performance index compared to the state-of-art robust approaches.
Imran Naseem, Roberto Togneri, Mohammed Bennamoun
ICPR2
2010 Sparse Representation for Speaker Identification
abstract
We address the closed-set problem of speaker identification by presenting a novel sparse representation classification algorithm. We propose to develop an over complete dictionary using the GMM mean super vector kernel for all the training utterances. A given test utterance corresponds to only a small fraction of the whole training database. We therefore propose to represent a given test utterance as a linear combination of all the training utterances, thereby generating a naturally sparse representation. Using this sparsity, the unknown vector of coefficients is computed via l1-minimization which is also the sparsest solution. Ideally, the vector of coefficients so obtained has nonzero entries representing the class index of the given test utterance. Experiments have been conducted on the standard TIMIT database and a comparison with the state-of-art speaker identification algorithms yields a favorable performance index for the proposed algorithm.
Imran Naseem, Roberto Togneri, Mohammed Bennamoun
ICPR2
2010 A feature extraction method for automatic speech recognition based on the cochlear nucleus
abstract
Motivated by the human auditory system, a feature extraction method for automatic speech recognition (ASR) based on the differential processing strategy of the AVCN, PVCN and the DCN of the cochlear nucleus is proposed. The method utilizes a zero-crossing with peak amplitudes (ZCPA) auditory model as synchrony detector to discriminate the low frequency formants. It utilizes the mean rate information in the synapse processing to capture the very rapidly changing dynamic nature of speech. Additionally, a temporal companding method is utilized for spectral enhancement through two-tone suppression. We propose to separate synchrony detection from synaptic processing as observed in the parallel processing methodology in the cochlear nucleus. HMM recognition using isolated digits showed improved recognition rates in clean and in nonstationary noise conditions than the existing auditory model. Index Terms: Speech recognition, zero-crossings, auditory model, cochlear nucleus, hidden Markov model.
Serajul Haque, Roberto Togneri
INTERSPEECH2
2010 Linear Regression for Face Recognition
abstract
In this paper, we present a novel approach of face identification by formulating the pattern recognition problem in terms of linear regression. Using a fundamental concept that patterns from a single-object class lie on a linear subspace, we develop a linear model representing a probe image as a linear combination of class-specific galleries. The inverse problem is solved using the least-squares method and the decision is ruled in favor of the class with the minimum reconstruction error. The proposed Linear Regression Classification (LRC) algorithm falls in the category of nearest subspace classification. The algorithm is extensively evaluated on several standard databases under a number of exemplary evaluation protocols reported in the face recognition literature. A comparative study with state-of-the-art algorithms clearly reflects the efficacy of the proposed approach. For the problem of contiguous occlusion, we propose a Modular LRC approach, introducing a novel Distance-based Evidence Fusion (DEF) algorithm. The proposed methodology achieves the best results ever reported for the challenging problem of scarf occlusion.
Imran Naseem, Roberto Togneri, Mohammed Bennamoun
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 A novel fuzzy clustering algorithm using observation weighting and context information for reverberant blind speech separation
Marco Kühne, Roberto Togneri, Sven Nordholm
Signal Process.2
2009 A criterion for the enhancement of time-frequency masks in missing data recognition
abstract
Despite their effectiveness for robust speech processing, missing data techniques are vulnerable to errors in the classification of the input speech signal's time-frequency points. A direct method for the removal of these mask errors is through the top-down optimization of the estimated mask, however this requires a measure to evaluate the mask quality without a priori noise knowledge. In this paper we propose the normalized likelihood confidence as such a criterion for robust speaker recognition. In this approach the accuracy with which an estimated mask classifies time-frequency points as corrupt or reliable is related to its likelihood score confidence. This is based on the conceptual effect of binary mask errors on the model likelihood distributions produced by accumulated marginalization densities. Experimental results confirm a relationship between the normalized likelihood distance and the accuracy of the time-frequency mask produced by various estimation strategies.
Daniel Pullella, Roberto Togneri
ICASSP2
2009 Face identification using linear regression
abstract
In this paper we present a novel approach of face identification by formulating the pattern recognition problem in terms of linear regression. Using a fundamental concept that patterns from a single object class lie on a linear subspace, we develop a linear model representing a probe image as a linear combination of class specific galleries. The inverse problem is solved using the least squares method and the decision is ruled in favor of the class with the minimum reconstruction error. The algorithm is extensively evaluated using two standard databases, a comparative study with the benchmark algorithms clearly reflects the efficacy of the proposed approach.
Imran Naseem, Roberto Togneri, Mohammed Bennamoun
ICIP2
2009 A stochastic version of Expectation Maximization algorithm for better estimation of Hidden Markov Model
Md. Shamsul Huda, John Yearwood, Roberto Togneri
Pattern Recognit. Lett.3
2009 Perceptual features for automatic speech recognition in noisy environments
Serajul Haque, Roberto Togneri, Anthony Zaknich
Speech Commun.2
2009 Robust Source Localization in Reverberant Environments Based on Weighted Fuzzy Clustering
abstract
Successful localization of sound sources in reverberant enclosures is an important prerequisite for many spatial signal processing algorithms. We investigate the use of a weighted fuzzyc-means cluster algorithm for robust source localization using location cues extracted from a microphone array. In order to increase the algorithm's robustness against sound reflections, we incorporate observation weights to emphasize reliable cues over unreliable ones. The weights are computed from local feature statistics around sound onsets because it is known that these regions are least affected by reverberation. Experimental results illustrate the superiority of the method when compared with standard fuzzy clustering. The proposed algorithm successfully located two speech sources for a range of angular separations in room environments with reverberation times of up to 600 ms.
Marco Kühne, Roberto Togneri, Sven Nordholm
IEEE Signal Process. Lett.2
2009 A Constraint-Based Evolutionary Learning Approach to the Expectation Maximization for Optimal Estimation of the Hidden Markov Model for Speech Signal Modeling
abstract
This paper attempts to overcome the tendency of the expectation-maximization (EM) algorithm to locate a local rather than global maximum when applied to estimate the hidden Markov model (HMM) parameters in speech signal modeling. We propose a hybrid algorithm for estimation of the HMM in automatic speech recognition (ASR) using a constraint-based evolutionary algorithm (EA) and EM, the CEL-EM. The novelty of our hybrid algorithm (CEL-EM) is that it is applicable for estimation of the constraint-based models with many constraints and large numbers of parameters (which use EM) like HMM. Two constraint-based versions of the CEL-EM with different fusion strategies have been proposed using a constraint-based EA and the EM for better estimation of HMM in ASR. The first one uses a traditional constraint-handling mechanism of EA. The other version transforms a constrained optimization problem into an unconstrained problem using Lagrange multipliers. Fusion strategies for the CEL-EM use a staged-fusion approach where EM has been plugged with the EA periodically after the execution of EA for a specific period of time to maintain the global sampling capabilities of EA in the hybrid algorithm. A variable initialization approach (VIA) has been proposed using a variable segmentation to provide a better initialization for EA in the CEL-EM. Experimental results on the TIMIT speech corpus show that CEL-EM obtains higher recognition accuracies than the traditional EM algorithm as well as a top-standard EM (VIA-EM, constructed by applying the VIA to EM).
Md. Shamsul Huda, John Yearwood, Roberto Togneri
IEEE Trans. Syst. Man Cybern. Part B3
2008 Towards the use of full covariance models for missing data speaker recognition
abstract
This work investigates the use of missing data techniques for noise robust speaker identification. Most previous work in this field relies on the diagonal covariance assumption in modeling speaker specific characteristics via Gaussian mixture models. This paper proposes the use of full covariance models that can capture linear correlations among feature components. This is of importance for missing data marginalization techniques as they depend on spectral rather than cepstral feature representations. Bounded and complete marginalization schemes are investigated both with diagonal and full covariance mixture models. Speaker identification experiments using stationary and non-stationary noise confirm that full covariance models are indeed superior compared to diagonal models.
Marco Kühne, Daniel Pullella, Roberto Togneri, Sven Nordholm
ICASSP3
2008 Robust speaker identification using combined feature selection and missing data recognition
abstract
Missing data techniques have been recently applied to speaker recognition to increase performance in noisy environments. The drawback of these techniques is the vulnerability of the recognizer to errors in the classification of time-frequency points as corrupt or reliable. In this paper we propose the combination of missing data processing and feature selection to reduce these errors. The formation of a set of speaker discriminative features allows time-frequency reliability masks to be refined via the removal of the non-discriminative frequency sub-bands. The reduced set is selected dynamically using multi-condition training and an estimate of the global SNR allowing for efficient top-down processing. Experimental results show that the combined technique achieves significant improvement over traditional bottom-up processing thus demonstrating the validity of the approach.
Daniel Pullella, Marco Kühne, Roberto Togneri
ICASSP3
2008 Adaptive beamforming and soft missing data decoding for robust speech recognition in reverberant environments
abstract
Abstract This paper presents a novel approach to combine microphonearray processing and robust speech recognition for reverberantmulti-speaker environments. Spatial cues are extracted from amicrophone array andautomaticallyclusteredtoestimate local-ization masks in the time-frequency domain. The localizationmasks are then used to blindly design adaptive filters in order toenhancethesourcesignalspriortomissingdataspeechrecogni-tion. A novel evidence model better exploiting the informationprovided by the source separation stage is proposed. Recogni-tion experiments demonstrate the effectiveness of the schemewhen compared to traditional microphone array enhancementand a related binaural separation model. Index Terms : missing data speech recognition, microphone ar-ray processing, adaptive beamforming 1. Introduction A key requirement for automatic speech recognition (ASR)technology to be employed in everyday situations is robust-ness to multiple competing speakers in reverberant enclosures.While today’s ASR systems still fall short in comparison withhuman listeners some promising success has been achieved indealing with multiple speakers in anechoic conditions. For ex-ample, the work reported in [1] utilizes binaural localizationcuessuchasinterauraltimeandintensitydifferencestoestimatetime-frequency (TF) masks for missing data speech recognition(MD-ASR).However,inreverberantenclosuresthelocalizationcues become increasingly unreliable making an accurate esti-mation of the TF mask considerably more difficult. Further-more when the speech models are trained on anechoic data thespectral features employed in MD-ASR are adversely affectedby the room reflections.Previous attempts to handle multisource reverberant en-vironments include dereverberation filters [2], adaptation ofspeech models to the room [3], an inhibition mechanism to em-phasizesoundonsets[4]andfeatureenhancementusinganane-choic speech prior based reconstruction technique [5].This paper proposes an alternative approach by extend-ing the two-channel system developed in [6] to the multi-microphone case (see Figure 1a). The system consists of asource separation stage coupled with a missing data decoderthrough the probabilistic concept of evidence models [7]. Forthe source separation stage a fuzzy clustering approach is em-ployed using direction of arrival (DOA) values of the two outersensors in the array. The clustering produces a source DOAestimate and a TF mask marking the dominant TF points foreach source. While in [6] the correlation between adjacent TFpoints wasonly utilized duringmask post-processing inthispa-per the neighborhood information is integrated during the clus-tering process itself. This biases the solution towards homoge-nous masks and helps greatly to reduce the effect of noise visi-ble in the localization cues under reverberant conditions.Motivated by [8] the masks and source DOAs are then uti-lized to blindly design multiple adaptive beamformers capableofenhancingthespeechsourcespriortorecognition. Thisworkproposes a novel evidence model incorporating both localiza-tion information and feature uncertainty. The system operatesinanunsupervised manneranddependsneither ontrainingdatafor source localization nor a priori learned echoic speech mod-els adapted to the reverberant environment.The remainder of this paper is as follows: Section 2 de-scribes the source separation stage in more detail. Section 3discusses the missing data recognizer and the evidence modelparameter estimation. Section 4 reports the evaluation resultsand compares these with a related binaural model. The papercloses in Section 5 with an outlook into future work.
Marco Kühne, Roberto Togneri, Sven Nordholm
INTERSPEECH2
2007 A Temporal Auditory Model with Adaptation for Automatic Speech Recognition
abstract
Rapid and short-term adaptation are dynamic mechanisms of human auditory system. An auditory model based on zero-crossings with peak amplitudes (ZCPA) was used as a front-end for automatic speech recognition (ASR) with the perceptual property of adaptation as determined by psychoacoustic observations. The model performance was evaluated on the isolated digits (TIDIGITS) database using continuous density HMM recognizer in additive noise environment. Experimental results indicate that the ASR performance of the ZCPA may be improved with adaptation over the static baseline performance in white Gaussian and factory noise. The perceptual front-end was also evaluated with dynamic (delta and delta-delta) features added to the adaptation. It was observed that adaptation with dynamic features performed better in factory, babble and car noise over a wide range of SNR values.
Serajul Haque, Roberto Togneri, Anthony Zaknich
ICASSP (4)2
2007 Mel-Spectrographic Mask Estimation for Missing Data Speech Recognition using Short-Time-Fourier-Transform Ratio Estimators
abstract
This paper adopts the framework of DUET, a recently proposed blind source separation (BSS) method, for speech recognition. Based on the attenuation and delay estimation in stereo signals spectrographic masks are designed to extract a target speaker from a mixture containing multiple speech sources. Instead of using these masks for resynthesis we avoid source reconstruction and propose to combine the source separation with a missing data speech recognizer. The obtained results for connected digit experiments in a multi-speaker environment demonstrate the validity of the approach.
Marco Kühne, Roberto Togneri, Sven Nordholm
ICASSP (4)2
2007 Smooth soft mel-spectrographic masks based on blind sparse source separation
abstract
Abstract This paper investigates the use of DUET, a recently proposedblind source separation method, as front-end for missing dataspeech recognition. Based on the attenuation and delay estima-tion in stereo signals soft time-frequency masks are designedto extract a target speaker from a mixture containing multiplespeech sources. A postprocessing step is introduced in order toremove isolated mask points that can cause insertion errors inthe speech decoder. The results for connected digit experimentsin a multi-speaker environment demonstrate that the proposedsoft masks closely match the performance of the oracle maskdesigned with a priori knowledge of the source spectra. Index Terms : speech recognition, missing data, attenuationand delay estimation 1. Introduction The concept of time-frequency (TF) masking has recently at-tracted some interest in the field of blind signal separation (BSS)[1, 2]. Demixing via TF-masks has the potential to separatemixtures with more sources than sensors as it does not rely onmatrix inversion. Instead the TF-plane is partitioned into dis-joint regions each assigned to a particular source. The sourcesare then recovered by converting each region back into the timedomain. It seems promising to use BSS systems as front-endsfor automatic speech recognition (ASR). In [3] we have pro-posed such a combination using a BSS technique called DUETand a missing data (MD) speech recognizer. The proposedsystem uses DUET to estimate TF-masks in the sparse Short-Time-Fourier-Transform (STFT) domain before converting thehigh STFT frequency resolution to a perceptual mel-frequencyscale suitable for ASR. In this way we can avoid source re-construction and directly exploit the spectrographic masks forMD-ASR. This paper extends our previous work in two regards.Firstly, we replace binary masks with soft masks which havebeen proven to be beneficial for both speech recognition andspeech enhancement. Several studies [2, 4] have reported thatbinary TF-masking can lead to audible unnatural sound arti-facts that can be avoided to some degree by soft masks. Theadvantages of soft masks in MD-ASR are even more evidentas marginalization approaches based on soft decisions consis-tently outperformed hard masks [5]. Secondly, we show thata simple mask postprocessing can lead to substantial recogni-tion improvements. A two-dimensional (2-D) median filter wasapplied to reduce the influence of outliers visible as scattered
Marco Kühne, Roberto Togneri, Sven Nordholm
INTERSPEECH2
2007 A structured speech model parameterized by recursive dynamics and neural networks
abstract
We present in this paper an overview of the Hidden Dynamic Model (HDM) paradigm, exemplifying parametric construction of structure-based speech models that can be used for recog-nition purposes. We explore a general class of the HDM that uses recursive, autoregression functions to represent the hid-den speech dynamics, and uses neural networks to represent the functional relationship between the hidden and observed speech vectors. This type of state-space formulation of the HDM is re-viewed in terms of model construction, a parameter estimation technique, and a decoding method. We also present some typ-ical experimental results on the use of this type of HDMs for phonetic recognition and for automatic vocal tract resonance tracking. We further provide analyses on the computational complexity (for decoding) and the parameter size of the HDM in comparison with the HMM. Finally, we discuss several key issues related to future exploration of the HDM paradigm. Index Terms: hidden dynamic model, recursive form of dynam-ics, neural network, nonlinear mapping, formant tracking, pho-netic recognition
Roberto Togneri, Li Deng 0001
INTERSPEECH1
2007 Feature and distribution normalization schemes for statistical mismatch reduction in reverberant speech recognition
abstract
Reverberant noise has been a major concern in speech recognition systems. Many speech recognition systems, even with state-of-art features, fail to respond to reverberant effects and the recognition rate deteriorates. This paper explores the significance of normalization strategies in reducing statistical mismatches for robust speech recognition in reverberant environment. Most normalization works focused only on ambient noise and have yet been experimented on reverberant noise. In addition, we propose a new approach for the odd order cepstral moment normalization which is computationally more efficient and reduces the convergence rate in the algorithm. The proposed method is experimentally justified and corroborated by the performance of other normalization schemes. The results emphasize the significance of reducing statistical mismatches in feature space for reverberant speech recognition.
A. M. Toh, Roberto Togneri, Sven Nordholm
INTERSPEECH2
2006 Prosodic features for a maximum entropy language model
abstract
A statistical language model attempts to characterise the patterns present in a natural language as a probability distribution defined over word sequences. Typically, they are trained using word co-occurrence statistics from a large sample of text. In some language modelling applications, such as automatic speech recognition (ASR), the availability of acoustic data provides an additional source of knowledge. This contains, amongst other things, the melodic and rhythmic aspects of speech referred to as prosody. Although prosody has been found to be an important factor in human speech recognition, its use in ASR has been limited. The goal of this research is to investigate how prosodic information can be employed to improve the language modelling component of a continuous speech recognition system. Because prosodic features are largely suprasegmental, operating over units larger than the phonetic segment, the language model is an appropriate place to incorporate such information. The prosodic features and standard language model features are combined under the maximum entropy framework, which provides an elegant solution to modelling information obtained from multiple, differing knowledge sources. We derive features for the model based on perceptually transcribed Tones and Break Indices (ToBI) labels, and analyse their contribution to the word recognition task. While ToBI has a solid foundation in linguistic theory, the need for human transcribers conflicts with the statistical model's requirement for a large quantity of training data. We therefore also examine the applicability of features which can be automatically extracted from the speech signal. We develop representations of an utterance's prosodic context using fundamental frequency, energy and duration features, which can be directly incorporated into the model without the need for manual labelling. Dimensionality reduction techniques are also explored with the aim of reducing the computational costs associated with training a maximum entropy model. Experiments on a prosodically transcribed corpus show that small but statistically significant reductions to perplexity and word error rates can be obtained by using both manually transcribed and automatically extracted features.
Oscar Chan, Roberto Togneri
INTERSPEECH2
2006 Automatic English stop consonants classification using wavelet analysis and hidden Markov models
abstract
This paper compares wavelet and STFT analysis for a speakerindependent stop classification task using the TIMIT database. In the designed experiment the HMM classifier had to assign each test token to one of the following stop classes [d,g,b,t,k,p,dx]. On 6332 stops the wavelet features obtained an overall accuracy of 86 % which corresponds to a 14 % relative error reduction compared to the STFT baseline system. Furthermore an analysis of the HMM misclassifications revealed that voiced stops were highly confused with their voiceless unaspirated counterparts.
Marco Kühne, Roberto Togneri
INTERSPEECH2
2006 A state-space model with neural-network prediction for recovering vocal tract resonances in fluent speech from Mel-cepstral coefficients
Roberto Togneri, Li Deng 0001
Speech Commun.1
2006 Statistical voice activity detection using low-variance spectrum estimation and an adaptive threshold
abstract
Traditionally, voice activity detection algorithms are based on any combination of general speech properties such as temporal energy variations, periodicity, and spectrum. This paper describes a novel statistical method for voice activity detection using a signal-to-noise ratio measure. The method employs a low-variance spectrum estimate and determines an optimal threshold based on the estimated noise statistics. A possible implementation is presented and evaluated over a large test set and compared to current modern standardized algorithms. The evaluations indicate promising results with the proposed scheme being comparable or favorable over the whole test set.
Alan Davis, Sven Nordholm, Roberto Togneri
IEEE Trans. Speech Audio Process.3
2004 Spatio-temporal processing for distant speech recognition
abstract
A new subband based front-end processor for speech recognition is presented. It integrates both spatial and temporal signal processing methods to enhance noisy signals as a means to reduce the mismatch problem in speech recognition. The approach makes use of the popular blind signal separation (BSS) to spatially separate the target signal from the interference. Due to the multipath/reverberant environment, BSS has its fundamental limitation in the separation quality. To overcome that, an adaptive noise canceller (ANC) is employed to perform further interference reduction. Experimental results show that even in an adverse environment, the proposed structure improves the word recognition rate (WRR) by 70% for the connected digit recognition task.
Siow Yong Low, Roberto Togneri, Sven Nordholm
ICASSP (1)2
2004 Use of neural network mapping and extended kalman filter to recover vocal tract resonances from the MFCC parameters of speech
abstract
In this paper, we present a state-space formulation of a neuralnetwork-based hidden dynamic model of speech whose parameters are trained using an approximate EM algorithm. The training makes use of the results of an off-the-shelf formant tracker (during the vowel segments) to simplify the complex sufficient statistics that would be required in the exact EM algorithm. The trained model, consisting of the state equation for the target-directed vocal tract resonance (VTR) dynamics on all classes of speech sounds (including consonant closure) and the observation equation for mapping from the VTR to acoustic measurement, is then used to recover the unobserved VTR based on Extended Kalman Filter. The results demonstrate accurate estimation of the VTRs, especially those during rapic consonant-vowel or vowel-consonant transitions and during consonant closure when the acoustic measurement alone provides weak or no information to infer the VTR values.
Li Deng 0001, Roberto Togneri
INTERSPEECH2
2004 Convolutive blind signal separation with post-processing
abstract
A new subband based speech enhancement scheme is presented. It integrates spatial and temporal signal processing methods to enhance speech signals in a noisy environment. The approach makes use of the popular blind signal separation (BSS) to spatially separate the target signal from the interference. Due to the multipath/reverberant environment, BSS has its fundamental limitation in its separation quality. To overcome that, an adaptive noise canceller (ANC) is employed to perform further interference reduction. The reference for the ANC in this case is simply the interference dominant output from the BSS. A higher order statistical method is proposed for the selection of the reference signal. This post processing acts as a spectral decorrelator and experimental results show that even in under-determined (more sources than elements) case, the structure offers impressive enhancement capability. Further, a remarkable improvement in recognition rate is registered when tested in automatic speech recognition (ASR).
Siow Yong Low, Sven Nordholm, Roberto Togneri
IEEE Trans. Speech Audio Process.3
2001 An EKF-based algorithm for learning statistical hidden dynamic model parameters for phonetic recognition
abstract
Presents a parameter estimation algorithm based on the extended Kalman filter (EKF) for the statistical coarticulatory hidden dynamic model (HDM). We show how the EKF parameter estimation algorithm unifies and simplifies the estimation of both the state and parameter vectors. Experiments based on N-best rescoring demonstrate superior performance of the (context-independent) HDM over a triphone baseline HMM in the TIMIT phonetic recognition task. We also show that the HDM is capable of generating speech vectors close to those from the corresponding real data.
Roberto Togneri, Li Deng 0001
ICASSP1
2001 Parameter estimation of a target-directed dynamic system model with switching states
Roberto Togneri, Jeff Z. Ma, Li Deng 0001
Signal Process.1
2000 A robust speech understanding system using conceptual relational grammar
abstract
We describe a robust speech understanding system based on our newly developed approach to spoken language processing. We show that a robust NLU system can be rapidly developed using a relatively simple speech recognizer to provide sufficient information for database retrieval by spoken language. Our experimental system consists of three components: a speech recognizer based on HMM, a natural language parser based on conceptual relational grammar and a data retrieval system based on the ATIS database. With the use of the robust parsing strategy, database query tasks can be successfully performed. 1.
Jiping Sun, Roberto Togneri, Li Deng 0001
INTERSPEECH2
1998 Speech recognition using the probabilistic neural network
abstract
A novel technique for speaker independent automated speech recognition is proposed. We take a segment model approach to Automated Speech Recognition (ASR), considering the trajectory of an utterance in vector space, then classify using a modified Probabilistic Neural Network (PNN) and maximum likelihood rule. The system performs favourably with established techniques. Our system achieves in excess of 94 % with isolated digit recognition, 88 % with isolated alphabetic letters, and 83% with the confusable /e / set. A favourable compromise between recognition accuracy and computer memory and speech can also be reached by performing clustering on the training data for the PNN. 1.
Raymond Low, Roberto Togneri
ICSLP2
1998 Parallel program analysis on workstation clusters: Memory utilisation and load balancing
abstract
Using a workstation cluster for parallel program development requires consideration of various factors to optimise the mapping of the algorithm to the characteristics of the environment. In this paper we present a new analysis and verification of well-known ideas in parallel programming research of specific importance to both the use and design of workstation cluster computing systems. We define a new performance measure related to memory resource utilisation and show how redundant memory usage can lead to poor memory utilisation of the cluster. We also present analytical and experimental evidence that the pool-of-tasks paradigm can lead to significantly improved speedup over series–parallel algorithms, especially when considering equivalent computational and communication requirements. The effect of load balancing on the series–parallel and pool-of-tasks algorithms is examined, and our analysis and experimental results confirm not only that the pool-of-tasks algorithms are more robust to load imbalances but that the effect of the imbalance is mitigated when more workstations are used. © 1998 John Wiley & Sons, Ltd.
Roberto Togneri
Concurr. Pract. Exp.1
1997 Parallel program analysis on workstation clusters: speedup profiling and latency hiding
abstract
Parallel programming on workstation clusters is subject to many factors and problems which determine the potential success or failure of any individual implementation. The most obvious problems are the difficulty in developing parallel algorithms and the high communication latency which may render such algorithms inefficient. In an attempt to address some of these issues we propose a strategy for estimating the potential speedup of a parallel program based on computation and communication profiling. We show that our proposed strategy yields accurate estimates of the speedup. We also propose a complete communication model so that the speedup can be estimated under different programming inputs and show that moderately accurate estimates can be obtained. High communication latency is the major problem with workstation cluster computing. We attempt to examine this problem from the system level point of view and show experimentally that latency hiding can allow almost full utilisation of the CPU resource even though individual programs may suffer from high communication latencies. © 1997 John Wiley & Sons, Ltd.
Roberto Togneri
Concurr. Pract. Exp.1
1997 Phoneme-based vector quantization in a discrete HMM speech recognizer
abstract
The quantization distortion of vector quantization (VQ) is a key element that affects the performance of a discrete hidden Markov modeling (DHMM) system. Many researchers have realized this problem and tried to use integrated feature or multiple codebook in their systems to offset the disadvantage of the conventional VQ. However the computational complexity of those systems is then increased. Investigations have shown that the speech signal space consists of finite clusters that represent phoneme data sets from male and female speakers and reveal Gaussian distributions. We propose an alternative VQ method in which the phoneme is treated as a cluster in the speech space and a Gaussian model is estimated for each phoneme. A Gaussian mixture model (GMM) is generated by the expectation-maximization (EM) algorithm for the whole speech space and used as a codebook in which each code word is a Gaussian model and represents a certain cluster. An input utterance would be classified as a certain phoneme or a set of phonemes only when the phoneme or phonemes gave highest likelihood. A typical discrete HMM system was used for both phoneme and isolated word recognition. The results show that the phoneme-based Gaussian modeling vector quantization classifies the speech space more effectively and significant improvements in the performance of the DHMM system have been achieved.
Roberto Togneri, Michael D. Alder
IEEE Trans. Speech Audio Process.2
1994 Using Gaussian mixture modeling in speech recognition
abstract
The paper describes a speaker-independent isolated word recognition system which uses a well known technique, the combination of vector quantization with hidden Markov modeling. The conventional vector quantization algorithm is substituted by a statistical clustering algorithm, the expectation-maximization algorithm, in this system. Based on the investigation of the data space, the phonemes were manually extracted from the training data and were used to generate the Gaussians in a code book in which each code word is a Gaussian rather than a centroid vector of the data class. Word-based hidden Markov modeling was then performed. Two English isolated digits data bases were investigated and the 12 Mel-spaced filter bank coefficients employed as the input feature. Compared with the conventional discrete HMM, the present system obtained a significant improvement of recognition accuracy.>
Michael D. Alder, Roberto Togneri
ICASSP (1)3
1990 An English language speech database at the University of Western Australia
abstract
The authors present a report on the content and status of a major speech database collection effort. The goal is to collect a useful set of speech material from a very large number of speakers. These speakers are drawn from a wide cross-section of the local community with a variety of ethnic and educational backgrounds. Speech materials include isolated digits and numbers, vowels and voiced phonemes, connected digits, and phonetically balanced sentences. Speech signals are encoded into 16-bit pulse code modulation (PCM) format and stored on Betamax format video tapes. In the seven months since this project started, speech from 100 speakers has been collected. A statistical breakdown of the backgrounds of the speakers is presented.>
Edmund M.-K. Lai, G. A. Carrijo, R. Bennett, Roberto Togneri, Michael D. Alder, Yianni Attikiouzel
ICASSP4
1990 Kohonen's algorithm for the numerical parametrisation of manifolds
Michael D. Alder, Roberto Togneri, Edmund M.-K. Lai, Yianni Attikiouzel
Pattern Recognit. Lett.2