Ian McLoughlin 0001

dblp:50/170 · also Ian Vince McLoughlin · DBLP profile ↗
← Back
114ranked-venue papers
16as first author
39since 2021 · last 2025
0000-0001-7111-2008ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 77 · 6 first-author · 31 since 2021Artificial intelligence and machine learning · 52 · 8 first-author · 19 since 2021Systems, architecture and hardware · 10 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Computer networks · 4Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 A Dual Modular Clock Algorithm for Cognitive Radio-based Emergency Response Network
abstract
Natural and human induced calamities can cause loss of lives and service disruption which requires immediate response setup for first responders and victims. In the aftermath of a disaster first 48 hours are crucial to provide temporary response network for responders to coordinate and conduct search and rescue operation in case of fully or partially destroyed telecommunication infrastructure until more stable help arrives. However, the problem remains the little or no information about the environment in which initial setup needs to operate. A cognitive radio in this regard has a potential to establish initial link in an unknown environment to provide rapid deployment with efficient spectrum usage. A rendezvous protocol can be used to establish initial link, but unpredictable and unknown environment can increase the rendezvous time due to the existence of Primary Radio (PR) activity. The existing protocols have limitations of index based channel selection, inefficient use of spectrum, time slot duration, hopping rate and rendezvous cycle. This paper proposes a dual modular clock-based rendezvous algorithm (DMCA), which takes advantage of prime and non-prime channel numbers to enhance spectrum utilization and incurs less rendezvous time. In the worst-case scenario of asymmetric, which involves only two similar channels (m = 2) with high PR activity, our DMCA protocol achieves 69%, 44%, and 21% less rendezvous time compare to traditional protocols proposed for unknown environments.
Saim Ghafoor, Saritha Unnikrishnan, Eoghan Furey, Ian McLoughlin 0001
CCNC5
2025 Analysis and Reconstruction of Laryngectomised Speech using GAN Models
abstract
Communication through whispering is crucial for individuals who have undergone laryngectomy, a surgical procedure involving the partial or total removal of the voice box or larynx. Whispering, which lacks fundamental frequency, often results in hushed and unintelligible speech for these patients, necessitating prosthetics or treatments. These treatments pose a risk of getting infections, and the resulting speech sounds unnatural.Recently, in a computational framework, deep learning-based methods have been developed for a whisper-to-speech conversion that could be useful for laryngectomy patients as well. In this paper, we analyse GAN-based methods for the reconstruction of laryngectomised speech, using a laryngectomy dataset that we collected in New Zealand. Based on this analysis, we propose modified architectures and loss functions specifically tailored to laryngectomy patients. The proposed models are trained on both the wTIMIT and laryngectomy patient datasets. Evaluations show improved performance in both spectral features and speech intelligibility compared to existing models.
Raj Patel, Hamid R. Sharifzadeh, Maryam Erfanian, Ian McLoughlin 0001, Jacqueline E. Allen
COMPSAC4
2025 PNP-RKD: A Positive-Negative Pair based Relational Knowledge Distillation Method for Cross-Domain Speaker Verification
abstract
Existing deep embedding learning based speaker verification (SV) methods suffer from performance degradation under domain shift conditions. This can be alleviated through unsupervised domain adaptation (UDA) techniques. While UDA improves global statistical consistency across domains, discriminative information may be overlooked or misaligned in the process. To combat this, we propose PNP-RKD, a relational knowledge distillation method that utilizes positive and negative pairs from both the source and target domains within a multitask learning framework. Two auxiliary tasks are conducted separately in the source and target domains to support PNP-RKD. Embeddings are learned in a supervised fashion from the labeled source domain, providing a robust foundation of prior knowledge. For the unlabeled target domain, we apply contrastive learning based on swapped prediction, a key component that enhances noise robustness and improves the quality of learned prototypes. More importantly, it facilitates reliable sampling in PNP-RKD, thereby enhancing the alignment of discriminative knowledge across domains. Extensive experiments conducted on the NIST SRE16 and SRE18 datasets demonstrate the superior performance of the proposed PNP-RKD method, achieving EERs of 6.83% and 8.28%, respectively.
Qing Gu 0002, Yan Song 0001, Nan Jiang 0022, Pengfei Cai, Ian McLoughlin 0001
ICASSP5
2025 Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
abstract
A significant challenge in sound event detection (SED) is the effective utilization of unlabeled data, given the limited availability of labeled data due to high annotation costs. Semi-supervised algorithms rely on labeled data to learn from unlabeled data, and the performance is constrained by the quality and size of the former. In this paper, we introduce the Prototype based Masked Audio Model (PMAM) algorithm for self-supervised representation learning in SED, to better exploit unlabeled data. Specifically, semantically rich frame-level pseudo labels are constructed from a Gaussian mixture model (GMM) based prototypical distribution modeling. These pseudo labels supervise the leaning of a Transformer-based masked audio model, in which binary cross-entropy loss is employed instead of the widely used InfoNCE loss, to provide independent loss contributions from different prototypes, which is important in real scenarios in which multiple labels may apply to unsupervised data frames. A final stage of fine-tuning with just a small amount of labeled data yields a very high performing SED model. On like-for-like tests using the DESED task, our method achieves a PSDS1 score of 62.5%, surpassing current state-of-the-art models and demonstrating the superiority of the proposed technique.
Pengfei Cai, Yan Song 0001, Nan Jiang 0022, Qing Gu 0002, Ian McLoughlin 0001
ICASSP5
2025 An Effective Anomalous Sound Detection Method Based on Global and Local Attribute Mining
Nan Jiang 0022, Yan Song 0001, Qing Gu 0002, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH6
2025 Finetune Large Pre-Trained Model Based on Frequency-Wise Multi-Query Attention Pooling for Anomalous Sound Detection
Nan Jiang 0022, Yan Song 0001, Qing Gu 0002, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH6
2025 A Domain Robust Pre-Training Method with Local Prototypes for Speaker Verification
Qing Gu 0002, Ian McLoughlin 0001
INTERSPEECH6
2025 Automated evaluation of children's speech fluency for low-resource languages
Nur Afiqah Abdul Latiff, Justin Kan, Rong Tong, Donny Soh, Xiaoxiao Miao, Ian McLoughlin 0001
INTERSPEECH7
2025 LSPnet: an ultra-low bitrate hybrid neural codec
Ian McLoughlin 0001, Xiaoxiao Miao, A. S. Madhukumar
INTERSPEECH2
2025 Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
abstract
Most existing sound event detection (SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting audio-language models, their performance is still far from satisfactory due to the lack of fine-grained alignment and cross-modal feature fusion. In this work, we propose the Detect Any Sound Model (DASM), a query-based framework for open-vocabulary SED guided by multi-modal queries. DASM formulates SED as a frame-level retrieval task, where audio features are matched against query vectors derived from text or audio prompts. To support this formulation, DASM introduces a dual-stream decoder that explicitly decouples event recognition and temporal localization: a cross-modality event decoder performs query-feature fusion and determines the presence of sound events at the clip-level, while a context network models temporal dependencies for frame-level localization. Additionally, an inference-time attention masking strategy is proposed to leverage semantic relations between base and novel classes, substantially enhancing generalization to novel classes. Experiments on the AudioSet Strong dataset demonstrate that DASM effectively balances localization accuracy with generalization to novel classes, outperforming CLAP-based methods in open-vocabulary setting (+ 7.8 PSDS) and the baseline in the closed-set setting (+ 6.9 PSDS). Furthermore, in cross-dataset zero-shot evaluation on DESED, DASM achieves a PSDS1 score of 42.2, even exceeding the supervised CRNN baseline. The project page is available at https://cai525.github.io/Transformer4SED/demo_page/DASM/.
Pengfei Cai, Yan Song 0001, Qing Gu 0002, Nan Jiang 0022, Ian McLoughlin 0001
ACM Multimedia6
2025 Adapting general disentanglement-based speaker anonymization for enhanced emotion preservation
abstract
A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for example, emotion, needs to be preserved to a greater extent. Two strategies for this are examined. First, we show that integrating emotion embeddings from a pre-trained emotion encoder can help preserve emotional cues, even though this approach slightly compromises privacy protection. Alternatively, we propose an emotion compensation strategy as a post-processing step applied to anonymized speaker embeddings. This conceals the original speaker’s identity and reintroduces the emotional traits lost during speaker embedding anonymization. Specifically, we model the emotion attribute using support vector machines to learn separate boundaries for each emotion. During inference, the original speaker embedding is processed in two ways: one, by an emotion indicator to predict emotion and select the emotion-matched SVM accurately; and two, by a speaker anonymizer to conceal speaker characteristics. The anonymized speaker embedding is then modified along the corresponding SVM boundary towards an enhanced emotional direction to save the emotional cues. The proposed strategies are also expected to be useful for adapting a general disentanglement-based speaker anonymization system to preserve other target paralinguistic attributes, with potential for a range of downstream tasks 2 .
Xiaoxiao Miao, Xin Wang 0037, Natalia A. Tomashenko, Cheng Lock Donny Soh, Ian McLoughlin 0001
Comput. Speech Lang.6
2024 Meta Representation Learning Method for Robust Speaker Verification in Unseen Domains
abstract
This paper presents a meta representation learning method for robust speaker verification (SV) in unseen domains. It is known that the existing embedding learning based SV systems may suffer from domain mismatch issues. To address this, we propose an episodic training procedure to compensate domain mismatch conditions at runtime. Specifically, episodes are constructed with domain balanced episodic sampling from two different domains, and a new domain alignment (DA) module is added besides the feature extractor (FE) and classifier to existing network structures. In each episodic training iteration, FE and DA modules are optimized separately with different objectives to improve the robustness of learning. Besides, a cross-domain inter-class alignment (CDICA) loss is proposed for improving the domain generalization ability. Experimental results on CNCeleb and VoxCeleb benchmarks demonstrate significant performance gains for unseen domains in SV.
Jian-Tao Zhang, Yan Song 0001, Wu Guo, Hao-Yu Song, Ian McLoughlin 0001
ICASSP6
2024 MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
Pengfei Cai, Yan Song 0001, Ian McLoughlin 0001
INTERSPEECH5
2024 An Effective Local Prototypical Mapping Network for Speech Emotion Recognition
Yuxuan Xi, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH5
2024 DB-PMAE: Dual-Branch Prototypical Masked AutoEncoder with locality for domain robust speaker verification
Wei-Lin Xie, Yu-Xuan Xi, Yan Song 0001, Jian-Tao Zhang, Hao-Yu Song, Ian McLoughlin 0001
INTERSPEECH6
2024 DNN-Based Speech Re-Reverb for VR Rooms
abstract
Reverberation is a physical phenomena caused by room impulse reflection from walls, ceilings and other plane surfaces. When occurring in speech, it contributes an additive echo to the speech which is proportional to the geometry of the environment in which that speech was recorded. Reverberation is not just noise though, as the echo timing in speech affects its intelligibility as well as the listener's interpretation of the playback environment. It gives listeners a sense of the size, shape and materials of the room it is in. For communications scenarios where playback reverberation differs significantly from the visual display (e.g. speech recorded in a small room but played back in a cavernous virtual reality (VR) space), a psychoacoustic dissonance arises between visual and auditory perception. This paper presents an AI-based algorithm to remove source room reverberation impulse effects from speech. We then add reverberation to match the target playback environment. To assess effectiveness, we simulate a set of recording room geometries, and perform psychoacoustic assessment to estimate the effect on a human listener.
Ian McLoughlin 0001, Jeannie Su Ann Lee, Indriyati Atmosukarto, Zhongqiang Ding
TENCON1
2023 An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies
abstract
In this paper, we propose an effective anomalous sound detection (ASD) method based on representation learning with simulated anomalies. Recently, ASD systems have used Outlier Exposure (OE) strategy to achieve promising performance in DCASE challenges. These exploit deep Convolutional Neural Networks (CNNs) to learn discriminative representations by treating the normal samples from different classes as pseudo anomalies. However, since the anomalous sounds occur only rarely, are diversely distributed and are unseen during training, the OE capability of representations learned from normal samples may be limited. To address this issue, we propose a statistics exchange (StEx) method, which constructs simulated anomalies to improve the effectiveness of representation learning via OE strategy. Specifically, the first and second order statistics are extracted from time or frequency axis of input spectrograms, and the simulated anomaly is then obtained by exchanging the statistics of spectrograms from different classes. Furthermore, an out-of-distribution (OOD) metric is introduced as an importance measure to qualitatively analyze OE capability, which enables appropriate simulated anomaly selection for ASD. Extensive experiments on the DCASE2021 challenge task2 development dataset verify the effectiveness of representation learning with simulated anomaly for OE based ASD.
Yan Song 0001, Zhu Zhuo, Yu-Hong Li, Ian McLoughlin 0001
ICASSP7
2023 Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification
abstract
Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to collect in practice, existing domain adaptation methods are hard to ensure the feature consistency of the same class but different domains. To this end, we propose a cross-domain data generation method to obtain a domain-invariant ASV system. Inspired by voice conversion (VC) task, a StarGAN based generative model first learns cross-domain mappings from SSMC data, and then generates missing domain data for all speakers, thus increasing the intra-class diversity of the training set. Considering the difference between ASV and VC task, we renovate the corresponding training objectives and network structure to make the adaptation task-specific. Evaluations on achieve a relative performance improvement of about 5-8% over the baseline in terms of minDCF and EER, outperforming the CNSRC winner’s system of the equivalent scale.
Hang-Rui Hu, Yan Song 0001, Jian-Tao Zhang, Li-Rong Dai 0001, Ian McLoughlin 0001, Zhu Zhuo, Yu-Hong Li
ICASSP5
2023 AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer
abstract
In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where they help mitigate a lack of sufficient real annotated data. However, mainly due to differences between the AT and SED tasks, it is suboptimal to directly utilize outputs from a pretrained AST model. Hence the proposed AST-SED adopts an encoder-decoder architecture to enable effective and efficient fine-tuning without needing to redesign or retrain the AST model. Specifically, the Frequency-wise Transformer Encoder (FTE) consists of transformers with self attention along the frequency axis to address multiple overlapped audio events issue in a single clip. The Local Gated Recurrent Units Decoder (LGD) consists of nearest-neighbor interpolation (NNI) and Bidirectional Gated Recurrent Units (Bi-GRU) to compensate for temporal resolution loss in the pretrained AST model output. Experimental results on DCASE2022 task4 development set have demonstrated the superiority of the proposed AST-SED with FTE-LGD architecture. Specifically, the Event-Based F1-score (EB-F1) of 59.60% and Polyphonic Sound detection Score scenario1 (PSDS1) of 0.5140 significantly outperform CRNN and other pretrained AST-based systems.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP4
2023 Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
abstract
In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together with original normal samples, are used for supervised contrastive representative learning in a multi-task framework. Besides cross-entropy loss between classes, contrastive loss is used to separate PAE output and original samples within each class. GeCo aims to better capture context information among frames, thanks to the self-attention mechanism for PAE model. Furthermore, GeCo combines generative and contrastive learning from which we aim to yield more effective and informative representations, compared to existing methods. Extensive experiments have been conducted on the DCASE2020 Task2 development dataset, showing that GeCo outperforms state-of-the-art generative and discriminative methods.
Xiao-Min Zeng, Yan Song 0001, Zhu Zhuo, Yu-Hong Li, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP8
2023 Fine-tuning Audio Spectrogram Transformer with Task-aware Adapters for Sound Event Detection
abstract
In this paper, we present a task-aware fine-tuning method to transfer Patchout faSt Spectrogram Transformer (PaSST) model to sound event detection (SED) task. Pretrained PaSST has shown significant performance on audio tagging (AT) and SED tasks, but it is not optimal to fine-tune the model from a single layer as the local and semantic information have not been well exploited. To address this, we first introduce task-aware adapters including SED-adapter and AT-adapter to fine-tune PaSST for SED and AT task respectively, and then propose task-aware fine-tuning to combine local information from shallower layer with semantic information from deeper layer, based on task-aware adapters. Besides, we propose the self-distillated mean teacher (SdMT) to train a robust student model with soft pseudo labels from teacher. Experiments are conducted on DCASE2022 task4 development set, the EB-F1 of 64.85% and PSDS1 of 0.5548 are achieved which outperform previous state-of-the-art systems.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH3
2023 Robust Prototype Learning for Anomalous Sound Detection
abstract
In this paper, we present a robust prototype learning framework for anomalous sound detection (ASD), where prototypical loss is exploited to measure the similarity between samples and prototypes. We show that existing generative and discriminative based ASD methods can be unified into this framework from the perspective of prototypical learning. For ASD in recent DCASE challenges, extensions related to imbalanced learning are proposed to improve the robustness of prototypes learned from source and target domains. Specifically, balanced sampling and multiple-prototype expansion (MPE) strategies are proposed to address imbalances across attributes of source and target domains. Furthermore, a novel negative-prototype expansion (NPE) method is used to construct pseudo-anomalies to learn a more compact and effective embedding space for normal sounds. Evaluation on the DCASE2022 Task2 development dataset demonstrates the validity of the proposed prototype learning framework.
Xiao-Min Zeng, Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH3
2022 Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift
abstract
In this paper, a self-supervised representation learning method is proposed for anomalous sound detection (ASD). ASD has received much research attention in recent DCASE challenges. It aims to identify whether a sound emitted from a machine is anomalous or not, given only normal sound data. This is a challenging task due to highly variable time-frequency characteristics of sounds from different machine types, and the fact that many attributes affect machine without being anomalous. This is especially true for domain shift tasks, where only a few training sound clips are available. From the perspective of self-supervised learning, each given sound clip can be considered as a transformation of an original clean sound, where the attribute of each clip may indicate different supervision signals. We propose a unified representation learning framework, equipped with a time-frequency attention mechanism, to perform ASD for different machine types and attributes. For domain shift, a centre imprinting method, which directly sets centres for target domain attributes, is presented. This provides immediate good representation and an initialization for further fine-tuning. Evaluation on DCASE2021 ASD task demonstrates the effectiveness of the proposed method.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP4
2022 Domain Robust Deep Embedding Learning for Speaker Recognition
abstract
This paper presents a domain robust deep embedding learning method for speaker verification (SV) tasks. Most recent methods utilize deep neural networks (DNN) to learn compact and discriminative speaker embeddings from large-scale labeled datasets such as VoxCeleb and the NIST SRE corpus. Despite the success of exiting methods, performance may degrade significantly for new target datasets, mainly due to the distribution discrepancy between training and test domains. Moreover, how corpora are collected, and the languages they contain differ, leading to them spanning multiple, perhaps mismatched, latent domains. To address this, a multi-task end-to-end framework is proposed to learn speaker embeddings from both labeled source and unlabeled target datasets. Motivated by label smoothing, a smoothed knowledge distillation (SKD) based self-supervised learning method is designed to exploit latent structural information from the unlabeled target domain. Furthermore, a domain-aware batch normalization (DABN) module aims to reduce the cross-domain distribution discrepancy, while a domain-agnostic instance normalization (DAIN) module aims to learn features that are robust to within-domain variance. Evaluation on NIST SRE16 demonstrates significant performance gains.
Hang-Rui Hu, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP5
2022 Frontend Attributes Disentanglement for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) with limited size dataset is a challenging task, since a spoken utterance contains various disturbing attributes besides emotion, including speaker, content, and language. However, due to a close relationship between speaker and emotion attributes, simply fine-tuning a linear model is enough to obtain a good SER performance on the utterance-level embeddings (i.e., i-vector and x-vectors) extracted from the pre-trained speaker recognition (SR) frontends. In this paper, we aim to perform frontend attributes disentanglement (AD) for SER task, using a pre-trained SR model. Specifically, the AD module consists of attribute normalization (AN) and attribute reconstruction (AR) phases. The AN filters out the variation information using instance normalization (IN), and AR reconstructs the emotion-relevant features from the residual space to ensure high emotion discrimination. For better disentanglement, a dual space loss is then designed to encourage the separability of emotion-relevant and emotion-irrelevant spaces. To introduce the long-range contextual information for emotion related reconstruction, a time-frequency (TF) attention is further proposed. Different from the style disentanglement of the extracted x-vectors, the proposed AD module can be applied on frontend feature extractor. Experiments on IEMOCAP benchmark demonstrate the effectiveness of the proposed method.
Yuxuan Xi, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP4
2022 Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
abstract
Transformers have recently dominated the ASR field. Although able to yield good performance, they involve an autoregressive (AR) decoder to generate tokens one by one, which is computationally inefficient. To speed up inference, non-autoregressive (NAR) methods, e.g. single-step NAR, were designed, to enable parallel generation. However, due to an independence assumption within the output tokens, performance of single-step NAR is inferior to that of AR models, especially with a large-scale corpus. There are two challenges to improving single-step NAR: Firstly to accurately predict the number of output tokens and extract hidden variables; secondly, to enhance modeling of interdependence between output tokens. To tackle both challenges, we propose a fast and accurate parallel transformer, termed Paraformer. This utilizes a continuous integrate-and-fire based predictor to predict the number of tokens and generate hidden variables. A glancing language model sampler then generates semantic embeddings to enhance the NAR decoder's ability to model context interdependence. Finally, we design a strategy to generate negative samples for minimum word error rate training to further improve performance. Experiments using the AISHELL-1, AISHELL-2 benchmark, and an industrial-level 20,000 hour task demonstrate that the proposed Paraformer can attain comparable performance to the state-of-the-art AR transformer, with over 10x speedup.
Zhifu Gao, Shiliang Zhang, Ian McLoughlin 0001, Zhijie Yan
INTERSPEECH3
2022 Class-Aware Distribution Alignment based Unsupervised Domain Adaptation for Speaker Verification
abstract
Existing speaker verification (SV) systems usually suffer from significant performance degradation when applied to a new domain that lies outside the training distribution. Given the unlabeled target-domain dataset, most Unsupervised Domain Adaptation (UDA) methods aim to minimize the distribution divergence between different domains. However, global distribution alignment strategies fail to consider the latent speaker label information and can hardly guarantee the feature discriminative capability in target domain. In this paper, we propose a novel UDA approach called WBDA (Within-class and Between-class Distribution Alignment), which aims to transfer the class-aware information (i.e., within- and between-class distributions) learned from the well-labeled source-domain to unlabeled target-domain. Motivated by the recent progress of self-supervised contrastive learning, the positive and negative pairs are constructed separately for source and target domains, from which the within- and between-class distribution can be estimated. And the SV system can then be learned by jointly optimizing the cross-domain class-aware distribution discrepancy loss and source-domain classification loss in an end-to-end manner. Evaluations on NIST SRE16 and SRE18 achieve a relative performance improvement of about 43.7% and 26.2% over the baseline in terms of Equal Error Rate (EER) separately, significantly outperforming the previous adaption methods based on global distribution alignment.
Hang-Rui Hu, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
INTERSPEECH4
2022 Hyperspectral Brain Tissue Classification using a Fast and Compact 3D CNN Approach
abstract
Glioblastoma (GB) is a malignant brain tumor and requires surgical resection. Although complete resection of GB improves prognosis, supratotal resection may cause neurological abnormalities. Therefore, intraoperative tissue classification techniques are needed to delineate infected tumor regions to remove reoccurrences. To delineate the affected regions, surgeons mostly rely on traditional magnetic resonance imaging (MRI) which often lacks accuracy and precision due to the brain-shift phenomenon. Hyperspectral Imaging (HSI) is a noninvasive advanced optical technique and has the potential to classify tissue cells accurately. However, HSI tumor classification is challenging due to overlapping regions, high interclass similarity, and homogeneous information. Additionally, HSI models using 2D Convolutional Neural Network (CNN) models works with spectral information eliminating spatial features and 3D followed by 2D hybrid model lacks abstract level spatial information. Therefore, in this study, we have used a minimal layer 3D CNN model to classify the GB tumor region from normal tissues using an intraoperative VivoHSI dataset. The HSI data have normal tissue (NT), tumor tissue (TT), hypervascularized tissue or blood vessels (BV), and background (BG) tissue cells. The proposed 3D CNN model consists of only two 3D layers using limited training samples (20%), which are further divided into 50% for training and 50% for validation and blind tested (80%) on the rest of the data. This study outperformed then state-of-the-art hybrid architecture by achieving an overall accuracy of 99.99%.
Hamail Ayaz, David Tormey, Ian McLoughlin 0001, Muhammad Ahmad 0002, Saritha Unnikrishnan
IPAS3
2022 Investigating the Cognitive Response of Brake Lights in Initiating Braking Action Using EEG
abstract
Half of all road accidents result from either lack of driver attention or from maintaining insufficient separation between vehicles. Collision from the rear, in particular, has been identified as the most common class of accident in the UK, and its influencing factors have been widely studied for many years. Rear-mounted stop lamps, illuminated when braking, are the primary mechanism to alert following drivers to the need to reduce speed or brake. This paper develops a novel brain response approach to measuring subject reaction to different brake light designs. A variety of off-the-shelf brake light assemblies are tested in a physical simulated driving environment to assess the cognitive reaction times of 22 subjects. Eight pairs of LED-based and two pairs of incandescent bulb-based brake light assemblies are used and electroencephalogram (EEG) data recorded. Channel Pz is utilised to extract the P3 component evoked during the decision making process that occurs in the brain when a participant decides to lift their foot from the accelerator and depress the brake. EEG analysis shows that both incandescent bulb-based lights are statistically slower to evoke cognitive responses than all tested LED-based lights. Between the LED designs, differences are evident, but not statistically significant, attributed to the significant amount of movement artifact in the EEG signal.
Ramaswamy Palaniappan, Surej Mouli, Howard Bowman, Ian McLoughlin 0001
IEEE Trans. Intell. Transp. Syst.4
2021 An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification
abstract
In this paper, we present an effective end-to-end deep embedding learning method based on Dense-Residual networks, which combine the advantages of a densely connected convolutional network (DenseNet) and a residual network (ResNet), for speaker verification (SV). Unlike a model ensemble strategy which merges the results of multiple systems, the proposed Dense-Residual networks perform feature fusion on every basic DenseR building block. Specifically, two types of DenseR blocks are designed. A sequential-DenseR block is constructed by densely connecting stacked basic units in a residual block of ResNet. A parallel-DenseR comprises split and concatenation operations on residual and dense components via corresponding skip connections. These building blocks are stacked into deep networks to exploit the complementary information with different receptive field sizes and growth rates. Extensive experiments have been conducted on the VoxCeleb1 dataset to evaluate the proposed methods. The SV performance achieved by the proposed Dense-Residual networks is shown to outperform corresponding ResNet, DenseNet or fusions of them, with similar model complexity, by a significant margin.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
ICASSP3
2021 Self-Attention Generative Adversarial Network for Speech Enhancement
abstract
Existing generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input. To remedy this issue, we propose a self-attention layer adapted from non-local attention, coupled with the convolutional and deconvolutional layers of a speech enhancement GAN (SEGAN) using raw signal input. Further, we empirically study the effect of placing the self-attention layer at the (de)convolutional layers with varying layer indices as well as at all of them when memory allows. Our experiments show that introducing self-attention to SEGAN leads to consistent improvement across the objective evaluation metrics of enhancement performance. Furthermore, applying at different (de)convolutional layers does not significantly alter performance, suggesting that it can be conveniently applied at the highest-level (de)convolutional layer with the smallest memory overhead1.
Huy Phan, Oliver Y. Chén, Philipp Koch, Ngoc Q. K. Duong, Ian McLoughlin 0001, Alfred Mertins
ICASSP6
2021 Multi-View Audio And Music Classification
abstract
We propose in this work a multi-view learning approach for audio and music classification. Considering four typical low-level representations (i.e. different views) commonly used for audio and music recognition tasks, the proposed multi-view network consists of four subnetworks, each handling one input types. The learned embedding in the subnetworks are then concatenated to form the multi-view embedding for classification similar to a simple concatenation network. However, apart from the joint classification branch, the network also maintains four classification branches on the single-view embedding of the subnetworks. A novel method is then proposed to keep track of the learning behavior on the classification branches and adapt their weights to proportionally blend their gradients for network training. The weights are adapted in such a way that learning on a branch that is generalizing well will be encouraged whereas learning on a branch that is overfitting will be slowed down. Experiments on three different audio and music classification tasks show that the proposed multi-view network not only outperforms the single-view baselines but also is superior to the multi-view baselines based on concatenation and late fusion.
Huy Phan, Oliver Y. Chén, Lam Dang Pham, Philipp Koch, Ian McLoughlin 0001, Alfred Mertins
ICASSP6
2021 An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
abstract
This paper presents an improved mean teacher (MT) based method for large-scale weakly labeled semi-supervised sound event detection (SED), by focusing on learning a better student model. Two main improvements are proposed based on the authors’ previous perturbation based MT method. Firstly, an event-aware module is de-signed to allow multiple branches with different kernel sizes to be fused via an attention mechanism. By inserting this module after the convolutional layer, each neuron can adaptively adjust its receptive field to suit different sound events. Secondly, instead of using the teacher model to provide a consistency cost term, we propose using a stochastic inference of unlabeled examples to generate high quality pseudo-targets by averaging multiple predictions from the perturbed student model. MixUp of both labeled and unlabeled data is further exploited to improve the effectiveness of student model. Finally, the teacher model can be obtained via exponential moving average (EMA) of the student model, which generates final predictions for SED during inference. Experiments on the DCASE2018 task4 dataset demonstrate the ability of the proposed method. Specifically, an F1-score of 42.1% is achieved, significantly outperforming the 32.4% achieved by the winning system, or the 39.3% by the previous perturbation based method.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
ICASSP3
2021 Extremely Low Footprint End-to-End ASR System for Smart Device
abstract
Recently, end-to-end (E2E) speech recognition has become popular, since it can integrate the acoustic, pronunciation and language models into a single neural network, which outperforms conventional models. Among E2E approaches, attention-based models, e.g. Transformer, have emerged as being superior. Such models have opened the door to deployment of ASR on smart devices, however they still suffer from requiring a large number of model parameters. We propose an extremely low footprint E2E ASR system for smart devices, to achieve the goal of satisfying resource constraints without sacrificing recognition accuracy. We design cross-layer weight sharing to improve parameter efficiency and further exploit model compression methods including sparsification and quantization, to reduce memory storage and boost decoding efficiency. We evaluate our approaches on the public AISHELL-1 and AISHELL-2 benchmarks. On the AISHELL-2 task, the proposed method achieves more than 10× compression (model size reduces from 248 to 24MB), at the cost of only minor performance loss (CER reduces from 6.49% to 6.92%).
Zhifu Gao, Yiwu Yao, Shiliang Zhang, Ian McLoughlin 0001
Interspeech6
2021 A Weight Moving Average Based Alternate Decoupled Learning Algorithm for Long-Tailed Language Identification
abstract
Language identification (LID) research has made tremendous progress in recent years, especially with the introduction of deep learning techniques. However, for real-world applications where the distribution of different language data is highly imbalanced, the performance of existing LID systems is still far from satisfactory. This raises the challenge of long-tailed LID. In this paper, we propose an effective weight moving average (WMA) based alternate decoupled learning algorithm, termed WADCL, for long-tailed LID. The system is divided into two components, a frontend feature extractor and a backend classifier. These are then alternately learned in an end-to-end manner using different sampling schemes to alleviate the distribution mismatch between training and test datasets. Furthermore, our WMA method aims to mitigate the side-effects of re-sampling schemes, by fusing the model parameters learned along the trajectory of stochastic gradient descent (SGD) optimization. To validate the effectiveness of the proposed WADCL algorithm, we evaluate and compare several systems over a language dataset constructed to match a long-tailed distribution based on real world application [1]. The experimental results from the long-tailed language dataset demonstrate that the proposed algorithm is able to achieve significant performance gains over existing state-of-the-art x-vector based LID methods.
Lin Liu 0017, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
Interspeech5
2021 An Effective Mutual Mean Teaching Based Domain Adaptation Method for Sound Event Detection
abstract
In this paper, we present a novel mutual mean teaching based domain adaptation (MMT-DA) method for sound event detection (SED) task, which can effectively exploit synthetic data to improve the SED performance. Existing methods simply treat the synthetic data as strongly-labeled data in semi-supervised learning (SSL) framework. Benefiting from the strong labels of synthetic data, superior SED performance can be achieved. However, a distribution mismatch between synthetic and real data raises an evident challenge for domain adaptation (DA). In MMT-DA, convolutional recurrent neural networks (CRNN) learned from different datasets (i.e. total data:real+synthetic, and real data) are exploited for DA. Specifically, mean teacher method using CRNN is employed for utilizing the unlabeled real data. To compensate the domain diversity, an additional domain classifier with gradient reverse layer(GRL) is used for training a mean teacher for total data. The student CRNNs are mutually taught using the soft predictions of unlabeled data obtained from different teachers. Furthermore, a strip pooling based attention module is exploited to model the inter-dependencies between channels and time-frequency dimensions to exploit the structure information. Experimental results on Task4 of DCASE2020 demonstrate the ability of the proposed method, achieving 52.0% F1-score on the validation dataset, which outperforms the winning system’s 50.6%.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
Interspeech4
2021 D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition
Xiaoxiao Miao, Ian McLoughlin 0001, Pengyuan Zhang
Neural Networks2
2021 Multi-Granularity Sequence Alignment Mapping for Encoder-Decoder Based End-to-End ASR
abstract
Encoder-decoder based automatic speech recognition (ASR) methods are increasingly popular due to their simplified processing stages and low reliance on prior knowledge. Conventional encoder-decoder based approaches usually learn a sequence-to-sequence mapping function from the source speech to target units (e.g., subwords, characters) in an end-to-end manner. However, it is still unclear how to choose the optimal target unit, or granularity of multiple units. In general, as increasing the information available for learning sequence-to-sequence mapping functions can improve modeling effectiveness, we therefore propose a multi-granularity sequence alignment (MGSA) approach. This aims to enhance cross-sequence interactions between different granularity units for both modeling and inference stages in the encoder-decoder based ASR. Specifically, a decoder module is designed to generate multi-granularity sequence predictions. We then exploit the latent alignment mapping among units having different levels of granularity, by utilizing the decoded multi-level sequences as input for model prediction. The cross-sequence interaction can also be employed to re-calibrate output probabilities in the proposed post-inference algorithm. Experimental results on both WSJ-80 hrs and Switchboard-300 hrs datasets show the superiority of the proposed method compared to traditional multi-task methods as well as to single granularity baseline systems.
Jie Zhang 0042, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 CNN-MoE Based Framework for Classification of Respiratory Anomalies and Lung Disease Detection
abstract
This paper presents and explores a robust deep learning framework for auscultation analysis. This aims to classify anomalies in respiratory cycles and detect diseases, from respiratory sound recordings. The framework begins with front-end feature extraction that transforms input sound into a spectrogram representation. Then, a back-end deep learning network is used to classify the spectrogram features into categories of respiratory anomaly cycles or diseases. Experiments, conducted over the ICBHI benchmark dataset of respiratory sounds, confirm three main contributions towards respiratory-sound analysis. Firstly, we carry out an extensive exploration of the effect of spectrogram types, spectral-time resolution, overlapping/non-overlapping windows, and data augmentation on final prediction accuracy. This leads us to propose a novel deep learning system, built on the proposed framework, which outperforms current state-of-the-art methods. Finally, we apply a Teacher-Student scheme to achieve a trade-off between model performance and model complexity which holds promise for building real-time applications.
Lam Pham, Huy Phan, Ramaswamy Palaniappan, Alfred Mertins, Ian McLoughlin 0001
IEEE J. Biomed. Health Informatics5
2020 An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation
abstract
Despite the significant progress of deep learning based speech separation methods, it remains challenging to extract and track the speech from target speakers, especially in a single-channel multiple speaker situation. Previously, the authors proposed a source-aware context network to exploit the temporal context in mixtures and estimated sources for online speech separation. In this paper, we propose a speaker-aware approach based on the source-aware context network structure, in which the speaker information is explicitly modeled by an auxiliary speaker identification branch. Then speech separation and speaker tracking can be jointly optimized by multi-task learning. Furthermore, we study the effectiveness of time-domain representation by proposing a raw sparse waveform encoder to preserve discriminative information. Experimental results on the WSJ0-2mix benchmark show that the proposed system significantly improves Signal-to-Distortion Ratio (SDR) performance.
Yan Song 0001, Zengxi Li, Ian McLoughlin 0001, Li-Rong Dai 0001
ICASSP4
2020 Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
abstract
Weakly labeled semi-supervised learning methods have recently drawn increasing attention from the research community for sound event detection tasks. Due to the weakness of the labelling, neural networks are often designed to perform sound event detection (SED) and audio tagging (AT) at the same time. In this paper, we propose a task-aware mean teacher method using a convolutional recurrent neural network (CRNN) with multi-branch structure to solve the SED and AT tasks differently. Specifically, a branch with coarse-level temporal resolution is designed for the AT task, while a branch with fine-level temporal resolution is designed for the SED task. The mean teacher based semi-supervised learning method is first adopted to improve the performance of the coarse-level AT branch by exploiting unlabeled data. Then the coarse-level AT branch is introduced as a teacher to guide the aggregated AT output of the fine-level SED branch, yielding an improvement in the SED performance. To further improve the AT and SED performance, information from multiple layers is exploited in the form of a multi-resolution feature. Experimental results on Task4 of the DCASE2018 challenge demonstrate the superiority of the proposed method, achieving 37.7% F1-score, which outperforms the winning system's 32.4%.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP4
2020 Deep Feature Embedding and Hierarchical Classification for Audio Scene Classification
abstract
In this work, we propose an approach that features deep feature embedding learning and hierarchical classification with triplet loss function for Acoustic Scene Classification (ASC). In the one hand, a deep convolutional neural network is firstly trained to learn a feature embedding from scene audio signals. Via the trained convolutional neural network, the learned embedding embeds an input into the embedding feature space and transforms it into a high-level feature vector for representation. In the other hand, in order to exploit the structure of the scene categories, the original scene classification problem is structured into a hierarchy where similar categories are grouped into meta-categories. Then, hierarchical classification is accomplished using deep neural network classifiers associated with triplet loss function. Our experiments show that the proposed system achieves good performance on both the DCASE 2018 Task 1A and 1B datasets, resulting in accuracy gains of 15.6% and 16.6% absolute over the DCASE 2018 baseline on Task 1A and 1B, respectively.
Lam Dang Pham, Ian McLoughlin 0001, Huy Phan, Ramaswamy Palaniappan, Alfred Mertins
IJCNN2
2020 SAN-M: Memory Equipped Self-Attention for End-to-End Speech Recognition
abstract
End-to-end speech recognition has become popular in recent years, since it can integrate the acoustic, pronunciation and language models into a single neural network. Among end-to-end approaches, attention-based methods have emerged as being superior. For example, Transformer, which adopts an encoder-decoder architecture. The key improvement introduced by Transformer is the utilization of self-attention instead of recurrent mechanisms, enabling both encoder and decoder to capture long-range dependencies with lower computational complexity. In this work, we propose boosting the self-attention ability with a DFSMN memory block, forming the proposed memory equipped self-attention (SAN-M) mechanism. Theoretical and empirical comparisons have been made to demonstrate the relevancy and complementarity between self-attention and the DFSMN memory block. Furthermore, the proposed SAN-M provides an efficient mechanism to integrate these two modules. We have evaluated our approach on the public AISHELL-1 benchmark and an industrial-level 20,000-hour Mandarin speech recognition task. On both tasks, SAN-M systems achieved much better performance than the self-attention based Transformer baseline system. Specially, it can achieve a CER of 6.46% on the AISHELL-1 task even without using any external LM, comfortably outperforming other state-of-the-art systems.
Zhifu Gao, Shiliang Zhang, Ian McLoughlin 0001
INTERSPEECH4
2020 An Effective Speaker Recognition Method Based on Joint Identification and Verification Supervisions
abstract
Deep embedding learning based speaker verification methods have attracted significant recent research interest due to their superior performance. Existing methods mainly focus on designing frame-level feature extraction structures, utterance-level aggregation methods and loss functions to learn discriminative speaker embeddings. The scores of verification trials are then computed using cosine distance or Probabilistic Linear Discriminative Analysis (PLDA) classifiers. This paper proposes an effective speaker recognition method which is based on joint identification and verification supervisions, inspired by multi-task learning frameworks. Specifically, a deep architecture with convolutional feature extractor, attentive pooling and two classifier branches is presented. The first, an identification branch, is trained with additive margin softmax loss (AM-Softmax) to classify the speaker identities. The second, a verification branch, trains a discriminator with binary cross entropy loss (BCE) to optimize a new triplet-based mutual information. To balance the two losses during different training stages, a ramp-up/ramp-down weighting scheme is employed. Furthermore, an attentive bilinear pooling method is proposed to improve the effectiveness of embeddings. Extensive experiments have been conducted on VoxCeleb1 to evaluate the proposed method, demonstrating results that relatively reduce the equal error rate (EER) by 22% compared to the baseline system using identification supervision only.
Yan Song 0001, Yiheng Jiang, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH4
2020 Automatic Assessment of Dysarthric Severity Level Using Audio-Video Cross-Modal Approach in Deep Learning
abstract
Dysarthria is a speech disorder disease that can have a significant impact on a person's daily life. Early detection of the disease can put the patient into therapy sessions more quickly. Researchers have established various approaches to detect the disease automatically. Traditional computational approaches commonly analysed acoustic features like Mel-Frequency Cepstral Coefficients (MFCC), Spectral Centroid, Linear Prediction Cepstral (LPC) coefficients and Perceptual Linear Prediction (PLP) from speech samples of patients to detect dysarthric speech characters like slow speech rate, short pauses, mis-articulated sounds, etc. Recent research has shown that some machine learning algorithms can also be deployed to extract speech features and detect the severity level automatically. \n \n In machine learning, feature extraction is a crucial step in dealing with classification and prediction problems. For different data formats, different well-established frameworks have been developed to extract and classify the corresponding features. For example, for an image data processing system, Convolution Neural Network (CNN) can provide the underlying network structure for the system to analyse the video data to obtain the visual features. In contrast, for audio data processing system, Natural Language Processing (NLP) algorithms can be applied to obtain acoustic features. Therefore, the selection of the framework to be used mainly depends on the modality of the input. As early steps in development of machine learning approaches for automatic assessment of dysarthric patients, classification systems based on audio features have been considered in literature; however, recent research efforts in other fields have shown that using an audio-video cross-modal framework can improve performance of the classification systems. \n \nIn this thesis, for the first time, an audio-video cross-modal framework is proposed using deep-learning algorithm that the network takes both audio and video data as input to detect severity levels of dysarthria. Within the deep-learning framework, we also propose two network architectures using audio-only or video-only input to detect dysarthria severity levels automatically. Comparing with current one-modality systems, the deep-learning framework yields satisfying results. More importantly, comparing with systems based only on audio data for automatic dysarthria severity level assessment, the audio-video deep-learning cross modal system proposed in this research can accelerate the training speed, improve accuracy and reduce the amount of required training data.
Han Tong, Hamid R. Sharifzadeh, Ian McLoughlin 0001
INTERSPEECH3
2020 Semi-Supervised End-to-End ASR via Teacher-Student Learning with Conditional Posterior Distribution
abstract
Encoder-decoder based methods have become popular for automatic speech recognition (ASR), thanks to their simplified processing stages and low reliance on prior knowledge. However, large amounts of acoustic data with paired transcriptions is generally required to train an effective encoder-decoder model, which is expensive, time-consuming to be collected and not always readily available. However unpaired speech data is abundant, hence several semi-supervised learning methods, such as teacher-student (T/S) learning and pseudo-labeling, have recently been proposed to utilize this potentially valuable resource. In this paper, a novel T/S learning with conditional posterior distribution for encoder-decoder based ASR is proposed. Specifically, the 1-best hypotheses and the conditional posterior distribution from the teacher are exploited to provide more effective supervision. Combined with model perturbation techniques, the proposed method reduces WER by 19.2% relatively on the LibriSpeech benchmark, compared with a system trained using only paired data. This outperforms previous reported 1-best hypothesis results on the same task.
Yan Song 0001, Jianshu Zhang 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH4
2020 An Effective Perturbation Based Semi-Supervised Learning Method for Sound Event Detection
abstract
Mean teacher based methods are increasingly achieving state-of-the-art performance for large-scale weakly labeled and unlabeled sound event detection (SED) tasks in recent DCASE challenges. By penalizing inconsistent predictions under different perturbations, mean teacher methods can exploit large-scale unlabeled data in a self-ensembling manner. In this paper, an effective perturbation based semi-supervised learning (SSL) method is proposed based on the mean teacher method. Specifically, a new independent component (IC) module is proposed to introduce perturbations for different convolutional layers, designed as a combination of batch normalization and dropblock operations. The proposed IC module can reduce correlation between neurons to improve performance. A global statistics pooling based attention module is further proposed to explicitly model inter-dependencies between the time-frequency domain and channels, using statistics information (e.g. mean, standard deviation, max) along different dimensions. This can provide an effective attention mechanism to adaptively re-calibrate the output feature map. Experimental results on Task 4 of the DCASE2018 challenge demonstrate the superiority of the proposed method, achieving about 39.8% F1-score, outperforming the previous winning system’s 32.4% by a significant margin.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
INTERSPEECH5
2020 Improving GANs for Speech Enhancement
abstract
Generative adversarial networks (GAN) have recently been shown to be efficient for speech enhancement. However, most, if not all, existing speech enhancement GANs (SEGAN) make use of a single generator to perform one-stage enhancement mapping. In this work, we propose to use multiple generators that are chained to perform multi-stage enhancement mapping, which gradually refines the noisy input signals in a stage-wise fashion. Furthermore, we study two scenarios: (1) the generators share their parameters and (2) the generators' parameters are independent. The former constrains the generators to learn a common mapping that is iteratively applied at all enhancement stages and results in a small model footprint. On the contrary, the latter allows the generators to flexibly learn different enhancement mappings at different stages of the network at the cost of an increased model size. We demonstrate that the proposed multi-stage enhancement approach outperforms the one-stage SEGAN baseline, where the independent generators lead to more favorable results than the tied generators. The source code is available at http://github.com/pquochuy/idsegan.
Huy Phan, Ian McLoughlin 0001, Lam Dang Pham, Oliver Y. Chén, Philipp Koch, Maarten De Vos, Alfred Mertins
IEEE Signal Process. Lett.2
2020 Glottal Flow Synthesis for Whisper-to-Speech Conversion
abstract
Whisper-to-speech conversion is motivated by laryngeal disorders, in which malfunction of the vocal folds leads to loss of voicing. Many patients with laryngeal disorders can still produce functional whispers, since these are characterised by the absence of vocal fold vibration. Whispers therefore constitute a common ground for speech rehabilitation across many kinds of laryngeal disorder. Whisper-to-speech conversion involves recreating natural-sounding speech from recorded whispers, and is a non-invasive and non-surgical rehabilitation that can maintain a natural method of speaking, unlike the existing methods of rehabilitation. This article proposes a new rule-based method for whisper-to-speech conversion that replaces the noisy whisper sound source with a synthesised speech-like harmonic source, while maintaining the vocal tract component unaltered. In particular, a novel glottal source generator is developed in which whisper information is used to parameterise the excitation through a high-quality glottis model. Evaluation of the system against the standard pulse train excitation method reveals significantly improved performance. Since our method is glottis-based, it is potentially compatible with the many existing vocal tract component adaptation systems.
Olivier Perrotin, Ian McLoughlin 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 A Spectral Glottal Flow Model for Source-filter Separation of Speech
abstract
The estimation of glottal flow from a speech waveform is an essential technique used in speech analysis and parameterisation. Significant research effort has been addressed at separating the first vocal tract resonance from the glottal formant (the low-frequency resonance that describes the open-phase of the vocal fold vibration), but few methods are capable of estimating the high-frequency spectral tilt, characteristic of the closing phase of the vocal fold vibration (which is crucial to the perception of vocal effort). This paper proposes an improved Iterative Adaptive Inverse Filtering (IAIF) method based on a Glottal Flow Model, which we call GFM-IAIF. The proposed method models the wide-band glottis response, incorporating both glottal formant and spectral tilt characteristics. Evaluation against IAIF and recently proposed IOP-IAIF shows that, while GFM-IAIF maintains good performance on vocal tract modelling, it significantly improves the glottis model. This ensures that timbral variations associated to voice quality can be correctly attributed and described.
Olivier Perrotin, Ian McLoughlin 0001
ICASSP2
2019 Unifying Isolated and Overlapping Audio Event Detection with Multi-label Multi-task Convolutional Recurrent Neural Networks
abstract
We propose a multi-label multi-task framework based on a convolutional recurrent neural network to unify detection of isolated and overlapping audio events. The framework leverages the power of convolutional recurrent neural network architectures; convolutional layers learn effective features over which higher recurrent layers perform sequential modelling. Furthermore, the output layer is designed to handle arbitrary degrees of event overlap. At each time step in the recurrent output sequence, an output triple is dedicated to each event category of interest to jointly model event occurrence and temporal boundaries. That is, the network jointly determines whether an event of this category occurs, and when it occurs, by estimating onset and offset positions at each recurrent time step. We then introduce three sequential losses for network training: multi-label classification loss, distance estimation loss, and confidence loss. We demonstrate good generalization on two datasets: ITC-Irst for isolated audio event detection, and TUT-SED-Synthetic-2016 for overlapping audio event detection.
Huy Phan, Oliver Y. Chén, Philipp Koch, Lam Dang Pham, Ian McLoughlin 0001, Alfred Mertins, Maarten De Vos
ICASSP5
2019 A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification
abstract
Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challenges. The introduction of GLU and temporal attention-based localization mechanisms plays an important role for both AT and SED tasks. In this paper, we propose a novel region based attention method to further boost the representation power of the existing GLU based CRNN. Specifically, we insert a feature selection (FS) structure after each GLU to create what we term a GLU-F. block, to exploit channel relationships. Furthermore, we extract region features (or the prototypes of certain sound events) from multi-scale sliding windows over higher convolutional layers, which are fed into an attention-based recurrent neural network to model their context information for AT and SED tasks. To evaluate the proposed region based attention method, we conduct extensive experiments on SED and AT tasks in DCASE2017. We achieve 59.5% and 60.1% AT F1-score, 51.3% and 55.1% SED F1-score for development and evaluation sets respectively, significantly outperforming state-of-the-art results.
Yan Song 0001, Wu Guo, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP5
2019 Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Yiheng Jiang, Li-Rong Dai 0001
INTERSPEECH3
2019 An Effective Deep Embedding Learning Architecture for Speaker Verification
Yiheng Jiang, Yan Song 0001, Ian McLoughlin 0001, Zhifu Gao, Li-Rong Dai 0001
INTERSPEECH3
2019 A New Time-Frequency Attention Mechanism for TDNN and CNN-LSTM-TDNN, with Application to Language Identification
Xiaoxiao Miao, Ian McLoughlin 0001, Yonghong Yan 0002
INTERSPEECH2
2019 GFM-Voc: A Real-Time Voice Quality Modification System
Olivier Perrotin, Ian McLoughlin 0001
INTERSPEECH2
2019 A Robust Framework for Acoustic Scene Classification
abstract
Acoustic scene classification (ASC) using front-end timefrequency features and back-end neural network classifiers has demonstrated good performance in recent years. However a profusion of systems has arisen to suit different tasks and datasets, utilising different feature and classifier types. This paper aims at a robust framework that can explore and utilise a range of different time-frequency features and neural networks, either singly or merged, to achieve good classification performance. In particular, we exploit three different types of frontend time-frequency feature; log energy Mel filter, Gammatone filter and constant Q transform. At the back-end we evaluate effective a two-stage model that exploits a Convolutional Neural Network for pre-trained feature extraction, followed by Deep Neural Network classifiers as a post-trained feature adaptation model and classifier. We also explore the use of a data augmentation technique for these features that effectively generates a variety of intermediate data, reinforcing model learning abilities, particularly for marginal cases. We assess performance on the DCASE2016 dataset, demonstrating good classification accuracies exceeding 90%, significantly outperforming the DCASE2016 baseline and highly competitive compared to state-of-the-art systems.
Lam Dang Pham, Ian McLoughlin 0001, Huy Phan, Ramaswamy Palaniappan
INTERSPEECH2
2019 Spatio-Temporal Attention Pooling for Audio Scene Classification
abstract
Acoustic scenes are rich and redundant in their content. In this work, we present a spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network to learn from patterns that are discriminative while suppressing those that are irrelevant for acoustic scene classification. The convolutional layers in this network learn invariant features from time-frequency input. The bidirectional recurrent layers are then able to encode the temporal dynamics of the resulting convolutional features. Afterwards, a two-dimensional attention mask is formed via the outer product of the spatial and temporal attention vectors learned from two designated attention layers to weigh and pool the recurrent output into a final feature vector for classification. The network is trained with between-class examples generated from between-class data augmentation. Experiments demonstrate that the proposed method not only outperforms a strong convolutional neural network baseline but also sets new state-of-the-art performance on the LITIS Rouen dataset.
Huy Phan, Oliver Y. Chén, Lam Dang Pham, Philipp Koch, Maarten De Vos, Ian McLoughlin 0001, Alfred Mertins
INTERSPEECH6
2019 Listening and Grouping: An Online Autoregressive Approach for Monaural Speech Separation
abstract
This paper proposes an autoregressive approach to harness the power of deep learning for multi-speaker monaural speech separation. It exploits a causal temporal context in both mixture and past estimated separated signals and performs online separation that is compatible with real-time applications. The approach adopts a learned listening and grouping architecture motivated by computational auditory scene analysis, with a grouping stage that effectively addresses the label permutation problem at both frame and segment levels. Experimental results on the WSJ0-2mix benchmark show that the new approach can achieve better signal-to-distortion ratio and perceptual evaluation of speech quality scores than most of the state-of-the-art methods for both closed-set and open-set evaluations, even methods that exploit whole-utterance statistics for separation. It achieves this while requiring fewer model parameters.
Zengxi Li, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Source-Aware Context Network for Single-Channel Multi-Speaker Speech Separation
abstract
Deep learning based approaches have achieved promising performance in speaker-dependent single-channel multispeaker speech separation. However, partly due to the label permutation problem, they may encounter difficulties in speaker-independent conditions. Recent methods address this problem by some assignment operations. Different from them, we propose a novel source-aware context network, which explicitly inputs speech sources as well as mixture signal. By exploiting the temporal dependency and continuity of the same source signal, the permutation order of outputs can be easily determined without any additional post-processing. Furthermore, a Multi-time-step Prediction Training strategy is proposed to address the mismatch between training and inference stages. Experimental results on benchmark WSJ0-2mix dataset revealed that our network achieved comparable or better results than state-of-the-art methods in both closed-set and open-set conditions, in terms of Signal-to-Distortion Ratio (SDR) improvement.
Zengxi Li, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP4
2018 Enabling Early Audio Event Detection with Neural Networks
abstract
This paper presents a methodology for early detection of audio events from audio streams. Early detection is the ability to infer an ongoing event during its initial stage. The proposed system consists of a novel inference step coupled with dual parallel tailored-loss deep neural networks (DNNs). The DNNs share a similar architecture except for their loss functions, i.e. weighted loss and multitask loss, which are designed to efficiently cope with issues common to audio event detection. The inference step is newly introduced to make use of the network outputs for recognizing ongoing events. The monotonicity of the detection function is required for reliable early detection, and will also be proved. Experiments on the ITC-Irst database show that the proposed system achieves state-of-the-art detection performance. Furthermore, even partial events are sufficient to achieve good performance similar to that obtained when an entire event is observed, enabling early event detection.
Huy Phan, Philipp Koch, Ian McLoughlin 0001, Alfred Mertins
ICASSP3
2018 An Improved Deep Embedding Learning Method for Short Duration Speaker Verification
abstract
This paper presents an improved deep embedding learning method based on convolutional neural networks (CNN) for short-duration speaker verification (SV). Existing deep learning-based SV methods generally extract frontend embeddings from a feed-forward deep neural network, in which the long-term speaker characteristics are captured via a pooling operation over the input speech. The extracted embeddings are then scored via a backend model, such as Probabilistic Linear Discriminative Analysis (PLDA). Two improvements are proposed for frontend embedding learning based on the CNN structure: (1) Motivated by the WaveNet for speech synthesis, dilated filters are designed to achieve a tradeoff between computational efficiency and receptive-filter size; and (2) A novel cross-convolutional-layer pooling method is exploited to capture $1^{st}$-order statistics for modelling long-term speaker characteristics. Specifically, the activations of one convolutional layer are aggregated with the guidance of the feature maps from the successive layer. To evaluate the effectiveness of our proposed methods, extensive experiments are conducted on the modified female portion of NIST SRE 2010 evaluations, with conditions ranging from 10s-10s to 5s-4s. Excellent performance has been achieved on each evaluation condition, significantly outperforming existing SV systems using i-vector and d-vector embeddings.
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH3
2018 An Attention Pooling Based Representation Learning Method for Speech Emotion Recognition
abstract
This paper proposes an attention pooling based representation learning method for speech emotion recognition (SER). The emotional representation is learned in an end-to-end fashion by applying a deep convolutional neural network (CNN) directly to spectrograms extracted from speech utterances. Motivated by the success of GoogleNet, two groups of filters with different shapes are designed to capture both temporal and frequency domain context information from the input spectrogram. The learned features are concatenated and fed into the subsequent convolutional layers. To learn the final emotional representation, a novel attention pooling method is further proposed. Compared with the existing pooling methods, such as max-pooling and average-pooling, the proposed attention pooling can effectively incorporate class-agnostic bottom-up, and class-specific top-down, attention maps. We conduct extensive evaluations on benchmark IEMOCAP data to assess the effectiveness of the proposed representation. Results demonstrate a recognition performance of 71.8% weighted accuracy (WA) and 68% unweighted accuracy (UA) over four emotions, which outperforms the state-of-the-art method by about 3% absolute for WA and 4% for UA.
Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH3
2018 Early Detection of Continuous and Partial Audio Events Using CNN
abstract
Sound event detection is an extension of the static auditory classification task into continuous environments, where performance depends jointly upon the detection of overlapping events and their correct classification. Several approaches have been published to date which either develop novel classifiers or employ well-trained static classifiers with a detection front-end. This paper takes the latter approach, by combining a proven CNN classifier acting on spectrogram image features, with time-frequency shaped energy detection that identifies seed regions within the spectrogram that are characteristic of auditory energy events. Furthermore, the shape detector is optimised to allow early detection of events as they are developing. Since some sound events naturally have longer durations than others, waiting until completion of entire events before classification may not be practical in a deployed system. The early detection capability of the system is thus evaluated for the classification of partial events. Performance for continuous event detection is shown to be good, with accuracy being maintained well when detecting partial events. © 2018 International Speech Communication Association. All rights reserved.
Ian McLoughlin 0001, Yan Song 0001, Lam Dang Pham, Ramaswamy Palaniappan, Huy Phan, Yue Lang
INTERSPEECH1
2018 Acoustic Modeling with Densely Connected Residual Network for Multichannel Speech Recognition
abstract
Motivated by recent advances in computer vision research, this paper proposes a novel acoustic model called Densely Connected Residual Network (DenseRNet) for multichannel speech recognition. This combines the strength of both DenseNet and ResNet. It adopts the basic "building blocks" of ResNet with different convolutional layers, receptive field sizes and growth rates as basic components that are densely connected to form so-called denseR blocks. By concatenating the feature maps of all preceding layers as inputs, DenseRNet can not only strengthen gradient back-propagation for the vanishing-gradient problem, but also exploit multi-resolution feature maps. Preliminary experimental results on CHiME-3 have shown that DenseRNet achieves a word error rate (WER) of 7.58% on beamforming-enhanced speech with six channel real test data by cross entropy criteria training while WER is 10.23% for the official baseline. Besides, additional experimental results are also presented to demonstrate that DenseRNet exhibits the robustness to beamforming-enhanced speech as well as near and far-field speech.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH4
2018 Improved Conditional Generative Adversarial Net Classification For Spoken Language Recognition
abstract
Recent research on generative adversarial nets (GAN) for language identification (LID) has shown promising results. In this paper, we further exploit the latent abilities of GAN networks to firstly combine them with deep neural network (DNN)-based i-vector approaches and then to improve the LID model using conditional generative adversarial net (cGAN) classification. First, phoneme dependent deep bottleneck features (DBF) combined with output posteriors of a pre-trained DNN for automatic speech recognition (ASR) are used to extract i-vectors in the normal way. These i-vectors are then classified using cGAN, and we show an effective method within the cGAN to optimize parameters by combining both language identification and verification signals as supervision. Results show firstly that cGAN methods can significantly outperform DBF DNN i-vector methods where 49-dimensional i-vectors are used, but not where 600-dimensional vectors are used. Secondly, training a cGAN discriminator network for direct classification has further benefit for low dimensional i-vectors as well as short utterances with high dimensional i-vectors. However, incorporating a dedicated discriminator network output layer for classification and optimizing both classification and verification loss brings benefits in all test cases.
Xiaoxiao Miao, Ian McLoughlin 0001, Shengyu Yao, Yonghong Yan 0002
SLT2
2018 LID-Senones and Their Statistics for Language Identification
abstract
Recent research on end-to-end training structures for language identification has raised the possibility that intermediate language-sensitive feature units exist which are analogous to phonetically sensitive senones in automatic speech recognition systems. Termed language identification (LID)-senones, the statistics derived from these feature units have been shown to be beneficial in discriminating between languages, particularly for short utterances. This paper examines the evidence for the existence of LID-senones before designing and evaluating LID systems based on low- and high-level statistics of LID-senones with both generative and discriminative models. For the standard NIST LRE 2009 task on 23 languages, LID-senone-based systems are shown to outperform state-of-the-art deep neural network/i-vector methods both when LID-senones are used directly for classification and when LID-senone statistics are used for i-vector formation.
Ma Jin, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Fisher vector based CNN architecture for image classification
abstract
In this paper, we tackle the representation learning problem for small scale fine-grained object recognition and scene classification tasks. Conventional bag of features(BoF) methods exploit hand-crafted frontend local features, and learn the representations via various machine learning techniques. Convolutional neural networks(CNN) directly learn the representation from raw images and benefit from joint optimization of network parameters in an end-to-end manner. However, the performance of existing representation learning methods is still unsatisfactory for the small-scale recognition tasks. To address this issue, we present a FV coding based CNN(FV-CNN) architecture. FV-CNN has three main advantages in that firstly it is able to exploit activations from the intermediate convolutional layer and a probabilistic discriminative model to derive the FV coding. Secondly, it takes advantage of the end-to-end back-propagation of the gradients to jointly optimize the whole learning process. Finally, it can learn a compact representation. When evaluated on benchmark datasets of fine grain object recognition (Caltech-CUB200), and scene classification (MIT67), accuracies of 88.0% and 82.2% are achieved.
Yan Song 0001, Peiseng Wang, Xinhai Hong, Ian McLoughlin 0001
ICIP4
2017 End-to-End Language Identification Using High-Order Utterance Representation with Bilinear Pooling
abstract
A key problem in spoken language identification (LID) is how to design effective representations which are specific to language information. Recent advances in deep neural networks have led to significant improvements in results, with deep end-to-end methods proving effective. This paper proposes a novel network which aims to model an effective representation for high (first and second)-order statistics of LID-senones, defined as being LID analogues of senones in speech recognition. The high-order information extracted through bilinear pooling is robust to speakers, channels and background noise. Evaluation with NIST LRE 2009 shows improved performance compared to current state-of-the-art DBF/i-vector systems, achieving over 33% and 20% relative equal error rate (EER) improvement for 3s and 10s utterances and over 40% relative Cavg improvement for all durations.
Ma Jin, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH3
2016 Compact convolutional neural network transfer learning for small-scale image classification
abstract
Transfer learning methods have demonstrated state-of-the-art performance on various small-scale image classification tasks. This is generally achieved by exploiting the information from an ImageNet convolution neural network (ImageNet CNN). However, the transferred CNN model is generally with high computational complexity and storage requirement. It raises the issue for real-world applications, especially for some portable devices like phones and tablets without high-performance GPUs. Several approximation methods have been proposed to reduce the complexity by reconstructing the linear or non-linear filters (responses) in convolutional layers with a series of small ones., In this paper, we present a compact CNN transfer learning method for small-scale image classification. Specifically, it can be decomposed into fine-tuning and joint learning stages. In fine-tuning stage, a high-performance target CNN is trained by transferring information from the ImageNet CNN. In joint learning stage, a compact target CNN is optimized based on ground-truth labels, jointly with the predictions of the high-performance target CNN. The experimental results on CIFAR-10 and MIT Indoor Scene demonstrate the effectiveness and efficiency of our proposed method.
Zengxi Li, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
ICASSP3
2016 Learning compact structural representations for audio events using regressor banks
abstract
We introduce a new learned descriptor for audio signals which is efficient for event representation. The entries of the descriptor are produced by evaluating a set of regressors on the input signal. The regressors are class-specific and trained using the random regression forests framework. Given an input signal, each regressor estimates the onset and offset positions of the target event. The estimation confidence scores output by a regressor are then used to quantify how the target event aligns with the temporal structure of the corresponding category. Our proposed descriptor has two advantages. First, it is compact, i.e. the dimensionality of the descriptor is equal to the number of event classes. Second, we show that even simple linear classification models, trained on our descriptor, yield better accuracies on audio event classification task than not only the nonlinear baselines but also the state-of-the-art results.
Huy Phan, Marco Maaß, Lars Hertel, Radoslaw Mazur, Ian McLoughlin 0001, Alfred Mertins
ICASSP5
2016 Robust Sound Event Detection in Continuous Audio Environments
abstract
Sound event detection in real world environments has attracted significant research interest recently because of it's applications in popular fields such as machine hearing and automated surveillance, as well as in sound scene understanding. This paper considers continuous robust sound event detection, which means multiple overlapped sound events in different types of interfering noise. First, a standard evaluation task is outlined based upon existing testing data sets for the sound event classification of isolated sounds. This paper then proposes and evaluates the use of spectrogram image features employing an energy detector to segment sound events, before developing a novel segmentation method making use of a Bayesian inference criteria. At the back end, a convolutional neural network is used to classify detected regions, and this combination is compared to several alternative approaches. The proposed method is shown capable of achieving very good performance compared with current state-of-the-art techniques.
Haomin Zhang, Ian McLoughlin 0001, Yan Song 0001
INTERSPEECH2
2016 Image classification with CNN-based Fisher vector coding
abstract
Fisher vector coding methods have been demonstrated to be effective for image classification. With the help of convolutional neural networks (CNN), several Fisher vector coding methods have shown state-of-the-art performance by adopting the activations of a single fully-connected layer as region features. These methods generally exploit a diagonal Gaussian mixture model (GMM) to describe the generative process of region features. However, it is difficult to model the complex distribution of high-dimensional feature space with a limited number of Gaussians obtained by unsupervised learning. Simply increasing the number of Gaussians turns out to be inefficient and computationally impractical. To address this issue, we re-interpret a pre-trained CNN as the probabilistic discriminative model, and present a CNN based Fisher vector coding method, termed CNN-FVC. Specifically, activations of the intermediate fully-connected and output soft-max layers are exploited to derive the posteriors, mean and covariance parameters for Fisher vector coding implicitly. To further improve the efficiency, we convert the pre-trained CNN to a fully convolutional one to extract the region features. Extensive experiments have been conducted on two standard scene benchmarks (i.e. SUN397 and MIT67) to evaluate the effectiveness of the proposed method. Classification accuracies of 60.7% and 82.1% are achieved on the SUN397 and MIT67 benchmarks respectively, outperforming previous state-of-the-art approaches. Furthermore, the method is complementary to GMM-FVC methods, allowing a simple fusion scheme to further improve performance to 61.1% and 83.1% respectively.
Yan Song 0001, Xinhai Hong, Ian McLoughlin 0001, Li-Rong Dai 0001
VCIP3
2016 Efficient Integer Frequency Offset Estimation Architecture for Enhanced OFDM Synchronization
abstract
In orthogonal frequency-division multiplexing (OFDM) systems, integer frequency offset (IFO) causes a circular shift of the subcarrier indices in the frequency domain. The IFO can be mitigated through strict RF front-end design, which tends to be expensive, or by strictly limiting mobility and channel agility, which constrains operating scenarios. The IFO is, therefore, often estimated and removed at baseband, allowing implementations to benefit from the relaxed RF front-end specifications and to be tolerant to both Doppler shift and multistandard channel selection. This paper proposes a novel architecture for the IFO estimation which achieves reduced power consumption and lower computational cost than contemporary methods, while achieving excellent estimation performance, close to theoretically achievable bounds. A pilot subsampling technique enables fourfold resource sharing to reduce the computational cost, while multiplierless computation yields further power reduction. Performance exceeds that of the conventional techniques, while being much more efficient. When implemented on field-programmable gate array for IEEE 802.16-2009, the dynamic power reductions of 78% are achieved. The architecture and method is applicable to other OFDM standards including IEEE 802.11 and IEEE 802.22.
Thinh Hung Pham, Suhaib A. Fahmy, Ian McLoughlin 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Multi-task deep neural network acoustic models with model adaptation using discriminative speaker identity for whisper recognition
abstract
This paper presents a study on large vocabulary continuous whisper automatic recognition (wLVCSR). wLVCSR provides the ability to use ASR equipment in public places without concern for disturbing others or leaking private information. However the task of wLVCSR is much more challenging than normal LVCSR due to the absence of pitch which not only causes the signal to noise ratio (SNR) of whispers to be much lower than normal speech but also leads to flatness and formant shifts in whisper spectra. Furthermore, the amount of whisper data available for training is much less than for normal speech. In this paper, multi-task deep neural network (DNN) acoustic models are deployed to solve these problems. Moreover, model adaptation is performed on the multi-task DNN to normalize speaker and environmental variability in whispers based on discriminative speaker identity information. On a Mandarin whisper dictation task, with 55 hours of whisper data, the proposed SI multi-task DNN model can achieve 56.7% character error rate (CER) improvement over a baseline Gaussian Mixture Model (GMM), discriminatively trained only using the whisper data. Besides, the CER of the proposed model for normal speech can reach 15.2%, which is close to the performance of a state-of-the-art DNN trained with one thousand hours of speech data. From this baseline, the model-adapted DNN gains a further 10.9% CER reduction over the generic model.
Ian McLoughlin 0001, Cong Liu 0006, Shaofei Xue, Si Wei
ICASSP2
2015 Improved language identification using deep bottleneck network
abstract
Effective representation plays an important role in automatic spoken language identification (LID). Recently, several representations that employ a pre-trained deep neural network (DNN) as the front-end feature extractor, have achieved state-of-the-art performance. However the performance is still far from satisfactory for dialect and short-duration utterance identification tasks, due to the deficiency of existing representations. To address this issue, this paper proposes the improved representations to exploit the information extracted from different layers of the DNN structure. This is conceptually motivated by regarding the DNN as a bridge between low-level acoustic input and high-level phonetic output features. Specifically, we employ deep bottleneck network (DBN), a DNN with an internal bottleneck layer acting as a feature extractor. We extract representations from two layers of this single network, i.e. DBN-TopLayer and DBN-MidLayer. Evaluations on the NIST LRE2009 dataset, as well as the more specific dialect recognition task, show that each representation can achieve an incremental performance gain. Furthermore, a simple fusion of the representations is shown to exceed current state-of-the-art performance.
Yan Song 0001, Ruilian Cui, Xinhai Hong, Ian McLoughlin 0001, Jiong Shi, Li-Rong Dai 0001
ICASSP4
2015 Robust sound event recognition using convolutional neural networks
abstract
Traditional sound event recognition methods based on informative front end features such as MFCC, with back end sequencing methods such as HMM, tend to perform poorly in the presence of interfering acoustic noise. Since noise corruption may be unavoidable in practical situations, it is important to develop more robust features and classifiers. Recent advances in this field use powerful machine learning techniques with high dimensional input features such as spectrograms or auditory image. These improve robustness largely thanks to the discriminative capabilities of the back end classifiers. We extend this further by proposing novel features derived from spectrogram energy triggering, allied with the powerful classification capabilities of a convolutional neural network (CNN). The proposed method demonstrates excellent performance under noise-corrupted conditions when compared against state-of-the-art approaches on standard evaluation tasks. To the author's knowledge this in the first application of CNN in this field.
Haomin Zhang, Ian McLoughlin 0001, Yan Song 0001
ICASSP2
2015 Low frequency ultrasonic voice activity detection using convolutional neural networks
abstract
Low frequency ultrasonic mouth state detection uses reflected audio chirps from the face in the region of the mouth to determine lip state, whether open, closed or partially open. The chirps are located in a frequency range just above the threshold of human hearing and are thus both inaudible as well as unaffected by interfering speech, yet can be produced and sensed using inexpensive equipment. To determine mouth open or closed state, and hence form a measure of voice activity detection, this recently invented technique relies upon the difference in the reflected chirp caused by resonances introduced by the open or partially open mouth cavity. Voice activity is then inferred from lip state through patterns of mouth movement, in a similar way to video-based lip-reading technologies. This paper introduces a new metric based on spectrogram features extracted from the reflected chirp, with a convolutional neural network classification back-end, that yields excellent performance without needing the periodic resetting of the template closed-mouth reflection required by the original technique.
Ian McLoughlin 0001, Yan Song 0001
INTERSPEECH1
2015 Deep bottleneck network based i-vector representation for language identification
abstract
This paper presents a unified i-vector framework for language identification (LID) based on deep bottleneck networks (DBN) trained for automatic speech recognition (ASR). The framework covers both front-end feature extraction and back-end modeling stages.The output from different layers of a DBN are exploited to improve the effectiveness of the i-vector representation through incorporating a mixture of acoustic and phonetic information. Furthermore, a universal model is derived from the DBN with a LID corpus. This is a somewhat inverse process to the GMM-UBM method, in which the GMM of each language is mapped from a GMM-UBM. Evaluations on specific dialect recognition tasks show that the DBN based i-vector can achieve significant and consistent performance gains over conventional GMM-UBM and DNN based i-vector methods. The generalization capability of this framework is also evaluated using DBNs trained on Mandarin and English corpuses. Index Terms: Language Identification, Deep Neural Network, Deep Bottleneck Feature, i-vector representation
Yan Song 0001, Xinhai Hong, Bing Jiang, Ruilian Cui, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH5
2015 Deep Bottleneck Feature for Image Classification
abstract
Effective image representation plays an important role for image classification and retrieval. Bag-of-Features (BoF) is well known as an effective and robust visual representation. However, on large datasets, convolutional neural networks (CNN) tend to perform much better, aided by the availability of large amounts of training data. In this paper, we propose a bag of Deep Bottleneck Features (DBF) for image classification, effectively combining the strengths of a CNN within a BoF framework. The DBF features, obtained from a previously well-trained CNN, form a compact and low-dimensional representation of the original inputs, effective for even small datasets. We will demonstrate that the resulting BoDBF method has a very powerful and discriminative capability that is generalisable to other image classification tasks.
Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
ICMR2
2015 Robust Sound Event Classification Using Deep Neural Networks
abstract
The automatic recognition of sound events by computers is an important aspect of emerging applications such as automated surveillance, machine hearing and auditory scene understanding. Recent advances in machine learning, as well as in computational models of the human auditory system, have contributed to advances in this increasingly popular research field. Robust sound event classification, the ability to recognise sounds under real-world noisy conditions, is an especially challenging task. Classification methods translated from the speech recognition domain, using features such as mel-frequency cepstral coefficients, have been shown to perform reasonably well for the sound event classification task, although spectrogram-based or auditory image analysis techniques reportedly achieve superior performance in noise. This paper outlines a sound event classification framework that compares auditory image front end features with spectrogram image-based front end features, using support vector machine and deep neural network classifiers. Performance is evaluated on a standard robust classification task in different levels of corrupting noise, and with several system enhancements, and shown to compare very well with current state-of-the-art classification techniques.
Ian McLoughlin 0001, Haomin Zhang, Zhipeng Xie, Yan Song 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Square-rich fixed point polynomial evaluation on FPGAs
abstract
Polynomial evaluation is important across a wide range of application domains, so significant work has been done on accelerating its computation. The conventional algorithm, referred to as Horner's rule, involves the least number of steps but can lead to increased latency due to serial computation. Parallel evaluation algorithms such as Estrin's method have shorter latency than Horner's rule, but achieve this at the expense of large hardware overhead. This paper presents an efficient polynomial evaluation algorithm, which reforms the evaluation process to include an increased number of squaring steps. By using a squarer design that is more efficient than general multiplication, this can result in polynomial evaluation with a 57.9% latency reduction over Horner's rule and 14.6% over Estrin's method, while consuming less area than Horner's rule, when implemented on a Xilinx Virtex 6 FPGA. When applied in fixed point function evaluation, where precision requirements limit the rounding of operands, it still achieves a 52.4% performance gain compared to Horner's rule with only a 4% area overhead in evaluating 5th degree polynomials.
Simin Xu, Suhaib A. Fahmy, Ian McLoughlin 0001
FPGA3
2014 Efficient multi-standard cognitive radios on FPGAs
abstract
Cognitive radios that support multiple standards and modify operation depending on environmental conditions are becoming more important as the demand for higher bandwidth and efficient spectrum use increases. Traditional implementations in custom ASICs cannot support such flexibility, with standards changing at a faster pace, while software baseband implementations fail to achieve the performance required. Hence, FPGAs offer an ideal platform bringing together flexibility, performance, and efficiency. This work explores the possible techniques for designing multi-standard radios on FPGAs, and explores how partial reconfiguration can be leveraged in a way that is amenable for domain experts with minimal FPGA knowledge.
Thinh Hung Pham, Suhaib A. Fahmy, Ian McLoughlin 0001
FPL3
2014 Task-aware deep bottleneck features for spoken language identification
abstract
Recently, deep bottleneck features (DBF) extracted from a deep neural network (DNN) containing a narrow bottleneck lay-er, have been applied for language identification (LID), and yield significant performance improvement over state-of-the-art methods on NIST LRE 2009. However, the DNN is trained us-ing a large corpus of specific language which is not directly related to the LID task. More recently, lattice based discrimi-native training methods for extracting more targeted DBF were proposed for ASR. Inspired by this, this paper proposes to tune the post-trained DNN parameters using an LID-specific train-ing corpus, which may make the resulting DBF, termed a Dis-criminative DBF (D2BF), more discriminative and task-aware. Specifically, the maximum mutual information (MMI) criteri-on, with gradient descent, is applied to update the DNN param-eters of the bottleneck layer in an iterative fashion. We evaluate the performance of the proposed D2BF using different back-end models, including GMM-MMI and ivector, over the most con-fused 6-languages selected from NIST LRE 2009. The results show that the proposed D2BF is more appropriate and effective than the original DBF. Index Terms: language identification, deep bottleneck feature, deep neural network, discriminative training, Gaussian mixture model, maximum mutual information 1.
Bing Jiang, Yan Song 0001, Si Wei, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH4
2014 The use of low-frequency ultrasound for voice activity detection
Ian McLoughlin 0001
INTERSPEECH1
2014 Shaping Spectral Leakage for IEEE 802.11p Vehicular Communications
abstract
IEEE 802.11p is a recently defined standard for the physical (PHY) and medium access control (MAC) layers for Dedicated Short-Range Communications. Four Spectrum Emission Masks (SEMs) are specified in 802.11p that are much more stringent than those for current 802.11 systems. In addition, the guard interval in 802.11p has been lengthened by reducing the bandwidth to support vehicular communication (VC) channels, and this results in a narrowing of the frequency guard. This raises a significant challenge for filtering the spectrum of 802.11p signals to meet the specifications of the SEMs. We investigate state of the art pulse shaping and filtering techniques for 802.11p, before proposing a new method of shaping the 802.11p spectral leakage to meet the most stringent, class D, SEM specification. The proposed method, performed at baseband to relax the strict constraints of the radio frequency (RF) front-end, allows 802.11p systems to be implemented using commercial off-the-shelf (COTS) 802.11a RF hardware, resulting in reduced total system cost.
Thinh Hung Pham, Ian McLoughlin 0001, Suhaib A. Fahmy
VTC Spring2
2014 Classifying watermelon ripeness by analysing acoustic signals using mobile devices
Wei Zeng 0004, Stefan Müller Arisona, Ian McLoughlin 0001
Pers. Ubiquitous Comput.4
2014 Super-Audible Voice Activity Detection
abstract
In this paper, reflected sound of frequency just above the audible range is used to detect speech activity. The active signal used is inaudible to humans, readily generated by the typical audio circuitry and components found in mobile telephones, and is robust to background sounds such as nearby voices. In use, the system relies upon a wideband excitation signal emitted from a loudspeaker located near the lips, which reflects from the mouth region and is then captured by a nearby microphone. The state of the lip opening is evaluated periodically by tracking the resonance patterns in the reflected excitation signal. When the lips are open, deep and complex resonances are formed as energy propagates into and then reflects out from the open mouth and vocal tract, with resonance depth being related to the open lip area. When the lips are closed, these resonance patterns are absent. The presence of the resonances can thus serve as a low complexity detection measure. The technique is evaluated for multiple users in terms of sensitivity to source placement and sensor placement. Voice activity detection performance using this measure is further evaluated in the presence of realistic wideband acoustic background noise, as well as artificially added noise. The system is shown to be relatively insensitive to sensor placement, highly insensitive to background noise, and able to achieve greater than 90% voice activity detection accuracy. The technique is even suitable when a subject is whispering in the presence of much louder multi-speaker babble. The technique has potential for speech-based systems operating in high noise environments as well as in silent speech interfaces, whisper-input systems and voice prostheses for speech-impaired users.
Ian McLoughlin 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Efficient Large Integer Squarers on FPGA
abstract
This paper presents an optimised high throughput architecture for integer squaring on FPGAs. The approach reduces the number of DSP blocks required compared to a standard multiplier. Previous work has proposed the tiling method for double precision squaring, using the least number of DSP blocks so far. However that approach incurs a large overhead in terms of look-up table (LUT) consumption and has a complex and irregular structure that is not suitable for higher word size. The architecture proposed in this paper can reduce DSP block usage by an equivalent amount to the tiling method while incurring a much lower LUT overhead: 21.8% fewer LUTs for a 53-bit squarer. The architecture is mapped to a Xilinx Virtex 6 FPGA and evaluated for a wide range of operand word sizes, demonstrating its scalability and efficiency.
Simin Xu, Suhaib A. Fahmy, Ian McLoughlin 0001
FCCM3
2013 Exemplar based language recognition method for short-duration speech segments
abstract
This paper proposes a novel exemplar-based language recognition method for short duration speech segments. It is known that language identity is a kind of weak information that can be deduced from the speech content. For short duration speech segments, the limited content also leads to a large intra-language variability. To address this issue, we propose a new method. This borrows a vector quantization based representation from image classification methods, and constructs the exemplar space using the popular i-vector representation of short duration speech segments. A mapping function is then defined to build the new representation. To evaluate the effectiveness of our proposed method, we conduct extensive experiments on the NIST LRE2007 dataset. The experimental results demonstrate improved performance for short duration speech segments.
Meng-Ge Wang, Yan Song 0001, Bing Jiang, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP5
2013 Human mouth state detection using low frequency ultrasound
Farzaneh Ahmadi, Mousa Ahmadi, Ian McLoughlin 0001
INTERSPEECH3
2013 Reconstruction of continuous voiced speech from whispers
abstract
Whispers are an important secondary vocal communications mechanism, that can be necessary for communicating private information and which are an integral aspect of natural human- to-human dialogue. Furthermore, they may be the primary com- munications method of those suffering from certain forms of aphonia, such as laryngectomees. This paper considers the con- version of continuous whispers to natural-sounding speech, and proposes a new reconstruction method based upon the synthesis of individual formants as excitation source, followed by artifi- cial glottal modulation. Early results show that the proposed method can improve quality and intelligibility over the original whispers when evaluated using continuous speech. It requires neither apriori nor speaker-dependent information, is of rela- tively low-complexity and suitable for real-time processing.
Ian McLoughlin 0001, Yan Song 0001
INTERSPEECH1
2013 Performance analysis of adaptive modulation and transmit antenna selection with channel prediction errors and feedback delay
abstract
Rate‐adaptive modulation and transmit antenna selection require channel knowledge at the transmitter, typically achieved using a feedback communications path from receiver to transmitter. Feedback delay in such systems causes outdated channel knowledge at the receiver and thus suboptimal switching decisions. This study evaluates the effect of degraded switching and proposes a channel prediction scheme to mitigate against delay‐induced performance degradation, specifically when deployed in Rayleigh fading channels using maximum ratio combining at the receiver. Shannon capacity expressions are found and then closed form bit‐error‐rate and average spectral efficiency expressions derived – all derivations supporting arbitrary number of antennas at both receiver and transmitter. These expressions are developed to allow optimal switching boundaries to be determined for M‐quadrature amplitude modulation rate adaptation under different degrees of channel prediction error, system arrangement and number of antennas.
Shiva Prakash, Ian McLoughlin 0001
IET Commun.2
2013 Low-Power Correlation for IEEE 802.16 OFDM Synchronization on FPGA
abstract
This brief compares the use of multiplierless and DSP slice-based cross-correlation for IEEE 802.16d orthogonal frequency division multiplexing (OFDM) timing synchronization on Xilinx Virtex-6 and Spartan-6 field programmable gate arrays (FPGAs). The natural approach, given the availability of embedded DSP blocks on these FPGAs, would be to implement standard multiplier-based cross-correlation. However, this can consume a significant number of DSP blocks, which may not fit on low-power devices. Hence, we compare a DSP48E1 slice-based design to four different quantizations of multiplierless correlation in terms of resource utilization and power consumption. OFDM timing synchronization accuracy is evaluated for each system at different signal-to-noise ratios. Results show that even relatively coarse multiplierless coefficient quantization can yield accurate timing synchronization, and does so at high clock speeds. Multiplierless designs enjoy reduced power consumption over the DSP48E1 Slice-based design, and can be used where DSP Slice resources are insufficient, such as on low-power FPGA devices.
Thinh Hung Pham, Suhaib A. Fahmy, Ian McLoughlin 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Channel prediction in non-regenerative multi-antenna relay selection systems
abstract
The use of multiple antennas in two-hop amplify-and-forward relay selection is analysed, where the source, relay and destination are each equipped with multiple receive but single transmit antennas. Since relay switching is based upon feedback information which is delay-limited, channel power prediction is employed to mitigate against the effect of outdated channel state information being used to make switching decisions. During transmission, a source selects a best relay on the basis of predicted signal-to-noise ratio over all available links. A chosen relay then employs maximal ratio combining at its receiver, and applies a variable gain to the received signal before forwarding to the destination. Closed form outage probability and bit error rate solutions are found for arbitrary numbers of relays and receive antennas, and used to explore trade offs between number of relays and number of antennas compared with single antenna alternatives. To assess predictor performance in combatting switching delay, comparison is made to non-predictive systems.
Shiva Prakash, Ian McLoughlin 0001
IET Commun.2
2012 Nonintrusive Quality Assessment of Noise Suppressed Speech With Mel-Filtered Energies and Support Vector Regression
abstract
Objective speech quality assessment is a challenging task which aims to emulate human judgment in the complex and time consuming task of subjective assessment. It is difficult to perform in line with the human perception due the complex and nonlinear nature of the human auditory system. The challenge lies in representing speech signals using appropriate features and subsequently mapping these features into a quality score. This paper proposes a nonintrusive metric for the quality assessment of noise-suppressed speech. The originality of the proposed approach lies primarily in the use of Mel filter bank energies (FBEs) as features and the use of support vector regression (SVR) for feature mapping. We utilize the sensitivity of FBEs to noise in order to obtain an effective representation of speech towards quality assessment. In addition, the use of SVR exploits the advantages of kernels which allow the regression algorithm to learn complex data patterns via nonlinear transformation for an effective and generalized mapping of features into the quality score. Extensive experiments conducted using two third party databases with different noise-suppressed speech signals show the effectiveness of the proposed approach.
Manish Narwaria, Weisi Lin, Ian McLoughlin 0001, Sabu Emmanuel, Liang-Tien Chia
IEEE Trans. Speech Audio Process.3
2012 Fourier Transform-Based Scalable Image Quality Measure
abstract
We present a new image quality assessment (IQA) algorithm based on the phase and magnitude of the 2D (twodimensional) Discrete Fourier Transform (DFT). The basic idea is to compare the phase and magnitude of the reference and distorted images to compute the quality score. However, it is well known that the Human Visual Systems (HVSs) sensitivity to different frequency components is not the same. We accommodate this fact via a simple yet effective strategy of nonuniform binning of the frequency components. This process also leads to reduced space representation of the image thereby enabling the reduced-reference (RR) prospects of the proposed scheme. We employ linear regression to integrate the effects of the changes in phase and magnitude. In this way, the required weights are determined via proper training and hence more convincing and effective. Lastly, using the fact that phase usually conveys more information than magnitude, we use only the phase for RR quality assessment. This provides the crucial advantage of further reduction in the required amount of reference image information. The proposed method is therefore further scalable for RR scenarios. We report extensive experimental results using a total of 9 publicly available databases: 7 image (with a total of 3832 distorted images with diverse distortions) and 2 video databases (totally 228 distorted videos). These show that the proposed method is overall better than several of the existing fullreference (FR) algorithms and two RR algorithms. Additionally, there is a graceful degradation in prediction performance as the amount of reference image information is reduced thereby confirming its scalability prospects. To enable comparisons and future study, a Matlab implementation of the proposed algorithm is available at http://www.ntu.edu.sg/home/wslin/reduced_phase.rar.
Manish Narwaria, Weisi Lin, Ian McLoughlin 0001, Sabu Emmanuel, Liang-Tien Chia
IEEE Trans. Image Process.3
2011 Fast and accurate GCPS selection scheme for SAR image registration based on an improved Trajkovic corner detector
abstract
There is a wide range of corner detector available such as the Moravec, Forstner, Harris and Trajkovic detectors. Beside the most commonly used Harris detector, the Trajkovic detector was found to have a slightly lower detection accuracy, but at a much faster computational rate, which is more favorable when it comes to high speed applications. However, the Trajkovic cornerness measure was defined using intensity differencing for optical imagery, hence it is not suitable for synthetic aperture radar (SAR) imagery, which suffers from multiplicative speckle noise. Moreover, the Trajkovic detector has a tendency to respond too readily to diagonal edges, thus it only shows good results to certain types of corners. This paper introduces a novel ground control points (GCPs) scheme to improve the speed and accuracy for SAR image registration. The scheme involves a reformulation of the Trajkovic cornerness measure in the form of a SAR intensity ratio, a hybrid scheme of Trajkovic operators for higher processing speed, and a corner refinement stage for better detection accuracy. Experiments carried out on a pair of multi-date RadarSat-2 images had shown a significant higher accuracy than in the case of the original Trajkovic detector, while performing much faster compared to the Harris detector.
Quang Huy Nguyen 0001, Ian McLoughlin 0001, Timo Rolf Bretschneider
IGARSS2
2011 Performance of Dual-Hop Multi-Antenna Systems with Fixed Gain Amplify-and-Forward Relay Selection
abstract
The performance of amplify-and-forward relaying in Rayleigh channels is explored for a dual-hop transmission system in which a transmission source selects one of several relays based upon instantaneous SNR (signal to noise ratio) to transmit a data packet to a destination receiver. Both source and receiver devices have multiple antennas, employing beamforming for transmission and MRC (maximal ratio combining) for reception. The chosen relay performs fixed gain forwarding. Closed form solutions for several performance measures are derived for this practical system, and the system is studied in terms of outage probability and symbol error rate which are verified through simulation. Several implementation alternatives are explored to note performance trade-offs, particularly between number of antennas and number of relays.
Shiva Prakash, Ian McLoughlin 0001
IEEE Trans. Wirel. Commun.2
2010 Autoregressive modelling for linear prediction of ultrasonic speech
abstract
Ultrasonic speech is a novel research area with significant applications: as a speech-aid prosthesis for patients with voice box difficulties, silent speech interfaces, secure mode of communication in mobile phones and as a communication medium in high noise industrial environments. Feature extraction is a critical part of the ultrasonic speech system. Linear prediction analysis (LPA) has been recently proven to be viable for extracting features from the three dimensional ultrasonic propagation in the vocal tract (VT). A one-dimensional autoregressive model based on averaging the LP coefficients, analysed in different recording positions has been investigated by the authors to fit the LF ultrasonic resonances of the VT. To reach a state of maturity for the LPA of ultrasonic speech and in continuum of the previous work, this paper compares the application of two major conventional methods of averaging and least squares error - already applied in room acoustics - for deriving the coefficients in autoregressive modelling of ultrasonic speech.
Farzaneh Ahmadi, Ian McLoughlin 0001, Hamid R. Sharifzadeh
INTERSPEECH2
2010 Non-intrusive Speech Quality Assessment with Support Vector Regression
Manish Narwaria, Weisi Lin, Ian McLoughlin 0001, Sabu Emmanuel, Liang-Tien Chia
MMM3
2010 Vowel Intelligibility in Chinese
abstract
Conventional wisdom states that, since the average amplitude of vowel articulation significantly exceeds that for consonants, an assessment of spoken intelligibility in obscuring noise should primarily be limited by consonant confusion. Furthermore, in both English and Chinese, consonant discrimination is considered to be more important to overall intelligibility than that of vowels. In the unbounded case, the assumption that vowel confusion is less important than consonant confusion may well be true; however, at least two situations exist where the influence of vowel confusion may be greater. The first is where vocabulary-specific restrictions confine the structure of a particular spoken word to alternatives differing primarily in their vowel. The second is the prevalence of non-additive white Gaussian noise (AWGN) interference, particularly impulsive noise which obscures only the vowel portion of a word, and similarly is present as a nonlinear effect of many time-sliced processing algorithms. This paper explores the issue of vowel intelligibility for spoken Chinese, where the confusion characteristics are complicated through the influence of lexical tone carried by the vowel in consonant-vowel-consonant (CVC) structure utterances. Experimental evidence from multilistener intelligibility testing are presented to build toward an understanding of the characteristics of Mandarin Chinese vowel confusion in the presence of AWGN. Results are also isolated by carrier word consonants and in terms of the lexical tone overlaid upon tested vowels. In particular, several factors relating to issues such as vowel length, tone combination and the crucial influence of the /a/ (IPA ) phone are revealed.
Ian McLoughlin 0001
IEEE Trans. Speech Audio Process.1
2010 Reliability through redundant parallelism for micro-satellite computing
abstract
Spacecraft typically employ rare and expensive radiation-tolerant, radiation-hardened, or at least military qualified parts for computational and other mission critical subsystems. Reasons include reliability in the harsh environment of space, and systems compatibility or heritage with previous missions. The overriding reliability concern leads most satellite computing systems to be rather conservative in design, avoiding novel or commercial-off-the-shelf components. This article describes an alternative approach: an FPGA-arbitrated parallel architecture that allows unqualified commercial devices to be incorporated into a computational device with aggregate reliability figures similar to those of traditional space-qualified alternatives. Apart from the obvious cost benefits in moving to commercial-off-the-shelf devices, these are attractive in situations where lower power consumption and/or higher processing performance are required. The latter argument is particularly of major importance at a time when the gap between required and available processing capability in satellites is widening. An analysis compares the proposed architecture to typical alternatives, maintaining risk of failure to within required levels, and discusses key applications for the parallel architecture.
Ian McLoughlin 0001, Timo Rolf Bretschneider
ACM Trans. Embed. Comput. Syst.1
2009 Predictive Transmit Antenna Selection with Maximal Ratio Combining
abstract
Antenna selection has long been a pragmatic method for exploiting spatial diversity in wireless systems with lower complexity than space-time or MIMO coding, and potentially having reduced hardware cost due to the reduction in the number of RF chains required. Whilst receive antenna selection is perhaps more common, transmit antenna selection also has several advantages, particularly for hardware-costly transmit schemes such as those requiring linearisation. However transmit antenna selection (TAS) requires either channel knowledge, or receiver knowledge at the transmitter, typically achieved using data transmission in the reverse direction, and this implies a delay between the channel being sampled and being acted upon. This outdated channel knowledge degrades system performance. In this paper, the degradation is determined, and related to the channel characteristics. A prediction scheme is then applied to mitigate against this degradation for the case of a (2,1;2) TAS system, where one of two transmit antenna is selected to communicate with two receive antennae employing maximal ratio combining.
Shiva Prakash, Ian McLoughlin 0001
GLOBECOM2
2009 Hardware-accelerated Edge Detection for Polarimetric Synthetic Aperture Radar Data
abstract
From the literature review, there are two constant false alarm rate detectors for detecting edges in multi-look fully polarimetric synthetic aperture radar (POLSAR) imagery, namely the likelihood ratio edge detector and the Roy's largest eigenvalue-based edge detector. In the latter approach, one major restriction is the computation complexity, i.e. in the context of the chosen C language-based implementation. Thus, in this paper, a novel hardware-based architecture is presented to improve the processing time for the Roy's largest eigenvalue-based edge detection. The algorithm was implemented in a field-programmable gate array (FPGA) with an accelerated solution targeting data rates of up to 1 Gb/s. Its performance was examined using nine-look NASA/JPL C-band data and evaluated in terms of processing speed and accuracy as compared to the C language-based implementation on a personal computer (PC) with a Core¿ 2 Duo processor clocked at 2.2 GHz.
Quang Huy Nguyen 0001, Ken Yoong Lee, Myo Tun Aung, Timo Rolf Bretschneider, Ian McLoughlin 0001
IGARSS (4)5
2009 Channel prediction for mitigating feedback link issues in transmit antenna selection systems
abstract
Transmit antenna selection (TAS) in multiple-input-multiple output (MIMO) has several advantages, particularly in hardware-costly transmit schemes such as those requiring linearisation. However TAS requires either channel knowledge, or receiver knowledge at the transmitter, typically achieved using data transmission in the reverse direction. This process involves a reverse channel bandwidth cost, but also implies a delay between the channel sampling and the switching. To reduce reverse channel costs, the bandwidth may be rate constrained. Thus both delay and feedback rate will limit the performance of a TAS system. To mitigate against these performance limitations, in this paper a prediction scheme is applied to systems employing transmit antenna selection with maximal ratio combining.
Shiva Prakash, Ian McLoughlin 0001
PIMRC2
2008 Secure Embedded Systems: The Threat of Reverse Engineering
abstract
Companies releasing newly designed embedded products typically recoup the cost of development through initial sales, and thus are unlikely to welcome early competition based around rapid reverse engineering of their products. By contrast, competitors able to shorten time-to-market though reverse engineering will gain design cost and market share advantages. Reverse engineering for nefarious purposes appears to be commonplace, and has significant cost impact on industry sales and profitability. In the Embedded Systems MSc programme at Nanyang Technological University, we are aiming to raise awareness of the unique security issues related to the reverse engineering of embedded systems. This effort is largely through devoting 50% of the secure embedded systems course, ES6190 to reverse engineering (the remainder to traditional security concerns). This paper covers the reverse engineering problem scope, and our approach to raising awareness through the secure embedded systems course. A classification of hardware reverse engineering steps and mitigations is also presented for the first time, with an overview of a reverse engineering curriculum. Since the quantity of published literature related to the reverse-engineering of embedded systems lies somewhere between scarce and nonexistent, this paper presents a full overview of the topic before descussing educational aspects related to this.
Ian McLoughlin 0001
ICPADS1
2008 Line spectral pairs
Ian McLoughlin 0001
Signal Process.1
2008 Subjective Intelligibility Testing of Chinese Speech
abstract
This paper presents a complete methodology and rationale for the subjective intelligibility testing of Chinese speech. It replaces the combination of several previously published Chinese intelligibility tests which have been in use for almost a decade, with a single composite test procedure constructed from a foundation of subjective trials and auditory evidence. Since publication of the first elements of Chinese intelligibility test, several factors have come to light which prompted this overhaul. First, international testing has highlighted words used in the original test that are unsuitable for speakers of particular regional dialects. Second, recent evidence indicates that the assumptions of tonal confusion made during the definition of the original tonal intelligibility tests are not borne out by subjective evidence. Finally, words published in the original test disadvantaged speakers from Mainland China due to the use of full-form Chinese characters rather than the more ubiquitous simplified form characters. This paper presents experimental evidence of tone confusion in Chinese speech, and uses this data to create a replacement tone test. Word choice has been adjusted to find more neutral alternatives for particular regional dialect speakers. The basic speech and tone extension tests are now presented with simplified form characters to ensure accessibility by the greatest number of test subjects. Finally, this paper includes a description of the full intelligibility test.
Ian McLoughlin 0001
IEEE Trans. Speech Audio Process.1
2008 A Group Ring Construction of the Extended Binary Golay Code
abstract
We give a new construction of the extended binary Golay code. The code is constructed as a zero divisor code over the dihedral group of twenty-four elements, D24. The construction is algebraic, much like the construction of the (23,12,7) binary Golay code from a polynomial.
Ian McLoughlin 0001, Ted Hurley
IEEE Trans. Inf. Theory1
2007 Linux as a teaching aid for embedded systems
abstract
Linux, an operating system kernel with a heritage derived from the room-sized mainframes of the 1970s, has seen a trememdous amount of multidisciplinary and worldwide development which has resulted in its ability to operate on some of the smallest and lowest power microprocessors available today. With its ongoing penetration into mass market embedded systems such as smart phones, automotive and aeronautical equipment, multimedia entertainment devices and personal digital assistants, versions of Linux are now the operating system of choice for many embedded developers. However Linux tends to be overlooked in the more traditional embedded electronics education. This paper explores the characteristics of embedded Linux that affect the education of embedded systems, and describes a tested educational model for its teaching within a general embedded systems education course.
Ian McLoughlin 0001, Anton J. R. Aendenroomer
ICPADS1
2007 An Embedded Systems graduate education for Singapore
Ian McLoughlin 0001, Douglas L. Maskell, Thambipillai Srikanthan, Wooi-Boon Goh
ICPADS1
2006 Transmit Antenna Selection for UHF MIMO Linking
abstract
The use of multiple-input multiple-output (MIMO) configurations in wireless systems is becoming increasingly popular, due to the potential capacity enhancing properties. The use of such configurations, requires deployment of multiple RF chains and results in an increase in cost and complexity. This paper focuses on transmitter design, where cost is often dominated by the need to utilise linearized power amplifiers. Thus, the effect of reducing the number of active transmit antennae through selection can yield a significant cost advantage. Three criteria for transmit antenna selection are evaluated and characterised through the use of real channel data in the UHF band. The criteria are based only on the channel gains and have relatively low complexity; thus are suitable for practical purposes. This study examines the performance of the tested selection algorithms under a variety of operating conditions. The paper also considers the issue of delayed switch time due to reverse-link communication latency
Marjan A. Baghaie, Ian McLoughlin 0001, Philippa A. Martin, Kishore Mehrotra, Desmond P. Taylor
VTC Spring2
2002 Intelligibility evaluation of GSM coder for Mandarin speech using CDRT
Ian McLoughlin 0001, Zhongqiang Ding, Eng Chong Tan
Speech Commun.1