Bin Ma 0001

dblp:70/6176-1 · DBLP profile ↗
← Back
211ranked-venue papers
11as first author
39since 2021 · last 2025
0000-0002-9223-9654ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 188 · 8 first-author · 35 since 2021Artificial intelligence and machine learning · 121 · 4 first-author · 17 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
abstract
The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent representations and poor speech quality, especially in out-of-domain scenarios. In this work, we propose HiFi-SR, a unified network that leverages end-to-end adversarial training to achieve high-fidelity speech super-resolution. Our model features a unified transformer-convolutional generator designed to seamlessly handle both the prediction of latent representations and their conversion into time-domain waveforms. The transformer network serves as a powerful encoder, converting low-resolution mel-spectrograms into latent space representations, while the convolutional network upscales these representations into high-resolution waveforms. To enhance high-frequency fidelity, we incorporate a multi-band, multi-scale time-frequency discriminator, along with a multi-scale mel-reconstruction loss in the adversarial training process. HiFi-SR is versatile, capable of upscaling any input speech signal between 4 kHz and 32 kHz to a 48 kHz sampling rate. Experimental results demonstrate that HiFi-SR significantly outperforms existing speech SR methods across both objective metrics and ABX preference tests, for both in-domain and out-of-domain scenarios.
Shengkui Zhao, Kun Zhou 0003, Zexu Pan, Chong Zhang 0003, Bin Ma 0001
ICASSP6
2025 Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning
abstract
Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or spectral domains, leading to increased generation complexity and slower inference speeds. Additionally, these methods have primarily modelled clean speech distributions, with limited exploration of noise distributions, thereby constraining the discriminative capability of diffusion models for speech enhancement. To address these issues, we propose a novel approach that integrates a conditional latent diffusion model (cLDM) with dual-context learning (DCL). Our method utilizes a variational autoencoder (VAE) to compress mel-spectrograms into a low-dimensional latent space. We then apply cLDM to transform the latent representations of both clean speech and background noise into Gaussian noise by the DCL process, and a parameterized model is trained to reverse this process, conditioned on noisy latent representations and text embeddings. By operating in a lower-dimensional space, the latent representations reduce the complexity of the generation process, while the DCL process enhances the model’s ability to handle diverse and unseen noise environments. Our experiments demonstrate the strong performance of the proposed approach compared to existing diffusion-based methods, even with fewer iterative steps, and highlight the superior generalization capability of our models to out-of-domain noise datasets.
Shengkui Zhao, Zexu Pan, Kun Zhou 0003, Chong Zhang 0003, Bin Ma 0001
ICASSP6
2025 Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
abstract
Achieving high-fidelity audio compression while preserving perceptual quality across diverse audio types remains a significant challenge in Neural Audio Coding (NAC). This paper introduces MUFFIN, a fully convolutional NAC framework that leverages psychoacoustically guided multi-band frequency reconstruction. Central to MUFFIN is the Multi-Band Spectral Residual Vector Quantization (MBS-RVQ) mechanism, which quantizes latent speech across different frequency bands. This approach optimizes bitrate allocation and enhances fidelity based on psychoacoustic studies, achieving efficient compression with unique perceptual features that separate content from speaker attributes through distinct codebooks. MUFFIN integrates a transformer-inspired convolutional architecture with proposed modified snake activation functions to capture fine frequency details with greater precision. Extensive evaluations on diverse datasets (LibriTTS, IEMOCAP, GTZAN, BBC) demonstrate MUFFIN’s ability to consistently surpass existing performance in audio reconstruction across various domains. Notably, a high-compression variant achieves an impressive SOTA 12.5 kHz rate while preserving reconstruction quality. Furthermore, MUFFIN excels in downstream generative tasks, demonstrating its potential as a robust token representation for integration with large language models. These results establish MUFFIN as a groundbreaking advancement in NAC and as the first neural psychoacoustic coding system. Speech demos and codes are available at https://demos46.github.io/muffin/ and https://github.com/dianwen-ng/MUFFIN.
Dianwen Ng, Kun Zhou 0003, Yi-Wen Chao, Zhiwei Xiong, Bin Ma 0001, Chng Eng Siong
ICML5
2025 A-SMiLE: Affective Sparse Mixture-of-Experts Adapter with Multi-Task Learning for Spoken Dialogue Models
Yi-Wen Chao, Yizhou Peng, Dianwen Ng, Chongjia Ni, Bin Ma 0001, Chng Eng Siong
INTERSPEECH6
2025 Thinking Fast and Slow: Robust Speech Recognition via Deep Filter-Tuning
Dianwen Ng, Kun Zhou 0003, Bin Ma 0001, Chng Eng Siong
INTERSPEECH3
2025 Online Audio-Visual Autoregressive Speaker Extraction
Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang 0003, Kun Zhou 0003, Bin Ma 0001
INTERSPEECH7
2025 Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
Zexu Pan, Shengkui Zhao, Kun Zhou 0003, Chong Zhang 0003, Bin Ma 0001
INTERSPEECH7
2025 FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
Yizhou Peng, Yi-Wen Chao, Dianwen Ng, Chongjia Ni, Bin Ma 0001, Chng Eng Siong
INTERSPEECH6
2025 ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
Shengkui Zhao, Zexu Pan, Bin Ma 0001
INTERSPEECH3
2024 Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?
abstract
Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. However, not many people understand how and why this is so. In this study, we aim to deepen our understanding of this emerging method by investigating the role of soft prompts in automatic speech recognition (ASR). Our findings highlight their role as zero-shot learners in improving ASR performance while also exposing them to the risk of malicious modifications. Soft prompts aid generalization but are not obligatory for inference. We also identify two primary roles of soft prompts: content refinement and noise information enhancement, which enhances robustness against background noise. Additionally, we propose an effective modification on noise prompts to show that they are capable of zero-shot learning on adapting to out-of-distribution noise environments.
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Fabian Ritter Gutierrez, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001
ICASSP10
2024 SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance
abstract
Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, which comprise half a dual-path model’s parameters, contribute minimally to performance. Thus, we propose the Single-Path Global Modulation (SPGM) block to replace inter-blocks. SPGM is named after its structure consisting of a parameter-free global pooling module followed by a modulation module comprising only 2% of the model’s total parameters. The SPGM block allows all transformer layers in the model to be dedicated to local feature modelling, making the overall model single-path. SPGM achieves 22.1 dB SI-SDRi on WSJ0-2Mix and 20.4 dB SI-SDRi on Libri2Mix, exceeding the performance of Sepformer by 0.5 dB and 0.3 dB respectively and matches the performance of recent SOTA models with up to 8 times fewer parameters. Model and weights are available at huggingface.co/yipjiaqi/spgm
Jia Qi Yip, Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Dianwen Ng, Chng Eng Siong, Bin Ma 0001
ICASSP11
2024 MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
abstract
Our previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale recurrent patterns. In this paper, we introduce a novel hybrid model that provides the capabilities to model both long-range, coarse-scale dependencies and fine-scale recurrent patterns by integrating a recurrent module into the MossFormer framework. Instead of applying the recurrent neural networks (RNNs) that use traditional recurrent connections, we present a recurrent module based on a feedforward sequential memory network (FSMN), which is considered "RNN-free" recurrent network due to the ability to capture recurrent patterns without using recurrent connections. Our recurrent module mainly comprises an enhanced dilated FSMN block by using gated convolutional units (GCU) and dense connections. In addition, a bottleneck layer and an output layer are also added for controlling information flow. The recurrent module relies on linear projections and convolutions for seamless, parallel processing of the entire sequence. The integrated MossFormer2 hybrid model demonstrates remarkable enhancements over MossFormer and surpasses other state-of-the-art methods in WSJ0-2/3mix, Libri2Mix, and WHAM!/WHAMR! benchmarks.
Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Jia Qi Yip, Dianwen Ng, Bin Ma 0001
ICASSP10
2024 Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
Kun Zhou 0003, Shengkui Zhao, Chong Zhang 0003, Hao Wang 0199, Dianwen Ng, Chongjia Ni, Trung Hieu Nguyen 0001, Jia Qi Yip, Bin Ma 0001
INTERSPEECH10
2024 Towards Audio Codec-based Speech Separation
Jia Qi Yip, Shengkui Zhao, Dianwen Ng, Chng Eng Siong, Bin Ma 0001
INTERSPEECH5
2024 Tuning Large Language Model for Speech Recognition With Mixed-Scale Re-Tokenization
abstract
Large Language Models (LLMs) have proven successful across a spectrum of speech-related tasks, such as speech recognition, text-to-speech, and spoken language understanding. Recently, the use of discretized speech features has gained attention as an efficient and compatible alternative to continuous features for LLMs. This is mainly due to their reduced storage requirements and better alignment of these features with LLM's input space. However, the typical practice of freezing the speech encoder during training poses challenges in bridging the modality gap between speech and text. To address this, we propose to use a mixed-scale re-tokenization layer, integrating multiple granularities in discretized speech features directly within the LLM's input module. Our experimental results demonstrated that the proposed method can effectively enhance the performance of ASR in the setting of continuous learning of an LLM, highlighting the importance of a meticulously designed input module for the integration of discretized speech features with an LLM.
Chong Zhang 0003, Qian Chen 0003, Wen Wang 0001, Bin Ma 0001
IEEE Signal Process. Lett.5
2023 Auxiliary Pooling Layer For Spoken Language Understanding
abstract
End-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to encode semantic information. Existing works explore transferring knowledge from a pre-trained text model with different alignment losses at a fixed granularity. In this paper, we address the variable granularity in transferring knowledge from texts to speech representation via APLY, an auxiliary pooling layer, that fuses the global information with the adaptively encoded local context. We demonstrate the effectiveness of APLY on three benchmarks of spoken language understanding.
Trung Hieu Nguyen 0001, Jinjie Ni, Wen Wang 0001, Qian Chen 0003, Chong Zhang 0003, Bin Ma 0001
ICASSP7
2023 De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition
abstract
Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a single source, which tends to do poorly when noise is present during testing. Nonetheless, it is crucial to overcome the adverse influence of noise for real-world applications. In this work, we propose a novel training framework, called deHuBERT, for noise reduction encoding inspired by H. Barlow’s redundancy-reduction principle. The new framework improves the HuBERT training algorithm by introducing auxiliary losses that drive the self- and cross-correlation matrix between pairwise noise-distorted embeddings towards identity matrix. This encourages the model to produce noise- agnostic speech representations. With this method, we report improved robustness in noisy environments, including unseen noises, without impairing the performance on the clean set.
Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Jinjie Ni, Chong Zhang 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001
ICASSP10
2023 Contrastive Speech Mixup for Low-Resource Keyword Spotting
abstract
Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more person-alized, KWS models need to adapt quickly to smaller user samples. To tackle this challenge, we propose a contrastive speech mixup (CosMix) learning algorithm for low-resource KWS. CosMix introduces an auxiliary contrastive loss to the existing mixup augmentation technique to maximize the relative similarity between the original pre-mixed samples and the augmented samples. The goal is to inject enhancing constraints to guide the model towards simpler but richer content-based speech representations from two augmented views (i.e. noisy mixed and clean pre-mixed utterances). We conduct our experiments on the Google Speech Command dataset, where we trim the size of the training set to as small as 2.5 mins per keyword to simulate a low-resource condition. Our experimental results show a consistent improvement in the performance of multiple models, which exhibits the effectiveness of our method.
Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Chng Eng Siong, Bin Ma 0001
ICASSP9
2023 Adaptive Knowledge Distillation Between Text and Speech Pre-Trained Models
abstract
Learning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models that are pre-trained on rich sources of texts. The distillation process, however, is challenging due to the modal disparity between textual and speech embedding spaces. This paper studies metric-based distillation to align the embedding space of text and speech with only a small amount of data without modifying the model structure. Since the semantic and granularity gap between text and speech has been omitted in literature, which impairs the distillation, we propose the Prior-informed Adaptive knowledge Distillation (PAD) that adaptively leverages text/speech units of variable granularity and prior distributions to achieve better global and local alignments between text and speech pre-trained models. We evaluate on three spoken language understanding benchmarks to show that PAD is more effective in transferring linguistic knowledge than other metric-based distillation approaches.
Jinjie Ni, Wen Wang 0001, Qian Chen 0033, Dianwen Ng, Han Lei, Trung Hieu Nguyen 0001, Chong Zhang 0003, Bin Ma 0001, Erik Cambria
ICASSP9
2023 D2Former: A Fully Complex Dual-Path Dual-Decoder Conformer Network Using Joint Complex Masking and Complex Spectral Mapping for Monaural Speech Enhancement
abstract
Monaural speech enhancement has been widely studied using real networks in the time-frequency (TF) domain. However, the input and the target are naturally complex-valued in the TF domain, a fully complex network is highly desirable for effectively learning the feature representation and modelling the sequence in the complex domain. Moreover, phase, an important factor for perceptual quality of speech, has been proved learnable together with magnitude from noisy speech using complex masking or complex spectral mapping. Many recent studies focus on either complex masking or complex spectral mapping, ignoring their performance boundaries. To address above issues, we propose a fully complex dual-path dual-decoder conformer network (D2Former) using joint complex masking and complex spectral mapping for monaural speech enhancement. In D2Former, we extend the conformer network into the complex domain and form a dual-path complex TF self-attention architecture for effectively modelling the complex-valued TF sequence. We further boost the TF feature representation in the encoder and the decoders using a dual-path learning structure by exploiting complex dilated convolutions on time dependency and complex feedforward sequential memory networks (CFSMN) for frequency recurrence. In addition, we improve the performance boundaries of complex masking and complex spectral mapping by combining the strengths of the two training targets into a joint-learning framework. As a consequence, D2Former takes fully advantages of the complex-valued operations, the dual-path processing, and the joint-training targets. Compared to the previous models, D2Former achieves state-of-the-art results on the VoiceBank+Demand benchmark with the smallest model size of 0.87M parameters.
Shengkui Zhao, Bin Ma 0001
ICASSP2
2023 MossFormer: Pushing the Performance Limit of Monaural Speech Separation Using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions
abstract
Transformer based models have provided significant performance improvements in monaural speech separation. However, there is still a performance gap compared to a recent proposed upper bound. The major limitation of the current dual-path Transformer models is the inefficient modelling of long-range elemental interactions and local feature patterns. In this work, we achieve the upper bound by proposing a gated single-head transformer architecture with convolution-augmented joint self-attentions, named MossFormer (Monaural speech separation TransFormer). To effectively solve the indirect elemental interactions across chunks in the dual-path architecture, MossFormer employs a joint local and global self-attention architecture that simultaneously performs a full-computation self-attention on local chunks and a linearised low-cost self-attention over the full sequence. The joint attention enables MossFormer model full-sequence elemental interaction directly. In addition, we employ a powerful attentive gating mechanism with simplified single-head self-attentions. Besides the attentive long-range modelling, we also augment MossFormer with convolutions for the position-wise local pattern modelling. As a consequence, MossFormer significantly outperforms the previous models and achieves the state-of-the-art results on WSJ0-2/3mix and WHAM!/WHAMR! benchmarks. Our model achieves the SI-SDRi upper bound of 21.2 dB on WSJ0-3mix and only 0.3 dB below the upper bound of 23.1 dB on WSJ0-2mix.
Shengkui Zhao, Bin Ma 0001
ICASSP2
2023 Adapter-tuning with Effective Token-dependent Representation Shift for Automatic Speech Recognition
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Qian Chen 0003, Wen Wang 0001, Chng Eng Siong, Bin Ma 0001
INTERSPEECH11
2023 Small Footprint Multi-channel Network for Keyword Spotting with Centroid Based Awareness
Dianwen Ng, Yang Xiao 0019, Jia Qi Yip, Biao Tian 0002, Qiang Fu 0001, Chng Eng Siong, Bin Ma 0001
INTERSPEECH8
2023 Dual Acoustic Linguistic Self-supervised Representation Learning for Cross-Domain Speech Recognition
Dianwen Ng, Chong Zhang 0003, Xiao Fu 0001, Wei Xi 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001, Jizhong Zhao
INTERSPEECH10
2023 A Unified Recognition and Correction Model under Noisy and Accent Speech Conditions
Dianwen Ng, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong
INTERSPEECH9
2023 Dual-Memory Multi-Modal Learning for Continual Spoken Keyword Spotting with Confidence Selection and Diversity Enhancement
Dianwen Ng, Xizhe Li, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong
INTERSPEECH10
2023 ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention
Jia Qi Yip, Duc-Tuan Truong, Dianwen Ng, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001
INTERSPEECH10
2022 Multimodal Sentiment Analysis on Unaligned Sequences Via Holographic Embedding
abstract
Multimodal sentiment analysis is built on fusion of inputs from multiple modalities. However, at the core of existing fusion method is the dot product between a key vector and a query vector and relies on multiple neural network layers to model the high-order correlation. In this paper, we present a method based on holographic reduced representation which is a compressed version of the outer product to model facilitate higher-order fusion across multiple modality. Experiment shows that our proposal performs promisingly on benchmark multimodal sentiment analysis data sets with improved efficiency.
Bin Ma 0001
ICASSP2
2022 CPT: Cross-Modal Prefix-Tuning for Speech-To-Text Translation
abstract
Speech translation models benefit from adapting multilingual pretrained language models. However, such adaptation modifies the parameters in the pretrained model to favor a specific task. Prefix-tuning, as a lightweight adaptation technique, has recently emerged as an efficient adaptation method that significantly reduces the number of trainable parameters and has demonstrated great potential in low-resource settings. It inserts prefixes into the output of each layer of a pretrained model, without modifying its parameters. During training, only the parameters of prefixes are updated while the rest of the model are being frozen. In this paper, we improve the performance of speech translation in medium-/low-resource settings by a cross-modal prefix that bridges the gap between speech input and translation modules to reduce the information loss in the cascaded model. We show that the proposed cross-modal prefix-tuning is effective, robust and parameter-efficient for adapting a speech recognition and translation pipeline.
Trung Hieu Nguyen 0001, Bin Ma 0001
ICASSP3
2022 End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression
abstract
Echo and noise suppression is an integral part of a full-duplex communication system. Many recent acoustic echo cancellation (AEC) systems rely on a separate adaptive filtering module for linear echo suppression and a neural module for residual echo suppression. However, in practice, adaptive filtering modules require time to converge and remain susceptible to changes in the acoustic environment. This introduces unnecessary delays to AEC systems using this two-stage framework, despite neural modules already having the capability to suppress both linear and nonlinear echo components. In this paper, we exploit the offset-compensating property of complex time-frequency masks and propose an end-to-end complex-valued neural network architecture. The building block of the proposed model is a pseudocomplex extension of the densely-connected multidilated DenseNet (D3Net), resulting in a very small network of only 354K parameters. The architecture utilized the multi-resolution nature of the D3Net to eliminate the need for pooling, allowing feature extraction using large receptive fields without any loss of output resolution. We also propose a dual-mask technique for joint echo and noise suppression with simultaneous speech enhancement. Evaluation on both synthetic and real test sets demonstrated promising results across multiple energy-based metrics and perceptual proxies.
Karn Watcharasupat, Thi Ngoc Tho Nguyen, Woon-Seng Gan, Shengkui Zhao, Bin Ma 0001
ICASSP5
2022 M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
abstract
Recent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologies. Speaker diarization and multi-speaker automatic speech recognition in meeting scenarios have attracted much attention recently. However, the lack of large public meeting data has been a major obstacle for advancement of the field. Therefore, we make available the AliMeeting corpus, which consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone. Each meeting session is composed of 2-4 speakers with different speaker overlap ratio, recorded in meeting rooms with different size. Along with the dataset, we launch the ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) with two tracks, namely speaker diarization and multi-speaker ASR, aiming to provide a common testbed for meeting rich transcription and promote reproducible research in this field. In this paper we provide a detailed introduction of the AliMeeting dateset, challenge rules, evaluation methods and baseline systems.
Fan Yu 0002, Shiliang Zhang, Yihui Fu, Lei Xie 0001, Zhihao Du, Weilong Huang, Zhijie Yan, Bin Ma 0001, Hui Bu
ICASSP10
2022 Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
abstract
The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic speech recognition (ASR) (track 2). Along with the challenge, we released 120 hours of real-recorded Mandarin meeting speech data with manual annotation, including far-field data collected by 8-channel micro-phone array as well as near-field data collected by each participants’ headset microphone. We briefly describe the released dataset, track setups, baselines and summarize the challenge results and major techniques used in the submissions.
Fan Yu 0002, Shiliang Zhang, Yihui Fu, Zhihao Du, Weilong Huang, Lei Xie 0001, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong-Aik Lee, Zhijie Yan, Bin Ma 0001, Hui Bu
ICASSP14
2022 FRCRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement
abstract
Convolutional recurrent networks (CRN) integrating a convolutional encoder-decoder (CED) structure and a recurrent structure have achieved promising performance for monaural speech enhancement. However, feature representation across frequency context is highly constrained due to limited receptive fields in the convolutions of CED. In this paper, we propose a convolutional recurrent encoder-decoder (CRED) structure to boost feature representation along the frequency axis. The CRED applies frequency recurrence on 3D convolutional feature maps along the frequency axis following each convolution, therefore, it is capable of catching long-range frequency correlations and enhancing feature representations of speech inputs. The proposed frequency recurrence is realized efficiently using a feedforward sequential memory network (FSMN). Besides the CRED, we insert two stacked FSMN layers between the encoder and the decoder to model further temporal dynamics. We name the proposed framework as Frequency Recurrent CRN (FRCRN). We design FRCRN to predict complex Ideal Ratio Mask (cIRM) in complex-valued domain and optimize FRCRN using both time-frequency-domain and time-domain losses. Our proposed approach achieved state-of-the-art performance on wideband bench-mark datasets and achieved 2nd place for the real-time fullband track in terms of Mean Opinion Score (MOS) and Word Accuracy (WAcc) in the ICASSP 2022 Deep Noise Suppression (DNS) challenge.
Shengkui Zhao, Bin Ma 0001, Karn Watcharasupat, Woon-Seng Gan
ICASSP2
2022 Learning Disentangled Representations for Counterfactual Regression via Mutual Information Minimization
abstract
Learning individual-level treatment effect is a fundamental problem in causal inference and has received increasing attention in many areas, especially in the user growth area which concerns many internet companies. Recently, disentangled representation learning methods that decompose covariates into three latent factors, including instrumental, confounding and adjustment factors, have witnessed great success in treatment effect estimation. However, it remains an open problem how to learn the underlying disentangled factors precisely. Specifically, previous methods fail to obtain independent disentangled factors, which is a necessary condition for identifying treatment effect. In this paper, we propose Disentangled Representations for Counterfactual Regression via Mutual Information Minimization (MIM-DRCFR), which uses a multi-task learning framework to share information when learning the latent factors and incorporates MI minimization learning criteria to ensure the independence of these factors. Extensive experiments including public benchmarks and real-world industrial user growth datasets demonstrate that our method performs much better than state-of-the-art methods.
Mingyuan Cheng, Xinru Liao, Quan Liu 0008, Bin Ma 0001, Jian Xu 0015, Bo Zheng 0007
SIGIR4
2021 Heterogeneous Graph Neural Networks for Large-Scale Bid Keyword Matching
abstract
Digital advertising is a critical part of many e-commerce platforms such as Taobao and Amazon. While in recent years a lot of attention has been drawn to the consumer side including canonical problems like ctr/cvr prediction, the advertiser side, which directly serves advertisers by providing them with marketing tools, is now playing a more and more important role. When speaking of sponsored search, bid keyword recommendation is the fundamental service. This paper addresses the problem of keyword matching, the primary step of keyword recommendation. Existing methods for keyword matching merely consider modeling relevance based on a single type of relation among ads and keywords, such as query clicks or text similarity, which neglects rich heterogeneous interactions hidden behind them. To fill this gap, the keyword matching problem faces several challenges including: 1) how to learn enriched and robust embeddings from complex interactions among various types of objects; 2) how to conduct high-quality matching for new ads that usually lack sufficient data.
Zongtao Liu, Bin Ma 0001, Quan Liu 0008, Jian Xu 0015, Bo Zheng 0007
CIKM2
2021 A Unified Speaker Adaptation Approach for ASR
abstract
Transformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mismatch between training and test data. Further finetuning a trained model with target speaker data is the most natural approach for adaptation, but it takes a lot of compute and may cause catastrophic forgetting to the existing speakers. In this work, we propose a unified speaker adaptation approach consisting of feature adaptation and model adaptation. For feature adaptation, we employ a speaker-aware persistent memory model which generalizes better to unseen test speakers by making use of speaker i-vectors to form a persistent memory. For model adaptation, we use a novel gradual pruning method to adapt to target speakers without changing the model architecture, which to the best of our knowledge, has never been explored in ASR. Specifically, we gradually prune less contributing parameters on model encoder to a certain sparsity level, and use the pruned parameters for adaptation, while freezing the unpruned parameters to keep the original model performance. We conduct experiments on the Librispeech dataset. Our proposed approach brings relative 2.74-6.52% word error rate (WER) reduction on general speaker adaptation. On target speaker adaptation, our method outperforms the baseline with up to 20.58% relative WER reduction, and surpasses the finetuning method by up to relative 2.54%. Besides, with extremely low-resource adaptation data (e.g., 1 utterance), our method could improve the WER by relative 6.53% with only a few epochs of training.
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001
EMNLP (1)6
2021 Preventing Early Endpointing for Online Automatic Speech Recognition
abstract
With the recent development of end-to-end models in speech recognition, there have been more interests in adapting these models for online speech recognition. However, using end-to-end models for online speech recognition is known to suffer from an early endpointing problem, which brings in many deletion errors. In this paper, we propose to address the early endpointing problem from the gradient perspective. Specifically, we leverage on the recently proposed ScaleGrad technique, which was proposed to mitigate the text degeneration issue. Different from ScaleGrad, we adapt it to discourage the early generation of the end-of-sentence () token. A scaling term is added to directly maneuver the gradient of the training loss to encourage the model to learn to keep generating non-tokens. Compared with previous approaches such as voice-activity-detection and end-of-query detection, the proposed method does not rely on various types of silence, and it also saves the trouble from obtaining the ground truth endpoint with forced alignment. Nevertheless, it can be jointly applied with other techniques. Experiments on AISHELL-1 dataset show that our model brings relative 5.4%-10.1% CER reductions over the baseline, and surpasses the unlikelihood training method which directly reduces the generation probability oftoken.
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001
ICASSP6
2021 Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
abstract
Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections, which heavily rely on the representation power of the complex-valued convolutional layers. In this paper, we propose a complex convolutional block attention module (CCBAM) to boost the representation power of the complex-valued convolutional layers by constructing more informative features. The CCBAM is a lightweight and general module which can be easily integrated into any complex-valued convolutional layers. We integrate CCBAM with the deep complex U-Net and CRN to enhance their performance for speech enhancement. We further propose a mixed loss function to jointly optimize the complex models in both time-frequency (TF) domain and time domain. By integrating CCBAM and the mixed loss, we form a new end-to-end (E2E) complex speech enhancement framework. Ablation experiments and objective evaluations show the superior performance of the proposed approaches.
Shengkui Zhao, Trung Hieu Nguyen 0001, Bin Ma 0001
ICASSP3
2021 Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram
abstract
Cross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to-speech (TTS) model, i.e., FastSpeech, and LPCNet neural vocoder to design a new cross-lingual VC framework named FastSpeech-VC. We address the mismatches of the phonetic set and the speech prosody by applying Phonetic PosteriorGrams (PPGs), which have been proved to bridge across speaker and language boundaries. Moreover, we add normalized logarithm-scale fundamental frequency (Log-F0) to further compensate for the prosodic mismatches and significantly improve naturalness. Our experiments on English and Mandarin languages demonstrate that with only mono-lingual corpus, the proposed FastSpeech-VC can achieve high quality converted speech with mean opinion score (MOS) close to the professional records while maintaining good speaker similarity. Compared to the baselines using Tacotron2 and Transformer TTS models, the FastSpeech-VC can achieve controllable converted speech rate and much faster inference speed. More importantly, the FastSpeech-VC can easily be adapted to a speaker with limited training utterances.
Shengkui Zhao, Hao Wang 0199, Trung Hieu Nguyen 0001, Bin Ma 0001
ICASSP4
2020 Independent Language Modeling Architecture for End-To-End ASR
abstract
The attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model (LM), is conditioned on the encoder output. This means that the acoustic encoder and the language model are entangled that doesn’t allow language model to be trained separately from external text data. To address this problem, in this work, we propose a new architecture that separates the decoder subnet from the encoder output. In this way, the decoupled subnet becomes an independently trainable LM subnet, which can easily be updated using the external text data. We study two strategies for updating the new architecture. Experimental results show that, 1) the independent LM architecture benefits from external text data, achieving 9.3% and 22.8% relative character and word error rate reduction on Mandarin HKUST and English NSC datasets respectively; 2) the proposed architecture works well with external LM and can be generalized to different amount of labelled data.
Van Tung Pham, Haihua Xu 0001, Yerbolat Khassanov, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001, Haizhou Li 0001
ICASSP7
2020 Speech Transformer with Speaker Aware Persistent Memory
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001
INTERSPEECH6
2020 Universal Speech Transformer
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001
INTERSPEECH6
2020 Cross Attention with Monotonic Alignment for Speech Transformer
Yingzhu Zhao, Chongjia Ni, Cheung-Chi Leung, Shafiq R. Joty, Chng Eng Siong, Bin Ma 0001
INTERSPEECH6
2020 Towards Natural Bilingual and Code-Switched Speech Synthesis Based on Mix of Monolingual Recordings and Cross-Lingual Voice Conversion
abstract
Recent state-of-the-art neural text-to-speech (TTS) synthesis models have dramatically improved intelligibility and naturalness of generated speech from text. However, building a good bilingual or code-switched TTS for a particular voice is still a challenge. The main reason is that it is not easy to obtain a bilingual corpus from a speaker who achieves native-level fluency in both languages. In this paper, we explore the use of Mandarin speech recordings from a Mandarin speaker, and English speech recordings from another English speaker to build high-quality bilingual and code-switched TTS for both speakers. A Tacotron2-based cross-lingual voice conversion system is employed to generate the Mandarin speaker's English speech and the English speaker's Mandarin speech, which show good naturalness and speaker similarity. The obtained bilingual data are then augmented with code-switched utterances synthesized using a Transformer model. With these data, three neural TTS models -- Tacotron2, Transformer and FastSpeech are applied for building bilingual and code-switched TTS. Subjective evaluation results show that all the three systems can produce (near-)native-level speech in both languages for each of the speaker.
Shengkui Zhao, Trung Hieu Nguyen 0001, Hao Wang 0199, Bin Ma 0001
INTERSPEECH4
2020 Fast Query-by-Example Speech Search Using Attention-Based Deep Binary Embeddings
abstract
State-of-the-art query-by-example (QbE) speech search approaches usually use recurrent neural network (RNN) based acoustic word embeddings (AWEs) to represent variable-length speech segments with fixed-dimensional vectors, and thus simple cosine distances can be measured over the embedded vectors of both the spoken query and the search content. In this paper, we aim to improve search accuracy and speed for the AWE-based QbE approach in low-resource scenario. First, multi-head self-attentive mechanism is introduced for learning a sequence of attention weights for all time steps of RNN outputs while attending to different positions of a speech segment. Second, as the real-valued AWEs suffer from substantial computation in similarity measure, a hashing layer is adopted for learning deep binary embeddings, and thus binary pattern matching can be directly used for fast QbE speech search. The proposed approach of self-attentive deep hashing network is effectively trained with three specifically-designed objectives: a penalization term, a triplet loss, and a quantization loss. Experiments show that our approach improves the relative search speed by 8 times and mean average precision (MAP) by 18.9%, as compared with the previous best real-valued embedding approach.
Yougen Yuan, Lei Xie 0001, Cheung-Chi Leung, Hongjie Chen 0001, Bin Ma 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Robust Audio-visual Speech Recognition Using Bimodal Dfsmn with Multi-condition Training and Dropout Regularization
abstract
Audio-visual speech recognition (AVSR) is thought to be one of the potential solutions for robust speech recognition, especially in noisy environments. Compared to audio only speech recognition, the major issues of AVSR include the lack of publicly available audio-visual corpora and the need of robust knowledge fusion of both speech and vision. In this work, based on the recently released NTCD-TIMIT audio-visual corpus, we address the challenges of AVSR through three aspects: 1) optimal integration of acoustic and visual information; 2) robust performance with multi-condition training; 3) robust modeling against missing visual information during decoding. We propose a bimodal-DFSMN to jointly learn feature fusion and acoustic modeling, and utilize a per-frame dropout approach to enhance the robustness of AVSR system against the missing of visual modality. In the experiments, we construct two setups based on the NTCD-TIMIT corpus that consists of 5 hours clean training data and 150 hours multi-condition training data, respectively. As a result, we achieve a phone error rate of 12.6% on clean test set and an average phone error rate of 26.2% on all test sets (clean, various SNRs, various noise types), which both dramatically improve the baseline performance in NTCD-TIMIT task.
Shiliang Zhang, Bin Ma 0001, Lei Xie 0001
ICASSP3
2019 Constrained Output Embeddings for End-to-End Code-Switching Speech Recognition with Only Monolingual Data
abstract
The lack of code-switch training data is one of the major concerns in the development of end-to-end code-switching automatic speech recognition (ASR) models. In this work, we propose a method to train an improved end-to-end code-switching ASR using only monolingual data. Our method encourages the distributions of output token embeddings of monolingual languages to be similar, and hence, promotes the ASR model to easily code-switch between languages. Specifically, we propose to use Jensen-Shannon divergence and cosine distance based constraints. The former will enforce output embeddings of monolingual languages to possess similar distributions, while the later simply brings the centroids of two distributions to be close to each other. Experimental results demonstrate high effectiveness of the proposed method, yielding up to 4.5% absolute mixed error rate improvement on Mandarin-English code-switching ASR task.
Yerbolat Khassanov, Haihua Xu 0001, Van Tung Pham, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001
INTERSPEECH7
2019 Towards Language-Universal Mandarin-English Speech Recognition
Shiliang Zhang, Bin Ma 0001, Lei Xie 0001
INTERSPEECH4
2019 Multi-Task Multi-Network Joint-Learning of Deep Residual Networks and Cycle-Consistency Generative Adversarial Networks for Robust Speech Recognition
Shengkui Zhao, Chongjia Ni, Rong Tong, Bin Ma 0001
INTERSPEECH4
2019 Fast Learning for Non-Parallel Many-to-Many Voice Conversion with Residual Star Generative Adversarial Networks
Shengkui Zhao, Trung Hieu Nguyen 0001, Hao Wang 0199, Bin Ma 0001
INTERSPEECH4
2018 Learning Acoustic Word Embeddings with Temporal Context for Query-by-Example Speech Search
abstract
We propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search.The temporal context includes the leading and trailing word sequences of a word.We assume that there exist spoken word pairs in the training database.We pad the word pairs with their original temporal context to form fixed-length speech segment pairs.We obtain the acoustic word embeddings through a deep convolutional neural network (CNN) which is trained on the speech segment pairs with a triplet loss.By shifting a fixed-length analysis window through the search content, we obtain a running sequence of embeddings.In this way, searching for the spoken query is equivalent to the matching of acoustic word embeddings.The experiments show that our proposed acoustic word embeddings learned with temporal context are effective in QbE speech search.They outperform the state-of-the-art frame-level feature representations and reduce run-time computation since no dynamic time warping is required in QbE speech search.We also find that it is important to have sufficient speech segment pairs to train the deep CNN for effective acoustic word embeddings.
Yougen Yuan, Cheung-Chi Leung, Lei Xie 0001, Hongjie Chen 0001, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH5
2017 Multilingual bottle-neck feature learning from untranscribed speech
abstract
We propose to learn a low-dimensional feature representation for multiple languages without access to their manual transcription. The multilingual features are extracted from a shared bottleneck layer of a multi-task learning deep neural network which is trained using un-supervised phoneme-like labels. The unsupervised phoneme-like labels are obtained from language-dependent Dirichlet process Gaussian mixture models (DPGMMs). Vocal tract length normalization (VTLN) is applied to mel-frequency cepstral coefficients to reduce talker variation when DPGMMs are trained. The proposed features are evaluated using the ABX phoneme discriminability test in the Zero Resource Speech Challenge 2017. In the experiments, we show that the proposed features perform well across different languages, and they consistently outperform our previously proposed DPGMM posteriorgrams which topped the performance in the same challenge in 2015.
Hongjie Chen 0001, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
ASRU4
2017 Extracting bottleneck features and word-like pairs from untranscribed speech for feature representation
abstract
We propose a framework to learn a frame-level speech representation in a scenario where no manual transcription is available. Our framework is based on pairwise learning using bottleneck features (BNFs). Initial frame-level features are extracted from a bottleneck-shaped multilingual deep neural network (DNN) which is trained with unsupervised phoneme-like labels. Word-like pairs are discovered in the untranscribed speech using the initial features, and frame alignment is performed on each word-like speech pair. The matching frame pairs are used as input-output to train another DNN with the mean square error (MSE) loss function. The final frame-level features are extracted from an internal hidden layer of MSE-based DNN. Our pairwise learned feature representation is evaluated on the ZeroSpeech 2017 challenge. The experiments show that pairwise learning improves phoneme discrimination in 10s and 120s test conditions. We find that it is important to use BNFs as initial features when pairwise learning is performed. With more word pairs obtained from the Switchboard corpus and its manual transcription, the phoneme discrimination of three languages in the evaluation data can further be improved despite data mismatch.
Yougen Yuan, Cheung-Chi Leung, Lei Xie 0001, Hongjie Chen 0001, Bin Ma 0001, Haizhou Li 0001
ASRU5
2017 Adaptation of PLDA for multi-source text-independent speaker verification
abstract
Probabilistic linear discriminant analysis (PLDA) is widely described as an effective model for text-independent speaker verification in the i-vector space. The PLDA scoring function is typically formulated as the likelihood ratio between the speaker-adapted and the universal PLDAs. In this case, the adaptation of PLDA was performed through the speaker factors. In this paper, we show that the channel factors of the PLDA could be equivalently exploited to deal with the multi-source conditions. In speaker verification, with the proposed method, a PLDAmodel trained on conversational telephone speech could be adequately adapted for interview-style microphone recordings. Experimental results on NIST SRE'08 and SRE'10 datasets confirm that the proposed method is effective, especially for the case whereby enrollment and test utterances were captured from different sources.
Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP3
2017 Efficient methods to train multilingual bottleneck feature extractors for low resource keyword search
abstract
Training a bottleneck feature (BNF) extractor with multilingual data has been common in low resource keyword search. In a low resource application, the amount of transcribed target language data is limited while there are usually plenty of multilingual data. In this paper, we investigated two methods to train efficient multilingual BNF extractors for low resource keyword search. One method is to use the target language data to update an existing BNF extractor, and another method is to combine the target language data to train a new multilingual BNF extractor from the start. In these two methods, we proposed to use long short-term memory recurrent neural network based language identification to select utterances in the multilingual training data that are acoustically close to the target language. Experiments on Swahili in the OpenKWS15 data demonstrated the efficiency of our proposed methods. The first method facilitates rapid system development, while both methods outperform using baseline BNF extractors in terms of accuracy.
Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Nancy F. Chen, Bin Ma 0001
ICASSP5
2017 Modification on LSA speech enhancement for speech recognition
abstract
Speech recognition performance deteriorates in face of unknown noise. Speech enhancement offers a solution by reducing the noise in speech at runtime. However, it also introduces artificial distortions to the speech signals. In this paper, we aim at reducing the artifacts that has adverse effects on speech recognition. With this motivation, we propose a modification scheme including smoothing adaptation to frame SNR and reestimation of a priori SNR for spectral-domain log-spectral-amplitude (LSA) speech enhancement. The experiments show that the proposed scheme of enhancement significantly improves the performance of the state-of-the-art speech recognition over the baseline speech enhancement.
Chang Huai You, Bin Ma 0001, Chongjia Ni
ICASSP2
2017 Pairwise learning using multi-lingual bottleneck features for low-resource query-by-example spoken term detection
abstract
We propose to use a feature representation obtained by pairwise learning in a low-resource language for query-by-example spoken term detection (QbE-STD). We assume that word pairs identified by humans are available in the low-resource target language. The word pairs are parameterized by a multi-lingual bottleneck feature (BNF) extractor that is trained using transcribed data in high-resource languages. The multi-lingual BNFs of the word pairs are used as an initial feature representation to train an autoencoder (AE). We extract features from an internal hidden layer of the pairwise trained AE to perform acoustic pattern matching for QbE-STD. Our experiments on the TIMIT and Switchboard corpora show that the pairwise learning brings 7.61% and 8.75% relative improvements in mean average precision (MAP) respectively over the initial feature representation.
Yougen Yuan, Cheung-Chi Leung, Lei Xie 0001, Hongjie Chen 0001, Bin Ma 0001, Haizhou Li 0001
ICASSP5
2017 The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016
abstract
18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017
Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah
INTERSPEECH22
2017 An Integrated Solution for Snoring Sound Classification Using Bhattacharyya Distance Based GMM Supervectors with SVM, Feature Selection with Random Forest and Spectrogram with CNN
Tin Lay Nwe, Tran Huy Dat, Wen Zheng Terence Ng, Bin Ma 0001
INTERSPEECH4
2017 Multi-Task Learning for Mispronunciation Detection on Singapore Children's Mandarin Speech
Rong Tong, Nancy F. Chen, Bin Ma 0001
INTERSPEECH3
2017 Filtering for Malice Through the Data Ocean: Large-Scale PHA Install Detection at the Communication Service Provider Level
Kai Chen 0012, Tongxin Li 0002, Bin Ma 0001, Peng Wang 0088, XiaoFeng Wang 0001, Peiyuan Zong
RAID3
2017 Spectral-domain speech enhancement for speech recognition
Chang Huai You, Bin Ma 0001
Speech Commun.2
2017 Modeling Latent Topics and Temporal Distance for Story Segmentation of Broadcast News
abstract
This paper studies a strategy to model latent topics and temporal distance of text blocks for story segmentation, that we call graph regularization in topic modeling or GRTM. We propose two novel approaches that consider both temporal distance and lexical similarity of text blocks, collectively referred to as data proximity, in learning latent topic representation, where a graph regularizer is involved to derive the latent topic representation while preserving data proximity. In the first approach, we extend the idea of Laplacian probabilistic latent semantic analysis (LapPLSA) by introducing a distance penalty function in the affinity matrix of a graph for latent topic estimation. The estimated latent topic distributions are used to replace the traditional term-frequency vectors as the data representation of the text blocks and to measure the cohesive strength between them. In the second approach, we perform Laplacian eigenmaps, which makes use of the graph regularizer for dimensionality reduction, on latent topic distributions estimated by conventional topic modeling. We conduct the experiments on the automatic speech recognition transcripts of the TDT2 English broadcast news corpus. The experiments show the proposed strategy outperforms the conventional techniques. LapPLSA performs the best with the highest F1-measure of 0.816. The effects of the penalty constant in the distance penalty function, the number of latent topics, and the size of training data on the segmentation performances are also studied.
Hongjie Chen 0001, Lei Xie 0001, Cheung-Chi Leung, Xiaoming Lu, Bin Ma 0001, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 Content-aware local variability vector for speaker verification with short utterance
abstract
I-vector has shown to be very effective in speaker verification with long-duration speech utterances. But when test utterances are of short duration, content mismatch between the enrollment and test utterances limit the performance of i-vector system. This paper proposes to extract local session variability vectors on different phonetic classes from the utterances instead of estimating the session variability across the whole utterance as i-vector does. Using the posteriors given by a deep neural network (DNN) trained for phone state classification, the local vectors represent the session variability contained in specific phonetic content. Our experiments show that the content-aware local vectors are better at coping with the content mismatch between training and test utterances of short durations for text-independent, text-constrained and text-dependent tasks.
Kong-Aik Lee, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP4
2016 Exemplar-inspired strategies for low-resource spoken keyword search in Swahili
abstract
We present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples.
Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001
ICASSP11
2016 Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword search
abstract
In this paper, we propose a cross-lingual deep neural network (DNN) based submodular unbiased data selection approach for low-resource keyword search (KWS). A small amount (e.g. one hour) of transcribed data is used to conduct cross-lingual transfer. The frame-level senone sequence activated by the cross-lingual DNN is used to represent each untranscribed speech utterance. The proposed submodular function considers utterance length normalization and the feature distribution matched to a development set. Experiments are conducted by selecting 9 hours of Tamil speech for the 2014 NIST Open Keyword Search Evaluation (OpenKWS14). The proposed data selection approach provides 35.8% relative actual term weighted value (ATWV) improvement over random selection on the OpenKWS14 Evalpartl data set. Further analysis of the experimental results shows that both utterance length normalization and the feature distribution estimated from a development set deployed in the submodular function can suppress the preference to select long utterances. The selected utterances can cover a more diverse range of tri-phones, words, and acoustic variations from a wider set of utterances. Moreover, the wider coverage of words also benefits the acquired linguistic knowledge, which also contributes to improving KWS performance.
Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Feng Rao, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
ICASSP8
2016 Approximate search of audio queries by using DTW with phone time boundary and data augmentation
abstract
Dynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DTW is sensitive to the mismatch of signal conditions between the query and the speech search data. To allow approximate search, we propose a partial template matching strategy using phone time boundary information generated by a phone recognizer. To have more invariant representation of audio signals, we use bottleneck features (BNF) as the input of DTW. The BNF network is trained from augmented data, which is generated by adding reverberation and additive noises to the clean training data. Experimental results on QUESST 2015 task shows the effectiveness of the proposed methods for QbE-STD when the queries and search data are both distorted by reverberation and noises.
Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Cheung-Chi Leung, Lei Wang 0020, Van Hai Do, Hang Lv 0001, Lei Xie 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
ICASSP10
2016 Discriminatively trained joint speaker and environment representations for adaptation of deep neural network acoustic models
abstract
A recent trend in normalization of factors extraneous to a speech recognition task has been to explicitly introduce features related to the unwanted variability in the training of Deep Neural Networks (DNN). Typically, this is done by either perturbing the training set with models of these extraneous factors such as vocal tract length and environmental noise or augmenting the conventional spectral features with auxiliary information such as i-vector, noise spectrum, etc. Another emerging approach is to derive low dimensional representations of the factors from the hidden layers of DNN and use it for normalization of the acoustic model. Almost all of these approaches focus on either speaker or environment normalization. In this paper we propose a novel approach for estimating a compact joint representation of speakers and environment by training a DNN, with a bottleneck layer, to classify the i-vector features into speaker and environment labels by Multi-Task Learning (MTL). Another novelty is to learn this compact representation while learning to map the i-vector of a noisy utterance into its corresponding clean speaker i-vector and noise-only i-vector. Experiments were conducted on an artificially noise-corrupted version of the WSJ corpus. The proposed compact joint speaker-environment representations show promising gains.
Maofan Yin, Sunil Sivadas, Kai Yu 0004, Bin Ma 0001
ICASSP4
2016 Unsupervised Bottleneck Features for Low-Resource Query-by-Example Spoken Term Detection
Hongjie Chen 0001, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2016 SingaKids-Mandarin: Speech Corpus of Singaporean Children Speaking Mandarin Chinese
Nancy F. Chen, Rong Tong, Darren Wee, Pei Xuan Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH5
2016 The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMS
abstract
Technical report for NIST LRE 2015 Workshop
Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier
INTERSPEECH18
2016 Toward High-Performance Language-Independent Query-by-Example Spoken Term Detection for MediaEval 2015: Post-Evaluation Analysis
Cheung-Chi Leung, Lei Wang 0020, Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Hang Lv 0001, Lei Xie 0001, Chongjia Ni, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH10
2016 Rapid Update of Multilingual Deep Neural Network for Low-Resource Keyword Search
Chongjia Ni, Lei Wang 0020, Cheung-Chi Leung, Feng Rao, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH6
2016 Context Aware Mispronunciation Detection for Mandarin Pronunciation Training
Rong Tong, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH3
2016 Joint Speaker and Lexical Modeling for Short-Term Characterization of Speaker
Guangsen Wang, Kong-Aik Lee, Trung Hieu Nguyen 0001, Hanwu Sun, Bin Ma 0001
INTERSPEECH5
2016 Learning Neural Network Representations Using Cross-Lingual Bottleneck Features with Word-Pair Information
Yougen Yuan, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2016 Large-scale characterization of non-native Mandarin Chinese spoken by speakers of European origin: Analysis on iCALL
abstract
In this work, we analyze phonetic and prosodic pronunciation patterns from iCALL, a speech corpus designed to evaluate Mandarin mispronunciations by non-native speakers of European origin and to address the lack of large-scale, non-native corpora with comprehensive annotations for applications in CAPT (computer-assisted pronunciation training). iCALL consists of 90,841 utterances from 305 speakers with a total duration of 142 hours. The speakers are from diverse linguistic backgrounds (spanning Germanic, Romance, and Slavic native languages). The read utterances are phonetically balanced with phonetic, tonal, and fluency annotations. Our findings on iCALL reveal that lexical tone errors are over six times more prevalent than phonetic errors, French speakers are twice as likely to mispronounce Tone 2, 3, 4 when compared to English speakers, native Romance language speakers are more likely to make de-aspiration and aspiration mistakes, and fluency scores correlate inversely with tone and phone error rate.
Nancy F. Chen, Darren Wee, Rong Tong, Bin Ma 0001, Haizhou Li 0001
Speech Commun.4
2015 Channel adaptation of plda for text-independent speaker verification
abstract
Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling channel variability in the i-vector space for text-independent speaker verification. Speaker verification is a binary hypothesis testing. Given a test segment, the verification score could be computed as the log-likelihood ratio between a speaker-adapted PLDA and the universal PLDA model. This work proposes to infer the channel factor specific to each test segment and to include the channel estimate in the PLDA models, which essentially shifts the scoring function to better match that of the test channel. We also explore the influence of covariance adaptation in both speaker and channel adaptations. Experimental results on NIST SRE'08 and SRE'10 dataset confirm that the proposed channel adaptation can be effective when the covariance is kept un-adapted, while the covariance adaptation is necessary in the speaker adaptation.
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP3
2015 Low-resource keyword search strategies for tamil
abstract
We propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological nature of Tamil, we present highlights of our current KWS system, including: (1) Submodular optimization data selection to maximize acoustic diversity through Gaussian component indexed N-grams; (2) Keywordaware language modeling; (3) Subword modeling of morphemes and homophones.
Nancy F. Chen, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Van Tung Pham, Haihua Xu 0001, Tze Siong Lau, Su Jun Leow, Boon Pang Lim, Cheung-Chi Leung, Lei Wang 0020, Chin-Hui Lee 0001, Alvina Goh, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001
ICASSP16
2015 Unsupervised data selection and word-morph mixed language model for tamil low-resource keyword search
abstract
This paper considers an unsupervised data selection problem for the training data of an acoustic model and the vocabulary coverage of a keyword search system in low-resource settings. We propose to use Gaussian component index based n-grams as acoustic features in a submodular function for unsupervised data selection. The submodular function provides a near-optimal solution in terms of the objective being optimized. Moreover, to further resolve the high out-of-vocabulary (OOV) rate for morphologically-rich languages like Tamil, word-morph mixed language modeling is also considered. Our experiments are conducted on the Tamil speech provided by the IAPRA Babel program for the 2014 NIST Open Keyword Search Evaluation (OpenKWS14). We show that the selection of data plays an important role to the word error rate of the speech recognition system and the actual term weighted value (ATWV) of the keyword search system. The 10 hours of speech selected from the full language pack (FLP) using the proposed algorithm provides a relative 23.2% and 20.7% ATWV improvement over two other data subsets, the 10-hour data from the limited language pack (LLP) defined by IARPA and the 10 hours of speech randomly selected from the FLP, respectively. The proposed algorithm also increases the vocabulary coverage, implicitly alleviating the OOV problem: The number of OOV search terms drops from 1,686 and 1,171 in the two baseline conditions to 972.
Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Nancy F. Chen, Bin Ma 0001
ICASSP5
2015 Submodular data selection with acoustic and phonetic features for automatic speech recognition
abstract
In this paper, we propose to use acoustic feature based submodular function optimization to select a subset of untranscribed data for manual transcription, and retrain the initial acoustic model with the additional transcribed data. The acoustic features are obtained from an unsupervised Gaussian mixture model. We also integrate the acoustic features with the phonetic features, which are obtained from an initial ASR system, in the submodular function. Submodular function optimization has been theoretically shown its near-optimal guarantee. We performed the experiments on 1000 hours of Mandarin mobile phone speech, in which 300 hours of initial data was for the training of an initial acoustic model. The experimental results show that the acoustic feature based approach, which does not rely on an initial ASR system, performs as well as the phonetic feature based approach. Moreover, there is complementary effect between the acoustic feature based and the phonetic feature based data selection. The submodular function with the combined features provides a relative 4.8% character error rate (CER) reduction over the corresponding ASR system using random selection. We also include the desired feature distribution obtained from a development set in a generalized function, but the improvement is insignificant.
Chongjia Ni, Lei Wang 0020, Cheung-Chi Leung, Bin Ma 0001
ICASSP6
2015 A new study of GMM-SVM system for text-dependent speaker recognition
abstract
This paper presents a new approach and the study of GMM-SVM system for text-dependent speaker recognition on scenario of the fixed pass-phrases. The uniform-split content-based GMM-SVM system is proposed and applied to text-dependent speaker evaluation. We conducted detailed study of the proposed method compared to the baseline GMM-SVM system on the RSR2015 database, which has been designed and collected for the evaluation of text-dependent speaker verification system. The experiment results show that the new approach can significantly reduce the detection error of the target-wrong error type (i.e., target speaker with wrong pass-phrase) while maintaining a low detection error for both imposter-correct and imposter-wrong error types (i.e., imposter with correct pass-phrase and imposter with wrong pass-phrase). We also show that score normalization could be applied with respect to the imposter-wrong distribution as opposed to the imposter-correct distribution.
Hanwu Sun, Kong-Aik Lee, Bin Ma 0001
ICASSP3
2015 Tokenizing fundamental frequency variation for Mandarin tone error detection
abstract
Tone error is commonly observed in tonal language acquisition. Correct tone production is especially challenging for native speakers of non-tonal languages. In this paper, we exploit the fundamental frequency variation (FFV) feature for Mandarin tone error detection. We propose to use FFV through two approaches: (1) Concatenating FFVs along side with standard speech recognition features; (2) Token FFV: Characterizing pitch variation with longer temporal context through GMM tokenization and n-gram language modeling. Our results show that tone error detection improves by incorporating FFV features and the two approaches are complementary to each other.
Rong Tong, Nancy F. Chen, Boon Pang Lim, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2015 Language independent query-by-example spoken term detection using N-best phone sequences and partial matching
abstract
In this paper, we propose a partial sequence matching based symbolic search (SS) method for the task of language independent query-by-example spoken term detection. One main drawback of conventional SS approach is the high miss rate for long queries. This is due to high variations in symbol representation of query and search audios, especially in language independent scenario. The successful matching of a query with its instances in search audio becomes exponentially more difficult as the query grows longer. To reduce miss rate, we propose a partial matching strategy, in which all partial phone sequences of a query are used to search for query instances. The partial matching is also suitable for real life applications where exact match is usually not necessary and word prefix, suffix, and order should not affect the search result. When applied to the QUESST 2014 task, results show the partial matching of phone sequences is able to reduce miss rate of long queries significantly compared with conventional full matching method. In addition, for the most challenging inexact matching queries (type 3), it also shows clear advantage over DTW-based methods.
Haihua Xu 0001, Lei Xie 0001, Cheung-Chi Leung, Hongjie Chen 0001, Jia Yu 0002, Hang Lv 0001, Lei Wang 0020, Su Jun Leow, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
ICASSP11
2015 Phone-centric local variability vector for text-constrained speaker verification
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
INTERSPEECH3
2015 Parallel inference of dirichlet process Gaussian mixture models for unsupervised acoustic modeling: a feasibility study
abstract
We adopt a Dirichlet process Gaussian mixture model (DPGMM) for unsupervised acoustic modeling and represent speech frames with Gaussian posteriorgrams. The model performs unsupervised clustering on untranscribed data, and each Gaussian component can be considered as a cluster of sounds from various speakers. The model infers its model complexity (i.e. the number of Gaussian components) from the data. For computation efficiency, we use a parallel sampler for the model inference. Our experiments are conducted on the corpus provided by the zero resource speech challenge. Experimental results show that the unsupervised DPGMM posteriorgrams obviously outperformMFCC, and perform comparably to the posteriorgrams derived from language-mismatched phoneme recognizers in terms of the error rate of ABX discrimination test. The error rates can be further reduced by the fusion of these two kinds of posteriorgrams.
Hongjie Chen 0001, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2015 iCALL corpus: Mandarin Chinese spoken by non-native speakers of European descent
Nancy F. Chen, Rong Tong, Darren Wee, Pei Xuan Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH5
2015 The reddots data collection for speaker recognition
abstract
de niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Kong-Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David A. van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma 0001, Haizhou Li 0001, Themos Stafylakis, Jahangir Alam 0001, Albert Swart, Javier Perez
INTERSPEECH10
2015 The reddots platform for mobile crowd-sourcing of speech data
Kong-Aik Lee, Guangsen Wang, Kam Pheng Ng, Hanwu Sun, Trung Hieu Nguyen 0001, Ngoc Thuy Huong Thai, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH7
2015 Topic modeling for conference analytics
abstract
This work presents our attempt to understand the research topics that characterize the papers submitted to a conference, by using topic modeling and data visualization techniques. We infer the latent topics from the abstracts of all the papers submitted to Interspeech2014 by means of Latent Dirichlet Allocation. Pertopic word distributions thus obtained are visualized through word clouds. We also compare the automatically inferred topics against the expert-defined topics (also known as tracks for Interspeech2014). The comparison is based on an information retrieval framework, where we use each latent topic as a query and each track as a document. For each latent topic, we retrieve a ranked list of tracks scored by the degree of word overlap. Each latent topic is associated with the top-scoring track. This analytic procedure was applied to all submissions to Interspeech2014 and sheds some interesting light in terms of providing an overview of topic categorization in the conference, popular versus unpopular topics, emerging topics and topic compositions. Such insights are potentially valuable for understanding the technical content of a field and planning the future development of its conference(s).
Pengfei Liu 0004, Shoaib Jameel, Wai Lam, Bin Ma 0001, Helen M. Meng
INTERSPEECH4
2015 Phonology-augmented statistical transliteration for low-resource languages
Hoang Gia Ngo, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2015 Stress level detection using double-layer subband filter
Tin Lay Nwe, Qianli Xu, Cuntai Guan, Bin Ma 0001
INTERSPEECH4
2015 Joint environment and speaker normalization using factored front-end CMLLR
Shakti Rath, Sunil Sivadas, Bin Ma 0001
INTERSPEECH3
2015 Investigation of parametric rectified linear units for noise robust speech recognition
Sunil Sivadas, Zhenzhou Wu, Bin Ma 0001
INTERSPEECH3
2015 Goodness of tone (GOT) for non-native Mandarin tone recognition
Rong Tong, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH3
2015 Acoustic Segment Modeling with Spectral Clustering Methods
abstract
This paper presents a study of spectral clustering-based approaches to acoustic segment modeling (ASM). ASM aims at finding the underlying phoneme-like speech units and building the corresponding acoustic models in the unsupervised setting, where no prior linguistic knowledge and manual transcriptions are available. A typical ASM process involves three stages, namely initial segmentation, segment labeling, and iterative modeling. This work focuses on the improvement of segment labeling. Specifically, we use posterior features as the segment representations, and apply spectral clustering algorithms on the posterior representations. We propose a Gaussian component clustering (GCC) approach and a segment clustering (SC) approach. GCC applies spectral clustering on a set of Gaussian components, and SC applies spectral clustering on a large number of speech segments. Moreover, to exploit the complementary information of different posterior representations, a multiview segment clustering (MSC) approach is proposed. MSC simultaneously utilizes multiple posterior representations to cluster speech segments. To address the computational problem of spectral clustering in dealing with large numbers of speech segments, we use inner product similarity graph and make reformulations to avoid the explicit computation of the affinity matrix and Laplacian matrix. We carried out two sets of experiments for evaluation. First, we evaluated the ASM accuracy on the OGI-MTS dataset, and it was shown that our approach could yield 18.7% relative purity improvement and 15.1% relative NMI improvement compared with the baseline approach. Second, we examined the performances of our approaches in the real application of zero-resource query-by-example spoken term detection on SWS2012 dataset, and it was shown that our approaches could provide consistent improvement on four different testing scenarios with three evaluation metrics.
Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Minimum divergence estimation of speaker prior in multi-session PLDA scoring
abstract
Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling speaker and channel variability in the i-vector space for text-independent speaker verification. This paper shows that the PLDA scoring function could be formulated as model comparison between an adapted PLDA model and the universal PLDA. Based on this formulation, we show that a more robust adaptation could be attained by adapting the PLDA model through the use of minimum divergence estimate of speaker prior in the latent subspace. Experimental results on NIST SRE'10 and SRE'12 dataset confirm that the proposed method is effective in handling multi-session task. Notably, it is free from the covariance shrinkage problem typically found in the standard multi-session PLDA scoring.
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP3
2014 Strategies for Vietnamese keyword search
abstract
We propose strategies for a state-of-the-art Vietnamese keyword search (KWS) system developed at the Institute for Infocomm Research (I2R). The KWS system exploits acoustic features characterizing creaky voice quality peculiar to lexical tones in Vietnamese, a minimal-resource transliteration framework to alleviate out-of-vocabulary issues from foreign loan words, and a proposed system combination scheme FusionX. We show that the proposed creaky voice quality features complement pitch-related features, reaching fusion gains of 17.7% relative (6.9% absolute). To the best of our knowledge, the proposed transliteration framework is the first reported rule-based system for Vietnamese; it outperforms statistical-approach baselines up to 14.93–36.73% relative on foreign loan word search tasks. Using FusionX to combine 3 sub-systems, the actual term-weighted value (ATWV) reaches 0.4742, exceeding the ATWV=0.3 benchmark for IARPA Babel participants in the NIST OpenKWSB Evaluation.
Nancy F. Chen, Sunil Sivadas, Boon Pang Lim, Hoang Gia Ngo, Haihua Xu 0001, Van Tung Pham, Bin Ma 0001, Haizhou Li 0001
ICASSP7
2014 Modelling the alternative hypothesis for text-dependent speaker verification
abstract
This paper describes text-dependent speaker verification as a task involving four classes of trials depending on whether the target speaker or an impostor pronounces the expected pass-phrase or not. These four classes are used to reformulate the log-likelihood ratio traditionally used in text-independent speaker verification. Three formulations of the alternative hypothesis are considered, leading to three new expressions of the verification score. Experiments performed on the publicly available RSR2015 database show a significant improvement compared to existing baseline scores. A relative gain up to 61% in term of minimum cost is achieved when considering that the alternative hypothesis is the union of three sub-hypotheses corresponding to the three existing classes of impostures.
Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
ICASSP3
2014 Imposture classification for text-dependent speaker verification
abstract
This work focuses on text-dependent speaker verification, where a user is required to chose and pronounce a customized pass-phrase to get authenticated. In this context, there are three types of impostures: an impostor pronouncing the correct pass-phrase, an impostor pronouncing a wrong pass-phrase and the most difficult one: an impostor playing back a recording of the target speaker pronouncing a wrong pass-phrase. Detecting and classifying different types of impostures can help to prevent future impostures of the same type. In this work, we first propose a new verification score to reject Playback impostures. This score allows a relative reduction of 90% of the equal error rate against Playback impostures while offering performance similar to the baseline text-dependent score against other types of impostures. As a second contribution, we show that the new score can be combined with an existing text-dependent verification score to improve the classification of the different types of impostures. The performance of the speaker verification engine for imposture classification is significantly improved with the Cllrdecreasing by at least 29% compared to the original system.
Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
ICASSP3
2014 Subspace Gaussian mixture model for computer-assisted language learning
abstract
In computer-assisted language learning (CALL), speech data from non-native speakers are usually insufficient for acoustic modeling. Subspace Gaussian Mixture Models (SGMM) have been effective in training automatic speech recognition (ASR) systems with limited amounts of training data. Therefore, in this work, we propose to use SGMM to improve the fluency assessment performance. In particular, the contributions of this work are: (i) The proposed SGMM acoustic model trained with native data outperforms the MMI-GMM/HMM baseline by 25% relative, (ii) when incorporating a small amount of non-native training data, the SGMM acoustic model further improves the performance of fluency assessment by 47% relative.
Rong Tong, Boon Pang Lim, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2014 Extended RSR2015 for text-dependent speaker verification over VHF channel
abstract
International audience
Anthony Larcher, Kong-Aik Lee, Pablo Luis Sordo Martinez, Trung Hieu Nguyen 0001, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH5
2014 A whispered Mandarin corpus for speech technology applications
Pei Xuan Lee, Darren Wee, Hilary Si Yin Toh, Boon Pang Lim, Nancy F. Chen, Bin Ma 0001
INTERSPEECH6
2014 A minimal-resource transliteration framework for vietnamese
Hoang Gia Ngo, Nancy F. Chen, Sunil Sivadas, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2014 On the use of Bhattacharyya based GMM distance and neural net features for identification of cognitive load levels
abstract
This paper presents a method for detecting cognitive load levels from speech. When speech is modulated by different lev-els of cognitive load, acoustic characteristics of speech change. In this paper, we measure acoustic distance of a stressed ut-terance from the baseline stress free speech using GMM-SVM kernel with Bhattacharyya based GMM distance. In addition, it is believed that airflow structure of speech production is non-linear. This motivates us to investigate better techniques to cap-ture nonlinear characteristic of stress information in acoustic features. Inspired by the recent success of neural networks for representation learning, we employ a single hidden layer feed forward network with non-linear activation to extract the fea-ture vectors. Furthermore, people have different reactions to a particular task load. This inter-speaker difference in stress re-sponses presents a major challenge for stress level detection. We use a bootstrapped training process to learn the stress re-sponse of a particular speaker. We perform experiments using data sets from Cognitive Load with Speech and EGG (CLSE) provided for the Cognitive Load Sub-Challenge of the INTER-SPEECH 2014 Computational Paralinguistics Challenge. The results show that the system with our proposed strategies per-forms well on validation and test sets. Index Terms: cognitive load, GMM-supervector, neural net features
Tin Lay Nwe, Trung Hieu Nguyen 0001, Bin Ma 0001
INTERSPEECH3
2014 The NIST SRE summed channel speaker recognition system
Hanwu Sun, Bin Ma 0001
INTERSPEECH2
2014 Virtual example for phonotactic language recognition
Rong Tong, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH2
2014 A graph-based Gaussian component clustering approach to unsupervised acoustic modeling
Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2014 Intrinsic spectral analysis based on temporal context features for query-by-example spoken term detection
abstract
We investigate the use of intrinsic spectral analysis (ISA) for query-by-example spoken term detection (QbE-STD). In the task, spoken queries and test utterances in an audio archive are converted to ISA features, and dynamic time warping is applied to match the feature sequence in each query with those in test utterances. Motivated by manifold learning, ISA has been pro-posed to recover from untranscribed utterances a set of nonlin-ear basis functions for the speech manifold, and shown with improved phonetic separability and inherent speaker indepen-dence. Due to the coarticulation phenomenon in speech, we propose to use temporal context information to obtain the ISA features. Gaussian posteriorgram, as an efficient acoustic rep-resentation usually used in QbE-STD, is considered a baseline feature. Experimental results on the TIMIT speech corpus show that the ISA features can provide a relative 13.5 % improvement in mean average precision over the baseline features, when the temporal context information is used. Index Terms: spoken term detection, intrinsic spectral analysis, Gaussian posteriorgram, dynamic time warping 1.
Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2014 How We Found These Vulnerabilities in Android Applications
Bin Ma 0001
SecureComm (2)1
2014 Text-dependent speaker verification: Classifiers, databases and RSR2015
abstract
The RSR2015 database, designed to evaluate text-dependent speaker verification systems under different durations and lexical constraints has been collected and released by the Human Language Technology (HLT) department at Institute for Infocomm Research (I2R) in Singapore. English speakers were recorded with a balanced diversity of accents commonly found in Singapore. More than 151 h of speech data were recorded using mobile devices. The pool of speakers consists of 300 participants (143 female and 157 male speakers) between 17 and 42 years old making the RSR2015 database one of the largest publicly available database targeted for text-dependent speaker verification. We provide evaluation protocol for each of the three parts of the database, together with the results of two speaker verification system: the HiLAM system, based on a three layer acoustic architecture, and an i-vector/PLDA system. We thus provide a reference evaluation scheme and a reference performance on RSR2015 database to the research community. The HiLAM outperforms the state-of-the-art i-vector system in most of the scenarios.
Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
Speech Commun.3
2013 Minimal-resource phonetic language models to summarize untranscribed speech
abstract
We propose to extract summary sentences from lexically untranscribed speech via phone tokenization. We use decoded phone sequences instead of words to train language models to infer semantically significant utterances. Phone tokens yield comparable results to words on the TDT-2 English corpus, yet require significantly less linguistic resources - no need for automatic speech recognition (ASR): (1) Using decoded phones of high phone error rate (78.7%) leads to comparable results to using ASR-decoded words. (2) Tokenizing English audio using a Czech phone recognizer leads to comparable results to using English words from closed-captions. These trends parallel those established in spoken language recognition and have practical significance: we can potentially summarize speech passages of resource-poor languages by leveraging existing tools developed on resource-rich languages.
Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
ICASSP2
2013 Speaker clustering using vector representation with long-term feature for lecture speech recognition
abstract
Speaker clustering has been widely adopted for clustering the speech data based on acoustic characteristics so that an unsupervised speaker normalization and speaker adaptive training can be applied for a better speech recognition performance. In this study, we present a vector space speaker clustering approach with long-term feature analysis. The supervector based on the GMM mean vectors is adopted to represent the characteristics of speakers. To achieve a robust representation, total variability subspace modeling, which has been successfully applied in speaker recognition for compensating channel and session variability over the GMM mean supervector, is used for speaker clustering. We apply a long-term feature analysis strategy to average short-time spectral features over a period of time to capture the speaker traits that are manifested over a speech segment longer than a spectral frame. Experiments conducted on lecture style speech show that this speaker clustering approach offers a better speech recognition performance.
Chien-Lin Huang, Chiori Hori, Hideki Kashioka, Bin Ma 0001
ICASSP4
2013 Joint analysis of vocal tract length and temporal information for robust speech recognition
abstract
This paper presents a joint analysis approach to address the acoustic feature normalization for robust speech recognition. The variations in acoustic environments and speakers are the major challenge for speech recognition. The conventional normalizations of these two variations are separately processed, applying the speaker normalization with an assumption of a noise free condition and applying the noise compensation with an assumption of speaker independency, and thus resulting in a suboptimal performance. The proposed joint analysis approach simultaneously considers the vocal tract length normalization and averaged temporal information of cepstral features. In a data-driven manner, the Gaussian mixture model is used to estimate the conditional parameters in the joint analysis. Experimental results show that the proposed approach achieves a substantial improvement.
Chien-Lin Huang, Chiori Hori, Hideki Kashioka, Bin Ma 0001
ICASSP4
2013 Phonetically-constrained PLDA modeling for text-dependent speaker verification with multiple short utterances
abstract
The importance of phonetic variability for short duration speaker verification is widely acknowledged. This paper assesses the performance of Probabilistic Linear Discriminant Analysis (PLDA) and i-vector normalization for a text-dependent verification task. We show that using a class definition based on both speaker and phonetic content significantly improves the performance of a state-of-the-art system. We also compare four models for computing the verification scores using multiple enrollment utterances and show that using PLDA intrinsic scoring obtains the best performance in this context. This study suggests that such scoring regime remains to be optimized.
Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
ICASSP3
2013 Broadcast news story segmentation using latent topics on data manifold
abstract
This paper proposes to use Laplacian Probabilistic Latent Semantic Analysis (LapPLSA) for broadcast news story segmentation. The latent topic distributions estimated by LapPLSA are used to replace term frequency vector as the representation of sentences and measure the cohesive strength between the sentences. Subword n-gram is used as the basic term unit in the computation. Dynamic Programming is used for story boundary detection. LapPLSA projects the data into a low-dimensional semantic topic representation while preserving the intrinsic local geometric structure of the data. The locality preserving property attempts to make the estimated latent topic distributions more robust to the noise from automatic speech recognition errors. Experiments are conducted on the ASR transcripts of TDT2 Mandarin broadcast news corpus. Our proposed approach is compared with other approaches which use dimensionality reduction technique with the locality preserving property, and two different topic modeling techniques. Experiment results show that our proposed approach provides the highest F1-measure of 0.8228, which significantly outperforms the best previous approaches.
Xiaoming Lu, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2013 Anti-model KL-SVM-NAP system for NIST SRE 2012 evaluation
abstract
This paper presents an anti-model based speaker recognition system for NIST SRE 2012 evaluation, which is one of subsystems in IIR SRE12 submission. We apply the anti-model approach for the SRE12 evaluation. The KL-SVM-NAP based speaker recognition system is adopted to evaluate the performance. We present detailed comparison study of the classical KL-SVM-NAP based speaker recognition system and anti-model based KL-SVM-NAP system for NIST 2012 speaker recognition evaluation. The results are reported on in-house pre-SRE12 development set and NIST SRE12 core task. The clear advantages of the anti-model approach over that the traditional KL-SVM-NAP approach are presented and discussed.
Hanwu Sun, Kong-Aik Lee, Bin Ma 0001
ICASSP3
2013 Using parallel tokenizers with DTW matrix combination for low-resource spoken term detection
abstract
Recently the posteriorgram-based template matching framework has been successfully applied to query-by-example spoken term detection tasks for low-resource languages. This framework employs a tokenizer to derive posteriorgrams, and applies dynamic time warping (DTW) to the posteriorgrams to locate the possible occurrences of a query term. Based on this framework, we propose to improve the detection performance by using multiple tokenizers with DTW distance matrix combination. The proposed approach uses multiple tokenizers in parallel as the front-end to generate different posteriorgram representations, and combines the distance matrices of the different posteriorgrams into a single matrix. DTW detection is then applied to the combined distance matrix. Lastly score post-processing techniques including pseudo-relevance feedback and score normalization are used for further improvement. Experiments were conducted on the spoken web search datasets of MediaEval 2011 and MediaEval 2012. Experimental results show that combining multiple tokenizers significantly outperforms the best single tokenizer, and that the DTW matrix combination method consistently outperforms the score combination method when more than three tokenizers are involved. Score post-processing techniques show further gains on top of using multiple tokenizers.
Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2013 A study on GMM-SVM with adaptive relevance factor and its comparison with i-vector and JFA for speaker recognition
abstract
Recently, joint factor analysis (JFA) and identity-vector (i-vector) represent the dominant techniques used for speaker recognition due to their superior performance. Developed relatively earlier, the Gaussian mixture model - support vector machine (GMM-SVM) with nuisance attribute projection (NAP) has gradually become less popular. However, when developing the relevance factor in maximum a posteriori (MAP) estimation of GMM to be adapted by application data in place of the conventional fixed value, it is noted that GMM-SVM demonstrates some advantages. In this paper, we conduct a comparative study between GMM-SVM with adaptive relevance factor and JFA/i-vector under the framework of Speaker Recognition Evaluation (SRE) formulated by the National Institute of Standards and Technology (NIST).
Chang Huai You, Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee
ICASSP3
2013 Large-scale characterization of Mandarin pronunciation errors made by native speakers of European languages
Nancy F. Chen, Vivaek Shivakumar, Mahesh Harikumar, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2013 Multi-session PLDA scoring of i-vector for partially open-set speaker detection
abstract
International audience
Kong-Aik Lee, Anthony Larcher, Chang Huai You, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2013 Improved unsupervised NAP training dataset design for speaker recognition
Hanwu Sun, Bin Ma 0001
INTERSPEECH2
2013 Unsupervised mining of acoustic subword units with segment-level Gaussian posteriorgrams
Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2013 I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verification
abstract
I4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort.
Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah
INTERSPEECH28
2013 Spoken Language Recognition: From Fundamentals to Practice
abstract
Spoken language recognition refers to the automatic process through which we determine or verify the identity of the language spoken in a speech sample. We study a computational framework that allows such a decision to be made in a quantitative manner. In recent decades, we have made tremendous progress in spoken language recognition, which benefited from technological breakthroughs in related areas, such as signal processing, pattern recognition, cognitive science, and machine learning. In this paper, we attempt to provide an introductory tutorial on the fundamentals of the theory and the state-of-the-art solutions, from both phonological and computational aspects. We also give a comprehensive review of current trends and future research directions using the language recognition evaluation (LRE) formulated by the National Institute of Standards and Technology (NIST) as the case studies.
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee
Proc. IEEE2
2013 Shifted-Delta MLP Features for Spoken Language Recognition
abstract
This letter presents our study of applying phoneme posterior features for spoken language recognition (SLR). In our work, phoneme posterior features are estimated from a multilayer perceptron (MLP) based phoneme recognizer, and are further processed through transformations including taking logarithm, PCA transformation, and appending shifted delta coefficients. The resulting shifted-delta MLP (SDMLP) features show similar distribution as conventional shifted-delta cepstral (SDC) features, and are more robust compared to the SDC features. Experiments on the NIST LRE2005 dataset show that the SDMLP features fit well with the state-of-the-art GMM-based SLR systems, and SDMLP features outperform SDC features significantly.
Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001
IEEE Signal Process. Lett.4
2013 Sparse Classifier Fusion for Speaker Verification
abstract
State-of-the-art speaker verification systems take advantage of a number of complementary base classifiers by fusing them to arrive at reliable verification decisions. In speaker verification, fusion is typically implemented as a weighted linear combination of the base classifier scores, where the combination weights are estimated using a logistic regression model. An alternative way for fusion is to use classifier ensemble selection, which can be seen as sparse regularization applied to logistic regression. Even though score fusion has been extensively studied in speaker verification, classifier ensemble selection is much less studied. In this study, we extensively study a sparse classifier fusion on a collection of twelve I4U spectral subsystems on the NIST 2008 and 2010 speaker recognition evaluation (SRE) corpora.
Ville Hautamäki, Tomi Kinnunen, Filip Sedlak, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
IEEE Trans. Speech Audio Process.5
2013 Spoken Language Recognition With Prosodic Features
abstract
Speech prosody is believed to carry much language-specific information that can be used for spoken language recognition (SLR). In the past, the use of prosodic features for SLR has been studied sporadically and the reported performances were considered unsatisfactory. In this paper, we exploit a wide range of prosodic attributes for large-scale SLR tasks. These attributes describe the multifaceted variations of F0, intensity and duration in different spoken languages. Prosodic attributes are modeled by the bag of n-grams approach with support vector machine (SVM) as in the conventional phonotactic SLR systems. Experimental results on OGI and NIST-LRE tasks showed that the use of proposed attributes gives significantly better SLR performance than those previously reported. The full feature set includes 87 prosodic attributes and redundancy among attributes may exist. Attributes are broken down into particular bigrams called bins. Four entropy-based feature selection metrics with different selection criteria are derived. Attributes can be selected by individual bins, or by attributes as batches of bins. It can also be done in a language-dependent or language-independent manner. By comparing different selection sizes and criteria, an optimal attribute subset comprising 5,000 bins is found by using a bin-level language-independent criterion. Feature selection reduces model size by 2.5 times and shortens the runtime by 6 times. The optimal subset of bins gives the lowest EER of 20.18% on NIST-LRE 2007 SLR task in a prosodic attribute model (PAM) system which exclusively modeled prosodic attributes. In a phonotactic-prosodic fusion SLR system, the detection cost, Cavgis 2.09%. The relative detection cost reduction is 23%.
Raymond W. M. Ng, Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001
IEEE Trans. Speech Audio Process.4
2012 An acoustic segment modeling approach to query-by-example spoken term detection
abstract
The framework of posteriorgram-based template matching has been shown to be successful for query-by-example spoken term detection (STD). This framework employs a tokenizer to convert query examples and test utterances into frame-level posteriorgrams, and applies dynamic time warping to match the query posteriorgrams with test posteriorgrams to locate possible occurrences of the query term. It is not trivial to design a reliable tokenizer due to heterogeneous test conditions and the limitation of training resources. This paper presents a study of using acoustic segment models (ASMs) as the tokenizer. ASMs can be obtained following an unsupervised iterative procedure without any training transcriptions. The STD performance of the ASM tokenizer is evaluated on Fisher Corpus with comparison to three alternative tokenizers. Experimental results show that the ASM tokenizer outperforms a conventional GMM tokenizer and a language-mismatched phoneme recognizer. In addition, the performance is significantly improved by applying unsupervised speaker normalization techniques.
Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2012 Acoustic TextTiling for story segmentation of spoken documents
abstract
We propose an acoustic TextTiling method based on segmental dynamic time warping for automatic story segmentation of spoken documents. Different from most of the existing methods using LVCSR transcripts, this method detects story boundaries directly from audio streams. In analogy to the cosine-based lexical similarity between two text blocks in a transcript, we define the acoustic similarity measure between two pseudo-sentences in an audio stream. Experiments on TDT2 Mandarin corpus show that acoustic TextTiling can achieve comparable performance to lexical TextTiling based on LVCSR transcripts. Moreover, we use MFCCs and Gaussian posteriorgrams as the acoustic representations in our experiments. Our experiments show that Gaussian posteriorgrams are more robust to perform segmentation for the stories each with multiple speakers.
Lilei Zheng, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2012 Ensemble Classifiers Using Unsupervised Data Selection for Speaker Recognition
Chien-Lin Huang, Chiori Hori, Hideki Kashioka, Bin Ma 0001
INTERSPEECH4
2012 PLDA Modeling in I-Vector and Supervector Space for Speaker Verification
abstract
In this paper, we advocate the use of uncompressed form of i-vector. We employ the probabilistic linear discriminant analysis (PLDA) to handle speaker and session variability for speaker verification task. An i-vector is a low-dimensional vector containing both speaker and channel information acquired from a speech segment. When PLDA is used on i-vector, dimension reduction is performed twice – first in the i-vector extraction process and second in the PLDA model. Keeping the full dimensionality of i-vector in the supervector space for PLDA modeling and scoring would avoid unnecessary loss of information. The drawback of using PLDA on uncompressed i-vector is the inversion of large matrices, which we show can be solved rather efficiently by portioning large matrix into smaller blocks. We also introduce the Gaussianized rank-norm, as an alternative to whitening, for feature normalization prior to PLDA modeling. Index Terms: speaker verification, i-vector, probabilistic LDA 1.
Kong-Aik Lee, Zhenmin Tang, Bin Ma 0001, Anthony Larcher, Haizhou Li 0001
INTERSPEECH4
2012 RSR2015: Database for Text-Dependent Speaker Verification using Multiple Pass-Phrases
abstract
International audience
Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH3
2012 Unsupervised NAP Training Data Design for Speaker Recognition
Hanwu Sun, Bin Ma 0001
INTERSPEECH2
2012 Effect of Relevance Factor of Maximum a posteriori Adaptation for GMM-SVM in Speaker and Language Recognition
Chang Huai You, Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee
INTERSPEECH3
2012 Discriminative feature extraction for speech recognition using continuous output codes
Omid Dehzangi, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
Pattern Recognit. Lett.2
2012 Speaker Clustering and Cluster Purification Methods for RT07 and RT09 Evaluation Meeting Data
abstract
This paper presents a design strategy for the speaker diarization system in the IIR submissions to the 2007 and 2009 NIST Rich Transcription Meeting Recognition Evaluations (RT07 and RT09) for the multiple distant microphone (MDM) condition. The system features two algorithms supporting two important steps in a diarization process. The first step is Initial Segmentation and Clustering (ISC), and the second one is cluster merging and purification. In the ISC step, we propose a histogram quantization and clustering technique based on time delay of arrival (TDOA) features by analyzing the correlation among the signals across multiple distant microphones. In the cluster merging and purification step, we further merge the speaker clusters using a Bayesian information criterion (BIC) to consolidate the clusters to arrive at one-cluster-per-speaker. The two steps work in tandem to form an integral process. We propose a novel Consensus Based Cluster Purification (CBCP) method that involves a technique to remove impure speaker segments in the speaker clusters before speaker modeling in the cluster purification process. The system reports a state-of-the-art performance of speaker diarization for RT07 and RT09 MDM condition with 7.47% and 8.77% Diarization error rates (DERs), respectively, for both overlapping and non-overlapping speech.
Tin Lay Nwe, Hanwu Sun, Bin Ma 0001, Haizhou Li 0001
IEEE Trans. Speech Audio Process.3
2011 Score fusion and calibration in multiple language detectors with large performance variation
abstract
In a large-scale language detection task, performance variation found between different component systems and different target languages has an adverse effect to the pooled error statistics. Special care has to be taken in score fusion and calibration. In this paper, we use a prosodic LID system to fuse with a phonotactic LID system using NIST Language Recognition Evaluation 2009 experimental data. Among four logistic regression models, the one which gives the lowest Cavg is chosen. We further explore our previously proposed calibration algorithm based on the minimum erroneous deviation criterion. The algorithm is made more robust by removing the predetermined list of target languages to be calibrated, as well as by adding an optimization constraint which enforces calibration in the data portion with a large performance variation. The fusion and calibration operations together bring a 33.9% relative Cavg reduction compared with the original result from a phonotactic LID system.
Raymond W. M. Ng, Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2011 Factored covariance modeling for text-independent speaker verification
abstract
Gaussian mixture models (GMMs) are commonly used to model the spectral distribution of speech signals for text-independent speaker verification. Mean vectors of the GMM, used in conjunction with support vector machine (SVM), have shown to be effective in characterizing speaker information. In addition to the mean vectors, covariance matrices capture the correlation between spectral features, which also represent some salient information about speaker identity. This paper investigates the use of local correlation between different dimensions of acoustic vector by using factor analysis and linear Gaussian model. Log-Euclidean inner product kernel is used to measure the similarity between two speech utterances in the form of covariance matrices. Experiments carried on NIST 2006 speaker verification tasks shows promising results.
Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001
ICASSP3
2011 Regularized Logistic Regression Fusion for Speaker Verification
abstract
Fusion of the base classifiers is seen as the way to achieve stateof-the art performance in the speaker verfication systems. Standard approach is to pose the fusion problem as the linear binary classification task. Most successful loss function in speaker verification fusion has been the weighted logistic regression popularized by the FoCal toolkit. However, it is known that optimizing logistic regression can overfit severely without appropriate regularization. In addition, subset classifier selection can be achieved by using an external 0/1 loss function on the best subset. In this work, we propose to use LASSO based regularization on the FoCal cost function to achive improved performance and classifier subset selection method integrated into one optimization task. Proposed method is able to achieve 51 % relative improvement in Actual DCF over the FoCal baseline. Index Terms: logistic regression, regularization, compressed sensing, linear fusion, speaker verification
Ville Hautamäki, Kong-Aik Lee, Tomi Kinnunen, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2011 Maximum Entropy Based Data Selection for Speaker Recognition
Chien-Lin Huang, Bin Ma 0001
INTERSPEECH2
2011 Speech Indexing Using Semantic Context Inference
abstract
Abstract This study presents a novel approach to spoken document retrieval based on semantic context inference for speech indexing. Each recognized term in a spoken document is mapped onto a semantic inference vector containing a bag of semantic terms through a semantic relation matrix. The semantic context inference vector is then constructed by summing up all the semantic inference vectors. Such a semantic term expansion and re-weighting make the semantic context inference vector a suitable representation for speech indexing. The experiments were conducted on 1550 anchor news stories collected from Mandarin Chinese broadcast news of 198 hours. The experimental results indicate that the proposed speech indexing using the semantic context inference contributes to a substantial performance improvement of spoken document retrieval. Index Terms : speech indexing, semantic context inference, spoken document retrieval 1. Introduction Speech is the most convenient way for the interaction of human-to-human and human-to-machine. The applications of spoken document retrieval in education, business and entertainment are rapidly growing. The recent attempts include multilingual oral history archives access [1], MIT lecture browsing [2], and the management of National Gallery consisting of speeches, news broadcasts and recordings [3], voice search about spoken dialog, call-routing systems [4], etc. All of them focus on retrieving the information to meet users' requirements. We know that it is not straightforward to directly compare the speech query with the spoken documents in the database. In order to construct an efficient and effective retrieval system, the state-of-the-art spoken document retrieval (SDR) technologies adopt the transcription obtained from automatic speech recognition for indexing. Vector space model [5] and probabilistic models (HMM [6], GMM [7], KL-divergence [8]), rely on certain similarity functions that assume a document is more likely to be relevant to a query if it contains more occurrences of query terms. The indexing techniques of text-based information retrieval have been widely adopted in spoken document retrieval. However, due to imperfect speech recognition results, out-of-vocabulary, and the ambiguity in homophone and word tokenization, conventional text-based indexing techniques are not always appropriate for spoken document retrieval. The transcription errors may cause undesired semantic and syntactic expression, thus result in an inadequate indexing. Several approaches have been proposed to address these problems with various indexing units such as word, sub-word, phone, and so on. The multi-level knowledge indexing approach considers three information sources including the speech transcription, keywords extracted from spoken documents, and hypernyms of the extracted keywords [9]. Hui et al. applied the
Chien-Lin Huang, Bin Ma 0001, Haizhou Li 0001, Chung-Hsien Wu 0001
INTERSPEECH2
2011 Joint Application of Speech and Speaker Recognition for Automation and Security in Smart Home
Kong-Aik Lee, Anthony Larcher, Helen Thai, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2011 Probabilistic Latent Semantic Analysis for Broadcast News Story Segmentation
abstract
This paper proposes to perform probabilistic latent semantic analysis (PLSA) for broadcast news (BN) story segmentation. PLSA exploits a deeper underlying relation among terms be-yond their occurrences thus conceptual matching can be em-ployed to replace literal term matching. Different from text seg-mentation, lexical based BN story segmentation has to be car-ried out over LVCSR transcripts, where the incorrect recogni-tion of out-of-vocabulary words inevitably impacts the seman-tic relation. We use phoneme subwords as the basic term units to address this problem. We integrate a cross entropy mea-surement with PLSA to depict lexical cohesion and compare its performance with the widely used cosine similarity metric. Furthermore, we evaluate two approaches, namely TextTiling and dynamic programming (DP), for story boundary identifica-tion. Experimental results show that the PLSA based methods bring a significant performance boost to story segmentation and the cross entropy based DP approach provides the best perfor-mance. Index Terms: story segmentation, probabilistic latent semantic analysis, cross entropy, dynamic programming, spoken docu-ment retrieval 1.
Mimi Lu, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2011 Study of Overlapped Speech Detection for NIST SRE Summed Channel Speaker Recognition
Hanwu Sun, Bin Ma 0001
INTERSPEECH2
2011 Target-Aware Lattice Rescoring for Dialect Recognition
Rong Tong, Bin Ma 0001, Haizhou Li 0001, Chng Eng Siong
INTERSPEECH2
2011 Speaker Verification With Feature-Space MAPLR Parameters
abstract
This paper studies a new technique that characterizes a speaker by the difference between the speaker and a cohort of background speakers in the form of feature-space maximum a posteriori linear regression (fMAPLR). The fMAPLR is a linear regression function that projects speaker dependent features to speaker independent ones, also known as an affine transform. It consists of two sets of parameters, bias vectors and transform matrices. The former, representing the first order information, is more robust than the latter, the second-order information. We propose a flexible tying scheme that allows the bias vectors and the matrices to be associated with different regression classes, such that both parameters are given sufficient statistics in a speaker verification task. We formulate a maximum a posteriori (MAP) algorithm for the estimation of feature transform parameters, that further alleviates the possible numerical problem. The fMAPLR parameters are then vectorized and compared via a support vector machine (SVM). We conduct the experiments on National Institute of Standards and Technology (NIST) 2006 and 2008 Speaker Recognition Evaluation databases. The experiments show that the proposed technique consistently outperforms the baseline Gaussian mixture model (GMM)-SVM speaker verification system.
Donglai Zhu, Bin Ma 0001, Haizhou Li 0001
IEEE Trans. Speech Audio Process.2
2010 Semi-supervised learning of language model using unsupervised topic model
abstract
We present a semi-supervised learning (SSL) method for building domain-specific language models (LMs) from general-domain data using probabilistic latent semantic analysis (PLSA). The proposed technique first performs topic decomposition (TD) on the combined dataset of domain-specific and general-domain data. Then it derives latent topic distribution of the interested domain, and derives domain-specific word n-gram counts with a PLSA style mixture model. Finally, it uses traditional n-gram modeling to construct domain-specific LMs from the domain-specific word n-gram counts. Experimental results show that this technique outperforms both states-of-the-art relative entropy text selection and traditional supervised training methods.
Shuanhu Bai, Chien-Lin Huang, Bin Ma 0001, Haizhou Li 0001
ICASSP3
2010 Error corrective classifier fusion for spoken Language Recognition
abstract
A number of effective classification algorithms have been developed for spoken language recognition, and it has been a common practice in the NIST Language Recognition Evaluations (LREs) that an information fusion is applied to boost the performance of the recognition system. This paper investigates the fusion of multiple output scores generated using different classifiers that complement to further reduce the classification error rate in spoken language recognition. We introduce a local performance metric to optimize the performance of the classifier fusion. The experiments are conducted on the 2009 NIST LRE corpus. The experimental results show that the proposed fusion effectively improves the performance over individual classifiers.
Omid Dehzangi, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
ICASSP2
2010 Prosodic attribute model for spoken language identification
abstract
Prosodic information is believed to carry language-specific information useful to spoken language recognition. Modeling prosodic features is a challenging problem, on which a wide diversity of approaches have been investigated. In this paper, a novel prosodic attribute model (PAM) is proposed to capture prosodic features with compact models. It models the language-specific co-occurrence statistics of a comprehensive set of prosodic features. When the prosodic LID system with PAM is evaluated in NIST Language Recognition Evaluations (LRE) 2007 and 2009, it demonstrates respectively 21% and 11% relative EER reduction compared to a phonotactic LID system. The contributions of prosodic features in detecting some of the target languages, including tonal languages, are even more substantial. It is also noted that most prosodic attributes in the comprehensive set are making positive contributions.
Raymond W. M. Ng, Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001
ICASSP4
2010 Speaker diarization system for RT07 and RT09 meeting room audio
abstract
This paper describes an improved speaker diarization system for the Single Distant Microphone (SDM) task in the 2007 and 2009 NIST Rich Transcription Meeting Recognition Evaluations. The system includes three main modules: front-end processing, initial speaker clustering and cluster purification/merging. The front-end processing involves the Wiener filtering for the targeted audio channels and a self-adaptation speech activity detection algorithm. A simple but effective energy based segmentation is applied to chunk the meeting data into small segments to construct the initial clusters. An enhanced purification algorithm is proposed to further improve the performance after the preliminary purification, and the BIC criterion is adopted for the cluster merging. The system achieves competitive overall DERs of 15.67% for RT07 SDM speaker diarization task and 17.34% for RT09 SDM speaker diarization task.
Hanwu Sun, Bin Ma 0001, Swe Zin Kalayar Khine, Haizhou Li 0001
ICASSP2
2010 Soft margin estimation of Gaussian mixture model parameters for spoken language recognition
abstract
This paper extends our previous work on large margin estimation (LME) of GMM parameters with extend Baum-Welch (EBW) for spoken language recognition. To overcome the problem in the LME that negative samples in the training set are not used in parameter estimation, we propose a soft margin estimation (SME) method in this paper. The soft margin is scaled by a loss function measuring the distance between a negative sample and the classification boundary. We formulate the constrained optimization of SME as an unconstrained optimization among both positive samples and negative samples using a penalty function, and update the GMM parameters with the EBW algorithm. Experiments on the NIST language recognition evaluation (LRE) 2007 task show that the SME method effectively improves the LME performance.
Donglai Zhu, Bin Ma 0001, Haizhou Li 0001
ICASSP2
2010 Voice conversion: From spoken vowels to singing vowels
abstract
In this paper, a voice conversion system that converts spoken vowels into singing vowels is proposed. Given the spoken vowels and their musical score, the system generates singing vowels. The system modifies the speech parameters of Fundamental frequency (F0), duration and spectral properties to produce singing voice. F0 contour is obtained using F0 fluctuation information from training singing voice and music score. Duration of each vowel of speech is stretched or shortened according to the length of the corresponding musical note. To transform speech spectrum to singing spectrum the following two approaches are employed. The first method employs spectral mean shifting and variance scaling method. And, the second approach uses weighted linear transformation method to transform speech to singing spectrum. The system is tested on the database including 75 speech and 30 singing voices sung using vowels. The results show that the proposed system is able to convert spoken vowels into singing vowels with a quality very close to the target singing voice.
Tin Lay Nwe, Minghui Dong, Paul Y. Chan, Bin Ma 0001, Haizhou Li 0001
ICME5
2010 Framewise Phone Classification Using Weighted Fuzzy Classification Rules
abstract
Our aim in this paper is to propose a rule-weight learning algorithm in fuzzy rule-based classifiers. The proposed algorithm is presented in two modes: first, all training examples are assumed to be equally important and the algorithm attempts to minimize the error-rate of the classifier on the training data by adjusting the weight of each fuzzy rule in the rule-base, and second, a weight is assigned to each training example as the cost of misclassification of it using the class distribution of its neighbors. Then, instead of minimizing the error-rate, the learning algorithm is modified to minimize the sum of costs for misclassified examples. Using six data sets from UCI-ML repository and the TIMIT speech corpus for frame wise phone classification, we show that our proposed algorithm considerably improves the prediction ability of the classifier.
Omid Dehzangi, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
ICPR2
2010 A study of term weighting in phonotactic approach to spoken language recognition
Sirinoot Boonsuk, Donglai Zhu, Bin Ma 0001, Atiwong Suchato, Proadpran Punyabukkana, Nattanun Thatphithakkul, Chai Wutiwiwatchai
INTERSPEECH3
2010 A discriminative performance metric for GMM-UBM speaker identification
Omid Dehzangi, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH2
2010 Approaching human listener accuracy with modern speaker verification
abstract
Being able to recognize people from their voice is a natural ability that we take for granted. Recent advances have shown significant improvement in automatic speaker recognition performance. Besides being able to process large amount of data in a fraction of time required by human, automatic systems are now able to deal with diverse channel effects. The goal of this paper is to examine how state-of-the-art automatic system performs in comparison with human listeners, and to investigate the strategy for human-assisted form of automatic speaker recognition, which is useful in forensic investigation. We set up an experimental protocol using data from the NIST SRE 2008 core set. A total of 36 listeners have participated in the listening experiments from three sites, namely Australia, Finland and Singapore. State-of-the-art automatic system achieved 20 % error rate, whereas fusion of human listeners achieved 22%. 1.
Ville Hautamäki, Tomi Kinnunen, Mohaddeseh Nosratighods, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH5
2010 Speaker characterization using long-term and temporal information
abstract
This paper presents new techniques for front-end analysis using long-term and temporal information for speaker recognition. We propose a long-term feature analysis strategy that averages short-time spectral features over a period of time in an effort to capture the speaker traits that are manifested over a speech segment longer than a spectral frame. We found that the moving averages of temporal information are effective in speaker recognition as well. The experiments on the 2008 NIST Speaker Recognition Evaluation dataset show the longterm and temporal information contribute to substantial EER reductions.
Chien-Lin Huang, Hanwu Sun, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH3
2010 Incorporating MAP estimation and covariance transform for SVM based speaker recognition
Cheung-Chi Leung, Donglai Zhu, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH4
2010 Effects of the phonological relevance in speaker verification
Yanhua Long, Li-Rong Dai 0001, Bin Ma 0001, Wu Guo
INTERSPEECH3
2010 Towards long-range prosodic attribute modeling for language recognition
abstract
As a high-level feature, prosody may be an effective feature when it is modeled over longer ranges than the typical range of a syllable. This paper is about language recognition with the high-level prosodic attributes. It studies two important issues of long-range modeling, namely the data scarcity handling method, and the model which properly describes prosodic boundary events. Illustrated by NIST language recognition evaluation (LRE) 2009, long-range modeling is shown to bring a 7.2% relative improvement to a prosodic language detector. Score fusion between the long-range prosodic system and a phonotactic system gives an EER of 3.07%. Exploiting boundary N -grams is the main contributing factor to global EER reduction, while different long-range prosodic modeling factors benefit the detection of different languages. Analysis reveals the evidence of language-specific long-range prosodic attributes, which sheds light on robust long-range modeling methods for language recognition. Index Terms: language recognition, prosody, long-range modeling
Raymond W. M. Ng, Cheung-Chi Leung, Ville Hautamäki, Tan Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH5
2010 Speaker diarization in meeting audio for single distant microphone
Tin Lay Nwe, Hanwu Sun, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH3
2010 The IIR NIST SRE 2008 and 2010 summed channel speaker recognition systems
abstract
This paper reports the IIR speaker recognition system for the summed channel evaluation tasks in the NIST SRE 2008 and 2010. The system includes three main modules: voice activity detection, speaker diarization and speaker recognition. The front-end process employs a voice activity detection algorithm for effective speech frame selection. The speaker diarization system that was developed for 2007 and 2009 NIST RT Evaluations is adopted for summed channel speech segmentation. A hybrid purifying and clustering algorithm is developed to segregate the summed channel speech by speakers. The GMM-SVM speaker recognition system is adopted to evaluate the performance with both MFCC and LPCC features. The system achieves an overall EER of 3.46% in the 1conv-summed task and 1.87% in the 8conv-summed task, respectively, where only all English trials are involved.
Hanwu Sun, Bin Ma 0001, Chien-Lin Huang, Trung Hieu Nguyen 0001, Haizhou Li 0001
INTERSPEECH2
2010 Selecting phonotactic features for language recognition
Rong Tong, Bin Ma 0001, Haizhou Li 0001, Chng Eng Siong
INTERSPEECH2
2010 The estimation and kernel metric of spectral correlation for text-independent speaker verification
Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH3
2010 Phoneme lattice based texttiling towards multilingual story segmentation
abstract
This paper proposes a phoneme lattice based TextTiling ap-proach towards multilingual story segmentation. The phoneme is the smallest segmental unit in a language and the number of phonemes in a language is usually far smaller than the number of words. Furthermore, many phonemes are shared by differ-ent languages. These properties make phonemes particularly appropriate for representing multilingual speech. As phoneme recognition is far from perfect, phoneme lattices, which carry much richer statistics than the 1-best hypotheses, are adopted in this paper as the input to the TextTiling approach. The term frequencies used in traditional TextTiling are replaced by the expected counts of phoneme n-gram units calculated from phoneme lattices. Experiments on TDT2 English and Mandarin corpora show that the phoneme lattice based TextTiling out-performs the phoneme 1-best based TextTiling and word based TextTiling in broadcast news story segmentation. Index Terms: story segmentation, topic detection and tracking, spoken document retrieval, phoneme lattice, speech processing. 1.
Lei Xie 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH3
2010 MAP estimation of subspace transform for speaker recognition
Donglai Zhu, Bin Ma 0001, Kong-Aik Lee, Cheung-Chi Leung, Haizhou Li 0001
INTERSPEECH2
2009 The I4U system in NIST 2008 speaker recognition evaluation
abstract
This paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU).
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin
ICASSP2
2009 Exploiting prosodic information for Speaker Recognition
abstract
In this paper, we study speaker characterization using prosodic supervectors with negative within-class covariance normalization (NWCCN) projection and speaker modeling with support vector regression (SVR). We also propose a segmental weight fusion (SWF) technique that combines acoustic and prosodic subsystems effectively, despite the big performance gap between the subsystems. We validate the effectiveness of our proposed techniques on the NIST 2006 Speaker Recognition Evaluation (SRE) in comparison with other prominent solutions. The experiments have reported competitive results of 17.72% Equal Error Rate for the prosodic subsystem alone and 4.50% for the fusion system on NIST 2006 SRE core test condition.
Yanhua Long, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Chng Eng Siong, Li-Rong Dai 0001
ICASSP2
2009 Evaluation of a fused FM and cepstral-based speaker recognition system on the NIST 2008 SRE
abstract
In this paper, the fusion of two speaker recognition subsystems, one based on Frequency Modulation (FM) and another on MFCC features, is reported. The motivation for their fusion was to improve the recognition accuracy across different types of channel variations, since the two features are believed to contain complementary information. It was found that the MFCC-based subsystem outperformed the FM-based subsystem on telephone conversations from NIST SRE-06 dataset, while the opposite was true for NIST SRE-08 telephone data. As a result, the FM-based subsystem performed as well as the MFCC-based subsystem and their fusion gave up to 23% relative improvement in terms of EER over the MFCC subsystem alone, when evaluated on the NIST 2008 core condition.
Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Bin Ma 0001, Haizhou Li 0001
ICASSP5
2009 Cross-validation of multiple language recognition systems using pseudo keys
abstract
In this paper, we present a pseudo-key analysis approach for cross-validation of language recognition systems before the ground truth (true key) becomes available. A state-of-the-art language recognition system typically employs multiple language recognition classifiers which are fused to form a mixture of experts. The individual classifiers are also called subsystems. To avoid the fused system from being brought down by some outlier classifiers, pseudo keys are designed to cross-examine the integrity of individual classifier candidates. The language recognition experiments are conducted on the NIST 2007 Language Recognition Evaluation (LRE) corpus using the subsystems in the primary submission from the Institute for Infocomm Research (IIR).
Hanwu Sun, Bin Ma 0001, Haizhou Li 0001
ICASSP2
2009 Joint map adaptation of feature transformation and Gaussian Mixture Model for speaker recognition
abstract
This paper extends our previous work on feature transformation-based support vector machines for speaker recognition by proposing a joint MAP adaptation of feature transformation (FT) and Gaussian Mixture Models (GMM) parameters. In the new approach, the prior probability density functions (PDFs) of FT and GMM parameters are jointly estimated using the background data under the maximum likelihood criteria. In this way, we derive a generic prior GMM that is more compact than the Universal Background Model due to the reduction of speaker variations. With the prior PDFs, we construct a supervector to characterize a speaker using FT and GMM parameters. We conducted experiments on NIST 2006 Speaker Recognition Evaluation (SRE06) data set. The results validated the effectiveness of the joint MAP adaptation approach.
Donglai Zhu, Bin Ma 0001, Haizhou Li 0001
ICASSP2
2009 Acoustic segment modeling for speaker recognition
abstract
We propose a speaker recognition system based on the acoustic segment modeling technique. It is assumed that the overall sound characteristics for speakers can be covered by a set of acoustic segment models (ASMs) while the ASMs are acoustically-motivated self-organized sound units without imposing any phonetic definitions. These acoustic segment models decode a spoken utterance into a string of segment units and the mean vectors of ASMs based on the unsupervised MAP adaptation are concatenated to represent the characteristics of the specific speaker. Support vector machines are thus applied on these high dimensional feature vectors for speaker recognition. We evaluate the proposed approach in the 2006 NIST speaker recognition evaluation core condition test trials.
Bin Ma 0001, Donglai Zhu, Haizhou Li 0001
ICME1
2009 Discriminative feature transformation using output coding for speech recognition
Omid Dehzangi, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH2
2009 Speaker diarization for meeting room audio
abstract
This paper describes a speaker diarization system in 2007 NIST Rich Transcription (RT07) Meeting Recognition Evaluation for the task of Multiple Distant Microphone (MDM) in meeting room scenarios. The system includes three major modules: data preparation, initial speaker clustering and cluster purification/merging. The data preparation consists of the raw data Wiener filtering and beamforming, Time Difference of Arrival estimate and speech activity detection. Based on the initial processed data, two-stage histogram quantization has been used to perform the initial speaker clustering. A modified purification strategy via high-order GMM clustering method is proposed. BIC criterion is applied for cluster merging. The system achieves a competitive overall DER of 8.31% for RT07 MDM speaker diarization task.
Hanwu Sun, Tin Lay Nwe, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH3
2009 Target-aware language models for spoken language recognition
Rong Tong, Bin Ma 0001, Haizhou Li 0001, Chng Eng Siong, Kong-Aik Lee
INTERSPEECH2
2009 Large margin estimation of Gaussian mixture model parameters with extended baum-welch for spoken language recognition
abstract
Discriminative training (DT) methods of acoustic models, such as SVM and MMI-training GMM, have been proved effective in spoken language recognition. In this paper we propose a DT method for GMM using the large margin (LM) estimation. Unlike traditional MMI or MCE methods, the LM estimation attempts to enhance the generalization ability of GMM to deal with new data that exhibits mismatch with training data. We define the multi-class separation margin as a function of GMM likelihoods, and derive update formulae of GMM parameters with the extended Baum-Welch algorithm. Results on the NIST language recognition evaluation (LRE) 2007 task show that the LM estimation achieves better performance and faster convergent speed than the MMI estimation.
Donglai Zhu, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH2
2009 A Target-Oriented Phonotactic Front-End for Spoken Language Recognition
abstract
This paper presents a strategy to optimize the phonotactic front-end for spoken language recognition. This is achieved by selecting a subset of phones from an existing phone recognizer's phone inventory such that only the phones that best discriminate each of the target languages are selected. Each such phone subset will be used to construct a target-oriented phone tokenizer (TOPT). In this study, we examine different approaches to construct such phone tokenizers for the front-end of a parallel phone recognizers followedbyvector space modeling (PPR-VSM) system. We show that the target-oriented phone tokenizers derived from language-specific phone recognizers are more effective than the original parallel phone recognizers. Our experimental results also show that the target-oriented phone tokenizers derived from universal phone recognizers achieve better performance than those derived from language-specific phone recognizers. Using the proposed target-oriented phone tokenizers as the phonotactic front-end, the language recognition system performance is significantly improved without the need for additional training samples. We achieve an equal error rate (EER) of 1.27%, 1.42% and 2.73% on the NIST 1996, 2003 and 2007 LRE databases respectively for 30-s closed-set tests. This system is one of the subsystems in IIR's submission to NIST 2007 LRE.
Rong Tong, Bin Ma 0001, Haizhou Li 0001, Chng Eng Siong
IEEE Trans. Speech Audio Process.2
2008 Target-oriented phone tokenizers for spoken language recognition
abstract
This paper presents a new strategy for designing the parallel phone recognizers for spoken language recognition. Given a collection of parallel phone recognizers, we select a subset of phones from each phone recognizer for each target language to construct a target-oriented phone tokenizer (TOPT). As a result, the collection of target-oriented phone tokenizers is more effective than the original parallel phone recognizers. This approach improves system performance significantly without requesting for additional transcribed training samples. We validate the effectiveness of the proposed strategy within the framework of the parallel phone recognizer followed by vector space modeling backend, or PPR-VSM. We achieve equal-error-rate of 2.21% and 3.65% on the 2003 and 2005 NIST LRE databases, respectively, for 30-second trials.
Rong Tong, Bin Ma 0001, Haizhou Li 0001, Chng Eng Siong
ICASSP2
2008 Discriminative learning for optimizing detection performance in spoken language recognition
abstract
We propose novel approaches for optimizing the detection performance in spoken language recognition. Two objective functions are designed to directly relate model parameters to two performance metrics of interest, the detection cost function and the area under the detection-error-tradeoff curve, respectively. Both metrics are approximated with differentiable functions of model parameters by using a smoothing function based on a class misclassification measure. The model parameters are optimized by using the generalized probabilistic descent algorithm. We conduct experiments on the NIST 2003 and 2005 Language Recognition Evaluation corpora. Results show that the proposed approaches effectively improve the performance over the maximum likelihood training approach.
Donglai Zhu, Haizhou Li 0001, Bin Ma 0001, Chin-Hui Lee 0001
ICASSP3
2008 Unsupervised pronunciation grammar growing using knowledge-based and data-driven approaches
abstract
This study presents a novel approach to unsupervised pronunciation grammar growing for non-native speech recognition. Unsupervised pronunciation grammar growing includes pronunciation variation graph construction and non-native grammar generation. Knowledge-based and data-driven approaches are considered for variation graph construction. The measurement of confidence and support is used for grammar selection. Experiments show that unsupervised pronunciation grammar growing is suitable for the improvement of non-native speech recognition.
Chien-Lin Huang, Chung-Hsien Wu 0001, Haizhou Li 0001, Chia-Hsin Hsieh, Bin Ma 0001
ICME5
2008 Fuzzy rule selection using Iterative Rule Learning for speech data classification
abstract
Fuzzy rule-based systems have been successfully used for pattern classification. These systems focus on generating a rule-base from numerical input data. The resulting rule-base can be applied on classification problems. However, we are faced with some challenges when generating and selecting the appropriate rules to create final rule-base. In this paper, a novel approach for rule selection is proposed. The proposed algorithm makes the use of Iterative Rule Learning (IRL) to reduce the search space of the classification problem in hand for rule-base extraction. The major element of our proposed approach is an evaluation metric which is able to accurately estimate the degree of cooperation of the candidate rule with current rules in the rule-base. Finally, fine-tuning of the selected rules is handled by employing a proposed rule-weighting mechanism. To evaluate the performance of the proposed scheme, TIMIT speech corpus was utilized for framewise classification of speech data. The results show the effectiveness of the proposed method while preserving the interpretability of the classification results.
Omid Dehzangi, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001
ICPR2
2008 Robust speaker verification using short-time frequency with long-time window and fusion of multi-resolutions
abstract
This study presents a novel approach of feature analysis to speaker verification. There are two main contributions in this paper. First, the feature analysis of short-time frequency with long-time window (SFLW) is a compact feature for the efficiency of speaker verification. The purpose of SFLW is to take account of short-time frequency characteristics and longtime resolution at the same time. Secondly, the fusion of multi-resolutions is used for the effectiveness of robust speaker verification. The speaker verification system can be further improved using multi-resolution features. The experimental results indicate that the proposed approaches not only speed up the processing time but also improve the performance of speaker verification.
Chien-Lin Huang, Bin Ma 0001, Chung-Hsien Wu 0001, Brian Kan-Wing Mak, Haizhou Li 0001
INTERSPEECH2
2008 Target-oriented phone selection from universal phone set for spoken language recognition
abstract
This paper studies target-oriented phone selection strategy for constructing phone tokenizers in the Parallel Phone Recognizers followed by Vector Space Model (PPR-VSM) paradigm of spoken language recognition. With this phone selection strategy, one derives a set of target-oriented phone tokenizers (TOPT), each having a subset of phones that have high discriminative ability for a target language. Two phone selection methods are proposed to derive such phone subsets from a phone recognizer. We show that the TOPTs derived from a universal phone recognizer (UPR) outperform those derived from language specific phone recognizers. The TOPT front-end derived from a UPR also consistently outperforms the UPR front-end without involving additional acoustic modeling. We achieve an equal error rates (EERs) of 1.33%, 1.75% and 2.80% on NIST 1996, 2003 and 2007 LRE databases respectively for 30 second closed-set tests by including multiple TOPTs in the PPR.
Rong Tong, Bin Ma 0001, Haizhou Li 0001, Chng Eng Siong
INTERSPEECH2
2008 Using MAP estimation of feature transformation for speaker recognition
abstract
We propose to use a new feature transformation (FT) function to construct supervectors of support vector machines for speaker recognition. Considering that estimation of bias vectors is more robust than that of transformation matrices, we define the FT function in a flexible form that transformation matrices and bias vectors are controlled by separate regression classes. Unlike the MLLR-based approach that needs a continuous speech recognition system, our FT function parameters are estimated based on a Gaussian mixture model (GMM). An iterative training procedure is used to achieve the maximum a posteriori estimation of the FT function parameters, which avoids the possible numerical problem caused by insufficient training data in the maximum likelihood estimation. Our approach is evaluated on the SRE2006 NIST evaluation and obtains better performance than a conventional SVM system based on GMM mean supervectors.
Donglai Zhu, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH2
2008 NIST 2007 Language Recognition Evaluation: From the Perspective of IIR
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Khe Chai Sim, Hanwu Sun, Rong Tong, Donglai Zhu, Chang Huai You
PACLIC2
2008 Optimizing the Performance of Spoken Language Recognition With Discriminative Training
abstract
The performance of spoken language recognition system is typically formulated to reflect the detection cost and the strategic decision points along the detection-error-tradeoff curve. We propose a performance metrics optimization (PMO) approach to optimizing the detection performance of Gaussian mixture model classifiers. We design the objective functions to directly relate the model parameters to the performance metrics of interest, i.e., the detection cost function and the area under the detection-error-tradeoff curve. Both metrics are approximated by differentiable functions of model parameters. In this way, the model parameters can be optimized with the generalized probabilistic descent algorithm, a typical discriminative training technique. We conduct the experiments on the NIST 2003 and 2005 Language Recognition Evaluation corpora. The experimental results show that the PMO approach effectively improves the performance over the maximum-likelihood training approach.
Donglai Zhu, Haizhou Li 0001, Bin Ma 0001, Chin-Hui Lee 0001
IEEE Trans. Speech Audio Process.3
2007 Effects of Device Mismatch, Language Mismatch and Environmental Mismatch on Speaker Verification
abstract
Device, language and environmental mismatch adversely affect speaker verification (SV) performance. We investigate such effects empirically based on the M3 (multibiometric, multilingual and multi-device) corpus (H. Meng et al., 2006). Device mismatch (among 3G phone, PocketPC and a desktop PC plug-in microphone) brings relative performance degradation of 523%; language mismatch (between English and Cantonese) brings 284% and environmental mismatch (between office environment and recording studio) brings 109%. In particular, verification with wide-band models on narrow-band test data outperforms narrow-band models on wide-band test data. The 3G phone's SV performance is generally low, but remains stable across environments. Additionally, durational variations within two-second utterances may cause a relative change of 633% in SV performance.
Bin Ma 0001, Helen M. Meng, Man-Wai Mak
ICASSP (4)1
2007 Discriminative Vector for Spoken Language Recognition
abstract
We propose a language recognition system based on discriminative vectors, in which parallel phone recognizers serve as the voice tokenization front-end followed by vector space modeling that effectively vectorizes phonotactic features, and the final classification is carried out based on the discriminative vectors. We design an ensemble of discriminative binary classifiers. The output values of these classifiers construct a discriminative vector, also referred to as output codes, to represent the high-dimensional phonotactic features. We achieve equal-error-rate of 1.95%, 3.02% and 4.9% on 1996, 2003 and 2005 NIST LRE databases, respectively, for 30-second trials.
Bin Ma 0001, Rong Tong, Haizhou Li 0001
ICASSP (4)1
2007 Spoken Language Recognition with Relevance Feedback
abstract
This paper applies relevance feedback technique in spoken language recognition task, in which we consider a test utterance as a test query. Assuming that we have a labeled multilingual corpus, we exploit the retrieved utterances from such a reference corpus to automatically augment the test query. Note that successful spoken language recognition relies on sufficient query data. The proposed method is especially effective for short query by expanding the query at a low cost. Experiments show that unsupervised relevance feedback reduces the relative equal-error-rate by 16.2%, 4.9% and 10.2% on NIST LRE 1996, 2003 and 2005 databases respectively for 3-second trials.
Rong Tong, Haizhou Li 0001, Bin Ma 0001, Chng Eng Siong, Siu-Yeung Cho
ICASSP (4)3
2007 A Generalized Feature Transformation Approach for Channel Robust Speaker Verification
abstract
In this paper we propose a generalized feature transformation approach to compensating for channel variation in speaker verification (SV) applications. Channel-dependent (CD) piecewise linear transformations are used for feature compensation. CD transformation parameters are estimated together with a channel-independent (CI) root Gaussian mixture model (GMM) from training data with a variety of channel conditions by using a maximum likelihood criterion. Experiments are conducted on the 2005 NIST Speaker Recognition Evaluation (SRE) corpus for several text-independent GMM-based SV systems. Experimental results show that the proposed approach achieves relative equal error rate (EER) reductions of 8.19% and 26.24% in comparison with a traditional feature mapping approach and a baseline system, respectively.
Donglai Zhu, Bin Ma 0001, Haizhou Li 0001, Qiang Huo
ICASSP (4)2
2007 Using direction of arrival estimate and acoustic feature information in speaker diarization
Chin-Wei Eugene Koh, Hanwu Sun, Tin Lay Nwe, Trung Hieu Nguyen 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001, Susanto Rahardja
INTERSPEECH5
2007 A Vector Space Modeling Approach to Spoken Language Identification
abstract
We propose a novel approach to automatic spoken language identification (LID) based on vector space modeling (VSM). It is assumed that the overall sound characteristics of all spoken languages can be covered by a universal collection of acoustic units, which can be characterized by the acoustic segment models (ASMs). A spoken utterance is then decoded into a sequence of ASM units. The ASM framework furthers the idea of language-independent phone models for LID by introducing an unsupervised learning procedure to circumvent the need for phonetic transcription. Analogous to representing a text document as a term vector, we convert a spoken utterance into a feature vector with its attributes representing the co-occurrence statistics of the acoustic units. As such, we can build a vector space classifier for LID. The proposed VSM approach leads to a discriminative classifier backend, which is demonstrated to give superior performance over likelihood-based n-gram language modeling (LM) backend for long utterances. We evaluated the proposed VSM framework on 1996 and 2003 NIST Language Recognition Evaluation (LRE) databases, achieving an equal error rate (EER) of 2.75% and 4.02% in the 1996 and 2003 LRE 30-s tasks, respectively, which represents one of the best results reported on these popular tasks
Haizhou Li 0001, Bin Ma 0001, Chin-Hui Lee 0001
IEEE Trans. Speech Audio Process.2
2007 Spoken Language Recognition Using Ensemble Classifiers
abstract
In this paper, we study a novel approach to spoken language recognition using an ensemble of binary classifiers. In this framework, we begin by representing a speech utterance with a high-dimensional feature vector such as the phonotactic characteristics or the polynomial expansion of cepstral features. A binary classifier can be built based on such feature vectors. We adopt a distributed output coding strategy in ensemble classifier design, where we decompose a multiclass language recognition problem into many binary classification tasks, each of which addresses a language recognition subtask by using a component classifier. Then, we combine the results of the component classifiers to form an output code as a hypothesized solution to the overall language recognition problem. In this way, we effectively project high-dimensional feature vectors into a tractable low-dimensional space, yet maintaining language discriminative characteristics of the spoken utterances. By fusing the output codes from both phonotactic features and cepstral features, we achieve equal-error-rates of 1.38% and 3.20% for 30-s trials on the 2003 and 2005 NIST language recognition evaluation databases.
Bin Ma 0001, Haizhou Li 0001, Rong Tong
IEEE Trans. Speech Audio Process.1
2006 Chinese Dialect Identification Using Tone Features Based on Pitch Flux
abstract
This paper presents a method to extract tone relevant features based on pitch flux from continuous speech signal. The autocorrelations of two adjacent frames are calculated and the covariance between them is estimated to extract multi-dimensional pitch flux features. These features, together with MFCCs, are modeled in a 2-stream GMM models, and are tested in a 3-dialect identification task for Chinese. The pitch flux features have shown to be very effective in identifying tonal languages with short speech segments. For the test speech segments of 3 seconds, 2-stream model achieves more than 30% error reduction over MFCC-based model.
Bin Ma 0001, Donglai Zhu, Rong Tong
ICASSP (1)1
2006 Integrating Acoustic, Prosodic and Phonotactic Features for Spoken Language Identification
abstract
The fundamental issue of the automatic language identification is to explore the effective discriminative cues for languages. This paper studies the fusion of five features at different level of abstraction for language identification, including spectrum, duration, pitch, n-gram phonotactic, and bag-of- sounds features. We build a system and report test results on NIST 1996 and 2003 LRE datasets. The system is also built to participate in NIST 2005 LRE. The experiment results show that different levels of information provide complementary language cues. The prosodic features are more effective for shorter utterances while the phonotactic features work better for longer utterances. For the task of 12 languages, the system with fusion of five features achieved 2.38% EER for 30-sec speech segments on NIST 1996 dataset.
Rong Tong, Bin Ma 0001, Donglai Zhu, Haizhou Li 0001, Chng Eng Siong
ICASSP (1)2
2006 Vector-based spoken language recognition using output coding
abstract
The vector-based spoken language recognition approach converts a spoken utterance into a high dimensional vector, also known as a bag-of-sounds vector, that consists of n-gram statistics of acoustic units. Dimensionality reduction would better prepare the bag-of-sounds vectors for classifier design. We propose projecting the bag-of-sounds vectors onto a low dimensional SVM output coding space, where each dimension represents a decision hyperplane between a pair of spoken languages. We also compare the performances of the output coding approach and the traditional low ranking approximation approach using latent semantic indexing (LSI) on the NIST 1996, 2003 and 2005 Language Recognition Evaluation (LRE) databases. The experiments show that the output coding approach consistently outperforms LSI with competitive results.
Haizhou Li 0001, Bin Ma 0001, Rong Tong
INTERSPEECH2
2006 Speaker cluster based GMM tokenization for speaker recognition
abstract
We present a speaker recognition system with multiple GMM tokenizers as the front-end, and vector space modeling as the back-end classifier. GMM tokenizer captures the acoustic and phonetic characteristics of a speaker from the speech without the need of phonetic transcription. To enhance the speaker characteristics coverage and provide more discriminative information, a speaker clustering algorithm is proposed to build multiple GMM tokenizers that are arranged in parallel. For an input utterance, each of the tokenizers outputs a token sequence, which is then represented by a vector of n-gram probabilities. Multiple vectors are concatenated to form a composite vector. Finally the Support Vector Machine (SVM) is used as the back-end classifier of the composite vectors. We use the 2002 NIST Speaker Recognition Evaluation (SRE) corpus for training GMM tokenizers and background modeling, and evaluate on the 2001 NIST SRE corpus. Index Terms: speaker recognition, speaker clustering, GMM tokenization
Bin Ma 0001, Donglai Zhu, Rong Tong, Haizhou Li 0001
INTERSPEECH1
2005 A Phonotactic Language Model for Spoken Language Identification
abstract
We have established a phonotactic language model as the solution to spoken language identification (LID). In this framework, we define a single set of acoustic tokens to represent the acoustic activities in the world's spoken languages. A voice tokenizer converts a spoken document into a text-like document of acoustic tokens. Thus a spoken document can be represented by a count vector of acoustic tokens and token n-grams in the vector space. We apply latent semantic analysis to the vectors, in the same way that it is applied in information retrieval, in order to capture salient phonotactics present in spoken documents. The vector space modeling of spoken utterances constitutes a paradigm shift in LID technology and has proven to be very successful. It presents a 12.4% error rate reduction over one of the best reported results on the 1996 NIST Language Recognition Evaluation database.
Haizhou Li 0001, Bin Ma 0001
ACL2
2005 Using Local & Global Phonotactic Features in Chinese Dialect Identification
abstract
Conventional techniques for spoken language identification use variants of phone similarity and language model scoring, which represent local phonetic constraints in spoken languages. We explore the identification of Chinese dialects which share the same written script and have similar sound systems and syllable structures. As such, local phonetic constraints do not provide enough discriminative information among dialects. We propose to use latent semantic analysis (LSA) to extract global features that represent the high-order statistics in the cooccurrence of sounds. Experiments show that we can achieve the best performance by combining acoustic, n-gram language modeling and LSA scores. An accuracy of 99.23% is achieved in 4-way classification tests using 20-second speech sessions.
Boon Pang Lim, Haizhou Li 0001, Bin Ma 0001
ICASSP (1)3
2005 A text categorization approach to automatic language identification
Bin Ma 0001, Haizhou Li 0001, Chin-Hui Lee 0001
INTERSPEECH2
2005 An acoustic segment modeling approach to automatic language identification
Bin Ma 0001, Haizhou Li 0001, Chin-Hui Lee 0001
INTERSPEECH1
2005 A phonotactic-semantic paradigm for automatic spoken document classification
abstract
We demonstrate a phonotactic-semantic paradigm for spoken document categorization. In this framework, we define a set of acoustic words instead of lexical words to represent acoustic activities in spoken languages. The strategy for acoustic vocabulary selection is studied by comparing different feature selection methods. With an appropriate acoustic vocabulary, a voice tokenizer converts a spoken document into a text-like document of acoustic words. Thus, a spoken document can be represented by a count vector, named a bag-of-sounds vector, which characterizes a spoken document's semantic domain. We study two phonotactic-semantic classifiers, the support vector machine classifier and the latent semantic analysis classifier, and their properties. The phonotactic-semantic framework constitutes a new paradigm in spoken document classification, as demonstrated by its success in the spoken language identification task. It achieves 18.2% error reduction over state-of-the-art benchmark performance on the 1996 NIST Language Recognition Evaluation database.
Bin Ma 0001, Haizhou Li 0001
SIGIR1
2004 English-Chinese bilingual text-independent speaker verification
abstract
This paper describes the development of a text-independent speaker verification (TISV) system for English and Chinese utterances. We have designed and collected a bilingual database that contains spoken responses and commands in short, medium and long durations. The TISV system uses Gaussian mixtures for speaker models. Our experiments indicate that language mismatch between enrolment and verification data leads to significant degradation in verification performance (between 40% to 49%). In order to maximize robustness towards language change in test utterances, speaker models were trained with utterances from both languages. Results indicate that this can effectively close the performance degradation gap due to language mismatch as mentioned above.
Bin Ma 0001, Helen M. Meng
ICASSP (5)1
2004 Fuzzy logic decision fusion in a multimodal biometric system
abstract
This paper presents a multi-biometric verification system that combines speaker verification, fingerprint verification with face identification. Their respective equal error rates (EER) are 4.3%, 5.1 % and the range of (5.1 % to 11.5%) for matched conditions in facial image capture. Fusion of the three by majority voting gave a relative improvement of 48 % over speaker verification (i.e. the best-performing biometric). Fusion by weighted average scores produced a further relative improvement of 52%. We propose the use of fuzzy logic decision fusion, in order to account for external conditions that affect verification performance. Examples include recording conditions of utterances for speaker verification, lighting and facial expressions in face identification and finger placement and pressure for fingerprint verification. The fuzzy logic framework incorporates some external factors relating to face and fingerprint verification and achieved an additional improvement of 19%. 1.
Chun Wai Lau, Bin Ma 0001, Helen M. Meng, Yiu Sang Moon, Yeung Yam
INTERSPEECH2
2002 Multilingual speech recognition with language identification
Bin Ma 0001, Cuntai Guan, Haizhou Li 0001, Chin-Hui Lee 0001
INTERSPEECH1
2001 Online adaptive learning of continuous-density hidden Markov models based on multiple-stream prior evolution and posterior pooling
abstract
We introduce a new adaptive Bayesian learning framework, called multiple-stream prior evolution and posterior pooling, for online adaptation of the continuous density hidden Markov model (CDHMM) parameters. Among three architectures we proposed for this framework, we study in detail a specific two stream system where linear transformations are applied to the mean vectors of the CDHMMs to control the evolution of their prior distribution. This new stream of prior distribution can be combined with another stream of prior distribution evolved without any constraints applied. In a series of speaker adaptation experiments on the task of continuous Mandarin speech recognition, we show that the new adaptation algorithm achieves a similar fast-adaptation performance as that of the incremental maximum likelihood linear regression (MLLR) in the case of small amount of adaptation data, while maintains the good asymptotic convergence property as that of our previously proposed quasi-Bayes adaptation algorithms.
Qiang Huo, Bin Ma 0001
IEEE Trans. Speech Audio Process.2
2000 Efficient ML training of CDHMM parameters based on prior evolution, posterior intervention and feedback
abstract
We present an efficient maximum likelihood (ML) training procedure for Gaussian mixture continuous density hidden Markov model (CDHMM) parameters. This procedure is proposed using the concept of approximate prior evolution, posterior intervention and feedback (PEPIF). In a series of experiments for training CDHMMs for a continuous Mandarin Chinese speech recognition task, the new PEPIF procedure achieves a 4-fold speed-up in terms of user CPU time over that of the Baum-Welch algorithm in producing models of given likelihood or recognition accuracy.
Qiang Hue, Nathan Smith, Bin Ma 0001
ICASSP3
2000 Robust speech recognition based on off-line elicitation of multiple priors and on-line adaptive prior fusion
Qiang Huo, Bin Ma 0001
INTERSPEECH2
1999 Irrelevant variability normalization in learning HMM state tying from data based on phonetic decision-tree
abstract
We propose to apply the concept of irrelevant variability normalization to the general problem of learning structure from data. Because of the problems of a diversified training data set and/or possible acoustic mismatches between training and testing conditions, the structure learned from the training data by using a maximum likelihood training method will not necessarily generalize well on mismatched tasks. We apply the above concept to the structural learning problem of phonetic decision-tree based hidden Markov model (HMM) state tying. We present a new method that integrates a linear-transformation based normalization mechanism into the decision-tree construction process to make the learned structure have a better modeling capability and generalizability. The viability and efficacy of the proposed method are confirmed in a series of experiments for continuous speech recognition of Mandarin Chinese.
Qiang Huo, Bin Ma 0001
ICASSP2
1999 On-line adaptive learning of CDHMM parameters based on multiple-stream prior evolution and posterior pooling
Qiang Huo, Bin Ma 0001
EUROSPEECH2