EDBT 2026 Demo / reviewers in the wild / expert
Chin-Hui Lee 0001
dblp:60/826-1
· DBLP profile ↗
214ranked-venue papers
1as first author
58since 2021 · last 2025
0000-0002-1892-2551ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 170 · 50 since 2021Artificial intelligence and machine learning · 125 · 27 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech EnhancementabstractIn this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly estimate the phase using a deep model, our idea is to exploit ad-hoc constraints to directly generate a consistent pair of magnitude and phase. Specifically, the proposed loss forces a set of complex numbers to be a consistent short-time Fourier transform (STFT) representation, i.e., to be the spectrogram of a real signal. Our approach thus avoids the difficulty of estimating the original phase, which is highly unstructured and sensitive to time shift. The influence of our proposed loss is first assessed on a PR task, experimentally demonstrating that our approach is viable. Next, we show its effectiveness on an SE task, using both the VB-DMD and WSJ0-CHiME3 data sets. On VB-DMD, our approach is competitive with conventional solutions. On the challenging WSJ0-CHiME3 set, the proposed framework compares favourably over those techniques that explicitly estimate the phase. Pin-Jui Ku, Chun-Wei Ho, Hao Yen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2025 | The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Shilong Wu, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001, Shinji Watanabe 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg |
INTERSPEECH | 5 |
| 2025 | Video Segmentation and Tokenization for Model-Based Video Scene ClassificationabstractIn this paper, we propose a novel approach for segmenting and tokenizing a video scene recording into a sequence of cascade units, known as visual segment units and modeled with visual segment models (VSMs) for video scene classification (VSC). Specifically, the proposed VSM framework takes deep visual features extracted from pre-trained encoders as inputs and models the temporal interactions between segment units by hidden Markov models. Next, we use unit co-occurrence statistics to introduce relationships between VSM units within a video scene recording. Furthermore, the VSM approach is extended to an acoustic-visual variant, subsequently integrating itself into a deep learning-based multi-modal scene classification system. This combination serves to further exploit the complementary nature of audio and video data. By incorporating a set of visual segment units into modeling a video scene class, it captures both inter-class similarity and intra-class diversity, facilitating improved scene classification, especially within categories prone to confusion. Extensive experimental results on a benchmark published by the DCASE (Detection and Classification of Acoustic Scenes and Events) 2021 Challenge show that the proposed framework can effectively handle the confusion issue among similar video scenes. In addition, our multi-modal integration system achieves state-of-the-art performance in the audio-visual scene classification task in the DCASE 2021 Challenge, thereby demonstrating the effectiveness of our proposed approach. Qing Wang 0008, Yajian Wang, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech RecognitionabstractAdvanced Audio- Visual Speech Recognition (AVSR) sys-tems have been observed to be sensitive to missing video frames, performing even worse than single-modality mod-els. While applying the common dropout techniques to the video modality enhances robustness to missing frames, it simultaneously results in a performance loss when dealing with complete data input. In this study, we delve into this contrasting phenomenon through the lens of modality bias and uncover that an excessive modality bias towards the audio modality induced by dropout constitutes the fun-damental cause. Next, we present the Modality Bias Hy-pothesis (MBH) to systematically describe the relationship between the modality bias and the robustness against missing modality in multimodal systems. Building on these findings, we propose a novel Multimodal Distribution Approxi-mation with Knowledge Distillation (MDA-KD)framework to reduce over-reliance on the audio modality, maintaining performance and robustness simultaneously. Finally, to address an entirely missing modality, we adopt adapters to dynamically switch decision strategies. The effective-ness of our proposed approach is evaluated through comprehensive experiments on the MISP2021 and MISP2022 datasets. Our code is available at https://github.com/dalision/ModalBiasAV5R. Yusheng Dai, Hang Chen 0001, Jun Du 0002, Ruoyu Wang 0029, Shihao Chen, Chin-Hui Lee 0001 |
CVPR | 7 |
| 2024 | A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech RecognitionabstractDeep learning (DL)-based speaker diarization methods have proven powerful performance comparing to traditional clustering-based methods for multi-talker speech diarization and recognition in farfield scenes. However, most DL-based approaches cannot utilize the spatial information well due to the poor robustness to unknown array topology and acoustic scenario. In this paper, a spatial long-term iterative mask estimation (SLT-IME) method is proposed to improve the performance of speaker diarization in various real-world acoustic scenarios. First, the complex angular central gaussian mixture model (cACGMM) with diarization results as initial values is used to estimate the presence probability of each speaker at each time-frequency bin, namely speaker masks, in a long-term chunk. Then, the speaker masks are converted to speaker activities according to the threshold, which deliver the diarization information of which speaker is active and when. Finally, the estimated speaker activity can also serve as the initial input for the diarization system, resulting in improved ASR performance. Experimental results on the CHiME-7 three datasets (CHiME-6, DiPCo, Mixer 6) show proposed method can improve diarization and recognition systems performance simultaneously. It also plays a key role in the ensemble system that achieves the best performance in the main track of CHiME-7 DASR Challenge. Yanhui Tu, Maokui He, Ruoyu Wang 0029, Shutong Niu, Lei Sun 0010, Zhongfu Ye, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 10 |
| 2024 | Improving Multi-Modal Emotion Recognition Using Entropy-Based Fusion and Pruning-Based Network Architecture OptimizationabstractIn this study, we aim to improve our recent hierarchical information fusion system for multi-modal emotion recognition challenge (MER 2023) in both efficiency and performance. Specifically, we extract robust acoustic and visual representations from pre-trained models and fuse them together in different structures. Then, an entropy-based fusion approach is proposed to obtain the final prediction of emotion and valence based on multi-label predictions of all different feature fusion structures. Furthermore, to reduce the network redundancy and improve the model generalization in low-resource multi-modal data conditions, we propose a novel approach for optimizing the network structure progressively based on structured pruning and learning-rate rewinding. When tested on the dataset of MER 2023, the optimized network structure with entropy-based fusion yields consistent and significant improvements, outperforming the champion system of the MER-MULTI sub-challenge. Jun Du 0002, Yusheng Dai, Chin-Hui Lee 0001, Yuling Ren |
ICASSP | 4 |
| 2024 | The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker ExtractionabstractPrevious Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompted a shift in focus towards the Audio-Visual Target Speaker Extraction (AVTSE) task for the MISP 2023 challenge in ICASSP 2024 Signal Processing Grand Challenges. Unlike existing audio-visual speech enhancement challenges primarily focused on simulation data, the MISP 2023 challenge uniquely explores how front-end speech processing, combined with visual clues, impacts back-end tasks in real-world scenarios. This pioneering effort aims to set the first benchmark for the AVTSE task, offering fresh insights into enhancing the accuracy of back-end speech recognition systems through AVTSE in challenging and real acoustic environments. This paper delivers a thorough overview of the task setting, dataset, and baseline system of the MISP 2023 challenge. It also includes an in-depth analysis of the challenges participants may encounter. The experimental results highlight the demanding nature of this task, and we look forward to the innovative solutions participants will bring forward. Shilong Wu, Hang Chen 0001, Yusheng Dai, Chenyue Zhang, Ruoyu Wang 0029, Hongbo Lan, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg, Zhongqiu Wang 0001, Jianqing Gao |
ICASSP | 9 |
| 2024 | Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence ArchitectureabstractWe propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates the strengths of memory-aware multi-speaker embedding (MA-MSE) and sequence-to-sequence (Seq2Seq) architecture, leading to improvement in both efficiency and performance. Next, we further decrease the memory occupation of decoding by incorporating input features fusion and then employ a multi-head attention mechanism to capture features at different levels. NSD-MS2S achieved a macro diarization error rate (DER) of 15.9% on the CHiME-7 EVAL set, which signifies a relative improvement of 49% over the official baseline system, and is the key technique for us to achieve the best performance for the main track of CHiME-7 DASR Challenge. Additionally, we introduce a deep interactive module (DIM) in MA-MSE module to better retrieve a cleaner and more discriminative multi-speaker embedding, enabling the current model to outperform the system we used in the CHiME-7 DASR Challenge. Our code is available at https://github.com/liyunlongaaa/NSD-MS2S. Gaobin Yang, Maokui He, Shutong Niu, Ruoyu Wang 0029, Yanyan Yue, Shuangqing Qian, Shilong Wu, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 9 |
| 2024 | Boosting End-to-End Multilingual Phoneme Recognition Through Exploiting Universal Speech Attributes ConstraintsabstractWe propose a first step toward multilingual end-to-end automatic speech recognition (ASR) by integrating knowledge about speech articulators. The key idea is to leverage a rich set of fundamental units that can be defined "universally" across all spoken languages, referred to as speech attributes, namely manner and place of articulation. Specifically, several deterministic attribute-to-phoneme mapping matrices are constructed based on the predefined set of universal attribute inventory, which projects the knowledge-rich articulatory attribute logits, into output phoneme logits. The mapping puts knowledge-based constraints to limit inconsistency with acoustic-phonetic evidence in the integrated prediction. Combined with phoneme recognition, our phone recognizer is able to infer from both attribute and phoneme information. The proposed joint multilingual model is evaluated through phoneme recognition. In multilingual experiments over 6 languages on benchmark datasets LibriSpeech and CommonVoice, we find that our proposed solution outperforms conventional multilingual approaches with a relative improvement of 6.85% on average, and it also demonstrates a much better performance compared to monolingual model. Further analysis conclusively demonstrates that the proposed solution eliminates phoneme predictions that are inconsistent with attributes. Hao Yen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2024 | Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic ScenariosabstractThis study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and multi-modal fusion. First, we propose a cross-modal teacher-student learning (TSL) framework to transfer information from an audio-only teacher model, trained on a rich collection of audio data with multiple data augmentation techniques, to an audiovisual student model trained with only a limited set of multimodal data. Next, we propose a two-stage audio-visual fusion strategy, consisting of an early feature fusion and a late video-guided decision fusion to exploit synergies between audio and video modalities. Finally, we introduce an innovative video pixel swapping (VPS) technique to extend an audio channel swapping (ACS) method to an audio-visual joint augmentation. Evaluation results on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge data set demonstrate significant improvements in SELD performances. Furthermore, our submission to the SELD task of the DCASE 2023 Challenge ranks first place by effectively integrating the proposed techniques into a model ensemble. Ya Jiang, Qing Wang 0008, Jun Du 0002, Maocheng Hu, Pengfei Hu 0006, Zeyan Liu, Shi Cheng 0001, Zhaoxu Nian, Mingqi Cai, Chin-Hui Lee 0001 |
ICME | 12 |
| 2024 | Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design
Hang Chen 0001, Jun Du 0002, Hongxiao Guo, Hui Bu, Jianxing Yang, Ming Li 0026, Chin-Hui Lee 0001 |
INTERSPEECH | 9 |
| 2024 | Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword RecognitionabstractWe propose a novel language-universal approach to end-to-end automatic spoken keyword recognition (SKR) leveraging upon (i) a self-supervised pre-trained model, and (ii) a set of universal speech attributes (manner and place of articulation).Specifically, Wav2Vec2.0 is used to generate robust speech representations, followed by a linear output layer to produce attribute sequences.A non-trainable pronunciation model then maps sequences of attributes into spoken keywords in a multilingual setting.Experiments on the Multilingual Spoken Words Corpus show comparable performances to character-and phoneme-based SKR in seen languages.The inclusion of domain adversarial training (DAT) improves the proposed framework, outperforming both character-and phoneme-based SKR approaches with 13.73% and 17.22% relative word error rate (WER) reduction in seen languages, and achieves 32.14% and 19.92% WER reduction for unseen languages in zero-shot settings. Hao Yen, Pin-Jui Ku, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2024 | Summary of Low-Resource Dysarthria Wake-Up Word Spotting ChallengeabstractIn recent years, the rapid advancement and widespread adoption of speech technology have made smart home systems a common feature in many households. However, individuals with dysarthria face difficulties using these technologies due to inconsistent speech patterns. This paper summarizes the Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge at SLT 2024, which aimed to develop effective voice wake-up systems for individuals with dysarthria. The challenge attracted 25 teams from 4 countries, with 7 teams submitting results and 5 providing detailed system descriptions. This paper presents an overview of the dataset, evaluation metrics, and key innovations from participating teams. Our findings highlight the potential of these systems to enhance the accessibility and usability of smart home technologies for individuals with dysarthria. The challenge results underscore the importance of developing specialized solutions to meet the unique needs of this user group. Hang Chen 0001, Jun Du 0002, Hongxiao Guo, Hui Bu, Ming Li 0026, Chin-Hui Lee 0001 |
SLT | 8 |
| 2024 | Optimizing Audio-Visual Speech Enhancement Using Multi-Level Distortion Measures for Audio-Visual Speech RecognitionabstractA multi-level distortion measure (MLDM) is proposed as an objective to optimize deep neural network-based speech enhancement (SE) in both audio-only and audio-visual scenarios. The aim is to achieve simultaneous performance improvements in speech quality, intelligibility, and recognition error reductions. Moreover, a comprehensive correlation analysis shows that these three evaluation metrics exhibit high Pearson correlation coefficient (PCC) values with three commonly used optimization objectives: the mean squared error between the ideal ratio and estimated magnitude masks, scale-invariant signal-to-noise ratio, and cross-entropy-guided measure. To further improve the performance, we leverage the complementarities of the three objectives and propose another correlated multi-level distortion measure (C-MLDM) defined as a weighted combination of MLDM and an average correlation measure based on the three PCCs. Experimental results on the TCD-TIMIT corpus corrupted by additive noise demonstrate that MLDM outperforms systems optimized with each objective in both audio-visual and audio-only scenarios, offering improved performances in all three metrics: speech quality, intelligibility, and recognition performance. C-MLDM also consistently outperforms MLDM in all test cases. Finally, the generalizability of both MLDM and C-MLDM is confirmed through extensive testing across diverse datasets, SE model architectures, and linguistic conditions. The source codes are publicly available.1 Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | A Variance-Preserving Interpolation Approach for Diffusion Models With Applications to Single Channel Speech Enhancement and RecognitionabstractIn this paper, we propose a variance-preserving interpolation framework to improve diffusion models for single-channel speech enhancement (SE) and automatic speech recognition (ASR). This new variance-preserving interpolation diffusion model (VPIDM) approach requires only 25 iterative steps and obviates the need for a corrector, an essential element in the existing variance-exploding interpolation diffusion model (VEIDM). Two notable distinctions between VPIDM and VEIDM are the scaling function of the mean of state variables and the constraint imposed on the variance relative to the mean's scale. We conduct a systematic exploration of the theoretical mechanism underlying VPIDM, and develop insights regarding VPIDM's applications in SE and ASR using VPIDM as a frontend. Our proposed approach, evaluated on two distinct data sets, demonstrates VPIDM's superior performances over conventional discriminative SE algorithms. Furthermore, we assess the performance of the proposed model under varying signal-to-noise ratio (SNR) levels. The investigation reveals VPIDM's improved robustness in target noise elimination when compared to VEIDM. Furthermore, utilizing the mid-outputs of both VPIDM and VEIDM results in enhanced ASR accuracies, thereby highlighting the practical efficacy of our proposed approach. Code and audio examples are available onlinehttps://github.com/zelokuo/VPIDM. Zilu Guo, Qing Wang 0008, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | Collaborative Viseme Subword and End-to-End Modeling for Word-Level Lip ReadingabstractWe propose a viseme subword modeling (VSM) approach to improve the generalizability and interpretability capabilities of deep neural network based lip reading. A comprehensive analysis of preliminary experimental results reveals the complementary nature of the conventional end-to-end (E2E) and proposed VSM frameworks, especially concerning speaker head movements. To increase lip reading accuracy, we propose hybrid viseme subwords and end-to-end modeling (HVSEM), which exploits the strengths of both approaches through multitask learning. As an extension to HVSEM, we also propose collaborative viseme subword and end-to-end modeling (CVSEM), which further explores the synergy between the VSM and E2E frameworks by integrating a state-mapped temporal mask (SMTM) into joint modeling. Experimental evaluations using different model backbones on both the LRW and LRW-1000 datasets confirm the superior performance and generalizability of the proposed frameworks. Specifically, VSM outperforms the baseline E2E framework, while HVSEM outperforms VSM in a hybrid combination of VSM and E2E modeling. Building on HVSEM, CVSEM further achieves impressive accuracies on 90.75% and 58.89%, setting new benchmarks for both datasets. Hang Chen 0001, Qing Wang 0008, Jun Du 0002, Genshun Wan, Shifu Xiong, Chin-Hui Lee 0001 |
IEEE Trans. Multim. | 8 |
| 2023 | Semi-Supervised Multi-Channel Speaker Diarization With Cross-Channel AttentionabstractMost neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised speaker diarization system to utilize large-scale multi-channel training data by generating pseudo-labels for unlabeled data. Furthermore, we introduce cross-channel attention into the Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding (NSD-MA-MSE) to learn channel contextual information of speaker embeddings better. Experimental results on the CHiME-7 Mixer6 dataset which only contains partial speakers’ labels of the training set, show that our system achieved 57.01% relative DER reduction compared to the clustering-based model on the development set. We further conducted experiments on the CHiME- 6 dataset to simulate the scenario of missing partial training set labels. When using 80% and 50% labeled training data, our system performs comparably to the results obtained using 100% labeled data for training. Shilong Wu, Jun Du 0002, Maokui He, Shutong Niu, Hang Chen 0001, Haitao Tang 0001, Chin-Hui Lee 0001 |
ASRU | 7 |
| 2023 | Summary on the Multimodal Information Based Speech Processing (MISP) 2022 ChallengeabstractThe Multimodal Information based Speech Processing (MISP) 2022 challenge aimed to enhance speech processing performance in harsh acoustic environments by leveraging additional modalities such as video or text. The challenge included two tracks: audio-visual speaker diarization (AVSD) and audio-visual diarization and recognition (AVDR). The training material was based on previous MISP 2021 recordings, but we have accurately synchronized audio and visual data. Additionally, a new evaluation set was provided. This paper gives an overview of the challenge setup, presents the results, and summarizes the effective techniques employed by the participants. We also analyze the current technical challenges and suggest directions for future research in AVSD and AVDR. Hang Chen 0001, Shilong Wu, Yusheng Dai, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 6 |
| 2023 | Incorporating Lip Features into Audio-Visual Multi-Speaker DOA Estimation by Gated FusionabstractThe audio-visual direction of arrival (DOA) estimation has demonstrated superior performance recently. In this paper, we present a novel audio-visual multi-speaker DOA estimation network, which for the first time incorporates multi-speaker lip features to adapt the complex overlapping and noisy scenarios. Firstly, we encode the multi-channel audio features, the reference angles and the lip Regions of Interest (RoIs) detected from the video respectively to acquire high-level representations. Then the multi-modal embeddings of audio, speaker angles and lips are fused by a tri-modal gated fusion module to balance their contributions to the output. The fused embedding is sent to the backend network to obtain the accurate DOA estimation with the combination of the predicted speaker angular vectors and the speaker activities. Experimental results show that our proposed approach can reduce the localization error by 73.48% compared to the previous work on the 2021 Multi-modal Information based Speech Processing (MISP) Challenge corpus. Meanwhile, the high accuracy and stability of localization results demonstrate the robustness of the proposed model in multi-speaker scenarios. Ya Jiang, Hang Chen 0001, Jun Du 0002, Qing Wang 0008, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2023 | An Experimental Study on Sound Event Localization and Detection Under Realistic Testing ConditionsabstractWe study four data augmentation (DA) techniques and two model architectures on realistic data for sound event localization and detection (SELD). First, based on ResNet-Conformer (RC), we compare the four DA approaches on the realistic DCASE 2022 SELD test set which is often not easy to handle due to room reverberations and audio overlaps in spontaneous recordings. Experimental results show that, except for audio channel swapping (ACS), the other three data augmentation methods that work well on the simulated SELD data set are no longer effective due to mismatches between simulated and realistic conditions. Next, using ACS-based augmentation, the two improved ResNet-Conformer networks further enhance SELD performances in realistic conditions. By incorporating these two sets of techniques, our overall system ranked the first place in SELD task of the DCASE 2022 Challenge. Shutong Niu, Jun Du 0002, Qing Wang 0008, Li Chai 0002, Huaxin Wu, Zhaoxu Nian, Lei Sun 0010, Chin-Hui Lee 0001 |
ICASSP | 10 |
| 2023 | Loss Function Design for DNN-Based Sound Event Localization and Detection on Low-Resource Realistic DataabstractThis study focuses on the design of a loss function for a deep neural network (DNN)-based model with two branches, which is used to solve sound event localization and detection (SELD) on low-resource realistic data. To this end, we employ a secondary network for audio classification, which provides global event information to the main network, enabling it to make robust SELD predictions. Furthermore, we suggest utilizing a momentum strategy for direction-of-arrival (DOA) estimation, taking advantage of the strong temporal consistency of sound events, thereby effectively reducing localization error. Lastly, we incorporate a regularization term into the loss function to alleviate the overfitting problem on the small dataset. We evaluate our proposed methods on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2022 Task 3 dataset, and the results demonstrate consistent improvements in SELD performance. In comparison to the baseline system, the proposed loss function yields significantly improved results for both localization and detection metrics on realistic data. Moreover, the proposed loss function demonstrates its ability to generalize across different network architectures, as evidenced by the consistent improvements achieved. Qing Wang 0008, Jun Du 0002, Zhaoxu Nian, Shutong Niu, Li Chai 0002, Huaxin Wu, Chin-Hui Lee 0001 |
ICASSP | 8 |
| 2023 | The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And RecognitionabstractThe Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two tracks: 1) audio-visual speaker diarization (AVSD), aiming to solve "who spoken when" using both audio and visual data; 2) a novel audio-visual diarization and recognition (AVDR) task that focuses on addressing "who spoken what when" with audio-visual speaker diarization results. Both tracks focus on the Chinese language, and use far-field audio and video in real home-tv scenarios: 2-6 people communicating each other with TV noise in the background. This paper introduces the dataset, track settings, and baselines of the MISP2022 challenge. Our analyses of experiments and examples indicate the good performance of AVDR baseline system, and the potential difficulties in this challenge due to, e.g., the far-field video quality, the presence of TV noise in the background, and the indistinguishable speakers. Shilong Wu, Hang Chen 0001, Maokui He, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 6 |
| 2023 | A Quantum Kernel Learning Approach to Acoustic Modeling for Spoken Command RecognitionabstractWe propose a quantum kernel learning (QKL) framework to address the inherent data sparsity issues often encountered in training large-scare acoustic models in low-resource scenarios. We project acoustic features based on classical-to-quantum feature encoding. Different from existing quantum convolution techniques, we utilize QKL with features in the quantum space to design kernel-based classifiers. Experimental results on challenging spoken command recognition tasks for a few low-resource languages, such as Arabic, Georgian, Chuvash, and Lithuanian, show that the proposed QKL-based hybrid approach attains good improvements over existing classical and quantum solutions. Chao-Han Huck Yang, Bo Li 0028, Yu Zhang 0033, Nanxin Chen, Tara N. Sainath, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 7 |
| 2023 | Incorporating Visual Information Reconstruction into Progressive Learning for Optimizing audio-visual Speech EnhancementabstractVideo information has been widely introduced to speech enhancement as its contribution at low signal-to-noise ratios (SNRs). Conventional audio-visual speech enhancement networks take noisy speech and video as input and learn features of clean speech directly. To reduce the large SNR gap between the learning target and input noisy speech, we propose a novel mask-based audio-visual progressive learning speech enhancement (AVPL) framework with visual information reconstruction (VIR) to increase SNRs gradually. Each stage of AVPL takes a concatenation of pre-trained visual embedding and the previous representation as input and predicts a mask with the intermediate representation of the current stage. To extract more visual information and deal with the performance distortion, the AVPL-VIR model reconstructs the visual embedding as it is fed in for each stage. Experiment on the TCD-TIMIT dataset shows that the progressive learning method significantly outperforms direct learning for both audio-only and audio-visual models. Moreover, by reconstructing video information, the VIR module provides a more accurate and comprehensive representation of the data, which in turn improves the performance of both AVDL and AVPL. Chenyue Zhang, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 6 |
| 2023 | Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion EncoderabstractIn recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in end-to-end frameworks with low-quality videos. Unmatching convergence rates and specialized input representations between audio-visual modalities are considered to cause the problem. In this paper, we propose two novel techniques to improve audio-visual speech recognition (AVSR) under a pre-training and fine-tuning training framework. First, we explore the correlation between lip shapes and syllable-level subword units in Mandarin through a frame-level subword unit classification task with visual streams as input. The fine-grained subword labels guide the network to capture temporal relationships between lip shapes and result in an accurate alignment between video and audio streams. Next, we propose an audio-guided Cross-Modal Fusion Encoder (CMFE) to utilize main training parameters for multiple cross-modal attention layers to make full use of modality complementarity. Experiments on the MISP2021-AVSR data set show the effectiveness of the two proposed techniques. Together, using only a relatively small amount of training data, the final system achieves better performances than state-of-the-art systems with more complex front-ends and back-ends. The code is released at1. Yusheng Dai, Hang Chen 0001, Jun Du 0002, Xiaofei Ding, Feijun Jiang, Chin-Hui Lee 0001 |
ICME | 7 |
| 2023 | Variance-Preserving-Based Interpolation Diffusion Models for Speech Enhancement
Zilu Guo, Jun Du 0002, Chin-Hui Lee 0001, Wenbin Zhang 0002 |
INTERSPEECH | 3 |
| 2023 | A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint ModelsabstractWe propose a multi-dimensional structured state space (S4) approach to speech enhancement. To better capture the spectral dependencies across the frequency axis, we focus on modifying the multi-dimensional S4 layer with whitening transformation to build new small-footprint models that also achieve good performance. We explore several S4-based deep architectures in time (T) and time-frequency (TF) domains. The 2-D S4 layer can be considered a particular convolutional layer with an infinite receptive field although it utilizes fewer parameters than a conventional convolutional layer. Evaluated on the VoiceBank-DEMAND data set, when compared with the conventional U-net model based on convolutional layers, the proposed TF-domain S4-based model is 78.6% smaller in size, yet it still achieves competitive results with a PESQ score of 3.15 with data augmentation. By increasing the model size, we can even reach a PESQ score of 3.18. Pin-Jui Ku, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2023 | Unsupervised Adaptation with Quality-Aware Masking to Improve Target-Speaker Voice Activity Detection for Speaker Diarization
Shutong Niu, Jun Du 0002, Maokui He, Chin-Hui Lee 0001, Baoxiang Li, Jiakui Li |
INTERSPEECH | 4 |
| 2023 | A Multiple-Teacher Pruning Based Self-Distillation (MT-PSD) Approach to Model Compression for Audio-Visual Wake Word Spotting
Jun Du 0002, Hengshun Zhou, Chin-Hui Lee 0001, Yuling Ren, Jiangjiang Zhao |
INTERSPEECH | 4 |
| 2023 | AD-TUNING: An Adaptive CHILD-TUNING Approach to Efficient Hyperparameter Optimization of Child Networks for Speech Processing Tasks in the SUPERB Benchmark
Gaobin Yang, Jun Du 0002, Maokui He, Shutong Niu, Baoxiang Li, Jiakui Li, Chin-Hui Lee 0001 |
INTERSPEECH | 7 |
| 2023 | Space-and-speaker-aware acoustic modeling with effective data augmentation for recognition of multi-array conversational speech
Li Chai 0002, Hang Chen 0001, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001 |
Speech Commun. | 5 |
| 2023 | Using iterative adaptation and dynamic mask for child speech extraction under real-world multilingual conditions
Shi Cheng 0001, Jun Du 0002, Shutong Niu, Alejandrina Cristià, Xin Wang 0037, Qing Wang 0008, Chin-Hui Lee 0001 |
Speech Commun. | 7 |
| 2023 | ANSD-MA-MSE: Adaptive Neural Speaker Diarization Using Memory-Aware Multi-Speaker EmbeddingabstractIn this paper, we propose a neural speaker diarization (NSD) network architecture consisting of three key components. First, a memory-aware multi-speaker embedding (MA-MSE) mechanism is proposed to facilitate a dynamical refinement of speaker embedding to reduce a potential data mismatch between the speaker embedding extraction and the NSD network. Next, a speaker selection procedure is introduced to handle situations where the detected number of speakers is different from the assumed speaker size in the NSD network. Finally, an adaptive procedure is proposed to improve the required prior information for the nonoverlap speech segments in a given utterance during each iteration. We call our proposed framework adaptive neural speaker diarization with memory-aware multi-speaker embedding (ANSD-MA-MSE). Our method improves diarization performance in realistic operating scenarios, such as adverse acoustic environments, domain mismatches, and a varying, rather than fixed, number of speakers. Having been tested on both the AMI corpus and the DIHARD-III evaluation sets, our proposed approach consistently outperforms other state-of-the-art techniques in diarization error rates, including the results reported by the best single-model system in the DIHARD-III challenge. Our code is publicly available athttps://github.com/Maokui-He/NSD-MA-MSE. Maokui He, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | QDM-SSD: Quality-Aware Dynamic Masking for Separation-Based Speaker DiarizationabstractWe improve iterative separation-based speaker diarization (ISSD) with quality-aware dynamic masking (QDM). We call the proposed framework QDM-SSD. Compared with ISSD, QDM-SSD enhances the simulated data used for model adaptation through QDM to alleviate the influence of errors in speaker priors. In addition to data quality purification, QDM-SSD also makes the adaptation data sparse by automatically adjusting speaker overlap ratios according to data quality. Furthermore, using a sliding window over the adaptation data, clean regions in speech segments can be better localized. Experiments on the two-speaker conversational telephone speech (CTS) corpus show that the proposed QDM-SSD framework can reduce the diarization error rate (DER) by 18.56% relatively compared with ISSD. Moreover, QDM-SSD is shown to generalize to other two-speaker non-conversation telephone speech data sets where ISSD fails to work. Finally, we demonstrate that QDM-SSD can serve as a front-end to improve the performances of back-end automatic speech recognition. Shutong Niu, Jun Du 0002, Lei Sun 0010, Yu Hu 0003, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | A Four-Stage Data Augmentation Approach to ResNet-Conformer Based Acoustic Modeling for Sound Event Localization and DetectionabstractIn this paper, we propose a novel four-stage data augmentation approach to ResNet-Conformer based acoustic modeling for sound event localization and detection (SELD). First, we explore two spatial augmentation techniques, namely audio channel swapping (ACS) and multi-channel simulation (MCS), to deal with data sparsity in SELD. ACS and MDS focus on augmenting the limited training data with expanding direction of arrival (DOA) representations such that the acoustic models trained with the augmented data are robust to localization variations of acoustic sources. Next, time-domain mixing (TDM) and time-frequency masking (TFM) are also investigated to deal with overlapping sound events and data diversity. Finally, ACS, MCS, TDM and TFM are combined in a step-by-step manner to form an effective four-stage data augmentation scheme. Tested on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 data set, our proposed augmentation approach greatly improves the system performance, ranking our submitted system in the first place in the SELD task of the DCASE 2020 Challenge. Furthermore, we employ a ResNet-Conformer architecture to model both global and local context dependencies of an audio sequence and win the first place in the DCASE 2022 SELD evaluations. Qing Wang 0008, Jun Du 0002, Huaxin Wu, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And ResultsabstractIn this paper we discuss the rational of the Multi-model Information based Speech Processing (MISP) Challenge, and provide a detailed description of the data recorded, the two evaluation tasks and the corresponding baselines, followed by a summary of submitted systems and evaluation results. The MISP Challenge aims at tack-ling speech processing tasks in different scenarios by introducing information about an additional modality (e.g., video, or text), which will hopefully lead to better environmental and speaker robustness in realistic applications. In the first MISP challenge, two bench-mark datasets recorded in a real-home TV room with two reproducible open-source baseline systems have been released to promote research in audio-visual wake word spotting (AVWWS) and audio-visual speech recognition (AVSR). To our knowledge, MISP is the first open evaluation challenge to tackle real-world issues of AVWWS and AVSR in the home TV scenario. Hang Chen 0001, Hengshun Zhou, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 4 |
| 2022 | The USTC-Ximalaya System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription (M2met) ChallengeabstractWe propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-Channel Multi-Party Meeting Transcription (M2MeT) challenge. These techniques are designed to handle multi-speaker conversations in real-world meeting scenarios with high speaker-overlap ratios and under heavy reverberant and noisy condition. First, for data preparation and augmentation in training TS-VAD models, speech data containing both real meetings and simulated indoor conversations are used. Second, in refining results obtained after TS-VAD based decoding, we perform a series of post-processing steps to improve the VAD results needed to reduce diarization error rates (DERs). Tested on the ALIMEETING corpus, the newly released Mandarin meeting dataset used in M2MeT, we demonstrate that our proposed system can decrease the DER by up to 66.55/60.59% relatively when compared with classical clustering based diarization on the Eval/Test set. Maokui He, Weilin Zhou, Jingjing Yin, Shutong Niu, Yuhang Cao, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 11 |
| 2022 | A Variational Bayesian Approach to Learning Latent Variables for Acoustic Knowledge TransferabstractWe propose a variational Bayesian (VB) approach to learning distributions of latent variables in deep neural network (DNN) models for cross-domain knowledge transfer, to address acoustic mismatches between training and testing conditions. Instead of carrying out point estimation in conventional maximum a posteriori estimation with a risk of having a curse of dimensionality in estimating a huge number of model parameters, we focus our attention on estimating a manageable number of latent variables of DNNs via a VB inference framework. To accomplish model transfer, knowledge learnt from a source domain is encoded in prior distributions of latent variables and optimally combined, in a Bayesian sense, with a small set of adaptation data from a target domain to approximate the corresponding posterior distributions. Experimental results on device adaptation in acoustic scene classification show that our proposed VB approach can obtain good improvements on target devices, and consistently outperforms 13 state-of-the-art knowledge transfer algorithms. Hu Hu, Sabato Marco Siniscalchi, Chao-Han Huck Yang, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2022 | Improving Separation-Based Speaker Diarization Via Iterative Model Refinement And Speaker Embedding Based Post-ProcessingabstractIn this paper, we propose an iterative separation-based speaker diarization (ISSD) approach to cope with the realistic data conditions. In the proposed ISSD, we iteratively generate adaptation data ac-cording to speaker priors and fine-tune the separation model, which leads to a gradual performance improvement. To further reduce some unavoidable speaker detection errors due to some undesirable prior errors using simple ISSD, we utilize speaker embedding information and propose two post-processing techniques, namely, speaker filtering and speaker recovery. We evaluate the diarization performance on the two-speaker conversational telephone speech (CTS) data set from DIHARD-III Challenge. When compared to state-of-the-art clustering-based speaker diarization (CSD) system, the proposed ISSD approach combined with the two post-processing schemes yields a 47.72 % and 46.97 % relative diarization error rate reduction on the development and evaluation sets, respectively. ISSD is also one key contributing factor to the best-performing system in DIHARD-III Challenge. Shutong Niu, Jun Du 0002, Lei Sun 0010, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2022 | A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network PruningabstractAudio-only based wake word spotting (WWS) is challenging under noisy conditions due to the environmental interference in signal transmission. In this paper, we investigate on designing a compact audio-visual WWS system by utilizing the visual information to alleviate the degradation. Specifically, in order to use visual information, we first encode the detected lips to fixed-size vectors with MobileNet and concatenate them with acoustic features followed by the fusion network for WWS. However, the audio-visual model based on neural network requires a large footprint and a high computational complexity. To meet the application requirements, we introduce a neural network pruning strategy via the lottery ticket hypothesis in an iterative fine-tuning manner (LTH-IF), to the single-modal and multi-modal models, respectively. Tested on our in-house corpus for audio-visual WWS in a home TV scene, the proposed audiovisual system achieves significant performance improvements over the single-modality (audio-only or video-only) system under different noisy conditions. Moreover, LTH-IF pruning can largely reduce the network parameters and computations with no degradation of WWS performance, leading to a potential product solution for the TV wake-up scenario. Hengshun Zhou, Jun Du 0002, Chao-Han Huck Yang, Shifu Xiong, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2022 | Audio-Visual Speech Recognition in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we present the updated Audio-Visual Speech Recognition (AVSR) corpus of MISP2021 challenge, a large-scale audio-visual Chinese conversational corpus consisting of 141h audio and video data collected by far/middle/near microphones and far/middle cameras in 34 real-home TV rooms. To our best knowledge, our corpus is the first distant multi-microphone conversational Chinese audio-visual corpus and the first large vocabulary continuous Chinese lip-reading dataset in the adverse home-tv scenario. Moreover, we make a deep analysis of the corpus and conduct a comprehensive ablation study of all audio and video data in the audio-only/video-only/audiovisual systems. Error analysis shows video modality supplement acoustic information degraded by noise to reduce deletion errors and provide discriminative information in overlapping speech to reduce substitution errors. Finally, we also design a set of experiments such as frontend, data augmentation and end-to-end models for providing the direction of potential future work. The corpus and the code are released to promote the research not only in speech area but also for the computer vision area and cross-disciplinary research. Hang Chen 0001, Jun Du 0002, Yusheng Dai, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen |
INTERSPEECH | 4 |
| 2022 | End-to-End Audio-Visual Neural Speaker Diarization
Maokui He, Jun Du 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2022 | Deep Segment Model for Acoustic Scene Classification
Yajian Wang, Jun Du 0002, Hang Chen 0001, Qing Wang 0008, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2022 | Audio-Visual Wake Word Spotting in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we describe and release publicly the audio-visual wake word spotting (WWS) database in the MISP2021 Challenge, which covers a range of scenarios of audio and video data collected by near-, mid-, and far-field microphone arrays, and cameras, to create a shared and publicly available database for WWS. The database and the code 2 are released, which will be a valuable addition to the community for promoting WWS research using multi-modality information in realistic and complex conditions. Moreover, we investigated the different data augmentation methods for single modalities on an end-to-end WWS network. A set of audio-visual fusion experiments and analysis were conducted to observe the assistance from visual information to acoustic information based on different audio and video field configurations. The results showed that the fusion system generally improves over the single-modality (audio- or video-only) system, especially under complex noisy conditions. Hengshun Zhou, Jun Du 0002, Gongzhen Zou, Zhaoxu Nian, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen, Shifu Xiong, Jianqing Gao |
INTERSPEECH | 5 |
| 2022 | An Experimental Study on Private Aggregation of Teacher Ensemble Learning for End-to-End Speech RecognitionabstractDifferential privacy (DP) is one data protection avenue to safeguard user information used for training deep models by imposing noisy distortion on privacy data. Such a noise perturbation often results in a severe performance degradation in automatic speech recognition (ASR) in order to meet a privacy budget ε. Private aggregation of teacher ensemble (PATE) utilizes ensemble probabilities to improve ASR accuracy when dealing with the noise effects controlled by small values of ε. We extend PATE learning to work with dynamic patterns, namely speech utterances, and perform a first experimental demonstration that it prevents acoustic data leakage in ASR training. We evaluate three end-to-end deep models, including LAS, hybrid CTC/attention, and RNN transducer, on the open-source LibriSpeech and TIMIT corpora. PATE learning-enhanced ASR models outperform the benchmark DP-SGD mechanisms, especially under strict DP budgets, giving relative word error rate reductions between 26.2% and 27.5% for an RNN transducer model evaluated with LibriSpeech. We also introduce a DP-preserving ASR solution for pretraining on public speech corpora. Chao-Han Huck Yang, I-Fan Chen, Andreas Stolcke, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
SLT | 5 |
| 2021 | A Two-Stage Approach to Device-Robust Acoustic Scene ClassificationabstractTo improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks (CNNs) is proposed. Our two-stage system leverages on an ad-hoc score combination based on two CNN classifiers: (i) the first CNN classifies acoustic inputs into one of three broad classes, and (ii) the second CNN classifies the same inputs into one of ten finergrained classes. Three different CNN architectures are explored to implement the two-stage classifiers, and a frequency sub-sampling scheme is investigated. Moreover, novel data augmentation schemes for ASC are also investigated. Evaluated on DCASE 2020 Task 1a, our results show that the proposed ASC system attains a state-of-the-art accuracy on the development set, where our best system, a two-stage fusion of CNN ensembles, delivers a 81.9% average accuracy among multi-device test data, and it obtains a significant improvement on unseen devices. Finally, neural saliency analysis with class activation mapping (CAM) gives new insights on the patterns learnt by our models. Hu Hu, Chao-Han Huck Yang, Xianjun Xia, Yajian Wang, Shutong Niu, Li Chai 0002, Juanjuan Li, Hongning Zhu, Sabato Marco Siniscalchi, Yannan Wang, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 16 |
| 2021 | A Progressive Learning Approach to Adaptive Noise and Speech Estimation for Speech Enhancement and Noisy Speech RecognitionabstractIn this paper, we propose a progressive learning-based adaptive noise and speech estimation (PL-ANSE) method for speech preprocessing in noisy speech recognition, leveraging upon a frame-level noise tracking capability of improved minima controlled recursive averaging (IMCRA) and an utterance-level deep progressive learning of nonlinear interactions between speech and noise. First, a bi-directional long short-term memory model is adopted at each network layer to learn progressive ratio masks (PRMs) as targets with progressively increasing signal-to-noise ratios. Then, the estimated PRMs at the utterance level are combined within a conventional speech enhancement algorithm at the frame level for speech enhancement. Finally, the enhanced speech based on multi-level information fusion is directly fed into a speech recognition system to improve the recognition performance. Experiments show that our proposed approach can achieve a relative word error rate (WER) reduction of 22.1% when compared to results attained with unprocessed noisy speech (from 23.84% to 18.57%) on the CHiME-4 single-channel real test data. Zhaoxu Nian, Yan-Hui Tu, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2021 | Speech Enhancement Autoencoder with Hierarchical Latent StructureabstractA new hierarchical convolutional neural network-based autoencoder architecture called SEHAE (Speech Enhancement Hierarchical AutoEncoder) is introduced, in which the latent representation is decomposed into several parts that correspond to different scales. The model consists of three functionally different components. First, a stack of encoders generates a set of latent vectors that contain information from an increasingly larger receptive field. Second, the decoders construct the clean speech in a stage-wise and additive fashion, starting from a learned initial vector. The third component, which we call funnel networks, is tasked with "knitting" together the outputs of the previous decoder and the encoder to compute latent vectors for the next decoder. Several options for initial vectors are explored. Experiments show that SEHAE achieves significant improvements for the considered speech quality and intelligibility measures, outperforming a denoising autoencoder and other step-wise models. Furthermore, its internal workings are investigated using the intermediate results from the decoders. Koen Oostermeijer, Jun Du 0002, Qing Wang 0008, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2021 | Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech RecognitionabstractWe propose a novel decentralized feature extraction approach in federated learning to address privacy-preservation issues for speech recognition. It is built upon a quantum convolutional neural network (QCNN) composed of a quantum circuit encoder for feature extraction, and a recurrent neural network (RNN) based end-to-end acoustic model (AM). To enhance model parameter protection in a decentralized architecture, an input speech is first up-streamed to a quantum computing server to extract Mel-spectrogram, and the corresponding convolutional features are encoded using a quantum circuit algorithm with random parameters. The encoded features are then down-streamed to the local RNN model for the final recognition. The proposed decentralized framework takes advantage of the quantum learning progress to secure models and to avoid privacy leakage attacks. Testing on the Google Speech Commands Dataset, the proposed QCNN encoder attains a competitive accuracy of 95.12% in a decentralized model, which is better than the previous architectures using centralized RNN models with convolutional features. We also conduct an in-depth study of different quantum circuit encoder architectures to provide insights into designing QCNN-based feature extractors. Neural saliency analyses demonstrate a correlation between the proposed QCNN features, class activation maps, and input spectrograms. We provide an implementation for future studies. Chao-Han Huck Yang, Jun Qi 0002, Samuel Yen-Chi Chen, Sabato Marco Siniscalchi, Xiaoli Ma, Chin-Hui Lee 0001 |
ICASSP | 7 |
| 2021 | Automatic Lip-Reading with Hierarchical Pyramidal Convolution and Self-Attention for Image Sequences with No Word Boundaries
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
Interspeech | 6 |
| 2021 | Scenario-Dependent Speaker Diarization for DIHARD-III Challenge
Jun Du 0002, Maokui He, Shutong Niu, Lei Sun 0010, Chin-Hui Lee 0001 |
Interspeech | 6 |
| 2021 | PATE-AAE: Incorporating Adversarial Autoencoder into Private Aggregation of Teacher Ensembles for Spoken Command ClassificationabstractWe propose using an adversarial autoencoder (AAE) to replace generative adversarial network (GAN) in the private aggregation of teacher ensembles (PATE), a solution for ensuring differential privacy in speech applications. The AAE architecture allows us to obtain good synthetic speech leveraging upon a discriminative training of latent vectors. Such synthetic speech is used to build a privacy-preserving classifier when non-sensitive data is not sufficiently available in the public domain. This classifier follows the PATE scheme that uses an ensemble of noisy outputs to label the synthetic samples and guarantee $\varepsilon$-differential privacy (DP) on its derived classifiers. Our proposed framework thus consists of an AAE-based generator and a PATE-based classifier (PATE-AAE). Evaluated on the Google Speech Commands Dataset Version II, the proposed PATE-AAE improves the average classification accuracy by +$2.11\%$ and +$6.60\%$, respectively, when compared with alternative privacy-preserving solutions, namely PATE-GAN and DP-GAN, while maintaining a strong level of privacy target at $\varepsilon$=0.01 with a fixed $δ$=10$^{-5}$. Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
Interspeech | 3 |
| 2021 | A Maximum Likelihood Approach to SNR-Progressive Learning Using Generalized Gaussian Distribution for LSTM-Based Speech Enhancement
Jun Du 0002, Li Chai 0002, Chin-Hui Lee 0001 |
Interspeech | 4 |
| 2021 | Audio-Visual Information Fusion Using Cross-Modal Teacher-Student Learning for Voice Activity Detection in Realistic Environments
Hengshun Zhou, Jun Du 0002, Hang Chen 0001, Zijun Jing, Shifu Xiong, Chin-Hui Lee 0001 |
Interspeech | 6 |
| 2021 | Acoustic Modeling for Multi-Array Conversational Speech Recognition in the Chime-6 ChallengeabstractThis paper presents our main contributions of acoustic modeling for multi-array multi-talker speech recognition in the CHiME-6 Challenge, exploring different strategies for acoustic data augmentation and neural network architectures. First, enhanced data from our front-end network preprocessing and spectral augmentation are investigated to be effective for improving speech recognition performance. Second, several neural network architectures are explored by different combinations of deep residual network (ResNet), factorized time delay neural network (TDNNF) and residual bidirectional long short-term memory (RBiLSTM). Finally, multiple acoustic models can be combined via minimum Bayes risk fusion. Compared with the official baseline acoustic model, the proposed solution can achieve a relatively word error rate reduction of 19% for the best single ASR system on the evaluation data, which is also one of main contributions to our top system for the Track 1 tasks of the CHiME-6 Challenge. Li Chai 0002, Jun Du 0002, Diyuan Liu, Yanhui Tu, Chin-Hui Lee 0001 |
SLT | 5 |
| 2021 | Correlating subword articulation with lip shapes for embedding aware audio-visual speech enhancement
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
Neural Networks | 6 |
| 2021 | A Cross-Entropy-Guided Measure (CEGM) for Assessing Speech Recognition Performance and Optimizing DNN-Based Speech EnhancementabstractA new cross-entropy-guided measure (CEGM) is proposed to indirectly assess accuracies of automatic speech recognition (ASR) of degraded speech with a speech enhancement front-end and without directly performing ASR experiments. The proposed CEGM is calculated in three steps, namely: (1) a low-level representations via feature extraction, (2) a high-level nonlinear mapping using an acoustic model, and (3) a final CEGM calculation between the high-level representations of clean and enhanced speech. Specifically, state posterior probabilities from outputs of conventional hybrid acoustic model of the target ASR system are adopted as the high-level representations and a cross-entropy criterion is used to calculate the CEGM. Due to CEGM's differentiability, it can also be used to replace the conventional minimum mean squared error (MMSE) criterion as an objective function for deep neural network (DNN)-based speech enhancement. Therefore, the front-end enhancement model can be optimized towards improving the accuracies of the back-end ASR system. Experiments on single-channel CHiME-4 Challenge show that CEGM yields consistently the highest correlations with word error rate (WER) which is often costly to calculate, and achieves the most accurate assessment of ASR performance when compared to the perceptual evaluation metrics commonly used for assessing speech enhancement performance. Furthermore, CEGM-optimized speech enhancement could effectively reduce the WER on the CHiME-4 real test set when compared to unprocessed noisy speech and enhanced speech obtained with MMSE-optimized enhancement for ASR systems with fixed multi-condition acoustic models in various deep architectures. Li Chai 0002, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Information Fusion in Attention Networks Using Adaptive and Multi-Level Factorized Bilinear Pooling for Audio-Visual Emotion RecognitionabstractMultimodal emotion recognition is a challenging task in emotion computing as it is quite difficult to extract discriminative features to identify the subtle differences in human emotions with abstract concept and multiple expressions. Moreover, how to fully utilize both audio and visual information is still an open problem. In this paper, we propose a novel multimodal fusion attention network for audio-visual emotion recognition based on adaptive and multi-level factorized bilinear pooling (FBP). First, for the audio stream, a fully convolutional network (FCN) equipped with 1-D attention mechanism and local response normalization is designed for speech emotion recognition. Next, a global FBP (G-FBP) approach is presented to perform audio-visual information fusion by integrating self-attention based video stream with the proposed audio stream. To improve G-FBP, an adaptive strategy (AG-FBP) to dynamically calculate the fusion weight of two modalities is devised based on the emotion-related representation vectors from the attention mechanism of respective modalities. Finally, to fully utilize the local emotion information, adaptive and multi-level FBP (AM-FBP) is introduced by combining both global-trunk and intra-trunk data in one recording on top of AG-FBP. Tested on the IEMOCAP corpus for speech emotion recognition with only audio stream, the new FCN method outperforms the state-of-the-art results with an accuracy of 71.40%. Moreover, validated on the AFEW database of EmotiW2019 sub-challenge and the IEMOCAP corpus for audio-visual emotion recognition, the proposed AM-FBP approach achieves the best accuracy of 63.09% and 75.49% respectively on the test set. Hengshun Zhou, Jun Du 0002, Qing Wang 0008, Qingfeng Liu, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | High-Resolution Attention Network with Acoustic Segment Model for Acoustic Scene ClassificationabstractThe spectral information of acoustic scenes is diverse and complex, which poses challenges for acoustic scene tasks. To improve the classification performance, a variety of convolutional neural networks (CNNs) are proposed to extract richer semantic information of scene utterances. However, the different regions of the features extracted from CNN-based encoder have different importance. In this paper, we propose a novel strategy for acoustic scene classification, namely high-resolution attention network with acoustic segment model (HRAN-ASM). In this approach, we utilize fully CNN to obtain high-level semantic information and then adopt two-stage attention strategy to select the relevant acoustic scene segments. Besides, the acoustic segment model (ASM) proposed in our recent work provides embedding vectors for this attention mechanism. The performance is evaluated on DCASE 2018 Task 1a, showing 70.5% good classification accuracy under single system and no data expansion, which is superior to CNN-based self-attention mechanism and highly competitive. Jun Du 0002, Hengshun Zhou, Yanhui Tu, Chin-Hui Lee 0001 |
ICASSP | 6 |
| 2020 | L-Vector: Neural Label Embedding for Domain AdaptationabstractWe propose a novel neural label embedding (NLE) scheme for the domain adaptation of a deep neural network (DNN) acoustic model with unpaired data samples from source and target domains. With NLE method, we distill the knowledge from a powerful source-domain DNN into a dictionary of label embeddings, or l-vectors, one for each senone class. Each l-vector is a representation of the senone-specific output distributions of the source-domain DNN and is learned to minimize the average L2, Kullback-Leibler (KL) or symmetric KL distance to the output vectors with the same label through simple averaging or standard back-propagation. During adaptation, the l-vectors serve as the soft targets to train the target-domain model with cross-entropy loss. Without parallel data constraint as in the teacher-student learning, NLE is specially suited for the situation where the paired target-domain data cannot be simulated from the source-domain data. We adapt a 6400 hours multi-conditional US English acoustic model to each of the 9 accented English (80 to 830 hours) and kids' speech (80 hours). NLE achieves up to 14.1% relative word error rate reduction over direct re-training with one-hot labels. Zhong Meng, Hu Hu, Jinyu Li 0001, Changliang Liu, Yan Huang 0028, Yifan Gong 0001, Chin-Hui Lee 0001 |
ICASSP | 7 |
| 2020 | A Maximum Likelihood Approach to Multi-Objective Learning Using Generalized Gaussian Distributions for Dnn-Based Speech EnhancementabstractThe multi-objective learning using minimum mean squared error criterion for DNN-based speech enhancement (MMSE-MOL-DNN) has been demonstrated to achieve better performance than single output DNN. However, one problem of MMSE-MOL-DNN is that the prediction error values on different targets have a very broad dynamic range, causing difficulty in DNN training. In this paper, we extend the maximum likelihood approach proposed in our previous work [1] to the multi-objective learning for DNN-based speech enhancement (ML-MOL-DNN) to achieve the automatic adjustment of the dynamic range of prediction error values on different targets. The conditional likelihood function to be maximized is derived under the generalized Gaussian distribution (GGD) error model. Moreover, the control of the dynamic range of the prediction error values on different targets is achieved by the scale factors in GGD. Furthermore, we propose a method to update the shape factors automatically utilizing the one-to-one mapping between the kurtosis and shape factor in GGD instead of manual adjustment. The experimental results show that our ML-MOL-DNN can achieve better performance than MMSE-MOL-DNN in terms of different objective measures. Shutong Niu, Jun Du 0002, Li Chai 0002, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2020 | Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train NetworkabstractWe propose a tensor-to-vector regression approach to multi-channel speech enhancement in order to address the issue of input size explosion and hidden-layer size expansion. The key idea is to cast the conventional deep neural network (DNN) based vector-to-vector regression formulation under a tensor-train network (TTN) framework. TTN is a recently emerged solution for compact representation of deep models with fully connected hidden layers. Thus TTN maintains DNN's expressive power yet involves a much smaller amount of trainable parameters. Furthermore, TTN can handle a multi-dimensional tensor input by design, which exactly matches the desired setting in multi-channel speech enhancement. We first provide a theoretical extension from DNN to TTN based regression. Next, we show that TTN can attain speech enhancement quality comparable with that for DNN but with much fewer parameters, e.g., a reduction from 27 million to only 5 million parameters is observed in a single-channel scenario. TTN also improves PESQ over DNN from 2.86 to 2.96 by slightly increasing the number of trainable parameters. Finally, in 8-channel conditions, a PESQ of 3.12 is achieved using 20 million parameters for TTN, whereas a DNN with 68 million parameters can only attain a PESQ of 3.06. Jun Qi 0002, Hu Hu, Yannan Wang, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 6 |
| 2020 | Progressive Multi-Target Network Based Speech Enhancement with Snr-Preselection for Robust Speaker DiarizationabstractIn this paper, we design a novel front-end processing system for speaker diarization under realistic conditions with challenging background noises. To cope with diversified environments, we first extend our perviously proposed progressive learning based speech enhancement model by adding multi-task learning in each intermediate layer. The corresponding progressive multi-target (PMT) in various layers includes both progressive ratio mask (PRM) and progressively enhanced log-power spectra (PELPS) with specified signal-to-noise ratios (SNRs). Speech distortions are commonly introduced during the front-end processing, which often deteriorate the back-end performance. However, the proposed speech enhancement model can be regarded as a bagging of models with multiple learning objectives, which provides flexibility for selecting the most appropriate output for robust speaker diarzation. In addition, a global SNR estimation is performed using the results of deep neural network (DNN) based speech activity detection (SAD) to decide whether the audio should be enhanced. We evaluate the speaker diarzation performance on the second DIHARD dataset which includes several different realistic conditions. Compared with the original data, experiments demonstrate that the enhanced data processed by our proposed method can effectively avoid the performance loss of every single domain, and achieve consistent improvements in most domains. Lei Sun 0010, Jun Du 0002, Xueyang Zhang, Tian Gao 0005, Chin-Hui Lee 0001 |
ICASSP | 6 |
| 2020 | Geometry Constrained Progressive Learning for Lstm-Based Speech EnhancementabstractIn our previous work, a progressive learning framework for long short-term memory (LSTM)-based speech enhancement was proposed to improve the performance in low SNR environment, where each LSTM layer is guided to learn an intermediate target with a specific SNR gain via the MMSE criterion. However, the constraint relationship among these targets is not considered in the objective function. In this paper, we incorporate two kinds of geometric constraints among these targets into the objective function to help LSTM achieve better training. One constraint is edge constraint and the other is the centroid constraint. In addition, we propose a method for constructing the intermediate targets online. It saves device storage space and alleviates the trouble of manually constructing intermediate targets. Experiment results demonstrate these geometric constraints can bring remarkable improvements in low SNR environments. Jun Du 0002, Li Chai 0002, Yannan Wang, Qing Wang 0008, Chin-Hui Lee 0001 |
ICASSP | 6 |
| 2020 | 2D-to-2D Mask Estimation for Speech Enhancement Based on Fully Convolutional Neural NetworkabstractIn recent years, the deep learning-based approaches are popular in the field of singe-channel speech enhancement. Convolutional neural networks (CNNs) are a standard component of many current speech enhancement system. In this study, we design a new Fully CNN (FCNN)-based regression model, which can directly achieve the 2-dimensional (2D) noisy lpg-power spectra (LPS) input to 2dimensional (2D) time-frequency mask output mapping, denoted as 2D-RFCNN. First, the whole 2D noisy LPS of one utterance is directly used as network input to make sure each convolutional filter can see more contextual information. Second, we only use the pooling operation on the frequency bin to ensure that the final dimension of frequency bin has a value of 1 and make the number of feature mapping same to frequency dimension, simultaneously. Finally, we also use the deep convolutional layers with a small size of filter, which is popularly used in speech recognition, for speech enhancement. Experiments of the CHiME-4 challenge task shows that our proposed 2D-RFCNN model not only improves the speech quality (PESQ) and intelligibility (STOI), but also reduces the recognition error rate on real test set. Yanhui Tu, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2020 | A Study of Child Speech Extraction Using Joint Speech Enhancement and Separation in Realistic ConditionsabstractIn this paper, we design a novel joint framework of speech enhancement and speech separation for child speech extraction in realistic conditions, targeting the problem of extracting child speech from daily conversations in BabyTrain mega corpus. To the best of our knowledge, it is the first discussion of a feasible method for child speech extraction in realistic conditions. First, we make detailed analysis of the BabyTrain mega corpus, which is recorded in adverse environments. We observe problems of background noises, reverberations and child speech that is partially obscured by adult speech (for instance due to speaker overlap but also imitation by the adult). Motivated by this, we conduct a joint framework of speech enhancement and speech separation for child speech extraction. To measure the extraction results in realistic conditions, we propose several objective measurements to evaluate the performance of the our system, which is different from those commonly used for simulation data. Compared with the unprocessed approach and classification approach, our proposed approach can yield the best performance among all subsets of BabyTrain. Xin Wang 0037, Jun Du 0002, Alejandrina Cristià, Lei Sun 0010, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2020 | A Cross-Task Transfer Learning Approach to Adapting Deep Speech Enhancement Models to Unseen Background Noise Using Paired Senone ClassifiersabstractWe propose an environment adaptation approach that improves deep speech enhancement models via minimizing the Kullback-Leibler divergence between posterior probabilities produced by a multi-condition senone classifier (teacher) fed with noisy speech features and a clean-condition senone classifier (student) fed with enhanced speech features to transfer an existing deep neural network (DNN) speech enhancer to specific noisy environments without using noisy/clean paired target waveforms needed in conventional DNN-based spectral regression. Our solution not only improves listening quality in the enhanced speech but also boosts noise robustness of existing automatic speech recognition (ASR) systems trained on clean data if employed as a pre-processing step before speech feature extraction. Experimental results show steady gains in objective quality measurements as a result of a teacher network producing adaptation targets for a student enhancement model to adjust its parameters in unseen noise conditions. The proposed technique is particularly advantageous in environments that are not handled effectively by the unadapted DNN-based enhancer, as we find that only very little data from a specific operating condition is required to yield good improvements. Finally, higher gains in speech quality directly translate to larger improvements in ASR. Sicheng Wang 0004, Wei Li 0119, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2020 | Characterizing Speech Adversarial Examples Using Self-Attention U-Net EnhancementabstractRecent studies have highlighted adversarial examples as ubiquitous threats to the deep neural network (DNN) based speech recognition systems. In this work, we present a U-Net based attention model, UNetAt, to enhance adversarial speech signals. Specifically, we evaluate the model performance by interpretable speech recognition metrics and discuss the model performance by the augmented adversarial training. Our experiments show that our proposed U-NetAtimproves the perceptual evaluation of speech quality (PESQ) from 1.13 to 2.78, speech transmission index (STI) from 0.65 to 0.75, shortterm objective intelligibility (STOI) from 0.83 to 0.96 on the task of speech enhancement with adversarial speech examples. We conduct experiments on the automatic speech recognition (ASR) task with adversarial audio attacks. We find that (i) temporal features learned by the attention network are capable of enhancing the robustness of DNN based ASR models; (ii) the generalization power of DNN based ASR model could be enhanced by applying adversarial training with an additive adversarial data augmentation. The ASR metric on word-error-rates (WERs) shows that there is an absolute 2.22 % decrease under gradient-based perturbation, and an absolute 2.03 % decrease, under evolutionary-optimized perturbation, which suggests that our enhancement models with adversarial training can further secure a resilient ASR system. Chao-Han Huck Yang, Jun Qi 0002, Xiaoli Ma, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2020 | Enhanced Adversarial Strategically-Timed Attacks Against Deep Reinforcement LearningabstractRecent deep neural networks based techniques, especially those equipped with the ability of self-adaptation in the system level such as deep reinforcement learning (DRL), are shown to possess many advantages of optimizing robot learning systems (e.g., autonomous navigation and continuous robot arm control.) However, the learning-based systems and the associated models may be threatened by the risks of intentionally adaptive (e.g., noisy sensor confusion) and adversarial perturbations from real-world scenarios. In this paper, we introduce timing-based adversarial strategies against a DRL-based navigation system by jamming in physical noise patterns on the selected time frames. To study the vulnerability of learning-based navigation systems, we propose two adversarial agent models: one refers to online learning; another one is based on evolutionary learning. Besides, three open-source robot learning and navigation control environments are employed to study the vulnerability under adversarial timing attacks. Our experimental results show that the adversarial timing attacks can lead to a significant performance drop, and also suggest the necessity of enhancing the robustness of robot learning systems. Chao-Han Huck Yang, Jun Qi 0002, I-Te Danny Hung, Chin-Hui Lee 0001, Xiaoli Ma |
ICASSP | 6 |
| 2020 | An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial UtterancesabstractIn this paper, we propose a sub-utterance unit selection framework to remove acoustic segments in audio recordings that carry little information for acoustic scene classification (ASC). Our approach is built upon a universal set of acoustic segment units covering the overall acoustic scene space. First, those units are modeled with acoustic segment models (ASMs) used to tokenize acoustic scene utterances into sequences of acoustic segment units. Next, paralleling the idea of stop words in information retrieval, stop ASMs are automatically detected. Finally, acoustic segments associated with the stop ASMs are blocked, because of their low indexing power in retrieval of most acoustic scenes. In contrast to building scene models with whole utterances, the ASM-removed sub-utterances, i.e., acoustic utterances without stop acoustic segments, are then used as inputs to the AlexNet-L back-end for final classification. On the DCASE 2018 dataset, scene classification accuracy increases from 68%, with whole utterances, to 72.1%, with segment selection. This represents a competitive accuracy without any data augmentation, and/or ensemble strategy. Moreover, our approach compares favourably to AlexNet-L with attention. Hu Hu, Sabato Marco Siniscalchi, Yannan Wang, Jun Du 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2020 | Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene ClassificationabstractIn this paper, we propose a domain adaptation framework to address the device mismatch issue in acoustic scene classification leveraging upon neural label embedding (NLE) and relational teacher student learning (RTSL). Taking into account the structural relationships between acoustic scene classes, our proposed framework captures such relationships which are intrinsically device-independent. In the training stage, transferable knowledge is condensed in NLE from the source domain. Next in the adaptation stage, a novel RTSL strategy is adopted to learn adapted target models without using paired source-target data often required in conventional teacher student learning. The proposed framework is evaluated on the DCASE 2018 Task1b data set. Experimental results based on AlexNet-L deep classification models confirm the effectiveness of our proposed approach for mismatch situations. NLE-alone adaptation compares favourably with the conventional device adaptation and teacher student based adaptation techniques. NLE with RTSL further improves the classification accuracy Hu Hu, Sabato Marco Siniscalchi, Yannan Wang, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2020 | Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech EnhancementabstractThis paper investigates different trade-offs between the number of model parameters and enhanced speech qualities by employing several deep tensor-to-vector regression models for speech enhancement. We find that a hybrid architecture, namely CNN-TT, is capable of maintaining a good quality performance with a reduced model parameter size. CNN-TT is composed of several convolutional layers at the bottom for feature extraction to improve speech quality and a tensor-train (TT) output layer on the top to reduce model parameters. We first derive a new upper bound on the generalization power of the convolutional neural network (CNN) based vector-to-vector regression models. Then, we provide experimental evidence on the Edinburgh noisy speech corpus to demonstrate that, in single-channel speech enhancement, CNN outperforms DNN at the expense of a small increment of model sizes. Besides, CNN-TT slightly outperforms the CNN counterpart by utilizing only 32% of the CNN model parameters. Besides, further performance improvement can be attained if the number of CNN-TT parameters is increased to 44% of the CNN model size. Finally, our experiments of multi-channel speech enhancement on a simulated noisy WSJ0 corpus demonstrate that our proposed hybrid CNN-TT architecture achieves better results than both DNN and CNN models in terms of better-enhanced speech qualities and smaller parameter sizes. Jun Qi 0002, Hu Hu, Yannan Wang, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2020 | A Space-and-Speaker-Aware Iterative Mask Estimation Approach to Multi-Channel Speech Recognition in the CHiME-6 Challenge
Yanhui Tu, Jun Du 0002, Lei Sun 0010, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2020 | A Noise-Aware Memory-Attention Network Architecture for Regression-Based Speech Enhancement
Jun Du 0002, Li Chai 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2020 | Using Speech Enhancement Preprocessing for Speech Emotion Recognition in Realistic Noisy Conditions
Hengshun Zhou, Jun Du 0002, Yanhui Tu, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2020 | On Mean Absolute Error for Deep Neural Network Based Vector-to-Vector RegressionabstractIn this paper, we exploit the properties of mean absolute error (MAE) as a loss function for the deep neural network (DNN) based vector-to-vector regression. The goal of this work is two-fold: (i) presenting performance bounds of MAE, and (ii) demonstrating new properties of MAE that make it more appropriate than mean squared error (MSE) as a loss function for DNN based vector-to-vector regression. First, we show that a generalized upper-bound for DNN-based vector-to-vector regression can be ensured by leveraging the known Lipschitz continuity property of MAE. Next, we derive a new generalized upper bound in the presence of additive noise. Finally, in contrast to conventional MSE commonly adopted to approximate Gaussian errors for regression, we show that MAE can be interpreted as an error modeled by Laplacian distribution. Speech enhancement experiments are conducted to corroborate our proposed theorems and validate the performance advantages of MAE over MSE for DNN based regression. Jun Qi 0002, Jun Du 0002, Sabato Marco Siniscalchi, Xiaoli Ma, Chin-Hui Lee 0001 |
IEEE Signal Process. Lett. | 5 |
| 2020 | A Multi-Target SNR-Progressive Learning Approach to Regression Based Speech EnhancementabstractWe propose a multi-target, signal-to-noise-ratio (SNR)-progressive learning (SNR-PL) framework for regression based speech enhancement (SE). At low SNR levels, it is often not easy to directly learn the complicated regression required in SE. We therefore decompose the original SE problem of mapping noisy to clean speech features, with a large SNR gap, into a series of sub-problems, each with a small SNR increment and presumably easier to learn. In our configurations, each hidden layer of the proposed regression neural network is guided to explicitly learn an intermediate target with a specified but small SNR gain. Tested on both deep neural network (DNN) and long short-term memory (LSTM) architectures, SNR-PL consistently outperforms the conventional “black box” DNN framework in terms of both objective measure superiority and network model compactness. Furthermore, with the best configured LSTM-based SNR-PL model, we often observe that the performance is easily saturated or even degraded when increasing the number of intermediate targets, due to the fact that useful information is lost in dimension reduction when involving more target layers. Accordingly, to address this information loss issue, we explore densely connected networks on top of the LSTM structure where the input and the preceding intermediate targets are concatenated together to learn the next target. Finally, to fully utilize the rich and complementary information of intermediate targets, a simple post-processing strategy is adopted to further improve the performance. Evaluated on the simulation speech data, experimental results in unseen noises cases demonstrate that the proposed approach consistently performs better than the conventional LSTM approach in terms of objective speech enhancement measures for speech intelligibility and quality. Furthermore, when evaluated on real data provided by the CHiME-4 Challenge for automatic speech recognition (ASR) of noisy microphone array speech, we show that the proposed approach with intermediate outputs can directly improve the ASR performance, while the conventional LSTM approach increases the word error rate. Yanhui Tu, Jun Du 0002, Tian Gao 0005, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Improving Audio-visual Speech Recognition Performance with Cross-modal Student-teacher TrainingabstractIn this paper, we propose a cross-modal student-teacher learning framework to make a full use of externally abundant acoustic data in addition to a given task-specific audio-visual training database for improving speech recognition performance under the low signal-to-noise-ratio (SNR) and acoustic mismatch conditions. First, a teacher model is trained with large-sized audio-only databases. Next, a student, namely a deep neural network (DNN) model, is trained on a small-sized audio-visual database to minimize the Kullback-Leibler (KL) divergence between its output and the posterior distribution of the teacher. We evaluate the proposed approach in both matched and mismatch acoustic conditions for phone recognition with the NTCD-TIMIT database. Compared to the DNN recognition system trained with the original audio-visual data only, the proposed solution reduces the phone error rate (PER) from 26.7% to 21.3% on a matched acoustic scenario. In the mismatch conditions, the PER is reduced from 47.9% to 42.9%. Moreover, we show that posteriors generated by the teacher contain environmental information, which enables our proposed student-teacher learning to work as an environmental-aware training and good PER reductions are observed in all SNR conditions. Wei Li 0119, Sicheng Wang 0004, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2019 | A Two-stage Single-channel Speaker-dependent Speech Separation Approach for Chime-5 ChallengeabstractIn this paper, we design a two-stage single-channel speaker-dependent speech separation approach for the CHiME-5 Challenge, targeting the problem of far-field and multi-talker conversational speech recognition in dinner party scenarios involving background noises, reverberations and overlapping speech. First, we make detailed analysis of the CHiME-5 data and observe problems of inaccurate human annotations and low-resource useable data for target speakers. Motivated by this, we conduct a first-stage speaker-dependent speech separation with a learning target for aggressive segregation to generate more and purer target speech data. Then a second-stage speaker-dependent speech separation with a new learning target is performed to obtain the final speech masks, which can be directly fed to back-end acoustic model. Compared with the official baseline, our proposed approach can yield an absolute word error rate reduction of 5.3%, namely from 81.3% to 76.0% in development test set. To the best of our knowledge, it is the first time to discuss a feasible method of single-channel speaker-dependent speech separation for such a challenging task although we make an assumption of oracle speaker diarization following the challenge rules. By integrating this crucial technique, our submitted systems achieved the first place of all four tasks in the CHiME-5 challenge. Lei Sun 0010, Jun Du 0002, Tian Gao 0005, Chin-Hui Lee 0001 |
ICASSP | 7 |
| 2019 | DNN Training Based on Classic Gain Function for Single-channel Speech Enhancement and RecognitionabstractFor conventional single-channel speech enhancement based on noise power spectrum, the speech gain function, which suppresses background noise at each time-frequency bin, is calculated by prior signal-to-noise-ratio (SNR). Hence, accurate prior SNR estimation is paramount for successful noise suppression. Accordingly, we have proposed a single-channel approach to combine conventional and deep learning techniques for speech enhancement and automatic speech recognition (ASR) recently. However, the combination process is at the testing stage, which is time-consuming with a complicated procedure. In this study, the gain function of classic speech enhancement will be utilized to optimize the ideal ratio mask based deep neural network (DNN-IRM) at the training stage, denoted as GF-DNN-IRM. And at the testing stage, the estimated IRM by GF-DNN-IRM model is directly used to generate enhanced speech without involving the conventional speech enhancement process. In addition, DNNs with less parameters in the causal processing mode are also discussed. Experiments of the CHiME-4 challenge task show that our proposed algorithm can achieve a relative word error rate reduction of 6.57% on RealData test set comparing to unprocessed speech without acoustic model retraining in causal mode, while the traditional DNN-IRM method fails to improve ASR performance in this case. Yanhui Tu, Jun Du 0002, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2019 | KL-Divergence Regularized Deep Neural Network Adaptation for Low-Resource Speaker-Dependent Speech Enhancement
Li Chai 0002, Jun Du 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2019 | A Cross-Entropy-Guided (CEG) Measure for Speech Enhancement Front-End Assessing Performances of Back-End Automatic Speech Recognition
Li Chai 0002, Jun Du 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2019 | A Hybrid Approach to Acoustic Scene Classification Based on Universal Acoustic Models
Jun Du 0002, Zi-Rui Wang, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2019 | Acoustic Model Ensembling Using Effective Data Augmentation for CHiME-5 Challenge
Li Chai 0002, Jun Du 0002, Diyuan Liu, Zhongfu Ye, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2019 | An iterative mask estimation approach to deep learning based multi-channel speech recognition
Yanhui Tu, Jun Du 0002, Lei Sun 0010, Hai-Kun Wang, Jingdong Chen, Chin-Hui Lee 0001 |
Speech Commun. | 7 |
| 2019 | Using Generalized Gaussian Distributions to Improve Regression Error Modeling for Deep Learning-Based Speech EnhancementabstractFrom a statistical perspective, the conventional minimum mean squared error (MMSE) criterion can be considered as the maximum likelihood (ML) solution under an assumed homoscedastic Gaussian error model. However, in this paper, a statistical analysis reveals the super-Gaussian and heteroscedastic properties of the prediction errors in nonlinear regression deep neural network (DNN)-based speech enhancement when estimating clean log-power spectral (LPS) components at DNN outputs with noisy LPS features in DNN input vectors. Accordingly, we propose treating all dimensions of the prediction error vector as statistically independent random variables and model them with generalized Gaussian distributions (GGDs). Then, the objective function with the GGD error model is derived according to the ML criterion. Experiments on the TIMIT corpus corrupted by simulated additive noises show consistent improvements of our proposed DNN framework over the conventional DNN framework in terms of various objective quality measures under 14 unseen noise types evaluated and at various signal-to-noise ratio levels. Furthermore, the ML optimization objective with GGD outperforms the conventional MMSE criterion, achieving improved generalization and robustness. Li Chai 0002, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Improving Mispronunciation Detection of Mandarin Tones for Non-Native Learners With Soft-Target Tone Labels and BLSTM-Based Deep Tone ModelsabstractWe investigate the effectiveness of soft-target tone labels and sequential context information for mispronunciation detection of Mandarin lexical tones pronounced by second language (L2) learners whose first language (L1) is of European origin. In conventional approaches, prosodic information (e.g., F0 and tone posteriors extracted from trained tone models) is used to calculate goodness of pronunciation (GOP) scores or train binary classifiers to verify pronunciation correctness. We propose three techniques to improve detection of mispronunciation of Mandarin tones for non-native learners. First, we extend our tone model from a deep neural network (DNN) to a bidirectional long short-term memory (BLSTM) network in order to more accurately model the high variability of non-native tone productions and the contextual information expressed in tone-level co-articulation. Second, we characterize ambiguous pronunciations where L2 learners' tone realizations are between two canonical tone categories by relaxing hard target labels to soft targets with probabilistic transcriptions. Third, segmental tone features fed into verifiers are extracted by a BLSTM to exploit sequential context information to improve mispronunciation detection. Compared to DNN-GOP trained with hard targets, the proposed BLSTM-GOP framework trained with soft targets reduces the tones' averaged equal error rate (ERR) from 7.58% to 5.83% and the averaged area under ROC curve (AUC) is increased from 97.85% to 98.31%. By utilizing BLSTM-based verifiers the EER further decreases to 5.16%, and the AUC is increased to 98.47%. Wei Li 0119, Nancy F. Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | A Theory on Deep Neural Network Based Vector-to-Vector Regression With an Illustration of Its Expressive Power in Speech EnhancementabstractThis paper focuses on a theoretical analysis of deep neural network (DNN) based functional approximation. Leveraging upon two classical theorems on universal approximation, an artificial neural network (ANN) with a single hidden layer of neurons is used. With modified ReLU and Sigmoid activation functions, we first generalize the related concepts to vector-to-vector regression. Then, we show that the width of the hidden layer of ANN is numerically related to the approximation of the regression function. Furthermore, we increase the number of hidden layers and show that the depth of the ANN-based regression function can enhance its expressive power. We illustrate this representation with recently-emerged DNN based speech enhancement. We first compare the expressive power by varying ANN structures and then test its related regression performance under different noisy conditions in various noise types and signal-to-noise-ratio levels. Experimental results verify our theoretical prediction that an ANN of a broader hidden layer and a deeper architecture can jointly ensure a closer approximation of the vector-to-vector regression functions in terms of the Euclidean distance between the log power spectra of noisy and expected clean speech. Moreover, a DNN with a broader width at the top hidden layer can improve the regression performance relative to those with a narrower width at the top hidden layers. Jun Qi 0002, Jun Du 0002, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Speech Enhancement Based on Teacher-Student Deep Learning Using Improved Speech Presence Probability for Noise-Robust Speech RecognitionabstractIn this paper, we propose a novel teacher-student learning framework for the preprocessing of a speech recognizer, leveraging the online noise tracking capabilities of improved minima controlled recursive averaging (IMCRA) and deep learning of nonlinear interactions between speech and noise. First, a teacher model with deep architectures is built to learn the target of ideal ratio masks (IRMs) using simulated training pairs of clean and noisy speech data. Next, a student model is trained to learn an improved speech presence probability by incorporating the estimated IRMs from the teacher model into the IMCRA approach. The student model can be compactly designed in a causal processing mode having no latency with the guidance of a complex and noncausal teacher model. Moreover, the clean speech requirement, which is difficult to meet in real-world adverse environments, can be relaxed for training the student model, implying that noisy speech data can be directly used to adapt the regression-based enhancement model to further improve speech recognition accuracies for noisy speech collected in such conditions. Experiments on the CHiME-4 challenge task show that our best student model with bidirectional gated recurrent units (BGRUs) can achieve a relative word error rate (WER) reduction of 18.85% for the real test set when compared to unprocessed system without acoustic model retraining. However, the traditional teacher model degrades the performance of the unprocessed system in this case. In addition, the student model with a deep neural network (DNN) in causal mode having no latency yields a relative WER reduction of 7.94% over the unprocessed system with 670 times less computing cycles when compared to the BGRU-equipped student model. Finally, the conventional speech enhancement and IRM-based deep learning method destroyed the ASR performance when the recognition system became more powerful. While our proposed approach could still improve the ASR performance even in the more powerful recognition system. Yanhui Tu, Jun Du 0002, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Densely Connected Progressive Learning for LSTM-Based Speech EnhancementabstractRecently, we proposed a novel progressive learning (PL) framework for deep neural network (DNN) based speech enhancement to improve the performance in low signal-to-noise ratio (SNR) environments. In this study, several new contributions are made to this framework. First, the advanced long short-term memory (LSTM) architecture is adopted to achieve better results, namely LSTM-PL, where each LSTM layer is guided to explicitly learn an intermediate target with a specific SNR gain. However, we observe that the performance of LSTM-PL architecture is easily degraded by increasing the number of intermediate targets due to the possible information loss when involving more target layers. Accordingly, we propose densely connected progressive learning in which the input and the estimations of intermediate targets are spliced together to learn the next target. This new structure can fully utilize the rich set of information from the multiple learning targets and alleviate the information loss problem. Experimental results demonstrate that the dense structure with deeper LSTM layers can yield significant gains of speech intelligibility measure for all noise types and levels. Moreover, the post-processing with more targets tends to achieve better performance. Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2018 | Improving Mandarin Tone Mispronunciation Detection for Non-Native Learners with Soft-Target Tone Labels and BLSTM-Based Deep ModelsabstractWe propose three techniques to improve mispronunciation detection of Mandarin tones of second language (L2) learners using tone-based extended recognition network (ERN). First, we extend our model from deep neural network (DNN) to bidirectionallon-short-term memory (BLSTM) in order to model tone-level co-articulation influenced by a broader temporal context (e.g., two or three consecutive Mandarin syllables). Second, we relax the hard labels to characterize the situations when a single tone class label is not enough because L2 learners' pronunciations are often between two canonical tone categories. Therefore, soft targets (a probabilistic transcription) are proposed for acoustic model training in place of conventional hard targets (one-hot targets). Third, we average tone scores produced by BLSTM models trained with hard and soft targets to seek the complementarity from modeling at the tone-target levels. Compared to our previous system based on the DNN-trained ERNs, the BLSTM-trained system with soft targets reduces the equal error rate (ERR) from 5.77% to 4.86%, and system combination decreases EER further to 4.34%, achieving a 24.78% relative error reduction. Wei Li 0119, Nancy F. Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2018 | A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch ConditionsabstractIn this study, we investigate on the effects of deep learning based speech enhancement as a preprocessor to speaker diarization in quite challenging realistic environments involving the background noises, reverberations and overlapping speech. To improve the generalization capability, the advanced long short-term memory (LSTM) architecture with the novel design of hidden layers via densely connected progressive learning and output layer via multiple-target learning is proposed for preprocessing. We build the deep model using synthesized training data pairs generated from WSJO reading-style speech and more than 100 noise types. Surprisingly, this proposed preprocessor demonstrates a strong generalization capability to speaker di-arization with the realistic noisy speech in highly mismatched conditions, in terms of the speaking style, interferences, and the interaction between them. Tested on three challenging tasks, namely AMI, ADOS, and SeedLings, the state-of-the-art diarization system with the novel LSTM-based speech preprocessor can yield consistent and significant reductions of diarization error rate (DER) over the systems using unprocessed noisy speech and traditional enhancement methods. Lei Sun 0010, Jun Du 0002, Tian Gao 0005, Yu-Ding Lu, Yu Tsao 0001, Chin-Hui Lee 0001, Neville Ryant |
ICASSP | 6 |
| 2018 | Error Modeling via Asymmetric Laplace Distribution for Deep Neural Network Based Single-Channel Speech Enhancement
Li Chai 0002, Jun Du 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2018 | Speaker Diarization with Enhancing Speech for the First DIHARD Challenge
Lei Sun 0010, Jun Du 0002, Xueyang Zhang, Chin-Hui Lee 0001 |
INTERSPEECH | 7 |
| 2018 | A Multiobjective Learning and Ensembling Approach to High-Performance Speech Enhancement With Compact Neural Network ArchitecturesabstractIn this study, we propose a novel deep neural network (DNN) architecture for speech enhancement (SE) via a multiobjective learning and ensembling (MOLE) framework to achieve a compact and lowlatency design, while maintaining good performance in quality evaluations. MOLE follows the boosting concept when combining weak models into a strong classifier and consists of two compact DNNs. The first, called the multiobjective learning DNN (MOL-DNN), takes multiple features, such as log-power spectra (LPS), mel-frequency cepstral coefficients (MFCCs) and Gammatone frequency cepstral coefficients (GFCCs) to predict a multiobjective set that includes clean speech feature, dynamic noise feature, and ideal ratio mask (IRM). The second, called the multiobjective ensembling DNN (MOE-DNN), takes the learned features from MOL-DNN as inputs and separately predicts clean LPS and IRM, clean MFCC and IRM, and clean GFCC and IRM using three sets of weak regression functions. Finally, a postprocessing operation can be applied to the estimated clean features by leveraging the multiple targets learned from both the MOL-DNN and the MOE-DNN. On speech corrupted by 15 noise types not seen in model training the SE results show that the MOLE approach, which features a small model size and low run-time latency, can achieve consistent improvements over both DNN- and long short-term memory (LSTM)-based techniques in terms of all the objective metrics evaluated in this study for all three cases (the input contexts contain 1-frame, 4-frame and 7-frame instances). The 1-frame MOLE-based SE system outperforms the DNN-based SE system with a 7-frame input expansion at a 3-frame delay and also achieves better performance than the LSTM-based SE system with 4-frame, no delay expansion by including only 3 previous frames, and with 170 times less processing latency. Qing Wang 0008, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | A transfer learning and progressive stacking approach to reducing deep model sizes with an application to speech enhancementabstractLeveraging upon transfer learning, we distill the knowledge in a conventional wide and deep neural network (DNN) into a narrower yet deeper model with fewer parameters and comparable system performance for speech enhancement. We present three transfer-learning solutions to accomplish our goal. First, the knowledge embedded in the form of the output values of a high-performance DNN is used to guide the training of a smaller DNN model in sequential transfer learning. In the second multi-task transfer learning solution, the smaller DNN is trained to learn the output value of the larger DNN, and the speech enhancement task in parallel. Finally, a progressive stacking transfer learning is accomplished through multi-task learning, and DNN stacking. Our experimental evidences demonstrate 5 times parameter reduction while maintaining similar enhancement performance with the proposed framework. Sicheng Wang 0004, Kehuang Li, Zhen Huang 0001, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2017 | Joint Training of Multi-Channel-Condition Dereverberation and Acoustic Modeling of Microphone Array Speech for Robust Distant Speech RecognitionabstractWe propose a novel data utilization strategy, called multichannel-condition learning, leveraging upon complementary information captured in microphone array speech to jointly train dereverberation and acoustic deep neural network (DNN) models for robust distant speech recognition. Experimental results, with a single automatic speech recognition (ASR) system, on the REVERB2014 simulated evaluation data show that, on 1-channel testing, the baseline joint training scheme attains a word error rate (WER) of 7.47%, reduced from 8.72% for separate training. The proposed multi-channel-condition learning scheme has been experimented on different channel data combinations and usage showing many interesting implications. Finally, training on all 8-channel data and with DNN-based language model rescoring, a state-of-the-art WER of 4.05% is achieved. We anticipate an even lower WER when combining more top ASR systems. Fengpei Ge, Kehuang Li, Bo Wu 0011, Sabato Marco Siniscalchi, Yonghong Yan 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2017 | Improving Mispronunciation Detection for Non-Native Learners with Multisource Information and LSTM-Based Deep ModelsabstractIn this paper, we utilize manner and place of articulation features and deep neural network models (DNNs) with long short-term memory (LSTM) to improve the detection performance of phonetic mispronunciations produced by second language learners. First, we show that speech attribute scores are complementary to conventional phone scores, so they can be concatenated as features to improve a baseline system based only on phone information. Next, pronunciation representation, usually calculated by frame-level averaging in a DNN, is now learned by LSTM, which directly uses sequential context information to embed a sequence of pronunciation scores into a pronunciation vector to improve the perfonnance of subsequent mispronunciation detectors. Finally, when both proposed techniques are incorporated into the baseline phone-based GOP (goodness of pronunciation) classifier system trained on the same data, the integrated system reduces the false acceptance rate (FAR) and false rejection rate (FRR) by 37.90% and 38.44% (relative), respectively, from the baseline system. Wei Li 0119, Nancy F. Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2017 | On Design of Robust Deep Models for CHiME-4 Multi-Channel Speech Recognition with Multiple Configurations of Array Microphones
Yanhui Tu, Jun Du 0002, Lei Sun 0010, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2017 | A Maximum Likelihood Approach to Deep Neural Network Based Nonlinear Spectral Mapping for Single-Channel Speech Separation
Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2017 | An information fusion framework with multi-channel feature concatenation and multi-perspective system combination for the deep-learning-based robust recognition of microphone array speech
Yanhui Tu, Jun Du 0002, Qing Wang 0008, Xiao Bao, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
Comput. Speech Lang. | 6 |
| 2017 | Hierarchical Bayesian combination of plug-in maximum a posteriori decoders in deep neural networks-based speech recognition and speaker adaptation
Zhen Huang 0001, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
Pattern Recognit. Lett. | 3 |
| 2017 | A unified DNN approach to speaker-dependent simultaneous speech enhancement and speech separation in low SNR environments
Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
Speech Commun. | 4 |
| 2017 | Bayesian Unsupervised Batch and Online Speaker Adaptation of Activation Function Parameters in Deep Models for Automatic Speech RecognitionabstractWe present a Bayesian framework to obtain maximum a posteriori (MAP) estimation of a small set of hidden activation function parameters in context-dependent-deep neural network-hidden markov model (CD-DNN-HMM)-based automatic speech recognition (ASR) systems. When applied to speaker adaptation, we aim at transfer learning from a well-trained deep model for a “general” usage to a “personalized” model geared toward a particular talker by using a collection of speaker-specific data. To make the framework applicable to practical situations, we perform adaptation in an unsupervised manner assuming that the transcriptions of the adaptation utterances are not readily available to the ASR system. We conduct a series of comprehensive batch adaptation experiments on the Switchboard ASR task and show that the proposed approach is effective even with CD-DNN-HMM built with discriminative sequential training. Indeed, MAP speaker adaptation reduces the word error rate (WER) to 20.1% from an initial 21.9% on the full NIST 2000 Hub5 benchmark test set. Moreover, MAP speaker adaptation compares favorably with other techniques evaluated on the same speech tasks. We also demonstrate its complementarity to other approaches by applying MAP adaptation to CD-DNN-HMM trained with speaker adaptive features generated through constrained maximum likelihood linear regression and further reduces the WER to 18.6%. Leveraging upon the intrinsic recursive nature in Bayesian adaptation and mitigating possible system constraints on batch learning, we also proposed an incremental approach to unsupervised online speaker adaptation by simultaneously updating the hyperparameters of the approximate posterior densities and the DNN parameters sequentially. The advantage of such a sequential learning algorithm over a batch method is not necessarily in the final performance, but in computational efficiency and reduced storage needs, without having to wait for all the data to be processed. So far, the experimental results are promising. Zhen Huang 0001, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | A Gender Mixture Detection Approach to Unsupervised Single-Channel Speech Separation Based on Deep Neural NetworksabstractWe propose an unsupervised speech separation framework for mixtures of two unseen speakers in a single-channel setting based on deep neural networks (DNNs). We rely on a key assumption that two speakers could be well segregated if they are not too similar to each other. A dissimilarity measure between two speakers is first proposed to characterize the separation ability between competing speakers. We then show that speakers with the same or different genders can often be separated if two speaker clusters, with large enough distances between them, for each gender group could be established, resulting in four speaker clusters. Next, a DNN-based gender mixture detection algorithm is proposed to determine whether the two speakers in the mixture are females, males, or from different genders. This detector is based on a newly proposed DNN architecture with four outputs, two of them representing the female speaker clusters and the other two characterizing the male groups. Finally, we propose to construct three independent speech separation DNN systems, one for each of the female-female, male-male, and female-male mixture situations. Each DNN gives dual outputs, one representing the target speaker group and the other characterizing the interfering speaker cluster. Trained and tested on the speech separation challenge corpus, our experimental results indicate that the proposed DNN-based approach achieves large performance gains over the state-of-the-art unsupervised techniques without using any specific knowledge about the mixed target and interfering speakers being segregated. Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | A Reverberation-Time-Aware Approach to Speech Dereverberation Based on Deep Neural NetworksabstractA reverberation-time-aware deep-neural-network (DNN)-based speech dereverberation framework is proposed to handle a wide range of reverberation times. There are three key steps in designing a robust system. First, in contrast to sigmoid activation and min-max normalization in state-of-the-art algorithms, a linear activation function at the output layer and global mean-variance normalization of target features are adopted to learn the complicated nonlinear mapping function from reverberant to anechoic speech and to improve the restoration of the low-frequency and intermediate-frequency contents. Next, two key design parameters, namely, frame shift size in speech framing and acoustic context window size at the DNN input, are investigated to show that RT60-dependent parameters are needed in the DNN training stage in order to optimize the system performance in diverse reverberant environments. Finally, the reverberation time is estimated to select the proper frame shift and context window sizes for feature extraction before feeding the log-power spectrum features to the trained DNNs for speech dereverberation. Our experimental results indicate that the proposed framework outperforms the conventional DNNs without taking the reverberation time into account, while achieving a performance only slightly worse than the oracle cases with known reverberation times even for extremely weak and severe reverberant conditions. It also generalizes well to unseen room sizes, loudspeaker and microphone positions, and recorded room impulse responses. Bo Wu 0011, Kehuang Li, Minglei Yang 0001, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Exemplar-inspired strategies for low-resource spoken keyword search in SwahiliabstractWe present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples. Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 9 |
| 2016 | Improving non-native mispronunciation detection and enriching diagnostic feedback with DNN-based speech attribute modelingabstractWe propose the use of speech attributes, such as voicing and aspiration, to address two key research issues in computer assisted pronunciation training (CAPT) for L2 learners, namely detecting mispronunciation and providing diagnostic feedback. To improve the performance we focus on mispronunciations occurred at the segmental and sub-segmental levels. In this study, speech attributes scores are first used to measure the pronunciation quality at a sub-segmental level, such as manner and place of articulation. These speech attribute scores are integrated by neural network classifiers to generate segmental pronunciation scores. Compared with the conventional phone-based GOP (Goodness of Pronunciation) system we implement with our dataset, the proposed framework reduces the equal error rate by 8.78% relative. Moreover, it attains comparable results to phone-based classifier approach to mispronunciation detection while providing comprehensive feedback, including segmental and sub-segmental diagnostic information, to help L2 learners. Wei Li 0119, Sabato Marco Siniscalchi, Nancy F. Chen, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2016 | An experimental study on joint modeling of mixed-bandwidth data via deep neural networks for robust speech recognitionabstractWe propose joint modeling strategies leveraging upon large-scale mixed-band training speech for recognition of both narrowband and wideband data based on deep neural networks (DNNs). We utilize conventional down-sampling and up-sampling schemes to go between narrowband and wideband data. We also explore DNN-based speech bandwidth expansion (BWE) to map some acoustic features from narrowband to wideband speech. By arranging narrowband and wideband features at the input or the output level of BWE-DNN, and combining down-sampling and up-sampling data, different DNNs can be established. Our experiments on a Mandarin speech recognition task show that the hybrid DNNs for joint modeling of mixed-band speech yield significant performance gains over both the narrowband and wideband speech models, well-trained separately, with a relative character error rate reduction of 7.9% and 3.9% on narrowband and wideband data, respectively. Furthermore, the proposed strategies also consistently outperform other conventional DNN-based methods. Jianqing Gao, Jun Du 0002, Changqing Kong, Huaifang Lu, Enhong Chen, Chin-Hui Lee 0001 |
IJCNN | 6 |
| 2016 | SNR-Based Progressive Learning of Deep Neural Network for Speech Enhancement
Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2016 | Detecting Mispronunciations of L2 Learners and Providing Corrective Feedback Using Knowledge-Guided and Data-Driven Decision TreesabstractWe propose a novel decision tree based framework to detect phonetic mispronunciations produced by L2 learners caused by using inaccurate speech attributes, such as manner and place of articulation. Compared with conventional score-based CAPT (computer assisted pronunciation training) systems, our proposed framework has three advantages: (1) each mispronunciation in a tree can be interpreted and communicated to the L2 learners by traversing the corresponding path from a leaf node to the root node; (2) corrective feedback based on speech attribute features, which are directly used to describe how consonants and vowels are produced using related articulators, can be provided to the L2 learners; and (3) by building the phone-dependent decision tree, the relative importance of the speech attribute features of a target phone can be automatically learned and used to distinguish itself from other phones. This information can provide L2 learners speech attribute feedback that is ranked in order of importance. In addition to the abovementioned advantages, experimental results confirm that the proposed approach can detect most pronunciation errors and provide accurate diagnostic feedback Wei Li 0119, Kehuang Li, Sabato Marco Siniscalchi, Nancy F. Chen, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2016 | An Iterative Phase Recovery Framework with Phase Mask for Spectral Mapping with an Application to Speech Enhancement
Kehuang Li, Bo Wu 0011, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2016 | A unified approach to transfer learning of deep neural networks with applications to speaker adaptation in automatic speech recognition
Zhen Huang 0001, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
Neurocomputing | 3 |
| 2016 | i-Vector Modeling of Speech Attributes for Automatic Foreign Accent RecognitionabstractWe propose a unified approach to automatic foreign accent recognition. It takes advantage of recent technology advances in both linguistics and acoustics based modeling techniques in automatic speech recognition (ASR) while overcoming the issue of a lack of a large set of transcribed data often required in designing state-of-the-art ASR systems. The key idea lies in defining a common set of fundamental units “universally” across all spoken accents such that any given spoken utterance can be transcribed with this set of “accent-universal” units. In this study, we adopt a set of units describing manner and place of articulation as speech attributes. These units exist in most spoken languages and they can be reliably modeled and extracted to represent foreign accent cues. We also propose an i-vector representation strategy to model the feature streams formed by concatenating these units. Testing on both the Finnish national foreign language certificate (FSD) corpus and the English NIST 2008 SRE corpus, the experimental results with the proposed approach demonstrate a significant system performance improvement with p-value over those with the conventional spectrum-based techniques. We observed up to a 15% relative error reduction over the already very strong i-vector accented recognition system when only manner information is used. Additional improvement is obtained by adding place of articulation clues along with context information. Furthermore, diagnostic information provided by the proposed approach can be useful to the designers to further enhance the system performance. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Tomi Kinnunen, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | A Regression Approach to Single-Channel Speech Separation Via High-Resolution Deep Neural NetworksabstractWe propose a novel data-driven approach to single-channel speech separation based on deep neural networks (DNNs) to directly model the highly nonlinear relationship between speech features of a mixed signal containing a target speaker and other interfering speakers. We focus our discussion on a semisupervised mode to separate speech of the target speaker from an unknown interfering speaker, which is more flexible than the conventional supervised mode with known information of both the target and interfering speakers. Two key issues are investigated. First, we propose a DNN architecture with dual outputs of the features of both the target and interfering speakers, which is shown to achieve a better generalization capability than that with output features of only the target speaker. Second, we propose using a set of multiple DNNs, each intending to be signal-noise-dependent (SND), to cope with the difficulty that one single general DNN could not well accommodate all the speaker mixing variabilities at different signal-to-noise ratio (SNR) levels. Experimental results on the speech separation challenge (SSC) data demonstrate that our proposed framework achieves better separation results than other conventional approaches in a supervised or semisupervised mode. SND-DNNs could also yield significant performance improvements over a general DNN for speech separation in low SNR cases. Furthermore, for automatic speech recognition (ASR) following speech separation, this purely front-end processing with a single set of speaker-independent ASR acoustic models, achieves a relative word error rate (WER) reduction of 11.6% over a state-of-the-art separation and recognition system where a complicated joint back-end decoding framework with multiple sets of speaker-dependent ASR acoustic models needs to be implemented. When speaker-adaptive ASR acoustic models for the target speakers are adopted for the enhanced signals, another 12.1% WER reduction over our best speaker-independent ASR system is achieved. Jun Du 0002, Yanhui Tu, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | An information fusion approach to recognizing microphone array speech in the CHiME-3 challenge based on a deep learning frameworkabstractWe present an information fusion approach to robust recognition of microphone array speech for the recently launched 3rd CHiME Challenge. It is based on a deep learning framework with a large neural network consisting of subnets with different architectures. Multiple knowledge sources are integrated via an early fusion of normalized noisy features with different beamforming techniques, speech enhanced features, speaker related features, and other auxiliary features concatenated as the input to each subnet, and a late fusion by combining the outputs of all subnets to produce one single output set. Our experiments demonstrate that all information sources are complementary in our proposed framework. Our best system achieves an average word error rate reduction of 68% from the officially released baseline results on the test set of real data. Jun Du 0002, Qing Wang 0008, Yanhui Tu, Xiao Bao, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
ASRU | 6 |
| 2015 | Low-resource keyword search strategies for tamilabstractWe propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological nature of Tamil, we present highlights of our current KWS system, including: (1) Submodular optimization data selection to maximize acoustic diversity through Gaussian component indexed N-grams; (2) Keywordaware language modeling; (3) Subword modeling of morphemes and homophones. Nancy F. Chen, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Van Tung Pham, Haihua Xu 0001, Tze Siong Lau, Su Jun Leow, Boon Pang Lim, Cheung-Chi Leung, Lei Wang 0020, Chin-Hui Lee 0001, Alvina Goh, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 13 |
| 2015 | A keyword-aware grammar framework for LVCSR-based spoken keyword searchabstractIn this paper, we proposed a method to realize the recently developed keyword-aware grammar for LVCSR-based keyword search using weight finite-state automata (WFSA). The approach creates a compact and deterministic grammar WFSA by inserting keyword paths to an existing n-gram WFSA. Tested on the evalpart1 data of the IARPA Babel OpenKWS13 Vietnamese and OpenKWS14 Tamil limited language pack tasks, the experimental results indicate the proposed keyword-aware framework achieves significant improvement, with about 50% relative actual term weighted value (ATWV) enhancement for both languages. Comparisons between the keyword-aware grammar and our previously proposed n-gram LM based approximation approach for the grammar also show that the KWS performances of these two realizations are complementary. I-Fan Chen, Chongjia Ni, Boon Pang Lim, Nancy F. Chen, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2015 | Joint training of front-end and back-end deep neural networks for robust speech recognitionabstractBased on the recently proposed speech pre-processing front-end with deep neural networks (DNNs), we first investigate different feature mapping directly from noisy speech via DNN for robust speech recognition. Next, we propose to jointly train a single DNN for both feature mapping and acoustic modeling. In the end, we show that the word error rate (WER) of the jointly trained system could be significantly reduced by the fusion of multiple DNN pre-processing systems which implies that features obtained from different domains of the DNN-enhanced speech signals are strongly complementary. Testing on the Aurora4 noisy speech recognition task our best system with multi-condition training can achieves an average WER of 10.3%, yielding a relative reduction of 16.3% over our previous DNN pre-processing only system with a WER of 12.3%. To the best of our knowledge, this represents the best published result on the Aurora4 task without using any adaptation techniques. Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2015 | Language-resource independent speech segmentation using cues from a spectrogram imageabstractIn this paper, we use image processing techniques on the speech spectrogram to perform speech phoneme segmentation. The proposed method relies solely on visual cues on the spectrogram, without the need for language-specific training data. The results are evaluated on the TIMIT corpus, and compared to other unsupervised speech segmentation techniques, with comparable results obtained. We also fuse the results with those obtained by hidden Markov models (HMM) and HMM-based forced alignment to investigate if image features can provide an additional feature representation for speech processing tasks. With the fusion, up to 10% absolute improvement in segmentation accuracy over the HMM baselines can be obtained. Results are promising and suggests a strong potential for image-based features applying to speech processing. Su Jun Leow, Chng Eng Siong, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2015 | Speech Separation based on signal-noise-dependent deep neural networks for robust speech recognitionabstractIn this paper, we propose a new signal-noise-dependent (SND) deep neural network (DNN) framework to further improve the separation and recognition performance of the recently developed technique for general DNN-based speech separation. We adopt a divide and conquer strategy to design the proposed SND-DNNs with higher resolutions that a single general DNN could not well accommodate for all the speaker mixing variabilities at different levels of signal-to-noise ratios (SNRs). In this study two kinds of SNR-dependent DNNs, namely positive and negative DNNs, are trained to cover the mixed speech signals with positive and negative SNR levels, respectively. At the separation stage, a first-pass separation using a general DNN can give an accurate SNR estimation for a model selection. Experimental results on the Speech Separation Challenge (SSC) task show that SND-DNNs could yield significant performance improvements for both speech separation and recognition over a general DNN. Furthermore, this purely front-end processing method achieves a relative word error rate reduction of 11.6% over a state-of-the-art recognition system where a complicated joint decoding framework needs to be implemented in the back-end. Yanhui Tu, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2015 | Rapid adaptation for deep neural networks through multi-task learningabstractWe propose a novel approach to addressing the adaptation effectiveness issue in parameter adaptation for deep neural network (DNN) based acoustic models for automatic speech recognition by adding one or more small auxiliary output layers modeling broad acoustic units, such as mono-phones or tied-state (often called senone) clusters. In scenarios with a limited amount of available adaptation data, most senones are usually rarely seen or not observed, and consequently the ability to model them in a new condition is often not fully exploited. With the original senone classification task as the primary task, and adding auxiliary mono-phone/senone-cluster classification as the secondary tasks, multi-task learning (MTL) is employed to adapt the DNN parameters. With the proposed MTL adaptation framework, we improve the learning ability of the original DNN structure, then enlarge the coverage of the acoustic space to deal with the unseen senone problem, and thus enhance the discrimination power of the adapted DNN models. Experimental results on the 20,000-word open vocabulary WSJ task demonstrate that the proposed framework consistently outperforms the conventional linear hidden layer adaptation schemes without MTL by providing 5.4% relative reduction in word error rate (WERR) with only 1 single adaptation utterance, and 10.7% WERR with 40 adaptation utterances against the un-adapted DNN models. Zhen Huang 0001, Jinyu Li 0001, Sabato Marco Siniscalchi, I-Fan Chen, Ji Wu 0002, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2015 | Maximum a posteriori adaptation of network parameters in deep modelsabstractWe present a Bayesian approach to adapting parameters of a well-trained context-dependent, deep-neural-network, hidden Markov model (CD-DNN-HMM) to improve automatic speech recognition performance. Given an abundance of DNN parameters but with only a limited amount of data, the effectiveness of the adapted DNN model can often be compromised. We formulate maximum a posteriori (MAP) adaptation of parameters of a specially designed CD-DNN-HMM with an augmented linear hidden networks connected to the output tied states, or senones, and compare it to feature space MAP linear regression previously proposed. Experimental evidences on the 20,000-word open vocabulary Wall Street Journal task demonstrate the feasibility of the proposed framework. In supervised adaptation, the proposed MAP adaptation approach provides more than 10% relative error reduction and consistently outperforms the conventional transformation based methods. Furthermore, we present an initial attempt to generate hierarchical priors to improve adaptation efficiency and effectiveness with limited adaptation data by exploiting similarities among senones. Zhen Huang 0001, Sabato Marco Siniscalchi, I-Fan Chen, Jinyu Li 0001, Jiadong Wu, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2015 | DNN-based speech bandwidth expansion and its application to adding high-frequency missing features for automatic speech recognition of narrowband speech
Kehuang Li, Zhen Huang 0001, Yong Xu 0004, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2015 | A universal VAD based on jointly trained deep neural networks
Qing Wang 0008, Jun Du 0002, Xiao Bao, Zi-Rui Wang, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2015 | High-resolution acoustic modeling and compact language modeling of language-universal speech attributes for spoken language identification
Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2015 | An entropy minimization framework for goal-driven dialogue management
Ji Wu 0002, Miao Li 0003, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2015 | Multi-objective learning and mask-based post-processing for deep neural network based speech enhancementabstractWe propose a multi-objective framework to learn both secondary targets not directly related to the intended task of speech enhancement (SE) and the primary target of the clean log-power spectra (LPS) features to be used directly for constructing the enhanced speech signals.In deep neural network (DNN) based SE we introduce an auxiliary structure to learn secondary continuous features, such as mel-frequency cepstral coefficients (MFCCs), and categorical information, such as the ideal binary mask (IBM), and integrate it into the original DNN architecture for joint optimization of all the parameters.This joint estimation scheme imposes additional constraints not available in the direct prediction of LPS, and potentially improves the learning of the primary target.Furthermore, the learned secondary information as a byproduct can be used for other purposes, e.g., the IBM-based post-processing in this work.A series of experiments show that joint LPS and MFCC learning improves the SE performance, and IBM-based post-processing further enhances listening quality of the reconstructed speech. Yong Xu 0004, Jun Du 0002, Zhen Huang 0001, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2015 | A Probabilistic Framework for Representing Dialog Systems and Entropy-Based Dialog Management Through Dynamic Stochastic State EvolutionabstractIn this paper, we present a probabilistic framework for goal-driven spoken dialog systems. A new dynamic stochastic state (DS-state) is then defined to characterize the goal set of a dialog state at different stages of the dialog process. Furthermore, an entropy minimization dialog management (EMDM) strategy is also proposed to combine with the DS-states to facilitate a robust and efficient solution in reaching a user's goals. A song-on-demand task, with a total of 38 117 songs and 12 attributes corresponding to each song, is used to test the performance of the proposed approach. In an ideal simulation, assuming no errors, the EMDM strategy is the most efficient goal-seeking method among all tested approaches, returning the correct song within 3.3 dialog turns on average. Furthermore, in a practical scenario, with top five candidates to handle the unavoidable automatic speech recognition (ASR) and natural language understanding (NLU) errors, the results show that only 61.7% of the dialog goals can be successfully obtained in 6.23 dialog turns on average when random questions are asked by the system, whereas if the proposed DS-states are updated with the top five candidates from the SLU output using the proposed EMDM strategy executed at every DS-state, then a 86.7% dialog success rate can be accomplished effectively within 5.17 dialog turns on average. We also demonstrate that entropy-based DM strategies are more efficient than non-entropy based DM. Moreover, using the goal set distributions in EMDM, the results are better than those without them, such as in sate-of-the-art database summary DM. Ji Wu 0002, Miao Li 0003, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | A Regression Approach to Speech Enhancement Based on Deep Neural NetworksabstractIn contrast to the conventional minimum mean square error (MMSE)-based noise reduction techniques, we propose a supervised method to enhance speech by means of finding a mapping function between noisy and clean speech signals based on deep neural networks (DNNs). In order to be able to handle a wide range of additive noises in real-world situations, a large training set that encompasses many possible combinations of speech and noise types, is first designed. A DNN architecture is then employed as a nonlinear regression function to ensure a powerful modeling capability. Several techniques have also been proposed to improve the DNN-based speech enhancement system, including global variance equalization to alleviate the over-smoothing problem of the regression model, and the dropout and noise-aware training strategies to further improve the generalization capability of DNNs to unseen noise conditions. Experimental results demonstrate that the proposed framework can achieve significant improvements in both objective and subjective measures over the conventional MMSE based technique. It is also interesting to observe that the proposed DNN approach can well suppress highly nonstationary noise, which is tough to handle in general. Furthermore, the resulting DNN model, trained with artificial synthesized data, is also effective in dealing with noisy speech data recorded in real-world scenarios without the generation of the annoying musical artifact commonly observed in conventional enhancement methods. Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Introducing attribute features to foreign accent recognitionabstractWe propose a hybrid approach to foreign accent recognition combining both phonotactic and spectral based systems by treating the problem as a spoken language recognition task. We extract speech attribute features that represent speech and acoustic cues reflecting foreign accents of a speaker to obtain feature streams that are modeled with the i-vector methodology. Testing on the Finnish Language Proficiency exam corpus, we find our proposed technique to achieve a significant performance improvement over the state-of-the-art systems using only spectral based features. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Tomi Kinnunen, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2014 | Attribute based lattice rescoring in spontaneous speech recognitionabstractIn this paper we extend attribute-based lattice rescoring to spontaneous speech recognition. This technique is based on two key features: (i) an attribute-based frontend, which consists of a bank of speech attribute detectors followed up by an evidence merger that generates confidence scores (e.g., sub-word posterior probabilities), and (ii) a rescoring module that integrates information generated by the frontend into an existing ASR engine through lattice rescoring. The speech attributes used in this work are phonetic features, such as frication and palatalization. Experimental results on the Switchboard part of the NIST 2000 Hub5 data set demonstrate that the proposed approach outperforms LVCSR systems based on Gaussian mixture model/ hidden Markov model (GMM/HMM) that does not use attribute related information. Furthermore, a small yet promising improvement is also observed when rescoring word-lattices generated by a state-of-the-art ASR system using deep neural networks. Different frontend configuration are investigated and tested. I-Fan Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2014 | An i-vector based descriptor for alphabetical gesture recognitionabstractAn i-vector approach to extracting features for video camera based gesture recognition is proposed. Conventional low-level raw features, such as position, speed, and acceleration, are low-dimensional feature representations which often suffer from measurement noise and thus are not highly discriminative. High-level features, such as Fourier descriptor, usually take a global transformation on the whole raw features of a gesture, but local statistical information is seldom considered. Moreover, compared with speech recordings, video cameras used to capture data are often at a low frame rate such that it is challenging for proper modeling and recognition. In this paper, we show that the proposed i-vector framework can handle both local statistical information and sparse trajectory representations more efficiently under the sparse data scenarios for an in-car hand-gesturing English letter recognition system. Experimental results confirm the effectiveness of the proposed i-vector features, which can reduce the letter error rate by as much as 36-44% relatively from the results obtained with the conventional location based raw features. You-Chi Cheng, Ville Hautamäki, Zhen Huang 0001, Kehuang Li, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2014 | Deep learning vector quantization for acoustic information retrievalabstractWe propose a novel deep learning vector quantization (DLVQ) algorithm based on deep neural networks (DNNs). Utilizing a strong representation power of this deep learning framework, with any vector quantization (VQ) method as an initializer, the proposed DLVQ technique is capable of learning a code-constrained codebook and thus improves over conventional VQ to be used in classification problems. Tested on an audio information retrieval task, the proposed DLVQ achieves a quite promising performance when it is initialized by the k-means VQ technique. A 10.5% relative gain in mean average precision (MAP) is obtained after fusing the k-means and DLVQ results together. Zhen Huang 0001, Chao Weng, Kehuang Li, You-Chi Cheng, Chin-Hui Lee 0001 |
ICASSP | 5 |
| 2014 | A maximal figure-of-merit learning approach to maximizing mean average precision with deep neural network based classifiersabstractWe propose a maximal figure-of-merit (MFoM) learning framework to directly maximize mean average precision (MAP) which is a key performance metric in many multi-class classification tasks. Conventional classifiers based on support vector machines cannot be easily adopted to optimize the MAP metric. On the other hand, classifiers based on deep neural networks (DNNs) have recently been shown to deliver a great discrimination capability in automatic speech recognition and image classification as well. However, DNNs are usually optimized with the minimum cross entropy criterion. In contrast to most conventional classification methods, our proposed approach can be formulated to embed DNNs and MAP into the objective function to be optimized during training. The combination of the proposed maximum MAP (MMAP) technique and DNNs introduces nonlinearity to the linear discriminant function (LDF) in order to increase the flexibility and discriminant power of the original MFoM-trained LDF based classifiers. Tested on both automatic image annotation and audio event classification, the experimental results show consistent improvements of MAP on both datasets when compared with other state-of-the-art classifiers without using MMAP. Kehuang Li, Zhen Huang 0001, You-Chi Cheng, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2014 | Dialect levelling in Finnish: a universal speech attribute approachabstractWe adopt automatic language recognition methods to study dialect levelling - a phenomenon that leads to reduced structural differences among dialects in a given spoken language. In terms of dialect characterisation, levelling is a nuisance variable that adversely affects recognition accuracy: The more similar two dialects are, the harder it is to set them apart. We address levelling in Finnish regional dialects using a new SAPU (Satakunta in Speech) corpus containing material from Satakunta (South-Western Finland) between 2007 and 2013. To define a compact and universal set of sound units to characterize dialects, we adopt speech attributes features, namely manner and place of articulation. It will be shown that speech attribute distributions can indeed characterise differences among dialects. Experiments with an i-vector system suggest that (1) the attribute features achieve higher dialect recognition accuracy and (2) they are less sensitive against age-related levelling in comparison to traditional spectral approach. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Elie Khoury 0001, Tommi Kurki, Tomi Kinnunen, Chin-Hui Lee 0001 |
INTERSPEECH | 7 |
| 2014 | A keyword-boosted sMBR criterion to enhance keyword search performance in deep neural network based acoustic modeling
I-Fan Chen, Nancy F. Chen, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2014 | Robust speech recognition with speech enhanced deep neural networks
Jun Du 0002, Qing Wang 0008, Tian Gao 0005, Yong Xu 0004, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2014 | Feature space maximum a posteriori linear regression for adaptation of deep neural networksabstractWe propose a feature space maximum a posteriori (MAP) linear regression framework to adapt parameters for context dependent deep neural network hidden Markov models (CD-DNN-HMMs). Due to the huge amount of parameters used in DNN acoustic models in large vocabulary continuous speech recognition, the problem of over-fitting can be severe in DNN adaptation, thus often impair the robustness of the adapted DNN model. Linear input network (LIN) as a straight-forward feature space adaptation method for DNN, similar to feature space maximum likelihood linear regression (fMLLR), can potentially suffer from the same robustness situation. The proposed adaptation framework is built based on MAP estimation of the LIN parameters by incorporating prior knowledge into the adaptation process. Experimental results on the Switchboard task show that against the speaker independent CD-DNN-HMM systems, LIN provides 4.28% relative word error rate reduction (WERR) and the proposed fMAPLIN method is able to provide further 1.15% (totally 5.43%) WERR on top of LIN. Zhen Huang 0001, Jinyu Li 0001, Sabato Marco Siniscalchi, I-Fan Chen, Chao Weng, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2014 | Beyond cross-entropy: towards better frame-level objective functions for deep neural network training in automatic speech recognitionabstractWe propose two approaches for improving the objective function for the deep neural network (DNN) frame-level training in large vocabulary continuous speech recognition (LVCSR). The DNNs used in LVCSR are often constructed with an output layer with softmax activation and the cross-entropy objective function is always employed in the frame-leveling training of DNNs. The pairing of softmax activation and crossentropy objective function contributes much in the success of DNN. The first approach developed in this paper improves the cross-entropy objective function by boosting the importance of the frames for which the DNN model has low target predictions (low target posterior probabilities) and the second one considers jointly minimizing the cross-entropy and maximizing the log posterior ratio between the target senone (tied-triphone states) and the most competing one. Experiments on Switchboard task demonstrate that the two proposed methods can provide 3.1% and 1.5% relative word error rate (WER) reduction , respectively, against the already very strong conventional crossentropy trained DNN system. Zhen Huang 0001, Jinyu Li 0001, Chao Weng, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2014 | Dynamic noise aware training for speech enhancement based on deep neural networksabstractWe propose three algorithms to address the mismatch problem in deep neural network (DNN) based speech enhancement. First, we investigate noise aware training by incorporating noise informationin the testutterance with anideal binary maskbased dynamic noise estimation approach to improve DNN’s speech separation ability from the noisy signal. Next, a set of more than 100 noise types is adopted to enrich the generalization capabilities of the DNN to unseen and non-stationary noise conditions. Finally, the quality of the enhanced speech can further be improved by global variance equalization. Empirical results show that each of the three proposed techniques contributes to the performance improvement. Compared to the conventional logarithmic minimum mean squared error speech enhancement method, our DNN system achieves 0.32 PESQ (perceptual evaluation of speech quality) improvement across six signal-tonoise ratio levels ranging from -5dB to 20dB on a test set with unknown noise types. We also observe that the combined strategies can well suppress highly non-stationary noise better than all the competing state-of-the-art techniques we have evaluated. Index Terms: Speech enhancement, deep neural networks, noise aware training, ideal binary mask, non-stationary noise Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2014 | An artificial neural network approach to automatic speech processing
Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
Neurocomputing | 3 |
| 2014 | An Experimental Study on Speech Enhancement Based on Deep Neural NetworksabstractThis letter presents a regression-based speech enhancement framework using deep neural networks (DNNs) with a multiple-layer deep architecture. In the DNN learning process, a large training set ensures a powerful modeling capability to estimate the complicated nonlinear mapping from observed noisy speech to desired clean signals. Acoustic context was found to improve the continuity of speech to be separated from the background noises successfully without the annoying musical artifact commonly observed in conventional speech enhancement algorithms. A series of pilot experiments were conducted under multi-condition training with more than 100 hours of simulated speech data, resulting in a good generalization capability even in mismatched testing conditions. When compared with the logarithmic minimum mean square error approach, the proposed DNN-based algorithm tends to achieve significant improvements in terms of various objective quality measures. Furthermore, in a subjective preference evaluation with 10 listeners, 76.35% of the subjects were found to prefer DNN-based enhanced speech to that obtained with other conventional technique. Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE Signal Process. Lett. | 4 |
| 2014 | A MAP-based Online Estimation Approach to Ensemble Speaker and Speaking Environment ModelingabstractAn ensemble speaker and speaking environment modeling (ESSEM) approach was recently developed. This ESSEM process consists of offline and online phases. The offline phase establishes an environment structure using speech data collected under a wide range of acoustic conditions, whereas the online phase estimates a set of acoustic models that matches the testing environment based on the established environment structure. Since the estimated acoustic models accurately characterize particular testing conditions, ESSEM can improve the speech recognition performance under adverse conditions. In this work, we propose two maximum a posteriori (MAP) based algorithms to improve the online estimation part of the original ESSEM framework. We first develop MAP-based environment structure adaptation to refine the original environment structure. Next, we propose to utilize the MAP criterion to estimate the mapping function of ESSEM and enhance the environment modeling capability. For the MAP estimation, three types of priors are derived; they are the clustered prior (CP), the sequential prior (SP), and the hierarchical prior (HP) densities. Since each prior density is able to characterize specific acoustic knowledge, we further derive a combination mechanism to integrate the three priors. Based on the experimental results on the Aurora-2 task, we verify that using the MAP-based online mapping function estimation can enable ESSEM to achieve better performance than using the maximum-likelihood (ML) based counterpart. Moreover, by using an integration of the online environment structuring adaptation and mapping function estimation, the proposed MAP-based ESSEM framework is found to provide the best performance. Compared with our baseline results, MAP-based ESSEM achieves an average word error rate reduction of 15.53% (5.41 to 4.57%) under 50 testing conditions at a signal-to-noise ratio (SNR) of 0 to 20 dB over the three standardized testing sets. Yu Tsao 0001, Shigeki Matsuda, Chiori Hori, Hideki Kashioka, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2013 | Knowledge integration for improving performance in LVCSRabstractThis paper presents a knowledge integration framework to improve performance in large vocabulary continuous speech recognition. Two types of knowledge sources, manner attribute and prosodic structure, are incorporated. For manner of articulation, six attribute detectors trained with an American English corpus (WSJ0) are utilized to rescore hypothesized phones in word lattices obtained by a baseline ASR system. For the prosodic structure, models trained with an unsupervised joint prosody labeling and modeling (PLM) technique using WSJ0 are used in lattice rescoring. Experimental results on the American English WSJ word recognition task of the Nov92 test set show that the proposed approach significantly outperforms the baseline system that does not use articulatory and prosodic information. The results also demonstrate the effectiveness and usefulness of the PLM technique in constructing prosodic models for American English ASR. Chen-Yu Chiang, Sabato Marco Siniscalchi, Sin-Horng Chen, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2013 | Minimax i-vector extractor for short duration speaker verificationabstractTotal variability modeling, based on i-vector extraction of converting a variable-length sequence of feature vectors into a fixed-length i-vector, is currently an adopted parametrization technique for state of-the-art speaker verification systems. However, when the number of the feature vectors is low, uncertainty in the i-vector representation as a point estimate of the linear-Gaussian model is understandably problematic. It is known that the zeroth and first order sufficient statistics, given the hyperparameters, completely characterize the extracted i-vectors. In this study we propose to use a minimax strategy to estimate the sufficient statistics in order to increase the robustness of the extracted i-vectors. We show by experiments that the proposed minimax technique can improve over the baseline system from 9.89 % to 7.99 % on the NIST SRE 2010 8conv-10sec task. Ville Hautamäki, You-Chi Cheng, Padmanabhan Rajan, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2013 | A blind segmentation approach to acoustic event detection based on i-vectorabstractWe propose a new blind segmentation approach to acous-tic event detection (AED) based on i-vectors. Conventional approaches to AED often required well-segmented data with non-overlapping boundaries for competing events. Inspired by block-based automatic image annotation in image retrieval tasks, we blindly segment audio streams into equal-length pieces, label the underlying observed acoustic events with mul-tiple categories and with no event boundary information, extract i-vector for them, and perform classification using support vec-tor machine and maximal figure-of-merit based classifiers. Ex-periments on various sets of audio data show promising results with an average of 8 % absolute gain in F1 over the conventional hidden Markov model based approach. An enhanced robustness at different noise levels is also observed. The key to the suc-cess lies in the enhanced discrimination power offered by the i-vector representation of the acoustic data. Index Terms: acoustic event detection, i-vector, blind segmen-tation, support vector machine, maximal figure-of-merit Zhen Huang 0001, You-Chi Cheng, Kehuang Li, Ville Hautamäki, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2013 | Universal attribute characterization of spoken languages for automatic spoken language recognition
Sabato Marco Siniscalchi, Jeremy Reed, Torbjørn Svendsen, Chin-Hui Lee 0001 |
Comput. Speech Lang. | 4 |
| 2013 | Model-based margin estimation for hidden Markov model learning and generalisationabstractRecently, speech scientists have been motivated by the great, success of building margin‐based classifiers, and have thus proposed novel methods to estimate continuous‐density hidden Markov model (HMM) for automatic speech recognition (ASR) according to the notion that the decision boundaries determined by the estimated HMMs attain the maximum classification margin as in learning support vector machines. Although a good performance has been observed, the margin used in the ASR community is often specified as a parameter that has no explicit relationship with the HMM parameters. The issues of how the margin is related to the HMM parameters and how it directly characterises the generalisation capability of HMM‐based classifiers have not been addressed so far in the community. In this study, the authors attempt to formulate the margin used in the soft margin estimation framework as a function of the HMM parameters. The key idea is to relate the standard distance‐based margin with the concept of divergence among competing HMM state Gaussian mixture model densities. Experimental results show that the proposed model‐based margin function is a good indication about the quality of HMMs on a given ASR task without the conventional needs of running experiments extensively using a separate set of test samples. Sabato Marco Siniscalchi, Jinyu Li 0001, Chin-Hui Lee 0001 |
IET Signal Process. | 3 |
| 2013 | Exploiting deep neural networks for detection-based speech recognition
Sabato Marco Siniscalchi, Dong Yu 0001, Li Deng 0001, Chin-Hui Lee 0001 |
Neurocomputing | 4 |
| 2013 | An Information-Extraction Approach to Speech Processing: Analysis, Detection, Verification, and RecognitionabstractThe field of automatic speech recognition (ASR) has enjoyed more than 30 years of technology advances due to the extensive utilization of the hidden Markov model (HMM) framework and a concentrated effort by the speech community to make available a vast amount of speech and language resources, known today as the Big Data Paradigm. State-of-the-art ASR systems achieve a high recognition accuracy for well-formed utterances of a variety of languages by decoding speech into the most likely sequence of words among all possible sentences represented by a finite-state network (FSN) approximation of all the knowledge sources required by the ASR task. However, the ASR problem is still far from being solved because not all information available in the speech knowledge hierarchy can be directly integrated into the FSN to improve the ASR performance and enhance system robustness. It is believed that some of the current issues of integrating various knowledge sources in top-down integrated search can be partially addressed by processing techniques that take advantage of the full set of acoustic and language information in speech. It has long been postulated that human speech recognition (HSR) determines the linguistic identity of a sound based on detected evidence that exists at various levels of the speech knowledge hierarchy, ranging from acoustic phonetics to syntax and semantics. This calls for a bottom-up attribute detection and knowledge integration framework that links speech processing with information extraction, by spotting speech cues with a bank of attribute detectors, weighting and combining acoustic evidence to form cognitive hypotheses, and verifying these theories until a consistent recognition decision can be reached. The recently proposed automatic speech attribute transcription (ASAT) framework is an attempt to mimic some HSR capabilities with asynchronous speech event detection followed by bottom-up knowledge integration and verification. In the last few years, ASAT has demonstrated good potential and has been applied to a variety of existing applications in speech processing and information extraction. Chin-Hui Lee 0001, Sabato Marco Siniscalchi |
Proc. IEEE | 1 |
| 2013 | Speech Recognition Using Long-Span Temporal Patterns in a Deep Network ModelabstractIn recent years, there has been a renewed interest in the use of artificial neural networks (ANNs) for speech applications, and it seems that a new trend to move the speech technology forward has begun. Two main contributions have triggered such a new trend: 1) a major advance has been made in training the weights in deep neural networks (DNNs), and a pre-trained deep neural network hidden Markov model (DNN-HMM) hybrid architecture has outperformed a conventional Gaussian mixture model hidden Markov model (GMM-HMM) automatic speech recognition (ASR) system on a challenging business search dataset, and 2) it has been shown that phoneme classification can be boosted by using a hierarchical structure of multi-layer perceptrons (MLPs) trained to model long-span temporal patterns with beneficial effects on language recognition tasks. In this work, we combine these two lines of research and demonstrate that word recognition accuracy can be significantly enhanced by arranging DNNs in a hierarchical structure to model long-term energy trajectories. The proposed solution has been evaluated on the 5000-word Wall Street Journal task, resulting in consistent and significant improvements in both phone and word recognition accuracy rates. We have also analyzed the effects of various modeling choices on the system performance, and several architectural solutions have been compared. Sabato Marco Siniscalchi, Dong Yu 0001, Li Deng 0001, Chin-Hui Lee 0001 |
IEEE Signal Process. Lett. | 4 |
| 2013 | Hermitian Polynomial for Speaker Adaptation of Connectionist Speech Recognition SystemsabstractModel adaptation techniques are an efficient way to reduce the mismatch that typically occurs between the training and test condition of any automatic speech recognition (ASR) system. This work addresses the problem of increased degradation in performance when moving from speaker-dependent (SD) to speaker-independent (SI) conditions for connectionist (or hybrid) hidden Markov model/artificial neural network (HMM/ANN) systems in the context of large vocabulary continuous speech recognition (LVCSR). Adapting hybrid HMM/ANN systems on a small amount of adaptation data has been proven to be a difficult task, and has been a limiting factor in the widespread deployment of hybrid techniques in operational ASR systems. Addressing the crucial issue of speaker adaptation (SA) for hybrid HMM/ANN system can thereby have a great impact on the connectionist paradigm, which will play a major role in the design of next-generation LVCSR considering the great success reported by deep neural networks - ANNs with many hidden layers that adopts the pre-training technique - on many speech tasks. Current adaptation techniques for ANNs based on injecting an adaptable linear transformation network connected to either the input, or the output layer are not effective especially with a small amount of adaptation data, e.g., a single adaptation utterance. In this paper, a novel solution is proposed to overcome those limits and make it robust to scarce adaptation resources. The key idea is to adapt the hidden activation functions rather than the network weights. The adoption of Hermitian activation functions makes this possible. Experimental results on an LVCSR task demonstrate the effectiveness of the proposed approach. Sabato Marco Siniscalchi, Jinyu Li 0001, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | A Bottom-Up Modular Search Approach to Large Vocabulary Continuous Speech RecognitionabstractA novel bottom-up decoding framework for large vocabulary continuous speech recognition (LVCSR) with a modular search strategy is presented. Weighted finite state machines (WFSMs) are utilized to accomplish stage-by-stage acoustic-to-linguistic mappings from low-level speech attributes to high-level linguistic units in a bottom-up manner. Probabilistic attribute and phone lattices are used as intermediate vehicles to facilitate knowledge integration at different levels of the speech knowledge hierarchy. The final decoded sentence is obtained by performing lexical access and applying syntactical constraints. Two key factors are critical to warrant a high recognition accuracy, namely: (i) generation of high-precision sets of competing hypotheses at every intermediate stage; and (ii) low-error pruning of unlikely theories to reduce input lattice sizes while maintaining high-quality hypotheses for the next layers of knowledge integration. The decoupled nature of the proposed techniques allows us to obtain recognition results at all stages, including attribute, phone and word levels, and enables an integration of various knowledge sources not easily done in the state-of-the-art hidden Markov model (HMM) systems based on top-down knowledge integration. Evaluation on the Nov92 test set of the 5000-word, Wall Street Journal task demonstrates that high-accuracy attribute and phone classification can be attained. As for word recognition, the proposed WFSM-based framework achieves encouraging word error rates. Finally, by combining attribute scores with the conventional HMM likelihood scores and re-ordering the N-best lists obtained from the word lattices generated with the proposed WFSM system, the word error rate (WER) can be further reduced. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Boosting attribute and phone estimation accuracies with deep neural networks for detection-based speech recognitionabstractGeneration of high-precision sub-phonetic attribute (also known as phonological features) and phone lattices is a key frontend component for detection-based bottom-up speech recognition. In this paper we employ deep neural networks (DNNs) to improve detection accuracy over conventional shallow MLPs (multi-layer perceptrons) with one hidden layer. A range of DNN architectures with five to seven hidden layers and up to 2048 hidden units per layer have been explored. Training on the SI84 and testing on the Nov92 WSJ data, the proposed DNNs achieve significant improvements over the shallow MLPs, producing greater than 90% frame-level attribute estimation accuracies for all 21 attributes tested for the full system. On the phone detection task, we also obtain excellent frame-level accuracy of 86.6%. With this level of high-precision detection of basic speech units we have opened the door to a new family of flexible speech recognition system design for both top-down and bottom-up, lattice-based search strategies and knowledge integration. Dong Yu 0001, Sabato Marco Siniscalchi, Li Deng 0001, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2012 | Per-Exemplar Fusion Learning for Video Retrieval and RecountingabstractWe propose a novel video retrieval framework based on an extension of per-exemplar learning [7]. Each training sample with multiple types of features (e.g., audio and visual) is regarded as an exemplar. For each exemplar, a localized per-exemplar distance function is learned and used to measure the similarity between itself and new test samples. Exemplars associate only with sufficiently similar test data, which accumulate to identify the data to be retrieved. In particular, for every exemplar, relevance of each feature type is discriminatively analyzed and the effect of less informative features is minimized during the fusion-based associations. In addition, we show that our framework can enable a rich set of recounting capabilities where the rationale for each retrieval result can be automatically described to users to aid their interaction with the system. We show that our system provides competitive retrieval accuracy against strong baseline methods, while adding the benefits of recounting. Ilseo Kim, Sangmin Oh, A. G. Amitha Perera, Chin-Hui Lee 0001 |
ICME | 4 |
| 2012 | Consumer-level multimedia event detection through unsupervised audio signal modelingabstractIn this work, a novel acoustic characterization approach to multimedia event detection (MED) task for unconstrained and unstructured consumer-level videos through audio signal modeling is proposed. The key idea is to characterize the acoustic space of interest with a set of fundamental acoustic units around which a set of acoustic segment models (ASMs) is built. A vector space modeling technique to address MED is here adopted, where an incoming audio signal is first decoded into a sequence of acoustic segments. Then, a feature vector is generated by using co-occurrence statistics of acoustic units, and the MED final decision is implemented with a vector space language classifier. Experimental evidence on the TRECVID2011 MED demonstrates the viability of the proposed approach. Furthermore, it better accounts for temporal dependencies than previously proposed MFCC bag-of-word approaches. Byungki Byun, Ilseo Kim, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2012 | Hermitian based Hidden Activation Functions for Adaptation of Hybrid HMM/ANN ModelsabstractThis work is concerned with speaker adaptation techniques for artificial neural network (ANN) implemented as feed-forward multi-layer perceptrons (MLPs) in the context of large vocabulary continuous speech recognition (LVCSR). Most successful speaker adaptation techniques for MLPs consist of augmenting the neural architecture with a linear transformation network connected to either the input or the output layer. The weights of this additional linear layer are learned during the adaptation phase while all of the other weights are kept frozen in order to avoid over-fitting. In doing so, the structure of the speaker-dependent (SD) and speaker-independent (Si) architecture differs and the number of adaptation parameters depends upon the dimension of either the input or output layers. We propose an alternative neural architecture for speaker-adaptation to overcome the limits of current approaches. This neural architecture adopts hidden activation functions that can be learned directly from the adaptation data. This adaptive capability of the hidden activation function is achieved through the use of orthonormal Hermite polynomials. Experimental evidence gathered on the Wall Street Journal Nov92 task demonstrates the viability of the proposed technique. Sabato Marco Siniscalchi, Jinyu Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2012 | Experiments on Cross-Language Attribute Detection and Phone Recognition With Minimal Target-Specific Training DataabstractA state-of-the-art automatic speech recognition (ASR) system can often achieve high accuracy for most spoken languages of interest if a large amount of speech material can be collected and used to train a set of language-specific acoustic phone models. However, designing good ASR systems with little or no language-specific speech data for resource-limited languages is still a challenging research topic. As a consequence, there has been an increasing interest in exploring knowledge sharing among a large number of languages so that a universal set of acoustic phone units can be defined to work for multiple or even for all languages. This work aims at demonstrating that a recently proposed automatic speech attribute transcription framework can play a key role in designing language-universal acoustic models by sharing speech units among all target languages at the acoustic phonetic attribute level. The language-universal acoustic models are evaluated through phone recognition. It will be shown that good cross-language attribute detection and continuous phone recognition performance can be accomplished for “unseen” languages using minimal training data from the target languages to be recognized. Furthermore, a phone-based background model (PBM) approach will be presented to improve attribute detection accuracies. Sabato Marco Siniscalchi, Dau-Cheng Lyu, Torbjørn Svendsen, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | A Bottom-Up Stepwise Knowledge-Integration Approach to Large Vocabulary Continuous Speech Recognition Using Weighted Finite State MachinesabstractA bottom-up, stepwise, knowledge integration framework is proposed to realize detection-based, large vocabulary continuous speech recognition (LVCSR) with a weighted finite state machine (WFSM). The WFSM framework offers a flexible architecture for different types of knowledge network compositions, each of them can be built and optimized independently. Speech attribute detectors are used as an intermediate block to obtain phoneme posterior probabilities over which a phoneme recognition network is designed. Lexical access and syntax knowledge integration over this phoneme network are then performed to deliver the decoded sentences. Experimental evidence illustrates that the proposed system outperforms several hybrid HMM/ANN systems with different configurations on the Wall Street Journal task while it is competitive with conventional LVCSR technology. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2010 | Experimental studies on continuous speech recognition using neural architectures with "adaptive" hidden activation functionsabstractThe choice of hidden non-linearity in a feed-forward multi-layer perceptron (MLP) architecture is crucial to obtain good generalization capability and better performance. Nonetheless, little attention has been paid to this aspect in the ASR field. In this work, we present some initial, yet promising, studies toward improving ASR performance by adopting hidden activation functions that can be automatically learned from the data and change shape during training. This adaptive capability is achieved through the use of orthonormal Hermite polynomials. The “adaptive” MLP is used in two neural architectures that generate phone posterior estimates, namely, a standalone configuration and a hierarchical structure. The posteriors are input to a hybrid phone recognition system with good results on the TIMIT corpus. A scheme for optimizing the contributions of high-accuracy neural architectures is also investigated, resulting in a relative improvement of ~9.0% over a non-optimized combination. Finally, initial experiments on the WSJ Nov92 task show that the proposed technique scales well up to large vocabulary continuous speech recognition (LVCSR) tasks. Sabato Marco Siniscalchi, Torbjørn Svendsen, Filippo Sorbello, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2010 | An acoustic segment model approach to incorporating temporal information into speaker modeling for text-independent speaker recognitionabstractWe propose an acoustic segment model (ASM) approach to incorporating temporal information into speaker modeling in text-independent speaker recognition. In training, the proposed framework first estimates a collection of ASM-based universal background models (UBMs). Multiple sets of speaker-specific ASMs are then obtained by adapting the ASM-based UBMs with speaker-specific enrollment data. A novel usage of language models of the ASM units is also proposed to characterize transitions among ASMs. In the testing phase the ASM sets for the claimed speaker and UBMs, along with a bigram ASM language model, are used to calculate detection scores for each given test utterance. We report on speaker recognition experiments using the NIST 2001 SRE database. The results clearly indicate that the proposed ASM-based method achieves a notable improvement over the GMM-based speaker modeling in which no temporal modeling is considered. Moreover, a further error reduction is obtained by integrating the language model, another inclusion of temporal properties made possibly by ASM based speaker modeling. Yu Tsao 0001, Hanwu Sun, Haizhou Li 0001, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2010 | Shrinkage model adaptation in automatic speech recognitionabstractInspired by the success of least absolute shrinkage and selection operator (LASSO) in statistical learning, we propose an regularized maximum likelihood linear regression (MLLR) to estimate models with only a limited set of adaptation data to improve accuracy for automatic speech recognition, by regularizing the standard MLLR objective function with an constraint. The so-called LASSO MLLR is a natural solution to the data insufficiency problem because the constraint regularizes some parameters to exactly 0 and reduces the number of free parameters to estimate. Tested on the 5k-WSJ0 task, the proposed LASSO MLLR gives significant word error rate reduction from the errors obtained with the standard MLLR in an utterance-by-utterance unsupervised adaptation scenario. 1. Jinyu Li 0001, Yu Tsao 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2010 | A particle filter feature compensation approach to robust speech recognition
Aleem Mushtaq, Yu Tsao 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2010 | Exploiting context-dependency and acoustic resolution of universal speech attribute models in spoken language recognitionabstractThis paper expands a previously proposed universal acoustic characterization approach to spoken language identification (LID) by studying different ways of modeling attributes to improve language recognition. The motivation is to describe any spoken language with a common set of fundamental units. Thus, a spoken utterance is first tokenized into a sequence of universal attributes. Then a vector space modeling approach delivers the final LID decision. Context-dependent attribute models are now used to better capture spectral and temporal characteristics. Also, an approach to expand the set of attributes to increase the acoustic resolution is studied. Our experiments show that the tokenization accuracy positively affects LID results by producing a 2.8% absolute improvement over our previous 30-second NIST 2003 performance. This result also compares favorably with the best results on the same task known by the authors when the tokenizers are trained on language-dependent OGI-TS data. Sabato Marco Siniscalchi, Jeremy Reed, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2010 | A Study on the Generalization Capability of Acoustic Models for Robust Speech RecognitionabstractIn this paper, we explore the generalization capability of acoustic model for improving speech recognition robustness against noise distortions. While generalization in statistical learning theory originally refers to the model's ability to generalize well on unseen testing data drawn from the same distribution as that of the training data, we show that good generalization capability is also desirable for mismatched cases. One way to obtain such general models is to use margin-based model training method, e.g., soft-margin estimation (SME), to enable some tolerance to acoustic mismatches without a detailed knowledge about the distortion mechanisms through enhancing margins between competing models. Experimental results on the Aurora-2 and Aurora-3 connected digit string recognition tasks demonstrate that, by improving the model's generalization capability through SME training, speech recognition performance can be significantly improved in both matched and low to medium mismatched testing cases with no language model constraints. Recognition results show that SME indeed performs better with than without mean and variance normalization, and therefore provides a complimentary benefit to conventional feature normalization techniques such that they can be combined to further improve the system performance. Although this study is focused on noisy speech recognition, we believe the proposed margin-based learning framework can be extended to dealing with different types of distortions and robustness issues in other machine learning applications. Jinyu Li 0001, Chng Eng Siong, Haizhou Li 0001, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 5 |
| 2009 | MAP estimation of online mapping parameters in ensemble speaker and speaking environment modelingabstractRecently, an ensemble speaker and speaking environment modeling (ESSEM) framework was proposed to enhance automatic speech recognition performance under adverse conditions. In the online phase of ESSEM, the prepared environment structure in the offline stage is transformed to a set of acoustic models for the target testing environment by using a mapping function. In the original ESSEM framework, the mapping function parameters are estimated based on a maximum likelihood (ML) criterion. In this study, we propose to use a maximum a posteriori (MAP) criterion to calculate the mapping function to avoid a possible over-fitting problem that can degrade the accuracy of environment characterization. For the MAP estimation, we also study two types of prior densities, namely, clustered prior and hierarchical prior, in this paper. On the Aurora-2 task using either type of prior densities, MAP-based ESSEM can achieve better performance than ML-based ESSEM, especially under low SNR conditions. When comparing to our best baseline results, the MAP-based ESSEM achieves a 14.97% (5.41% to 4.60%) word error rate reduction in average at a signal to noise ratio of 0 dB to 20 dB over the three testing sets. Yu Tsao 0001, Shigeki Matsuda, Satoshi Nakamura 0001, Chin-Hui Lee 0001 |
ASRU | 4 |
| 2009 | A study on hidden Markov model's generalization capability for speech recognitionabstractFrom statistical learning theory, the generalization capability of a model is the ability to generalize well on unseen test data which follow the same distribution as the training data. This paper investigates how generalization capability can also improve robustness when testing and training data are from different distributions in the context of speech recognition. Two discriminative training (DT) methods are used to train the hidden Markov model (HMM) for better generalization capability, namely the minimum classification error (MCE) and the soft-margin estimation (SME) methods. Results on Aurora-2 task show that both SME and MCE are effective in improving one of the measures of acoustic model's generalization capability, i.e. the margin of the model, with SME be moderately more effective. In addition, the better generalization capability translates into better robustness of speech recognition performance, even when there is significant mismatch between the training and testing data. We also applied the mean and variance normalization (MVN) to preprocess the data to reduce the training-testing mismatch. After MVN, MCE and SME perform even better as the generalization capability now is more closely related to robustness. The best performance on Aurora-2 is obtained from SME and about 28% relative error rate reduction is achieved over the MVN baseline system. Finally, we also use SME to demonstrate the potential of better generalization capability in improving robustness in more realistic noisy task using the Aurora-3 task, and significant improvements are obtained. Jinyu Li 0001, Chng Eng Siong, Haizhou Li 0001, Chin-Hui Lee 0001 |
ASRU | 5 |
| 2009 | A study on multilingual acoustic modeling for large vocabulary ASRabstractWe study key issues related to multilingual acoustic modeling for automatic speech recognition (ASR) through a series of large-scale ASR experiments. Our study explores shared structures embedded in a large collection of speech data spanning over a number of spoken languages in order to establish a common set of universal phone models that can be used for large vocabulary ASR of all the languages seen or unseen during training. Language-universal and language-adaptive models are compared with language-specific models, and the comparison results show that in many cases it is possible to build general-purpose language-universal and language-adaptive acoustic models that outperform language-specific ones if the set of shared units, the structure of shared states, and the shared acoustic-phonetic properties among different languages can be properly utilized. Specifically, our results demonstrate that when the context coverage is poor in language-specific training, we can use one tenth of the adaptation data to achieve equivalent performance in cross-lingual speech recognition. Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero, Chin-Hui Lee 0001 |
ICASSP | 6 |
| 2009 | A phonetic feature based lattice rescoring approach to LVCSRabstractLarge Vocabulary Continuous Speech Recognition (LVCSR) systems decode the input speech using diverse information sources, such as acoustic, lexical, and linguistic. Although most of the unreliable hypotheses are pruned during the recognition process, current state-of-the-art systems often make errors that are ldquounreasonablerdquo for human listeners. Several studies have shown that a proper integration of acoustic-phonetic information can be beneficial to reducing such errors. We have previously shown that high-accuracy phone recognition can be achieved if a bank of speech attribute detectors is used to compute a confidence score describing attribute activation levels that the current frame exhibits. In those experiments, the phone recognition system did not rely on the language model to follow their word sequence constraints, and the vocabulary was small. In this work, we extend our approach to LVCSR by introducing a second recognition step during which additional information not directly used during conventional log-likelihood based decoding is introduced. Experimental results show promising performance. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2009 | Ensemble speaker and speaking environment modeling approach with advanced online estimation processabstractRecently, we proposed an ensemble speaker and speaking environment modeling (ESSEM) framework to characterize speaker variability and speaking environments. In contrast to multi-style training, ESSEM uses single-style training to prepare multiple sets of environment-specific acoustic models. The ensemble of these acoustic models forms a prior structure of the environment for flexible prediction of unknown environment during testing. In this study, we present methods to further improve the precision for model characterization. We first study a weighted N-best information technique to well utilize the N-best transcription hypothesis in an unsupervised adaptation manner. Next, we introduce cohort selection and environment space adaptation techniques to online improve the resolution and coverage of the prior structure. With an integration of the proposed methods, we further improve the ESSEM performance over our previous study. On the Aurora-2 task, ESSEM achieves an average word error rate (WER) of 4.64%, corresponding to a 15.64% relative WER reduction over our best baseline result (5.50% to 4.64% WER) obtained with multi-condition training. Yu Tsao 0001, Jinyu Li 0001, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2009 | A study on soft margin estimation of linear regression parameters for speaker adaptationabstractWe formulate a framework for soft margin estimation-based linear regression (SMELR) and apply it to supervised speaker adaptation. Enhanced separation capability and increased discriminative ability are two key properties in margin-based discriminative training. For the adaptation process to be able to flexibly utilize any amount of data, we also propose a novel interpolation scheme to linearly combine the speaker independent (SI) and speaker adaptive SMELR (SMELR/SA) models. The two proposed SMELR algorithms were evaluated on a Japanese large vocabulary continuous speech recognition task. Both the SMELR and interpolated SI+SMELR/SA techniques showed improved speech adaptation performance in comparison with the well-known maximum likelihood linear regression (MLLR) method. We also found that the interpolation framework works even more effectively than SMELR when the amount of adaptation data is relatively small. Shigeki Matsuda, Yu Tsao 0001, Jinyu Li 0001, Satoshi Nakamura 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 5 |
| 2009 | Exploring universal attribute characterization of spoken languages for spoken language recognitionabstractWe propose a novel universal acoustic characterization approach to spoken language identification (LID), in which any spoken language is described with a common set of fundamental units defined “universally.” Specifically, manner and place of articulation form this unit inventory and are used to build a set of universal attribute models with data-driven techniques. Using the vector space modeling approaches to LID a spoken utterance is first decoded into a sequence of attributes. Then, a feature vector consisting of co-occurrence statistics of attribute units is created, and the final LID decision is implemented with a set of vector space language classifiers. Although the present study is just in its preliminary stage, promising results comparable to acoustically rich phone-based LID systems have already been obtained on the NIST 2003 LID task. The results provide clear insight for further performance improvements and encourage a continuing exploration of the proposed framework. Sabato Marco Siniscalchi, Jeremy Reed, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2009 | A study on integrating acoustic-phonetic information into lattice rescoring for automatic speech recognition
Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
Speech Commun. | 2 |
| 2009 | An Ensemble Speaker and Speaking Environment Modeling Approach to Robust Speech RecognitionabstractWe propose an ensemble speaker and speaking environment modeling (ESSEM) approach to characterizing environments in order to enhance performance robustness of automatic speech recognition systems under adverse conditions. The ESSEM process comprises two phases, the offline and the online. In the offline phase, we prepare an ensemble speaker and speaking environment space formed by a collection of super-vectors. Each super-vector consists of the entire set of means from all the Gaussian mixture components of a set of hidden Markov models that characterizes a particular environment. In the online phase, with the ensemble environment space prepared in the offline phase, we estimate the super-vector for a new testing environment based on a stochastic matching criterion. In this paper, we focus on methods for enhancing the construction and coverage of the environment space in the offline phase. We first demonstrate environment clustering and partitioning algorithms to structure the environment space well; then, we propose a minimum classification error training algorithm to enhance discrimination across environment super-vectors and therefore broaden the coverage of the ensemble environment space. We evaluate the proposed ESSEM framework on the Aurora2 connected digit recognition task. Experimental results verify that ESSEM provides clear improvement over a baseline system without environmental compensation. Moreover, the performance of ESSEM can be further enhanced by using well-structured environment spaces. Finally, we confirm that ESSEM gives the best overall performance with an environment space refined by an integration of all techniques. Yu Tsao 0001, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Toward a detector-based universal phone recognizerabstractIn recent research, we have proposed a high-accuracy bottom-up detection-based paradigm for continuous phone speech recognition. The key component of our system was a bank of articulatory detectors each of which computes a score describing an activation level of the specified speech phonetic features that the current frame exhibits. In this work, we present our first attempt at designing a universal phone recognizer using the detection-based approach. We show that our technique is intrinsically language independent since reliable articulatory detectors can be designed for diverse languages, and robust detection can be performed across languages. Moreover, a universal set of detectors is designed by sharing the training material available for several diverse languages. We further demonstrate that our approach makes it possible to decode new target languages by neither retraining nor applying acoustic adaptation techniques. We report phone recognition performance that compares favorably with the best results known by the authors on the OGI Multi-language Telephone Speech corpus. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2008 | Discriminative learning for optimizing detection performance in spoken language recognitionabstractWe propose novel approaches for optimizing the detection performance in spoken language recognition. Two objective functions are designed to directly relate model parameters to two performance metrics of interest, the detection cost function and the area under the detection-error-tradeoff curve, respectively. Both metrics are approximated with differentiable functions of model parameters by using a smoothing function based on a class misclassification measure. The model parameters are optimized by using the generalized probabilistic descent algorithm. We conduct experiments on the NIST 2003 and 2005 Language Recognition Evaluation corpora. Results show that the proposed approaches effectively improve the performance over the maximum likelihood training approach. Donglai Zhu, Haizhou Li 0001, Bin Ma 0001, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2008 | On a generalization of margin-based discriminative training to robust speech recognitionabstractRecently, there have been intensive studies of margin-based learning for automatic speech recognition (ASR). It is our believe that by securing a margin from the decision boundaries to the training samples, a correct decision can still be made if the mismatches between testing and training samples are well within the tolerance region specified by the margin. This nice property should be effective for robust ASR, where the testing condition is different from those in training. In this paper, we report on experiment results with soft margin estimation (SME) on the Aurora2 task and show that SME is very effective under clean training with more than 50% relative word error reductions in the clean, 20db, and 15db testing conditions, and still gives a slight improvement over conventional multi-condition training approaches. This demonstrates that the margin in SME can equip recognizers with a nice generalization property under adverse conditions. Jinyu Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2008 | Soft margin estimation with various separation levels for LVCSRabstractWe continue our previous work on soft margin estimation (SME) to large vocabulary continuous speech recognition (LVCSR) in two new aspects. The first is to formulate SME with different unit separation. SME methods focusing on string-, word-, and phone-level separation are defined. The second is to compare SME with all the popular conventional discriminative training (DT) methods, including maximum mutual information estimation (MMIE), minimum classification error (MCE), and minimum word/phone error (MWE/MPE). Tested on the 5k-word Wall Street Journal task, all the SME methods achieves a relative word error rate (WER) reduction from 17% to 25% over our baseline. Among them, phone-level SME obtains the best performance. Its performance is slightly better than that of MPE, and much better than those of other conventional DT methods. With the comprehensive comparison with conventional DT methods, SME demonstrates its success on LVCSR tasks. Jinyu Li 0001, Zhijie Yan, Chin-Hui Lee 0001, Renhua Wang |
INTERSPEECH | 3 |
| 2008 | Continuous phone recognition without target language training dataabstractDesigning an automatic speech recognition system with little or no language-specific training data is a challenging research topic because collecting abundant speech training data is not always an easy job for all possible languages of interest. According to our previous studied detection-based paradigm, we used a set of 21 acoustic phonetic attributes shared by five languages to perform Japanese phone recognition without using any Japanese speech training data. In this paper, we address the key issue of designing attribute-to-phone mapping models by two techniques: (1) a phone-based background model for each of the speech attribute detector to improve attribute detection; and (2) a data-driven clustering algorithm to group attribute-to-phone mapping rules of known languages to predict such rules for target phones in an unseen language. We report on experimental results of continuous Japanese phone recognition with the OGI Multilingual Speech Corpus and show that the proposed approach indeed decreases the false rejection rate of attribute detection, and improves the phone recognition accuracy Dau-Cheng Lyu, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2008 | A penalized logistic regression approach to detection based phone classificationabstractRecently, we have proposed a detection-based speech recognizer which has two main components: a bank of phonetic feature detectors implemented with hidden Markov models (HMMs), and an event merger. Each detector generates a score that pertains to some phonetic features, e.g. voicing. The merger combines all these scores to generate phone labels. The parameters of the detectors and the merger can be optimized either separately or jointly, and we showed that penalized logistic regression machine (PLRM) is a convenient tool for joint optimization. We validated our approach on a rescoring scheme. In this work, we tackle the phone classification problem and show that high level phone accuracy can be achieved without a direct modeling of the phones when PLRM is used. We also show that better results can be obtained by increasing the number of phonetic features, and that our method outperforms phone classifiers trained either by maximum likelihood estimation, or maximum mutual information Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2008 | Improving the ensemble speaker and speaking environment modeling approach by enhancing the precision of the online estimation process
Yu Tsao 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2008 | Optimizing the Performance of Spoken Language Recognition With Discriminative TrainingabstractThe performance of spoken language recognition system is typically formulated to reflect the detection cost and the strategic decision points along the detection-error-tradeoff curve. We propose a performance metrics optimization (PMO) approach to optimizing the detection performance of Gaussian mixture model classifiers. We design the objective functions to directly relate the model parameters to the performance metrics of interest, i.e., the detection cost function and the area under the detection-error-tradeoff curve. Both metrics are approximated by differentiable functions of model parameters. In this way, the model parameters can be optimized with the generalized probabilistic descent algorithm, a typical discriminative training technique. We conduct the experiments on the NIST 2003 and 2005 Language Recognition Evaluation corpora. The experimental results show that the PMO approach effectively improves the performance over the maximum-likelihood training approach. Donglai Zhu, Haizhou Li 0001, Bin Ma 0001, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | A study on soft margin estimation for LVCSRabstractWe extend our previous work on soft margin estimation (SME) to large vocabulary continuous speech recognition in two aspects. The first is to use the extended Baum-Welch method to replace the conventional generalized probabilistic descent algorithm for optimization. The second is to compare SME with minimum classification error (MCE) training with the same implementation details in order to show that it is indeed the margin component in the objective function with margin-based utterance and frame selection that contributes to the success of SME. Tested on the 5 k-word Wall Street Journal task, all the SME methods work better than MCE. The best SME approach achieves a relative word error rate reduction of about 19% over our best baseline performance. This enhancement can only be demonstrated because of our use of margin-based objective function and the extended Baum-Welch parameter optimization method. Jinyu Li 0001, Zhijie Yan, Chin-Hui Lee 0001, Renhua Wang |
ASRU | 3 |
| 2007 | Towards bottom-up continuous phone recognitionabstractWe present a novel approach to designing bottom-up automatic speech recognition (ASR) systems. The key component of the proposed approach is a bank of articulatory attribute detectors implemented using a set of feed-forward artificial neural networks (ANNs). Each detector computes a score describing an activation level of the specified speech attributes that the current frame exhibits. These cues are first combined by an event merger that provides some evidence about the presence of a higher level feature which is then verified by an evidence verifier to produce hypotheses at the phone or word level. We evaluate several configurations of our proposed system on a continuous phone recognition task. Experimental results on the TIMIT database show that the system achieves a phone error rate of 25% which is superior to results obtained with either hidden Markov model (HMM) or conditional random field (CRF) based recognizers. We believe the system's inherent flexibility and the ease of adding new detectors may provide further improvements. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
ASRU | 3 |
| 2007 | Two extensions to ensemble speaker and speaking environment modeling for robust automatic speech recognitionabstractRecently an ensemble speaker and speaking environment modeling (ESSEM) approach to characterizing unknown testing environments was studied for robust speech recognition. Each environment is modeled by a super-vector consisting of the entire set of mean vectors from all Gaussian densities of a set of HMMs for a particular environment. The super-vector for a new testing environment is then obtained by an affine transformation on the ensemble super-vectors. In this paper, we propose a minimum classification error training procedure to obtain discriminative ensemble elements, and a super-vector clustering technique to achieve refined ensemble structures. We test these two extentions to ESSEM on Aurora2. In a per-utterance unsupervised adaptation mode we achieved an average WER of 4.99% from OdB to 20 dB conditions with these two extentions when compared with a 5.51% WER obtained with the ML-trained gender-dependent baseline. To our knowledge this represents the best result reported in the literature on the Aurora2 connected digit recognition task. Yu Tsao 0001, Chin-Hui Lee 0001 |
ASRU | 2 |
| 2007 | Approximate Test Risk Minimization Through Soft Margin EstimationabstractIn a recent study, we proposed soft margin estimation (SME) to learn parameters of continuous density hidden Markov models (HMMs). Our earlier experiments with connect digit recognition have shown that SME offers great advantages over other state-of-the-art discriminative training methods. In this paper, we illustrate SME from a perspective of statistical learning theory and show that by including a margin in formulating the SME objective function it is capable of directly minimizing the approximate test risk, while most other training methods intent to minimize only the empirical risks. We test SME on the 5k-word Wall Street Journal task, and find the proposed approach achieves a relative word error rate reduction of about 10% over our best baseline results in different experimental configurations. We believe this is the first attempt to show the effectiveness of margin-based acoustic modeling for large vocabulary continuous speech recognition. We also expect further performance improvements in the future because the approximate test risk minimization principle offers a flexible and yet rigorous framework to facilitate easy incorporation of new margin-based optimization criteria into HMM training. Jinyu Li 0001, Sabato Marco Siniscalchi, Chin-Hui Lee 0001 |
ICASSP (4) | 3 |
| 2007 | High-Accuracy Phone Recognition By Combining High-Performance Lattice Generation and Knowledge Based RescoringabstractThis study is a result of a collaboration project between two groups, one from Brno University of Technology and the other from Georgia Institute of Technology (GT). Recently the Brno recognizer is known to outperform many state-of-the-art systems on phone recognition, while the GT knowledge-based lattice rescoring module has been shown to improve system performance on a number of speech recognition tasks. We believe a combination of the two system results in high-accuracy phone recognition. To integrate the two very different modules, we modify Brno's phone recognizer into a phone lattice hypothesizer to produce high-quality phone lattices, and feed them directly into the knowledge-based module to rescore the lattices. We test the combined system on the TIMIT continuous phone recognition task without retraining the individual subsystems, and we observe that the phone error rate was effectively reduced to 19.78% from 24.41% produced by the Brno phone recognizer. To the best of the authors' knowledge this result represents the lowest ever error rate reported on the TIMIT continuous phone recognition task. Sabato Marco Siniscalchi, Petr Schwarz, Chin-Hui Lee 0001 |
ICASSP (4) | 3 |
| 2007 | Boosting of Maximal Figure of Merit Classifiers for Automatic Image AnnotationabstractVisual information contained in a scene is very complex and can be represented with multiple features describing aspects of the entire information. In this paper we propose a boosting approach to automatic image annotation by building strong classifiers based on multiple collections of weak concept classifiers with each collection focused on a single visual feature. The weak classifiers are trained with a maximal figure-of-merit learning approach. By exploiting multiple features the boosting procedure allows to build classifiers able to pick the most discriminative feature for the specific annotation task. Filippo Vella, Chin-Hui Lee 0001, Salvatore Gaglio |
ICIP (2) | 2 |
| 2007 | Use of Generalized Pattern Model for Video AnnotationabstractThis paper proposes an integrated framework that combines intra-shot and temporal inter-shot sequence analysis based on visual features to find stable patterns for video annotation. At the shot level, we perform multi-stage kNN classification using the global visual features to identify good candidate shots containing the concept. At the sequence level, we aim to find patterns of shot sequences around candidate shots with consistent statistical characteristics and dynamics. We discretize the shot contents into fixed set of tokens, and transform the high dimensional continuous video streams into tractable token sequences. We then extend the soft matching model to reveal video sequence patterns and flexibly match the patterns around candidate shots. We combine both local shot matching method and generalized pattern model using both visual and text features. Experimental results on TRECVID2006 dataset demonstrate that the proposed approach is effective. Tat-Seng Chua, Lekha Chaisorn, Chin-Hui Lee 0001 |
ICME | 4 |
| 2007 | Soft margin feature extraction for automatic speech recognitionabstractWe propose a new discriminative learning framework, called soft margin feature extraction (SMFE), for jointly optimizing the parameters of transformation matrix for feature extraction and of hidden Markov models (HMMs) for acoustic modeling. SMFE extends our previous work of soft margin estimation (SME) to feature extraction. Tested on the TIDIGITS connected digit recognition task, the proposed approach achieves a string accuracy of 99.61%, much better than our previously reported SME results. To our knowledge, this is the first study on applying the margin-based method in joint optimization of feature extraction and acoustic modeling. The excellent performance of SMFE demonstrates the success of soft margin based method, which targets to obtain both high accuracy and good model generalization. Index Terms: discriminative feature extraction, hidden Markov model, margin, automatic speech recognition Jinyu Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2007 | An ensemble modeling approach to joint characterization of speaker and speaking environments
Yu Tsao 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2007 | Enhancing image annotation by integrating concept ontology and text-based bayesian learning modelabstractAutomatic image annotation (AIA) has been a hot research topic in recent years since it can be used to support concept-based image retrieval. However, most existing AIA models depend heavily on the availability of a large number of labeled training samples, which require significant human labeling efforts. In this paper, we propose a novel learning framework which integrates text-based Bayesian model (TBM) and concept ontology to effectively expand the training set of each concept class without the need of additional human labeling efforts or collecting additional training images from other data sources. The basic idea lies in exploiting the text information from training set to provide additional effective annotations for training images so that training data for each concept class can be augmented. In this study we employ Bayesian Hierarchical Multinomial Mixture Models (BHMMMs) as our baseline AIA model. By combining additional annotations obtained from TBM into each concept class in the training phase, the performance of BHMMMs can be significantly improved on Corel image dataset with 263 testing concepts as compared to the state-of-the-art AIA models under the same experimental configurations. Chin-Hui Lee 0001, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2007 | Fusion of Region and Image-Based Techniques for Automatic Image Annotation
Tat-Seng Chua, Chin-Hui Lee 0001 |
MMM (1) | 3 |
| 2007 | Report on the NSF-sponsored Human Language Technology Workshop on Industrial Centers
Mary P. Harper, Alex Acero, Srinivas Bangalore, Jordan Cohen, Barbara Cuthill, Carol Y. Espy-Wilson, Christiane Fellbaum, John Garofolo, Chin-Hui Lee 0001, Jim Lester, Andrew McCallum, Nelson Morgan, Michael Picheney, Joseph Picone, Lance Ramshaw, Jeffrey C. Reynar, Hadar Shemtov, Clare Voss |
MTSummit | 10 |
| 2007 | A Vector Space Modeling Approach to Spoken Language IdentificationabstractWe propose a novel approach to automatic spoken language identification (LID) based on vector space modeling (VSM). It is assumed that the overall sound characteristics of all spoken languages can be covered by a universal collection of acoustic units, which can be characterized by the acoustic segment models (ASMs). A spoken utterance is then decoded into a sequence of ASM units. The ASM framework furthers the idea of language-independent phone models for LID by introducing an unsupervised learning procedure to circumvent the need for phonetic transcription. Analogous to representing a text document as a term vector, we convert a spoken utterance into a feature vector with its attributes representing the co-occurrence statistics of the acoustic units. As such, we can build a vector space classifier for LID. The proposed VSM approach leads to a discriminative classifier backend, which is demonstrated to give superior performance over likelihood-based n-gram language modeling (LM) backend for long utterances. We evaluated the proposed VSM framework on 1996 and 2003 NIST Language Recognition Evaluation (LRE) databases, achieving an equal error rate (EER) of 2.75% and 4.02% in the 1996 and 2003 LRE 30-s tasks, respectively, which represents one of the best results reported on these popular tasks Haizhou Li 0001, Bin Ma 0001, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Approximate Test Risk Bound Minimization Through Soft Margin EstimationabstractInspired by the great success of margin-based classifiers, there is a trend to incorporate the margin concept into hidden Markov modeling for speech recognition. Several attempts based on margin maximization were proposed recently. In this paper, a new discriminative learning framework, called soft margin estimation (SME), is proposed for estimating the parameters of continuous-density hidden Markov models. The proposed method makes direct use of the successful ideas of soft margin in support vector machines to improve generalization capability and decision feedback learning in minimum classification error training to enhance model separation in classifier design. SME is illustrated from a perspective of statistical learning theory. By including a margin in formulating the SME objective function, SME is capable of directly minimizing an approximate test risk bound. Frame selection, utterance selection, and discriminative separation are unified into a single objective function that can be optimized using the generalized probabilistic descent algorithm. Tested on the TIDIGITS connected digit recognition task, the proposed SME approach achieves a string accuracy of 99.43%. On the 5 k-word Wall Street Journal task, SME obtains relative word error rate reductions of about 10% over our best baseline results in different experimental configurations. We believe this is the first attempt to show the effectiveness of margin-based acoustic modeling for large vocabulary continuous speech recognition in a hidden Markov model framework. Further improvements are expected because the approximate test risk bound minimization principle offers a flexible and rigorous framework to facilitate incorporation of new margin-based optimization criteria into hidden Markov model training. Jinyu Li 0001, Ming Yuan 0001, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Soft margin estimation of hidden Markov model parametersabstractWe propose a new discriminative learning framework, called soft margin estimation (SME), for estimating parameters of continuous density hidden Markov models. The proposed method makes direct usage of the successful ideas of soft margin in support vector machines to improve generalization capability, and of decision feedback learning in minimum classification error training to enhance model separation in classifier design. We attempt to incorporate frame selection, utterance selection and discriminative separation in a single unified objective function that can be optimized with the wellknown generalized probabilistic descent algorithm. We demonstrate the advantage of SME in theory and practice over other state-of-the-art techniques. Tested on a connected digit recognition task, the proposed SME approach achieves a string accuracy of 99.33%. To our knowledge, this is the best result ever reported on the TIDIGITS database. 1. Jinyu Li 0001, Ming Yuan 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2006 | A study on detection based automatic speech recognition
Chengyuan Ma, Yu Tsao 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2006 | A study on lattice rescoring with knowledge scores for automatic speech recognitionabstractWe study lattice rescoring with knowledge scores for automatic speech recognition. Frame-based log likelihood ratio is adopted as a score measure of the goodness-of-fit between a speech segment and the knowledge sources. We evaluate our approach in two different applications: phone recognition, and connected digit continuous recognition. By incorporating knowledge scores obtained from 15 attribute detectors for place and manner of articulation, we reduced phone error rate from 40.52% to 35.16% using monophone models. The error rate can be further reduced to 33.42% for triphone models. The same lattice rescoring algorithm is extended to connected digit recognition using the TIDIGITS database, and without using any digit-specific training data. We observed the digit error rate can be effectively reduced to 4.03% from 4.54% which was obtained with the conventional Viterbi decoding algorithm with no knowledge scores. Sabato Marco Siniscalchi, Jinyu Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2006 | A vector space approach to environment modeling for robust speech recognitionabstractWe propose a vector space approach to characterizing environments for robust speech recognition. We represent a given environment by a super-vector formed by concatenating all the mean vectors of the Gaussian mixture components of the state observation densities of all hidden Markov models trained in the particular environment. New environment super-vectors can now be obtained either by an interpolation method with a collection of super-vectors trained from many real or simulated environments or by a transformation performed on an anchor super-vector for a specific environment, such as a clean condition. At a 5dB signal-to-noise (SNR) level, both interpolation- and transformation-based approaches achieve a significant error rate reduction of close to 47% from a baseline system with cepstral mean subtraction (CMS) with only two adaptation utterances. When incorporating N-best information to perform unsupervised adaptation at 5dB SNR with the same two utterances, we achieve a relative error reduction of about 40%, close to that achieved in the supervised mode. Index Terms: acoustic modeling, environment adaptation Yu Tsao 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2006 | A maximal figure-of-merit (MFoM)-learning approach to robust classifier design for text categorizationabstractWe propose a maximal figure-of-merit (MFoM)-learning approach for robust classifier design, which directly optimizes performance metrics of interest for different target classifiers. The proposed approach, embedding the decision functions of classifiers and performance metrics into an overall training objective, learns the parameters of classifiers in a decision-feedback manner to effectively take into account both positive and negative training samples, thereby reducing the required size of positive training data. It has three desirable properties: (a) it is a performance metric, oriented learning; (b) the optimized metric is consistent in both training and evaluation sets; and (c) it is more robust and less sensitive to data variation, and can handle insufficient training data scenarios. We evaluate it on a text categorization task using the Reuters-21578 dataset. Training an F 1 -based binary tree classifier using MFoM, we observed significantly improved performance and enhanced robustness compared to the baseline and SVM, especially for categories with insufficient training samples. The generality for designing other metrics-based classifiers is also demonstrated by comparing precision, recall, and F 1 -based classifiers. The results clearly show consistency of performance between the training and evaluation stages for each classifier, and MFoM optimizes the chosen metric. Chin-Hui Lee 0001, Tat-Seng Chua |
ACM Trans. Inf. Syst. | 3 |
| 2005 | A Study on Knowledge Source Integration for Candidate Rescoring in Automatic Speech RecognitionabstractWe propose a rescoring framework for speech recognition that incorporates acoustic phonetic knowledge sources. The scores corresponding to all knowledge sources are generated from a collection of neural network based classifiers. Rescoring is then performed by combining different knowledge scores and they are used to reorder candidate strings provided by state-of-the-art HMM-based speech recognizers. We report on continuous phone recognition experiments using the TIMIT database. Our results indicate that classifying manners and places of articulation provides additional information in rescoring, and improved accuracies over our best baseline speech recognizers are achieved using both context-independent and context-dependent phone models. The same technique can be extended to lattice rescoring and large vocabulary continuous speech recognition. Jinyu Li 0001, Yu Tsao 0001, Chin-Hui Lee 0001 |
ICASSP (1) | 3 |
| 2005 | A text categorization approach to automatic language identification
Bin Ma 0001, Haizhou Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2005 | On designing and evaluating speech event detectorsabstractWe study issues related to designing speech event detectors for automatic speech recognition. Event detection is a critical component of a recently proposed automatic speech attribute transcription (ASAT) paradigm for speech research. Similar to keyword spotting and non-keyword rejection, a good detector needs to effectively detect speech attributes of interest while rejecting extraneous events. We compare frame and segment based detectors, study their properties in detecting manners of articulation, and propose new performance measures. We test these detectors on the TIMIT database with several evaluation criteria. Our results indicate that segment based detectors outperform frame based detectors in several key aspects of speech detector design. We also show that the performance can be significantly enhanced by incorporating discriminative training into designing speech event detectors. Jinyu Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2005 | An acoustic segment modeling approach to automatic language identification
Bin Ma 0001, Haizhou Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2005 | A study on separation between acoustic models and its applicationsabstractWe study separation between models of speech attributes. A good measure of separation usually serves as a key indicator of the discrimination power of these speech models because it can often be used to indirectly determine the performance of speech recognition and verification systems. In this study, we use a probabilistic distance, called generalized log likelihood ratio (GLLR), to measure the separation between a model of a target speech attribute and models of its competing attributes. We illustrate five applications to compare separations among models obtained over multiple levels of discrimination capabilities, at various degrees of acoustic definitions and resolutions, under mismatched training and testing conditions, and with different training criteria and speech parameters. We demonstrate that the well-known GLLR distance and its corresponding histograms also provide a good utility to qualitatively and quantitatively characterize the properties of trained models without performing large scale speech recognition and verification experiments. Yu Tsao 0001, Jinyu Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2004 | A hierarchical approach to story segmentation of large broadcast news video corpusabstractA multi-modal two-level framework for news story segmentation was proposed in Chaisorn et al. (2002). This paper presents our extended work scaled to cope with a large news video corpus used in TRECVID 2003 evaluation. We divided our system into two levels: the shot level that classifies input video shots into one of the predefined categories using a hybrid of heuristic and learning based approaches; and story level that performs story segmentation using the HMM framework based on the output of shot level and other temporal features. A heuristic rule-based technique is then employed to classify each detected story into "news" or "miscellaneous". We evaluated our system on over 120 hours of news video and showed that our system could achieve an accuracy of more than 77%. Our system came first in the TRECVID 2003 story segmentation task Lekha Chaisorn, Tat-Seng Chua, Chin-Hui Lee 0001, Qi Tian 0002 |
ICME | 3 |
| 2004 | A MFoM learning approach to robust multiclass multi-label text categorizationabstractProceedings, Twenty-First International Conference on Machine Learning, ICML 2004 Chin-Hui Lee 0001, Tat-Seng Chua |
ICML | 3 |
| 2003 | A maximal figure-of-merit learning approach to text categorizationabstractA novel maximal figure-of-merit (MFoM) learning approach to text categorization is proposed. Different from the conventional techniques, the proposed MFoM method attempts to integrate any performance metric of interest (e.g. accuracy, recall, precision, or F1 measure) into the design of any classifier. The corresponding classifier parameters are learned by optimizing an overall objective function of interest. To solve this highly nonlinear optimization problem, we use a generalized probabilistic descent algorithm. The MFoM learning framework is evaluated on the Reuters-21578 task with LSI-based feature extraction and a binary tree classifier. Experimental results indicate that the MFoM classifier gives improved F1 and enhanced robustness over the conventional one. It also outperforms the popular SVM method in micro-averaging F1. Other extensions to design discriminative multiple-category MFoM classifiers for application scenarios with new performance metrics could be envisioned too. Chin-Hui Lee 0001, Tat-Seng Chua |
SIGIR | 3 |
| 2003 | A Multi-Modal Approach to Story Segmentation for News Video
Lekha Chaisorn, Tat-Seng Chua, Chin-Hui Lee 0001 |
World Wide Web | 3 |
| 2002 | The segmentation of news video into story unitsabstractThe segmentation of news video into single-story semantic units is a challenging problem. This research proposes a two-level, multi-modal framework to tackle this problem. The video is analyzed at the shot and story unit (or scene) levels using a variety of features and techniques. At the shot level, we employ a decision tree to classify the shot into one of 13 predefined categories. At the scene level, we perform HMM (hidden Markov models) analysis to locate the story boundaries. We test the performance of our system using two days of news video obtained from the MediaCorp of Singapore. Our initial results indicate that we could achieve a high accuracy of over 95% for shot classification, and over 89% in F/sub 1/ measure on scene/story boundary detection. Lekha Chaisorn, Tat-Seng Chua, Chin-Hui Lee 0001 |
ICME (1) | 3 |
| 2002 | Weighted graph based decision tree optimization for high accuracy acoustic modeling
Jinsong Zhang 0001, Satoshi Nakamura 0001, Chin-Hui Lee 0001, Tat-Seng Chua |
INTERSPEECH | 4 |
| 2002 | Multilingual speech recognition with language identification
Bin Ma 0001, Cuntai Guan, Haizhou Li 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |