VLDB 2026 Research / reviewers in the wild / expert
Chenglin Xu
dblp:125/2814
· DBLP profile ↗
32ranked-venue papers
10as first author
18since 2021 · last 2025
0000-0002-1584-6282ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 21 · 7 first-author · 10 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InvoxSVC: Any-to-any Zero-shot Singing Voice Conversion with In-Context Learning in Latent Flow MatchingabstractRecent advancements in singing voice conversion (SVC) have focused on achieving zero-shot, any-to-any voice transformation capabilities. Many approaches attempt to modify voice characteristics by incorporating global timbre variables into acoustic models. However, these methods often depend heavily on the capabilities of timbre extractors and lack an understanding of temporal local information. This limitation poses challenges, particularly in replicating specific voice qualities such as those of children. To address this issue, we introduce InvoxSVC, a latent flow matching model (LFM) designed for rapid and precise singing voice conversion with a particular emphasis on capturing temporal local features. While reducing the residual timbral information in the source singing encoding through singer-guidance, InvoxSVC enhances the model’s ability to capture temporal nuances by integrating in-context learning during inference. Additionally, the model employs a pre-trained high-fidelity variational autoencoder (VAE) to improve waveform generation. In comparative evaluations, InvoxSVC outperforms the open-source project So-VITS-SVC in both objective and subjective assessments. Wangjin Zhou, Tianjiao Du, Wenhao Guan, Chenglin Xu, Yi Zhao 0006, Tatsuya Kawahara |
ICME | 5 |
| 2025 | Simple and Effective Content Encoder for Singing Voice Conversion via SSL-Embedding Dimension Reduction
Wangjin Zhou, Tianjiao Du, Chenglin Xu, Sheng Li 0010, Yi Zhao 0006, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2025 | A neural network approach for speech enhancement and noise-robust bandwidth extension
Chenglin Xu |
Comput. Speech Lang. | 2 |
| 2023 | KAQ: A Non-Intrusive Stacking Framework for Mean Opinion Score Prediction with Multi-Task LearningabstractNon-intrusive speech quality assessment aims to replace the time-consuming subjective evaluation metric, i.e., mean opinion score (MOS), by designing a trainable model to predict the MOS. In this work, we propose a non-intrusive stacking framework, named KAQ, to automatically achieve MOS prediction. KAQ benefits from the stacking algorithm to harness the capabilities of multiple well-performing machine learning models. And these models are trained to predict not only MOS but also intelligibility score via a multi-task learning. The experimental results show that KAQ achieves significantly better performance than the baselines in the 2023 VoiceMOS challenge and also wins second place in terms of mean square error at the utterance and system levels. Chenglin Xu, Xiguang Zheng |
ASRU | 1 |
| 2023 | Load-aware dynamic controller placement based on deep reinforcement learning in SDN-enabled mobile cloud-edge computing networks
Chenglin Xu, Cheng Xu 0001, Siqi Li 0005, Tao Li 0075 |
Comput. Networks | 1 |
| 2023 | Neural speech enhancement with unsupervised pre-training and mixture training
Chenglin Xu, Lei Xie 0001 |
Neural Networks | 2 |
| 2022 | Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech EnhancementabstractDeep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, personalized speech enhancement. As shown in the evaluation results, this is achieved by employing a multi-stage and multi-loss training architecture that incorporates the recently proposed two-step structure, ASR loss produced by a back-end ASR encoder, and the speaker extraction network. Lianwu Chen, Chenglin Xu, Xinlei Ren, Xiguang Zheng |
ICASSP | 2 |
| 2022 | L-SpEx: Localized Target Speaker ExtractionabstractSpeaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s location is known in advance or detected by an extra visual cue, e.g., face image or video. In this paper, we propose an end-to-end localized target speaker extraction on pure speech cues, that is called L-SpEx. Specifically, we design a speaker localizer driven by the target speaker’s embedding to extract the spatial features, including direction-of-arrival (DOA) of the target speaker and beamforming output. Then, the spatial cues and target speaker’s embedding are both used to form a top-down auditory attention to the target speaker. Experiments on the multi-channel reverberant dataset called MCLibri2Mix show that our L-SpEx approach significantly outperforms the baseline system. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2022 | Progressive Tandem Learning for Pattern Recognition With Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) have shown clear advantages over traditional artificial neural networks (ANNs) for low latency and high computational efficiency, due to their event-driven nature and sparse communication. However, the training of deep SNNs is not straightforward. In this paper, we propose a novel ANN-to-SNN conversion and layer-wise learning framework for rapid and efficient pattern recognition, which is referred to as progressive tandem learning. By studying the equivalence between ANNs and SNNs in the discrete representation space, a primitive network conversion method is introduced that takes full advantage of spike count to approximate the activation value of ANN neurons. To compensate for the approximation errors arising from the primitive network conversion, we further introduce a layer-wise learning method with an adaptive training scheduler to fine-tune the network weights. The progressive tandem learning framework also allows hardware constraints, such as limited weight precision and fan-in connections, to be progressively imposed during training. The SNNs thus trained have demonstrated remarkable classification and regression capabilities on large-scale object recognition, image reconstruction, and speech separation tasks, while requiring at least an order of magnitude reduced inference time and synaptic operations than other state-of-the-art SNN implementations. It, therefore, opens up a myriad of opportunities for pervasive mobile and embedded devices with a limited power budget. Jibin Wu, Chenglin Xu, Daquan Zhou, Malu Zhang, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Selective Listening by Synchronizing Speech With LipsabstractA speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accompanying video track. Visual cues are particularly useful when a pre-enrolled speech is not available. In this work, we don’t rely on the target speaker’s pre-enrolled speech, but rather use the target speaker’s face track as the speaker cue, that is referred to as the auxiliary reference, to form an attractor towards the target speaker. We advocate that the temporal synchronization between the speech and its accompanying lip movements is a direct and dominant audio-visual cue. Therefore, we propose a self-supervised pre-training strategy, to exploit the speech-lip synchronization cue for target speaker extraction, which allows us to leverage abundant unlabeled in-domain data. We transfer the knowledge from the pre-trained model to the attractor encoder of the speaker extraction network. We show that the proposed speaker extraction network outperforms various competitive baselines in terms of signal quality, perceptual quality, and intelligibility, achieving state-of-the-art performance. Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference SignalsabstractSpeaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted speech in early stages is used as the reference speech for late stages. For the first time, we use frame-level sequential speech embedding as the reference for target speaker. This is a departure from the traditional utterance-based speaker embedding reference. In addition, a signal fusion scheme is proposed to combine the decoded signals in multiple scales with automatically learned weights. Experiments on WSJ0-2mix and its noisy versions (WHAM! and WHAMR!) show that SpEx++ consistently outperforms other state-of-the-art baselines. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2021 | Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion RecognitionabstractConvolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spectro-temporal-channel (STC) attention module that is integrated with CNN to improve representation learning ability. Our module infers an attention map along the three dimensions, namely time, frequency, and CNN channel. Experiments are conducted on the IEMOCAP database to evaluate the effectiveness of the proposed representation learning method. The results demonstrate that the proposed method outperforms the traditional CNN method by an absolute increase of 3.13% in terms of F1 score. Lili Guo 0001, Longbiao Wang, Chenglin Xu, Jianwu Dang 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 3 |
| 2021 | Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial TrainingabstractNeural speech enhancement degrades significantly in face of unseen noise. To address such mismatch, we propose to learn noise-agnostic feature representations by disentanglement learning, which removes the unspecified noise factor, while keeping the specified factors of variation associated with the clean speech. Specifically, a discriminator module is introduced to distinguish the type of noises, which is referred to as the disentangler. With the adversarial training strategy, a gradient reversal layer seeks to disentangle the noise factor and remove it from the feature representation. Experiment results show that the proposed approach achieves 5.8% and 5.2% relative improvements over the best baseline in terms of perceptual evaluation of the speech quality (PESQ) and segmental signal-to-noise ratio (SSNR), respectively. The ablation study indicates that the proposed disentangler module is also effective in other encoder-decoder-like structures. Nana Hou, Chenglin Xu, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 2 |
| 2021 | Muse: Multi-Modal Target Speaker Extraction with Visual CuesabstractSpeaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between speech and lip movement also serves as an informative cue. Motivated by this idea, we study a novel technique to use speech-lip visual cues to extract reference target speech directly from mixture speech during inference time, without the need of pre-recorded reference speech. We propose a multi-modal speaker extraction network, named MuSE, that is conditioned only on a lip image sequence. MuSE not only outperforms other competitive baselines in terms of SI-SDR and PESQ, but also shows consistent improvement in cross-dataset evaluations. Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li 0001 |
ICASSP | 3 |
| 2021 | Universal Speaker Extraction in the Presence and Absence of Target Speakers for Speech of One and Two Talkers
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz |
Interspeech | 2 |
| 2021 | GlobalPhone Mix-To-Separate Out of 2: A Multilingual 2000 Speakers Mixtures Database for Speech Separation
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz |
Interspeech | 2 |
| 2021 | Neural Speaker Extraction with Speaker-Speech Cross-Attention Network
Wupeng Wang, Chenglin Xu, Meng Ge, Haizhou Li 0001 |
Interspeech | 2 |
| 2021 | Target Speaker Verification With Selective Auditory Attention for Single and Multi-Talker SpeechabstractSpeaker verification has been studied mostly under the single-talker condition. It is adversely affected in the presence of interference speakers. Inspired by the study on target speaker extraction, e.g., SpEx, we propose a unified speaker verification framework for both single- and multi-talker speech, that is able to pay selective auditory attention to the target speaker. This target speaker verification (tSV) framework jointly optimizes a speaker attention module and a speaker representation module via multi-task learning. We study four different target speaker embedding schemes under the tSV framework. The experimental results show that all four target speaker embedding schemes significantly outperform other competitive solutions for multi-talker speech. Notably, the best tSV speaker embedding scheme achieves 76.0% and 55.3% relative improvements over the baseline system on the WSJ0-2mix-extr and Libri2Mix corpora in terms of equal-error-rate for 2-talker speech, while the performance of tSV for single-talker speech is on par with that of traditional speaker verification system, that is trained and evaluated under the same single-talker condition. Chenglin Xu, Wei Rao 0002, Jibin Wu, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Time-Domain Neural Network Approach for Speech Bandwidth ExtensionabstractIn this paper, we study the time-domain neural network approach for speech bandwidth extension. We propose a network architecture, named multi-scale fusion neural network (MfNet), that gradually restores the low-frequency signal and predicts the high-frequency signal through the exchange of information across different scale representations. We propose a training scheme to optimize the network with a combination of perceptual loss and time-domain adversarial loss. Experiments show the proposed multi-scale fusion network consistently outperforms the competing methods in terms of perceptual evaluation of speech quality (PESQ), signal to distortion rate (SDR), signal to noise ratio (SNR), log-spectral distance (LSD) and word error rate (WER). More promisingly, the multi-scale fusion network requires only 10% of the parameters of the time-domain reference baseline. Chenglin Xu, Nana Hou, Lei Xie 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 2 |
| 2020 | SpEx+: A Complete Time Domain Speaker Extraction NetworkabstractSpeaker extraction aims to extract the target speech signal from a multi-talker environment given a target speaker's reference speech.We recently proposed a time-domain solution, SpEx, that avoids the phase estimation in frequency-domain approaches.Unfortunately, SpEx is not fully a time-domain solution since it performs time-domain speech encoding for speaker extraction, while taking frequency-domain speaker embedding as the reference.The size of the analysis window for timedomain and the size for frequency-domain input are also different.Such mismatch has an adverse effect on the system performance.To eliminate such mismatch, we propose a complete time-domain speaker extraction solution, that is called SpEx+.Specifically, we tie the weights of two identical speech encoder networks, one for the encoder-extractor-decoder pipeline, another as part of the speaker encoder.Experiments show that the SpEx+ achieves 0.8dB and 2.1dB SDR improvement over the state-of-the-art SpEx baseline, under different and same gender conditions on WSJ0-2mix-extr database respectively. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | Speaker and Phoneme-Aware Speech Bandwidth Extension with Residual Dual-Path NetworkabstractSpeech bandwidth extension aims to generate a wideband signal from a narrowband (low-band) input by predicting the missing high-frequency components. It is believed that the general knowledge about the speaker and phonetic content strengthens the prediction. In this paper, we propose to augment the low-band acoustic features with i-vector and phonetic posteriorgram (PPG), which represent speaker and phonetic content of the speech, respectively. We also propose a residual dual-path network (RDPN) as the core module to process the augmented features, which fully utilizes the utterance-level temporal continuity information and avoids gradient vanishing. Experiments show that the proposed method achieves 20.2% and 7.0% relative improvements over the best baseline in terms of log-spectral distortion (LSD) and signal-to-noise ratio (SNR), respectively. Furthermore, our method is 16 times more compact than the best baseline in terms of the number of parameters. Nana Hou, Chenglin Xu, Van Tung Pham, Joey Tianyi Zhou, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | Multi-Task Learning for End-to-End Noise-Robust Bandwidth ExtensionabstractBandwidth extension aims to reconstruct wideband speech signals from narrowband inputs to improve perceptual quality. Prior studies mostly perform bandwidth extension under the assumption that the narrowband signals are clean without noise. The use of such extension techniques is greatly limited in practice when signals are corrupted by noise. To alleviate such problem, we propose an end-to-end time-domain framework for noise-robust bandwidth extension, that jointly optimizes a mask-based speech enhancement and an ideal bandwidth extension module with multi-task learning. The proposed framework avoids decomposing the signals into magnitude and phase spectra, therefore, requires no phase estimation. Experimental results show that the proposed method achieves 14.3% and 15.8% relative improvements over the best baseline in terms of perceptual evaluation of speech quality (PESQ) and log-spectral distortion (LSD), respectively. Furthermore, our method is 3 times more compact than the best baseline in terms of the number of parameters. Nana Hou, Chenglin Xu, Joey Tianyi Zhou, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | SpEx: Multi-Scale Time Domain Speaker Extraction NetworkabstractSpeaker extraction aims to mimic humans' selective auditory attention by extracting a target speaker's voice from a multi-talker environment. It is common to perform the extraction in frequency-domain, and reconstruct the time-domain signal from the extracted magnitude and estimated phase spectra. However, such an approach is adversely affected by the inherent difficulty of phase estimation. Inspired by Conv-TasNet, we propose a time-domain speaker extraction network (SpEx) that converts the mixture speech into multi-scale embedding coefficients instead of decomposing the speech signal into magnitude and phase spectra. In this way, we avoid phase estimation. The SpEx network consists of four network components, namely speaker encoder, speech encoder, speaker extractor, and speech decoder. Specifically, the speech encoder converts the mixture speech into multi-scale embedding coefficients, the speaker encoder learns to represent the target speaker with a speaker embedding. The speaker extractor takes the multi-scale embedding coefficients and target speaker embedding as input and estimates a receptive mask. Finally, the speech decoder reconstructs the target speaker's speech from the masked embedding coefficients. We also propose a multi-task learning framework and a multi-scale embedding implementation. Experimental results show that the proposed SpEx achieves 37.3%, 37.7% and 15.0% relative improvements over the best baseline in terms of signal-to-distortion ratio (SDR), scale-invariant SDR (SI-SDR), and perceptual evaluation of speech quality (PESQ) under an open evaluation condition. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Time-Domain Speaker Extraction NetworkabstractSpeaker extraction is to extract a target speaker's voice from multi-talker speech. It simulates humans' cocktail party effect or the selective listening ability. The prior work mostly performs speaker extraction in frequency domain, then reconstructs the signal with some phase approximation. The inaccuracy of phase estimation is inherent to the frequency domain processing, that affects the quality of signal reconstruction. In this paper, we propose a time-domain speaker extraction network (TseNet) that doesn't decompose the speech signal into magnitude and phase spectrums, therefore, doesn't require phase estimation. The TseNet consists of a stack of dilated depthwise separable convolutional networks, that capture the long-range dependency of the speech signal with a manageable number of parameters. It is also conditioned on a reference voice from the target speaker, that is characterized by speaker i-vector, to perform the selective listening to the target speaker. Experiments show that the proposed TseNet achieves 16.3% and 7.0% relative improvements over the baseline in terms of signal-to-distortion ratio (SDR) and perceptual evaluation of speech quality (PESQ) under open evaluation condition. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
ASRU | 1 |
| 2019 | Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation LossabstractThe SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which doesn't calculate direct signal reconstruction error and consider the speech context. To address these problems, this paper proposes a magnitude and temporal spectrum approximation loss to estimate a phase sensitive mask for the target speaker with the speaker characteristics. Moreover, this paper explores a concatenation framework instead of the context adaptive deep neural network in the SBF method to encode a speaker embedding into the mask estimation network. Experimental results under open evaluation condition show that the proposed method achieves 70.4% and 17.7% relative improvement over the SBF baseline on signal-to-distortion ratio (SDR) and perceptual evaluation of speech quality (PESQ), respectively. A further analysis demonstrates 69.1% and 72.3% relative SDR improvements obtained by the proposed method for different and same gender mixtures. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 1 |
| 2019 | Target Speaker Extraction for Multi-Talker Speaker Verification
Wei Rao 0002, Chenglin Xu, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2018 | Single Channel Speech Separation with Constrained Utterance Level Permutation Invariant Training Using Grid LSTMabstractUtterance level permutation invariant training (uPIT) technique is a state-of-the-art deep learning architecture for speaker independent multi-talker separation. uPIT solves the label ambiguity problem by minimizing the mean square error (MSE) over all permutations between outputs and targets. However, uPIT may be sub-optimal at segmental level because the optimization is not calculated over the individual frames. In this paper, we propose a constrained uPIT (cuPIT) to solve this problem by computing a weighted MSE loss using dynamic information (i.e., delta and acceleration). The weighted loss ensures the temporal continuity of output frames with the same speaker. Inspired by the heuristics (i.e., vocal tract continuity) in computational auditory scene analysis, we then extend the model by adding a Grid LSTM layer, that we name it as cuPIT-Grid LSTM, to automatically learn both temporal and spectral patterns over the input magnitude spectrum simultaneously. The experimental results show 9.6% and 8.5% relative improvements on WSJ0-2mix dataset under both closed and open conditions comparing with the uPIT baseline. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 1 |
| 2018 | A Shifted Delta Coefficient Objective for Monaural Speech Separation Using Multi-task Learning
Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 47 |
| 2017 | Weighted Spatial Covariance Matrix Estimation for MUSIC Based TDOA Estimation of Speech Source
Chenglin Xu, Sining Sun, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 16 |
| 2014 | A deep neural network approach for sentence boundary detection in broadcast newsabstractThis paper presents a deep neural network (DNN) approach to sentence boundary detection in broadcast news. We extract prosodic and lexical features at each inter-word position in the transcripts and learn a sequential classifier to label these po-sitions as either boundary or non-boundary. This work is real-ized by a hybrid DNN-CRF (conditional random field) architec-ture. The DNN accepts prosodic feature inputs and non-linearly maps them into boundary/non-boundary posterior probability outputs. Subsequently, the posterior probabilities are combined with lexical features and the integrated features are modeled by a linear-chain CRF. The CRF finally labels the inter-word po-sitions as boundary or non-boundary by Viterbi decoding. Ex-periments show that, as compared with the state-of-the-art DT-CRF approach [1], the proposed DNN-CRF approach achieves 16.7 % and 4.1 % reduction in NIST boundary detection error in reference and speech recognition transcripts, respectively. Index Terms: sentence boundary detection, structural event de-tection, deep neural network, rich transcription 1. Chenglin Xu, Lei Xie 0001, Guangpu Huang, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 1 |