Thomas Fang Zheng

dblp:53/2843 · also Fang Zheng 0001 · DBLP profile ↗
← Back
102ranked-venue papers
13as first author
16since 2021 · last 2025
0000-0002-0249-4767ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 68 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-authorDatabases, data management, data science and information retrieval · 3Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
abstract
We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowdsourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing.
Turi Abu, Ying Shi 0001, Thomas Fang Zheng, Dong Wang 0013
ICASSP3
2025 ISL-MED: A General Iterative Self-Learning Framework for Speech Complex Emotion Detection
abstract
Speech complex emotion detection (SCED) aims to identify all emotion categories and their intensities in speech, which is crucial for understanding the speaker’s genuine intentions. A significant challenge inherent to SCED is the coarse nature of manual annotations, such as one-hot labels. These labels merely indicate the presence or absence of an emotion, thereby lacking the granularity required to guide models in capturing crucial intensity information. To overcome this limitation, we propose a novel framework: the Iterative Self-Learning-based Multiple Emotion Detector (ISL-MED). This framework utilizes an iterative self-learning approach to infer the latent complex emotion distribution directly from coarse one-hot labels. Specifically, ISL-MED employs multiple dedicated emotion detectors, each responsible for estimating the intensity component for a distinct emotion category. Notably, these detectors can be instantiated using various existing speech emotion recognition (SER) models, and their quantity can be flexibly configured based on task requirements. Furthermore, this paper proposes a data selection strategy based on Curriculum Learning and Human-Machine Consensus (HMCC). This strategy enhances model performance and accelerates convergence by systematically identifying and excluding highly ambiguous samples from the training set. We validated the effectiveness of ISL-MED on the complex emotions dataset CNSCED for speech complex emotion detection tasks and further evaluated its generalization capability on the IEMOCAP dataset for single emotion recognition tasks.
Xinxin Luo, Chang Feng, Hankiz Yilahun, Mingxing Xu, Askar Hamdulla, Thomas Fang Zheng
IJCNN7
2024 Enhancing Quantised End-to-End ASR Models Via Personalisation
abstract
Recent end-to-end automatic speech recognition (ASR) models have become increasingly larger, making them particularly challenging to be deployed on resource-constrained devices. Model quantisation is an effective solution that sometimes causes the word error rate (WER) to increase. In this paper, a novel strategy of personalisation for a quantised model (PQM) is proposed, which combines speaker adaptive training (SAT) with model quantisation to improve the performance of heavily compressed models. Specifically, PQM uses a 4-bit NormalFloat Quantisation (NF4) approach for model quantisation and low-rank adaptation (LoRA) for SAT. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size and 1% additional speaker-specific parameters, 15.1% and 23.3% relative WER reductions were achieved on quantised Whisper and Conformer-based attention-based encoder-decoder ASR models respectively, comparing to the original full precision models.
Qiuming Zhao, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng
ICASSP5
2024 Advancing Respiratory Sound Classification: Integration of Audio Spectrogram Transformer with ConnectMix and NEFTune Augmentation
Runze Huang, Mingxing Xu, Thomas Fang Zheng
ICONIP (10)3
2024 Emotional Atmosphere Soft Label for Emotion Recognition in Conversations
Chang Feng, Hankiz Yilahun, Mingxing Xu, Askar Hamdulla, Thomas Fang Zheng
ICONIP (3)6
2024 A Joint Noise Disentanglement and Adversarial Training Framework for Robust Speaker Verification
abstract
Automatic Speaker Verification (ASV) suffers from performance degradation in noisy conditions. To address this issue, we propose a novel adversarial learning framework that incorporates noise-disentanglement to establish a noise-independent speaker invariant embedding space. Specifically, the disentanglement module includes two encoders for separating speaker related and irrelevant information, respectively. The reconstruction module serves as a regularization term to constrain the noise. A feature-robust loss is also used to supervise the speaker encoder to learn noise-independent speaker embeddings without losing speaker information. In addition, adversarial training is introduced to discourage the speaker encoder from encoding acoustic condition information for achieving a speaker-invariant embedding space. Experiments on VoxCeleb1 indicate that the proposed method improves the performance of the speaker verification system under both clean and noisy conditions.
Xujiang Xing, Mingxing Xu, Thomas Fang Zheng
INTERSPEECH3
2024 SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
Qiuming Zhao, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng
INTERSPEECH5
2024 Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
Shuai Wang 0016, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng
INTERSPEECH7
2024 Hierarchical Multi-Path and Multi-Model Selection For Fake Speech Detection
abstract
The variety of spoofing algorithms used in generating speech poses obstacles to fake speech detection. Earlier methods have demonstrated complementary effects for detection. This paper proposes a novel hierarchical multi-path multi-model selection method for fake speech detection. It is designed to dynamically select and utilise the most suitable model from a set of complementary models. In our method, four basic detection models are incorporated, each offering partial but complementary detection abilities, to enhance balanced performance on diverse fake speech. The models are trained through a multi-path schema and the selection mechanism is structured hierarchically to improve the generalisation ability. Our method achieves an Equal Error Rate (EER) of 0.37% on the ASVspoof 2019 LA dataset, and outperforms other state-of-the-art method on the cross-domain and cross-dataset scenarios. A statistical analysis of EERs against thirteen unknown attacks reveals our method’s superiority, evidenced by the lowest standard deviation of 0.24, further underscoring our method’s robustness against a range of attacks.
Chang Feng, Guangzhi Sun, Shuai Wang 0016, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng
SLT8
2023 CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech Synthesis
abstract
Research on Video to Speech Synthesis (VTS) surges recently and the focus is gradually shifting from small-vocabulary short-phrase VTS to large-vocabulary continuous VTS (LVC-VTS). A large-scale dataset with sufficient speakers and utterances is a prerequisite for such research, and the database is certainly language dependent.In this paper, we introduce CN-CVS, a large-scale Mandarin continuous visual-speech dataset, to support LVC-VTS research. The dataset contains about 200k utterances from more than 2500 individuals, amounting to more than 300 hours of visual-speech data. We built a state-of-the-art VTS model with the new dataset and conducted preliminary studies. Our results show that models that achieve good performance on small vocabulary tasks may perform very poor on CN-CVS, indicating that continuous VTS is indeed a challenging task, and the main challenge comes from the unconstrained vocabulary. The dataset and baseline code can be downloaded for free from http://cncvs.cslt.org.
Chen Chen 0075, Dong Wang 0013, Thomas Fang Zheng
ICASSP3
2023 Random Cycle Loss and Its Application to Voice Conversion
abstract
Speech disentanglement aims to decompose independent causal factors of speech signals into separate codes. Perfect disentanglement benefits to a broad range of speech processing tasks. This paper presents a simple but effective disentanglement approach based on cycle consistency loss and random factor substitution. This leads to a novel random cycle (RC) loss that enforces analysis-and-resynthesis consistency, a main principle of reductionism. We theoretically demonstrate that the proposed RC loss can achieve independent codes if well optimized, which in turn leads to superior disentanglement when combined with information bottleneck (IB). Extensive simulation experiments were conducted to understand the properties of the RC loss, and experimental results on voice conversion further demonstrate the practical merit of the proposal. Source code and audio samples can be found on the webpage http://rc.cslt.org.
Dong Wang 0013, Lantian Li, Chen Chen 0075, Thomas Fang Zheng
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 CN-Celeb: Multi-genre speaker recognition
Lantian Li, Jiawen Kang 0002, Yunqi Cai, Ravichander Vipperla, Thomas Fang Zheng, Dong Wang 0013
Speech Commun.8
2021 Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification
abstract
Domain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly present an empirical study to show that simply adding cross-domain data does not help performance in conditions with enrollment-test mismatch. Careful analysis shows that this striking result is caused by the incoherent statistics between the enrollment and test conditions. Based on this analysis, we present a decoupled scoring approach that can maximally squeeze the value of cross-domain labels and obtain optimal verification scores in the enrollment-test mismatch condition. When the statistics are coherent, the new formulation falls back to the conventional PLDA. Experimental results on cross-channel test show that the proposed approach is highly effective and is a principal solution to domain mismatch.
Lantian Li, Yang Zhang 0052, Jiawen Kang 0002, Thomas Fang Zheng, Dong Wang 0013
ICASSP4
2021 Attack on Practical Speaker Verification System Using Universal Adversarial Perturbations
abstract
In authentication scenarios, applications of practical speaker verification systems usually require a person to read a dynamic authentication text. Previous studies played an audio adversarial example as a digital signal to perform physical attacks, which would be easily rejected by audio replay detection modules. This work shows that by playing our crafted adversarial perturbation as a separate source when the adversary is speaking, the practical speaker verification system will misjudge the adversary as a target speaker. A two-step algorithm is proposed to optimize the universal adversarial perturbation to be text-independent and has little effect on the authentication text recognition. We also estimated room impulse response (RIR) in the algorithm which allowed the perturbation to be effective after being played over the air. In the physical experiment, we achieved targeted attacks with success rate of 100%, while the word error rate (WER) on speech recognition was only increased by 3.55%. And recorded audios could pass replay detection for the live person speaking.
Shuning Zhao, Jianmin Li 0001, Xingliang Cheng, Thomas Fang Zheng, Xiaolin Hu 0001
ICASSP6
2021 Cross-Database Replay Detection in Terminal-Dependent Speaker Verification
Xingliang Cheng, Mingxing Xu, Thomas Fang Zheng
Interspeech3
2021 When Automatic Voice Disguise Meets Automatic Speaker Verification
abstract
The technique of transforming voices in order to hide the real identity of a speaker is called voice disguise, among which automatic voice disguise (AVD) by modifying the spectral and temporal characteristics of voices with miscellaneous algorithms are easily conducted with softwares accessible to the public. AVD has posed great threat to both human listening and automatic speaker verification (ASV). In this paper, we have found that ASV is not only a victim of AVD but could be a tool to beat some simple types of AVD. Firstly, three types of AVD, pitch scaling, vocal tract length normalization (VTLN) and voice conversion (VC), are introduced as representative methods. State-of-the-art ASV methods are subsequently utilized to objectively evaluate the impact of AVD on ASV by equal error rates (EER). Moreover, an approach to restore disguised voice to its original version is proposed by minimizing a function of ASV scores w.r.t. restoration parameters. Experiments are then conducted on disguised voices from Voxceleb, a dataset recorded in real-world noisy scenario. The results have shown that, for the voice disguise by pitch scaling, the proposed approach obtains an EER around 7% comparing to the 30% EER of a recently proposed baseline using the ratio of fundamental frequencies. The proposed approach generalizes well to restore the disguise with nonlinear frequency warping in VTLN by reducing its EER from 34.3% to 18.5%. However, it is difficult to restore the source speakers in VC by our approach, where more complex forms of restoration functions or other paralinguistic cues might be necessary to restore the nonlinear transform in VC. Finally, contrastive visualization on ASV features with and without restoration illustrate the role of the proposed approach in an intuitive way.
Linlin Zheng, Jiakang Li, Meng Sun 0001, Xiongwei Zhang, Thomas Fang Zheng
IEEE Trans. Inf. Forensics Secur.5
2020 ASR-Free Pronunciation Assessment
abstract
Most of the pronunciation assessment methods are based on local features derived from automatic speech recognition (ASR), e.g., the Goodness of Pronunciation (GOP) score. In this paper, we investigate an ASR-free scoring approach that is derived from the marginal distribution of raw speech signals. The hypothesis is that even if we have no knowledge of the language (so cannot recognize the phones/words), we can still tell how good a pronunciation is, by comparatively listening to some speech data from the target language. Our analysis shows that this new scoring approach provides an interesting correction for the phone-competition problem of GOP. Experimental results on the ERJ dataset demonstrated that combining the ASR-free score and GOP can achieve better performance than the GOP baseline.
Sitong Cheng, Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH6
2020 Domain-Invariant Speaker Vector Projection by Model-Agnostic Meta-Learning
abstract
Domain generalization remains a critical problem for speaker recognition, even with the state-of-the-art architectures based on deep neural nets. For example, a model trained on reading speech may largely fail when applied to scenarios of singing or movie. In this paper, we propose a domain-invariant projection to improve the generalizability of speaker vectors. This projection is a simple neural net and is trained following the Model-Agnostic Meta-Learning (MAML) principle, for which the objective is to classify speakers in one domain if it had been updated with speech data in another domain. We tested the proposed method on CNCeleb, a new dataset consisting of single-speaker multi-condition (SSMC) data. The results demonstrated that the MAML-based domain-invariant projection can produce more generalizable speaker vectors, and effectively improve the performance in unseen domains.
Jiawen Kang 0002, Lantian Li, Yunqi Cai, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH6
2020 Neural Discriminant Analysis for Deep Speaker Embedding
abstract
Probabilistic Linear Discriminant Analysis (PLDA) is a popular tool in open-set classification/verification tasks.However, the Gaussian assumption underlying PLDA prevents it from being applied to situations where the data is clearly non-Gaussian.In this paper, we present a novel nonlinear version of PLDA named as Neural Discriminant Analysis (NDA).This model employs an invertible deep neural network to transform a complex distribution to a simple Gaussian, so that the linear Gaussian model can be readily established in the transformed space.We tested this NDA model on a speaker recognition task where the deep speaker vectors (x-vectors) are presumably non-Gaussian.Experimental results on two datasets demonstrate that NDA consistently outperforms PLDA, by handling the non-Gaussian distributions of the x-vectors.
Lantian Li, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH3
2018 Full-Info Training for Deep Speaker Feature Learning
abstract
In recent studies, it has shown that speaker patterns can be learned from very short speech segments (e.g., 0.3 seconds) by a carefully designed convolutional & time-delay deep neural network (CT-DNN) model. By enforcing the model to discriminate the speakers in the training data, frame-level speaker features can be derived from the last hidden layer. In spite of its good performance, a potential problem of the present model is that it involves a parametric classifier, i.e., the last affine layer, which may consume some discriminative knowledge, thus leading to `information leak' for the feature learning. This paper presents a full-info training approach that discards the parametric classifier and enforces all the discriminative knowledge learned by the feature net. Our experiments on the Fisher database demonstrate that this new training scheme can produce more coherent features, leading to consistent and notable performance improvement on the speaker verification task.
Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng
ICASSP4
2018 Deep Factorization for Speech Signal
abstract
Various informative factors mixed in speech signals, leading to great difficulty when decoding any of the factors. An intuitive idea is to factorize each speech frame into individual informative factors, though it turns out to be highly difficult. Recently, we found that speaker traits, which were assumed to be long-term distributional properties, are actually short-time patterns, and can be learned by a carefully designed deep neural network (DNN). This discovery motivated a cascade deep factorization (CDF) framework that will be presented in this paper. The proposed framework infers speech factors in a sequential way, where factors previously inferred are used as conditional variables when inferring other factors. We will show that this approach can effectively factorize speech signals, and using these factors, the original speech spectrum can be recovered with a high accuracy. This factorization and reconstruction approach provides potential values for many speech processing tasks, e.g., speaker recognition and emotion recognition, as will be demonstrated in the paper.
Lantian Li, Dong Wang 0013, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Thomas Fang Zheng
ICASSP6
2018 Imbalance Learning-based Framework for Fear Recognition in the MediaEval Emotional Impact of Movies Task
Xingliang Cheng, Mingxing Xu, Thomas Fang Zheng
INTERSPEECH4
2017 Speaker segmentation using deep speaker vectors for fast speaker change scenarios
abstract
A novel speaker segmentation approach based on deep neural network is proposed and investigated. This approach uses deep speaker vectors (d-vectors) to represent speaker characteristics and to find speaker change points. The d-vector is a kind of frame-level speaker discriminative feature, whose discriminative training process corresponds to the goal of discriminating a speaker change point from a single speaker speech segment in a short time window. Following the traditional metric-based segmentation, each analysis window contains two sub-windows and is shifting along the audio stream to detect speaker change points, where the speaker characteristics are represented by the means of deep speaker vectors for all frames in each window. Experimental investigations conducted in fast speaker change scenarios show that the proposed method can detect speaker change points more quickly and more effectively than the commonly used segmentation methods.
Renyu Wang, Mingliang Gu, Lantian Li, Mingxing Xu, Thomas Fang Zheng
ICASSP5
2017 A Study on Replay Attack and Anti-Spoofing for Automatic Speaker Verification
abstract
For practical automatic speaker verification (ASV) systems, replay attack poses a true risk.By replaying a pre-recorded speech signal of the genuine speaker, ASV systems tend to be easily fooled.An effective replay detection method is therefore highly desirable.In this study, we investigate a major difficulty in replay detection: the over-fitting problem caused by variability factors in speech signal.An F-ratio probing tool is proposed and three variability factors are investigated using this tool: speaker identity, speech content and playback & recording device.The analysis shows that device is the most influential factor that contributes the highest over-fitting risk.A frequency warping approach is studied to alleviate the over-fitting problem, as verified on the ASV-spoof 2017 database.
Lantian Li, Yixiang Chen 0003, Dong Wang 0013, Thomas Fang Zheng
INTERSPEECH4
2017 Distributed representation learning for knowledge graphs with entity descriptions
Thomas Fang Zheng, Ralph Grishman
Pattern Recognit. Lett.3
2016 Learning Embedding Representations for Knowledge Inference on Imperfect and Incomplete Repositories
abstract
This paper considers the problem of knowledge inference on large-scale imperfect repositories with incomplete coverage by means of embedding entities and relations at the first attempt. We propose IIKE (Imperfect and Incomplete Knowledge Embedding), a probabilistic model which measures the probability of each belief, i.e., in large-scale knowledge bases such as NELL and Freebase, and our objective is to learn a better low-dimensional vector representation for each entity (h and t) and relation (r) in the process of minimizing the loss of fitting the corresponding confidence given by machine learning (NELL) or crowdsouring (Freebase), so that we can use ||h+r-t|| to assess the plausibility of a belief when conducting inference. We use subsets of those inexact knowledge bases to train our model and test the performances of link prediction and triplet classification on ground truth beliefs, respectively. The results of extensive experiments show that IIKE achieves significant improvement compared with the baseline and state-of-the-art approaches.
Thomas Fang Zheng
WI3
2016 Improving speaker verification performance against long-term speaker variability
Jun Wang 0073, Lantian Li, Thomas Fang Zheng, Frank K. Soong
Speech Commun.4
2016 Improving Short Utterance Speaker Recognition by Modeling Speech Unit Classes
abstract
Short utterance speaker recognition (SUSR) is highly challenging due to the limited enrollment and/or test data. We argue that the difficulty can be largely attributed to the mismatched prior distributions of the speech data used to train the universal background model (UBM) and those for enrollment and test. This paper presents a novel solution that distributes speech signals into a multitude of acoustic subregions that are defined by speech units, and models speakers within the subregions. To avoid data sparsity, a data-driven approach is proposed to cluster speech units into speech unit classes, based on which robust subregion models can be constructed. Further more, we propose a model synthesis approach based on maximum likelihood linear regression (MLLR) to deal with no-data speech unit classes. The experiments were conducted on a publicly available database SUD12. The results demonstrated that on a text-independent speaker recognition task where the test utterances are no longer than 2 seconds and mostly shorter than 0.5 seconds, the proposed subregion modeling offered a 21.51% relative reduction in equal error rate (EER), compared with the standard GMM-UBM baseline. In addition, with the model synthesis approach, the performance can be greatly improved in scenarios where no enrollment data are available for some speech unit classes.
Lantian Li, Dong Wang 0013, Thomas Fang Zheng
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Unseen Noise Estimation Using Separable Deep Auto Encoder for Speech Enhancement
abstract
Unseen noise estimation is a key yet challenging step to make a speech enhancement algorithm work in adverse environments. At worst, the only prior knowledge we know about the encountered noise is that it is different from the involved speech. Therefore, by subtracting the components which cannot be adequately represented by a well defined speech model, the noises can be estimated and removed. Given the good performance of deep learning in signal representation, a deep auto encoder (DAE) is employed in this work for accurately modeling the clean speech spectrum. In the subsequent stage of speech enhancement, an extra DAE is introduced to represent the residual part obtained by subtracting the estimated clean speech spectrum (by using the pre-trained DAE) from the noisy speech spectrum. By adjusting the estimated clean speech spectrum and the unknown parameters of the noise DAE, one can reach a stationary point to minimize the total reconstruction error of the noisy speech spectrum. The enhanced speech signal is thus obtained by transforming the estimated clean speech spectrum back into time domain. The above proposed technique is called separable deep auto encoder (SDAE). Given the under-determined nature of the above optimization problem, the clean speech reconstruction is confined in the convex hull spanned by a pre-trained speech dictionary. New learning algorithms are investigated to respect the non-negativity of the parameters in the SDAE. Experimental results on TIMIT with 20 noise types at various noise levels demonstrate the superiority of the proposed method over the conventional baselines.
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme, Thomas Fang Zheng
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Distant Supervision for Entity Linking
Thomas Fang Zheng
PACLIC3
2015 Document Representation with Statistical Word Senses in Cross-Lingual Document Clustering
abstract
Cross-lingual document clustering is the task of automatically organizing a large collection of multi-lingual documents into a few clusters, depending on their content or topic. It is well known that language barrier and translation ambiguity are two challenging issues for cross-lingual document representation. To this end, we propose to represent cross-lingual documents through statistical word senses, which are automatically discovered from a parallel corpus through a novel cross-lingual word sense induction model and a sense clustering method. In particular, the former consists in a sense-based vector space model and the latter leverages on a sense-based latent Dirichlet allocation. Evaluation on the benchmarking datasets shows that the proposed models outperform two state-of-the-art methods for cross-lingual document clustering.
Guoyu Tang, Yunqing Xia, Erik Cambria, Thomas Fang Zheng
Int. J. Pattern Recognit. Artif. Intell.5
2015 Statistical word sense aware topic models
Guoyu Tang, Yunqing Xia, Jun Sun 0024, Min Zhang 0005, Thomas Fang Zheng
Soft Comput.5
2015 Detection and reconstruction of clipped speech for speaker recognition
Fanhu Bie, Dong Wang 0013, Jun Wang 0073, Thomas Fang Zheng
Speech Commun.4
2014 Distant Supervision for Relation Extraction with Matrix Completion
abstract
The essence of distantly supervised relation extraction is that it is an incomplete multi-label classification problem with sparse and noisy features.To tackle the sparsity and noise challenges, we propose solving the classification problem using matrix completion on factorized matrix of minimized rank.We formulate relation classification as completing the unknown labels of testing items (entity pairs) in a sparse matrix that concatenates training and testing textual features with training labels.Our algorithmic framework is based on the assumption that the rank of item-byfeature and item-by-label joint matrix is low.We apply two optimization models to recover the underlying low-rank matrix leveraging the sparsity of feature-label matrix.The matrix completion problem is then solved by the fixed point continuation (FPC) algorithm, which can find the global optimum.Experiments on two widely used datasets with different dimensions of textual features demonstrate that our low-rank matrix completion approach significantly outperforms the baseline and the state-of-the-art methods.
Deli Zhao, Zhiyuan Liu 0001, Thomas Fang Zheng, Edward Y. Chang
ACL (1)5
2014 Mining the Personal Interests of Microbloggers via Exploiting Wikipedia Knowledge
Thomas Fang Zheng
CICLing (2)3
2014 Topic Models Incorporating Statistical Word Senses
Guoyu Tang, Yunqing Xia, Jun Sun 0024, Min Zhang 0005, Thomas Fang Zheng
CICLing (1)5
2014 Using Word Sense as a Latent Variable in LDA Can Improve Topic Modeling
abstract
Since proposed, LDA have been successfully used in modeling text documents. So far, words are the common features to induce latent topic, which are later used in document representation. Observation on documents indicates that the polysemous words can make the latent topics less discriminative, resulting in less accurate document representation. We thus argue that the semantically deterministic word senses can improve quality of the latent topics. In this work, we proposes a series of word sense aware LDA models which use word sense as an extra latent variable in topic induction. Preliminary experiments on benchmark datasets show that word sense can indeed improve topic modeling.
Yunqing Xia, Guoyu Tang, Erik Cambria, Thomas Fang Zheng
ICAART (1)5
2014 Clustering tweets usingWikipedia concepts
Guoyu Tang, Yunqing Xia, Weizhi Wang, Raymond Y. K. Lau, Thomas Fang Zheng
LREC5
2014 Transition-based Knowledge Graph Embedding with Relational Mapping Properties
Emily Chang, Thomas Fang Zheng
PACLIC4
2013 Sequential model adaptation for speaker verification
abstract
GMM-UBM-based speaker verification heavily relies on well-trained UBMs.In practice, it is not often easy to obtain a UBM that fully matches the acoustic channel in operation.In a previous study, we proposed to address this problem by a novel sequential UBM adaptation approach based on MAP.This work extends the study by applying the sequential approach to speaker model adaptation.In addition, we investigate a new feature-space sequential adaptation approach based on feature MAP linear regression (fMAPLR) and compare it with the previously proposed model-space MAP approach.We find that these two approaches are complementary and can be combined to deliver additional performance gains.The experiments conducted on a time-varying speech database demonstrate that the proposed MAP-fMAPLR approach leads to significant EER reduction with two mismatched UBMs (25% and 39% respectively).
Jun Wang 0073, Dong Wang 0013, Thomas Fang Zheng, Javier Tejedor
INTERSPEECH4
2013 Ranking Search Intents Underlying a Query
Yunqing Xia, Xiaoshi Zhong, Guoyu Tang, Thomas Fang Zheng, Qinan Hu, Sen Na, Yaohai Huang
NLDB6
2012 Content-Based Semantic Tag Ranking for Recommendation
abstract
Content-based social tagging recommendation, which considers the relationship between the tags and the descriptions contained in resources, is proposed to remedy the cold-start problem of collaborative filtering. There is such a common phenomenon that certain tag does not appear in the corresponding description, however, they do semantically relate with each other. State-of-the-art methods seldom consider this phenomenon and thus still need to be improved. In this paper, we propose a novel content-based social tag ranking scheme, aiming to recommend the semantic tags that the descriptions may not contain. The scheme firstly acquires the quantized semantic relationships between words with empirical methods, then constructs the weighted tag-digraph based on the descriptions and acquired quantized semantics, and finally performs a modified graph-based ranking algorithm to refine the score of each candidate tag for recommendation. Experimental results on both English and Chinese datasets show that the proposed scheme performs better than several state-of-the-art content-based methods.
Thomas Fang Zheng
Web Intelligence3
2011 Reliable accent specific unit generation with dynamic Gaussian mixture selection for multi-accent speech recognition
abstract
Multiple accents are often present in Mandarin speech, as most Chinese have learned Mandarin as a second language. We propose generating reliable accent specific unit together with dynamic Gaussian mixture selection for multi-accent speech recognition. Time alignment phoneme recognition is used to generate such unit and to model accent variations explicitly and accurately. Dynamic Gaussian mixture selection scheme builds a dynamical observation density for each specified frame in decoding, and leads to use Gaussian mixture component efficiently. This method increases the covering ability for a diversity of accent variations in multi-accent, and alleviates the performance degradation caused by pruned beam search without augmenting the model size. The effectiveness of this approach is evaluated on three typical Chinese accents Chuan, Yue and Wu. Our approach outperforms traditional acoustic model reconstruction approach significantly by 6.30%, 4.93% and 5.53%, respectively on Syllable Error Rate (SER) reduction, without degrading on standard speech.
Chao Zhang 0031, Yi Liu 0050, Yunqing Xia, Thomas Fang Zheng, Jesper Ø. Olsen, Jilei Tian
ICME4
2011 CLGVSM: Adapting Generalized Vector Space Model to Cross-lingual Document Clustering
Guoyu Tang, Yunqing Xia, Min Zhang 0005, Haizhou Li 0001, Thomas Fang Zheng
IJCNLP5
2010 Using phoneme recognition and text-dependent speaker verification to improve speaker segmentation for Chinese speech
abstract
Speaker segmentation is widely used in many tasks such as multi-speaker detection and speaker tracking. The segmentation performance depends on the performance of speaker verification (SV) between two short utterances to a large extent, so the improvement of the SV performance for short utterances would give the segmentation performance a great help. In this paper, a method based on phoneme recognition and text-dependent speaker recognition is proposed. During segmentation, a phoneme sequence is first recognized using a phoneme recognizer and then text-dependent speaker recognition based on dynamic time warping (DTW) is performed on the same phoneme in two adjacent windows. Experiments over Chinese Corpus Consortium (CCC) MSS database showed that better performance was achieved compared with the BIC method and the GLR method. Index Terms: speaker segmentation, phoneme recognition, text-dependent, short utterances
Thomas Fang Zheng
INTERSPEECH3
2009 A phrase-level piecewise linear scaling algorithm for melody match in Query-by-Humming systems
abstract
The Query-by-Humming (QBH) system allows users to retrieve songs by singing/humming. In this paper we propose a phraselevel piecewise linear scaling algorithm for melody match. Musical phrase boundaries are predicted for the query to split it to phrases. The boundaries of melody fragment corresponding to each phrase are allowed for adjusting in a limited scope. The algorithm employs Dynamic Programming and Recursive Alignment to search for the minimal piecewise matching cost upon Linear Scaling at phrase-level. Our experimental results on 5223 melody database show that the proposed algorithm outperforms traditional algorithms. The proposed algorithm gives significant improvements of 17.0%, 14.7% and 4.8% with respect to Linear Scaling, Dynamic Time Wrapping and Recursive Alignment in top-1 rate, respectively. The results show that the proposed algorithm is more efficient than the previous algorithms.
Wenxiao Cao, Danning Jiang, Yong Qin 0001, Thomas Fang Zheng, Yi Liu 0050
ICME5
2009 Effectiveness of n-gram fast match for query-by-humming systems
abstract
To achieve a good balance between matching accuracy and computation efficiency is a key challenge for query-by-humming (QBH) system. In this paper, we propose an approach of n-gram based fast match. Our n-gram method uses a robust statistical note transcription as well as error compensation method based on the analysis of frequent transcription errors. The effectiveness of our approach has been evaluated on a relatively large melody database with 5223 melodies. The experimental results show that when the searching space was reduced to only 10% of the whole size, 90% of the target melodies were preserved in the candidates, and 88% of the match accuracy of system was kept. Meanwhile, no obvious additional computation was applied.
Danning Jiang, Wenxiao Cao, Yong Qin 0001, Thomas Fang Zheng, Yi Liu 0050
ICME5
2008 State-dependent phoneme-based model merging for dialectal Chinese speech recognition
Linquan Liu, Thomas Fang Zheng, Wenhu Wu
Speech Commun.2
2007 State-dependent mixture tying with variable codebook size for accented speech recognition
abstract
In this paper, we propose a state-dependent tied mixture (SDTM) models with variable codebook size to improve the model robustness for accented phonetic variations while maintaining model discriminative ability. State tying and mixture tying are combined to generate SDTM models. Compared to a pure mixture tying system, the SDTM model uses state tying to reserve the state identity; compared to the sole state tying system, such model uses a small set of parameters to discard the overlapping mixture distributions for robust model estimation. The codebook size of SDTM model is varied according to the confusion probability of states. The more confusable a state is, the larger its codebook size gets for a higher degree of model resolution. The codebook size is governed by state level variation probability of accented phonetic confusions which can be automatically extracted by frame-to-state alignment based on the local model mismatch. The effectiveness of this approach is evaluated on Mandarin accented speech. Our method yields a significant 2.1%, 9.5% and 3.5% absolute word error rate reduction compared with state tying, mixture tying and state-based phonetic tied mixtures, respectively.
Thomas Fang Zheng, Yunqing Xia
ASRU2
2007 Session Variability Subspace Projection Based Model Compensation for Speaker Verification
abstract
In this paper, a session variability subspace projection (SVSP) based model compensation method for speaker verification is proposed. During the training phase the session variability is removed from speaker models by projection, while during the testing phase the session variability in a test utterance is used to compensate speaker models. Finally, the compensated speaker models and UBM are used to recognize the identity of the test utterance. Compared with the conventional GMM-UBM system, the relative equal error rate reduction of SVSP is 16.2% on the NIST 2006 single-side one conversation training, single-side one conversation test.
Thomas Fang Zheng, Wenhu Wu
ICASSP (4)2
2007 Emotion attribute projection for speaker recognition on emotional speech
abstract
Emotion is one of the important factors that cause the system performance degradation. By analyzing the similarity between channel effect and emotion effect on speaker recognition, an emotion compensation method called emotion attribute projection (EAP) is proposed to alleviate the intraspeaker emotion variability. The use of this method has achieved an equal error rate (EER) reduction of 11.7% with the EER reduced from 9.81% to 8.66%. When a linear fusion based on a GMM-UBM system with an EER of 9.38% and an SVM-EAP system with an EER of 8.66% is adopted, another EER reduction of 22.5% and 16.1% can be further achieved, respectively, and the final EER can be 7.27%. Index Terms: speaker recognition, emotional speech, emotion attribute projection, fusion
Huanjun Bao, Ming-Xing Xu, Thomas Fang Zheng
INTERSPEECH3
2007 Using a small development set to build a robust dialectal Chinese speech recognizer
abstract
To make full use of a small development data set to build a robust dialectal Chinese speech recognizer from a standard Chinese speech recognizer (based on Chinese Initial/Final, IF), a novel, simple but effective acoustic modeling method, named state-dependent phoneme-based model merging (SDPBMM), is proposed and evaluated, where a shared-state of standard tri-IF is merged with a state of dialectal mono-IF in terms of pronunciation variation modeling. Specifically, in order to deal with phonetic-level pronunciation variations in SDPBMM, distance-based pronunciation modeling is proposed based on a small dialectal Chinese data set. With a 40-minute Shanghai-dialectal Chinese data set, SDPBMM can achieve a significant syllable error rate (SER) reduction of 14.3 % for dialectal Chinese with almost no performance degradation for standard Chinese. Experimentally, SDPBMM can also outperform the maximum likelihood linear regression (MLLR) adaptation and the pooled retraining methods with relative SER reductions by 2.8 % and 10.6%, respectively. If SDPBMM is combined with the MLLR adaptation, another relative SER reduction of 3.3 % can be further achieved. Index Terms: dialectal Chinese, speech recognition, accented speech, pronunciation modeling, acoustic modeling
Linquan Liu, Thomas Fang Zheng, Makoto Akabane, Ruxin Chen, Wenhu Wu
INTERSPEECH2
2007 A Cohort-Based Speaker Model Synthesis for Mismatched Channels in Speaker Verification
abstract
Mismatch between enrollment and test data is one of the top performance degrading factors in speaker recognition applications. This mismatch is particularly true over public telephone networks, where input speech data is collected over different handsets and transmitted over different channels from one trial to the next. In this paper, a cohort-based speaker model synthesis (SMS) algorithm, designed for synthesizing robust speaker models without requiring channel-specific enrollment data, is proposed. This algorithm utilizes a priori knowledge of channels extracted from speaker-specific cohort sets to synthesize such speaker models. The cohort selection in the proposed new SMS can be either speaker-specific or Gaussian component based. Results on the China Criminal Police College (CCPC) speaker recognition corpus, which contains utterances from both landline and mobile channel, show the new algorithms yield significant speaker verification performance improvement over Htnorm and universal background model (UBM)-based speaker model synthesis.
Thomas Fang Zheng, Mingxing Xu, Frank K. Soong
IEEE Trans. Speech Audio Process.2
2006 Cohort-Based Speaker Model Synthesis for Channel Robust Speaker Recognition
abstract
Speaker recognition over a public telephone network involves various types of transmission channels and handsets, which leads to mismatched channels (between the enrolled models and the test utterances), and hence to a significant decline in the speaker recognition performance. In this paper a cohort-based speaker model synthesis algorithm, which aims at synthesizing speaker models for channels where no enrollment data is available is proposed. This algorithm applies a priori knowledge of channels extracted from speaker-specific cohort sets to synthesize speaker models. Results for the China Criminal Police College (CCPC) speaker recognition corpus, which contains utterances from both a landline and a mobile channel, show significant improvements over the HT-norm and UBM-based speaker model synthesis algorithms
Thomas Fang Zheng, Mingxing Xu
ICASSP (1)2
2006 Automatic initial/final generation for dialectal Chinese speech recognition
abstract
Phonetic differences always exist between any Chinese dialect and standard Chinese (Putonghua). In this paper, a method, named automatic dialect-specific Initial/Final (IF) generation, is proposed to deal with the issue of phonemic difference which can automatically produce the dialect-specific units based on model distance measure. A dialect-specific decision tree regrowing method is also proposed to cope with the tri-IF expansion due to the introduction of dialect-specific IFs (DIFs). In combination with a certain adaptation technique, the proposed methods can achieve a syllable error rate (SER) reduction of 18.5% for Shanghai-accented Chinese compared with the Putonghua-based baseline while the use of the DIF set only can lead to an SER reduction of 5.5%. Index Terms: speech recognition, dialectal Chinese speech recognition, phone set generation, acoustic distance measure
Linquan Liu, Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2006 Study on speaker verification on emotional speech
abstract
Besides background noise, channel effect and speaker’s health condition, emotion is another factor which may influence the performance of a speaker verification system. In this paper, the performance of a GMM-UBM based speaker verification system on emotional speech is studied. It is found that speech with various emotions aggravates the verification performance. Two reasons for the performance aggravation are analyzed, they are mismatched emotions between the speaker models and the test utterances, and the articulating styles of certain emotions which create intense intra-speaker vocal variability. In response to the first reason, an emotion-dependent score normalization method is proposed, which is borrowed from the idea of Hnorm. Index Terms: speaker verification, emotional speech
Thomas Fang Zheng, Ming-Xing Xu, Huanjun Bao
INTERSPEECH2
2006 A tree-based kernel selection approach to efficient Gaussian mixture model-universal background model based speaker identification
Thomas Fang Zheng, Zhanjiang Song, Frank K. Soong, Wenhu Wu
Speech Commun.2
2005 Combining Selection Tree with Observation Reordering Pruning for Efficient Speaker Identification Using GMM-UBM
abstract
In this paper a new method of reducing the computational load for Gaussian mixture model universal background model (GMM-UBM) based speaker identification is proposed. In order to speed up the selection of N-best Gaussian mixtures in a UBM, a selection tree (ST) structure as well as relevant operations is proposed. Combined with the existing observation reordering pruning (ORP) method which was proposed for rapid pruning of unlikely speaker model candidates, the proposed method achieves a much larger computation reduction factor than any single individual method. Experimental results show that a GMM-UBM system used in a conjunction with ST and ORP can speed up the computation by a factor of about 16 with an error rate increase of only about 1% compared with a baseline GMM-UBM system.
Thomas Fang Zheng, Zhanjiang Song, Wenhu Wu
ICASSP (1)2
2005 The predictive differential amplitude spectrum for robust speaker recognition in stationary noises
abstract
The performance of any speaker identification system degrades quite seriously when the acoustic conditions for testing mismatch those for training. In this paper, we propose a method to restore clean speech from noisy speech with two steps: 1) a predictive difference function is employed to estimate the differential amplitude spectrums (DAS) from both the left-side and right-side of the amplitude spectrum of the noisy speech, so as to eliminate the noise as precisely as possible, and 2) an average of the left-side and right-side integral DASs is taken as the estimated amplitude spectrum of the original clean speech. The spectrum in the traditional MFCC calculation is then replaced with this estimated amplitude and the extracted features based on this are referred to as predictive differential amplitude spectrum (PDAS) based cepstral coefficients (PDASCCs). We compare PDASCCs with cepstral mean subtraction (CMS) based, spectral subtraction (SS) based, and differential power spectrum (DPS) based cepstral coefficients at different noise levels. Experimental results show that the PDASCCs are more effective in enhancing the robustness of a speaker recognition system, and used with the CMS method the average error rate can be reduced by 7.5%.
Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2005 Modeling high-level information by using Gaussian mixture correlation for GMM-UBM based speaker recognition
Thomas Fang Zheng, Zhanjiang Song
INTERSPEECH2
2005 Real-time pitch tracking based on combined SMDSF
abstract
This paper presents a novel pitch tracking method in the time domain. Based on the difference function as used in YIN referred to as the sum magnitude difference square function (SMDSF) thereinafter -- we propose two modified types of SMDSFs, with several methods presented to calculate these SMDSFs efficiently and without bias by using the FFT algorithm. In pitch estimation, every type of SMDSF has its own estimation error characteristics. By analyzing these characteristics, we define a new function which combines the foresaid two types of SMDSFs to prevent estimation errors. A new, relatively accurate, and real-time pitch tracking algorithm is then proposed which does not need any extra preprocessing and post-processing. Experimental results show that this proposed algorithm can achieve remarkably good performance for pitch tracking.
Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2005 Rapidly developing spoken Chinese dialogue systems with the d-ear SDS SDK
abstract
Developing a spoken dialog system is typically timeconsuming, and must often be accomplished using difficult-tolearn professional technologies. Most existing toolkits use statistical semantic parsers and model a dialogue interaction as a finite-state network. However, for developing flexible spoken Chinese dialogue systems, these toolkits have several problems. A new toolkit named the d-Ear SDS SDK is introduced here. The SDK suggests a multi-session dialogue system framework with a powerful semantic parser specially designed for spoken Chinese understanding, and a powerful dialogue manager providing non-finite-state dialogue control. To set up a new dialogue system, the developer can customize all the system modules with domain-specific information and operations, using the d-Ear SDS Studio to save time. Using the SDK, we have built several dialogue systems with excellent performance in a very short time.
Thomas Fang Zheng, Michael Brasser, Zhanjiang Song
INTERSPEECH2
2004 Weighting observation vectors for robust speech recognition in noisy environments
abstract
In this paper, we propose a novel approach to robust speech recognition in noisy environments by discriminating the observation vectors. In conventional HMM-based speech recognition, all the observation vectors are treated with equal importance no matter how the corresponding speech segment is corrupted with noise. Our approach proposed here modifies the conventional decoder by weighting the likelihood scores for different observation vectors based on the signal to noise ratios (SNRs) of the corresponding speech frames when the probabilities of generating a sequence of observations are being calculated for some models. The proposed approach combined with spectral subtraction is evaluated with four different kinds of noises added to the clean speech. The experimental results show the superior performance of the proposed method over the method where only the spectral subtraction is applied, especially in the median SNR environments. 1.
Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2003 Using word confidence measure for OOV words detection in a spontaneous spoken dialog system
abstract
Developing a real-life spoken dialogue system must face with many practical issues, where the out-of-vocabulary (OOV) words problem is one of the key difficulties. This paper presents the OOV detection mechanism based on the word confidence scoring developed for the d-Ear Attendant system, a spontaneous spoken dialogue system. In the d-Ear Attendant system, an explicit filler model is originally used to detect the presence of OOV words [1]. Although this approach has a satisfactory OOV detection rate, it badly degrades the accuracy of in-vocabulary (IV) detection by 4.4 % absolutely (from 97 % to 92.6%). Such the degradation will not be acceptable in a practical system. By using a few commonly used acoustic confidence features and some new context confidence features, our confidence measure method not only is able to detect the word level speech recognition errors, but also has a good ability for OOV words detection with an acceptable false alarm rate. For example, with a false rejection rate of 2.5%, the false acceptance rate of 26 % is achieved. 1.
Thomas Fang Zheng, Mingxing Xu
INTERSPEECH3
2003 A Method to Build a Super Small but Practically Accurate Language Model for Handheld Devices
Genqing Wu, Thomas Fang Zheng
J. Comput. Sci. Technol.2
2002 Improved katz smoothing for language modeling in speech recogniton
abstract
In this paper, a new method is proposed to improve the canonical Katz back-off smoothing technique in language modeling. The process of Katz smoothing is detailedly analyzed and the global discounting parameters are selected for discounting. Further more, a modified version of the formula for discounting parameters is proposed, in which the discounting parameters are determined by not only the occurring counts of the n-gram units but also the low-order history frequencies. This modification makes the smoothing more reasonable for those n-gram units that have homophonic (same in pronunciation) histories. The new method is tested on a Chinese Pinyin-to-character (where Pinyin is the pronunciation string) conversion system and the results show that the improved method can achieve a surprising reduction both in perplexity and Chinese character error rate. 1.
Genqing Wu, Thomas Fang Zheng, Wenhu Wu, Mingxing Xu
INTERSPEECH2
2002 Reducing pronunciation lexicon confusion and using more data without phonetic transcription for pronunciation modeling
abstract
The multiple-pronunciation lexicon (MPL) is very important to model the pronunciation variations for spontaneous speech recognition. But the introduction of MPL brings out two problems. First, the MPL will increase the among-lexicon confusion and degrade the recognizer's performance. Second, the MPL needs more data with phonetic transcription so as to cover as many surface forms as possible. Accordingly, two solutions are proposed, they are the context-dependent weighting method and the iterative forced-alignment based transcription method. The use of them can compensate what the MPL causes and improve the overall performance. Experiments across a naturally spontaneous speech database show that the proposed methods are effective and better than other methods.
Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne
INTERSPEECH1
2002 Speech Detection in Non-Stationary Noise Based on the 1/f Process
Thomas Fang Zheng, Wenhu Wu
J. Comput. Sci. Technol.2
2002 Mandarin Pronunciation Modeling Based on CASS Corpus
Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne
J. Comput. Sci. Technol.1
2001 Automatic generation of pronunciation lexicons for Mandarin spontaneous speech
abstract
Pronunciation modeling for large vocabulary speech recognition attempts to improve recognition accuracy by identifying and modeling pronunciations that are not in the ASR systems pronunciation lexicon. Pronunciation variability in spontaneous Mandarin is studied using the newly created CASS corpus of phonetically annotated spontaneous speech. Pronunciation modeling techniques developed for English,are applied to this corpus to train pronunciation models which are then used for Mandarin broadcast news transcription.
William J. Byrne, Veera Venkataramani, Terri Kamm, Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, Yi Liu 0022, Umar Ruhi
ICASSP4
2001 Topic Forest: a plan-based dialog management structure
abstract
There are many task-oriented dialog systems, but few of them can cope with the issues such as the multiple-topic issue, the topic changing issue, information sharing among different topics, and the difference in importance for different information items. To provide efficient solutions, a plan-based dialog management structure named Topic Forest is proposed, which makes the mixed-initiative dialog control easier. The Topic Forest based reasoning engine with a certain strategy for both remembering and forgetting is also described. The reasoning strategy is designed to be domain-independent; therefore it makes the dialog management model easy to be ported to other different domains.
Thomas Fang Zheng, Mingxing Xu
ICASSP2
2001 A theme structure method for the ellipsis resolution
abstract
The purpose of this paper is to solve the contextual ellipsis problem that is popular in our Chinese spoken dialogue system named EasyNav. A Theme Structure is proposed to describe the attentional state. Its dynamic generation feature makes it suitable to model the topic transition in user-initiative dialogues. By studying the differences and the similarities between the ellipsis and the anaphora phenomena, we extend the resolution procedure and the theory from anaphora to ellipsis. The ellipsis resolution is now based on the semantic knowledge and the discourse factor other than the syntactic information. A Theme Structure Method proposed in this paper for the ellipsis resolution is uniform to not only all kinds of elliptical elements but also some particular ellipsis types such as the fragmental ellipsis and the default ellipsis. 1.
Yinfei Huang, Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2001 Design of a semantic parser with support to ellipsis resolution in a Chinese spoken language dialogue system
abstract
In this paper, a semantic parser with support to ellipsis resolution in a Chinese spoken language dialogue system is proposed. The grammar and parsing strategy of this parser is designed to address the characteristics of spoken language and to support the ellipsis resolution. Namely, it parses the user utterance with a domain-specific semantic grammar based on a template-filling approach. Syntactic constraints extracted by a Generalized LR parser are also used in the parsing process. With a paradigm of two-state bottom-up parsing and a scoring scheme, the ellipsis resolution module is integrated into the parser seamlessly. The parsing result is represented by a linked structure of semantic frames, which is convenient to both the parser and its successive components of the dialogue system. 1.
Thomas Fang Zheng, Yinfei Huang
INTERSPEECH2
2001 An MCE based classification tree using hierarchical feature-weighting in speech recognition
abstract
In this paper a hierarchical classification framework using the feature-weighting tree for the objective of applying diverse weighting to acoustic features is proposed for speech recognition. The hierarchical feature-weighting tree with a flexible structure complexity can be constructed optimally with the optimal splitting for the recognition confusion graph. Based on the minimum classification error principle, the subset-dependent training and the multi-level recognition method are proposed, where the feature weighting can be automatically trained without normalization in recognition. Both the mathematical analysis and the experimental results show that such a supervised hierarchical classification tree based on the feature weighting is efficient to reduce the speech recognition error. 1.
Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2001 An online incremental language model adaptation method
abstract
In this paper, an online incremental language model adaptation method is proposed, which is different from the traditional offline language model adaptation method. There are some problems in the online incremental adaptation. The first one is how to adjust the model parameters online and modify the model incrementally. The second one is how to induce new words and assign initial probabilities to the ngrams related to them. In our application for Chinese character input method editor, the language model is divided into two parts, corresponding to the background (generalpurpose) model and the user model, respectively. A modified maximum a posterior method is proposed for adapting the user model dynamically. Experiments are done to test the proposed method on an Chinese sentence input system and the results show that a satisfying word error rate reduction is obtained when the input articles are of similar topics.
Genqing Wu, Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2001 Robust parsing in spoken dialogue systems
abstract
The rule-based parsing is a prevalent method for the natural language understanding (NLU) and has been introduced in dialogue systems for spoken language processing (SLP). However, additional measures must be taken to cope with the severe spoken linguistic phenomena, such as garbage, repetition, ellipsis, word disordering, fragment and ill form, which frequently occur in the spoken language. We propose in this paper a robust parsing scheme, which integrates the following methods. Keywords are used as terminal symbols; hence the symbol set of the grammar is purely within the semantical category. The definition of the grammar is extended to accommodate four types of rules, called up-tying, by-passing, up-messing, and over-crossing respectively. An improved chart parser, named marionette, is designed to parse the semantic grammar instance. The robust parsing scheme has been adopted in an air traveling information service system, called EasyFlight, and has achieved a high performance when dealing with the spontaneous speech. 1.
Pengju Yan, Thomas Fang Zheng, Mingxing Xu
INTERSPEECH2
2001 Improved context-dependent acoustic modeling for continuous Chinese speech recognition
abstract
This paper describes the new framework of context-dependent (CD) Initial/Final (IF) acoustic modeling using the decision tree based state tying for continuous Chinese speech recognition. The Extended Initial/Final (XIF) set is chosen as the basic speech recognition unit (SRU) set according to the Chinese language characteristics, which outperforms the standard IF set. An adaptive mixture increasing strategy is applied when splitting the single Gaussian into mixed Gaussians in each tied state after the decision tree has been constructed. Our experimental results show that these two improvements are helpful to the acoustic modeling of Chinese speech recognition and that the CD XIF model outperforms the baseline syllable model over 30%. 1.
Thomas Fang Zheng, Chunhua Luo
INTERSPEECH2
2001 A two-layer lexical tree based beam search in continuous Chinese speech recognition
abstract
In this paper, an approach to continuous speech recognition based on a two-layer lexical tree is proposed. The search network is maintained by the two-layer lexical tree, in which the first layer reflects the word net and the phone net while the second layer the dynamic programming (DP). Because the acoustic information is tied in the second layer, the memory cost is so small that it has the ability to process some complicated applications, such as the use of cross-word context-dependent (CD) triphone models, the Chinese fuzzy syllable mapping and the pronunciation modeling. The search algorithm based on the two-layer lexical tree is also proposed, which is derived from the token-passing algorithm. Finally, an implementation of the two-layer lexical tree using the cross-word context-dependent triphone models is presented, and the experimental results show that the highly efficient decoding can be achieved without too much memory cost. 1.
Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2001 Modeling pronunciation variation using context-dependent weighting and b/s refined acoustic modeling
abstract
The pronunciation variability is an important issue that must be faced with when developing practical automatic spontaneous speech recognition systems. By studying the initial/final (IF) characteristics of Chinese language and developing the Bayesian equation, we propose the concepts of generalized initial/final (GIF) and generalized syllable (GS), the GIF modeling method and the IF-GIF modeling method, as well as the context-dependent pronunciation weighting method. By using these approaches, the IF-GIF modeling reduces the Chinese syllable error rate (SER) by 6.3 % and 4.2 % compared with the GIF modeling and IF modeling respectively when the language modeling, such as syllable or word N-gram, is not used. 1.
Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne
INTERSPEECH1
2001 Comparison of Different Implementations of MFCC
Thomas Fang Zheng, Zhanjiang Song
J. Comput. Sci. Technol.1
2000 Statistical knowledge based frame synchronous search strategies in continuous speech recognition
abstract
In this paper, we propose a novel and efficient search algorithm for the continuous speech recognition (CSR). The proposed algorithm is on the basis of the traditional frame synchronous search (FSS) algorithm. It makes full use of some statistical knowledge, such as the differential state dwelling distribution (DSDD), as one of the control factors for the state transition. It also incorporates some other rule-based knowledge, such as the pruning criterion based on the dynamic forward prediction and the lexical word search tree (WST), as the search constraint. Experimental result shows that the statistical knowledge based search strategies can improve the performance of the CSR system significantly with an increase of the accuracy by 36.6% compared with the baseline FSS. Also, EasyTalk, the Chinese CSR system based on it, has achieved higher recognition accuracy and decoding efficiency.
Zhanjiang Song, Thomas Fang Zheng, Wenhu Wu
ICASSP2
2000 Language understanding component for Chinese dialogue system
abstract
In this paper we present the design and the implement of the language understanding component of a Chinese spoken language dialogue system EasyNav. In pursuing the coherence with the goal of understanding, we design the structure of system with speech decoding and language understanding integrated closely. Thus the language understanding component need to be restrictive besides portable. Actually we implement a general syntactic parser and domain-specific semantic parser for the purpose. The grammar rules written for understanding are suitable for spoken language. The feature of spoken language also exists throughout the system. 1.
Yinfei Huang, Thomas Fang Zheng, Mingxing Xu, Pengju Yan, Wenhu Wu
INTERSPEECH2
2000 The phonetic labeling on read and spontaneous discourse corpora
abstract
Read and spontaneous discourses are two different but very significant speech styles to be investigated. So phonetic labeling on read and spontaneous discourse corpora are made one is ASCCD, a 10 hours read discourse corpus and the other is CASS, a 4 hours spontaneous discourse corpus. First the principles and conventions of transcription are presented. Then, these two speech styles are compared from phonetic and syntactic point of view, including the statistic results of different phonetic units got from the annotated corpora.
Guohua Sun, Wu Hua, Zhigang Yin, Yiqing Zu, Thomas Fang Zheng, Zhanjiang Song
INTERSPEECH7
2000 CASS: a phonetically transcribed corpus of mandarin spontaneous speech
Thomas Fang Zheng, William J. Byrne, Pascale Fung, Terri Kamm, Yi Liu 0022, Zhanjiang Song, Umar Ruhi, Veera Venkataramani
INTERSPEECH2
2000 An equivalent-class based MMI learning method for MGCPM
abstract
In this paper, we present an Equivalent-Class Based Maximum Mutual Information (ECB-MMI) learning method for our previously proposed Mixed Gaussian Continuous Probability Model (MGCPM). Similar to HMMs, the defined object function for MGCPM training considers the mutual information among different models so as to maximally separate the Speech Recognition Units (SRUs) in model space. Experimental result shows that for MGCPM the MMI training method can improve the recognition rate by 5 % compared to the traditional training method MLE (Maximum Likelihood Estimation). Because the computation amount of MMI algorithm is very large, we propose an N-Best strategy to find the corresponding equivalent class (EC) in order to reduce complexity. Our experimental result shows that this criterion works very well.
Chunhua Luo, Thomas Fang Zheng, Mingxing Xu
INTERSPEECH2
2000 A c/v segmentation method for Mandarin speech based on multiscale fractal dimension
abstract
This paper proposes a new algorithm for Mandarin speech Consonant and Vowel (C/V) segmentation based on the fractal theory. The new method focuses on searching the transient region between the Consonant and Vowel parts in a Mandarin syllable that in general is a concatenation of a consonant followed by a vowel. The Multiscale Fractal Dimension Set (MFD) stands for the fractal dimensions at multiple maximum resolutions of computation. Just using the r-variance of MFD (the degree of the difference from all elements of a MFD) to distinguish clearly between the stable phonemes and their transient region, the algorithm can directly search the speech frame with minimum r-variance of MFD as the C/V segmentation boundary. A result of 95.2% segmentation accuracy is obtained for clean test corpus, and 82.3% accuracy in noisy environment with the SNR of 10 dB. This shows that the new C/V segmentation algorithm is qualified for the task of continuous Mandarin speech recognition.
Thomas Fang Zheng, Wenhu Wu
INTERSPEECH2
2000 On enhancing katz-smoothing based back-off language model
Jian Wu 0034, Thomas Fang Zheng
INTERSPEECH2
2000 Reducing time-synchronous beam search effort using stage based look-ahead and language model rank based pruning
abstract
In this paper, we present an efficient look-ahead technique based on both the Language Model (LM) Look-Ahead and the Acoustic Model (AM) Look-Ahead, for the time-synchronous beam search in the large vocabulary speech recognition. In this so-call stage based look-ahead (SLA) technique, two predicting processes with different hypothesis evaluating criteria are organized by stages according to the different requirements for pruning the unlikely surviving hypotheses. Furthermore, in order to reduce the efforts for distributing the LM over the lexical tree more effectively, the LM Rank based Pruning (LMRP) is integrated with the extension of each new phoneme node. The recognition experiments performed on the 50k-word Mandarin Dictation task (Easytalk2000) show that a reduction by 10 percents in the search effort in comparison with the standard word-conditioned search using LM look-ahead only, and a reduction of 25 percents in the word error rates in comparison with the search algorithm without any look-ahead can be achieved. 1.
Jian Wu 0034, Thomas Fang Zheng
INTERSPEECH2
2000 Semi-continuous segmental probability modeling for continuous speech recognition
abstract
In this paper the design of semi-continuous segmental probability models (SCSPMs) in large vocabulary continuous speech recognition is presented. The tied Gaussian densities are trained using data from all states of all utterances while the mixture weights are estimated using data from the state being trained individually. The SCSPMs tie all the densities of all states from all Speech Recognition Units (SRUs) to form a shared pdf codebook, thus the number of Gaussian densities is greatly reduced. Several pruning methods are reviewed and then a new pruning criterion is proposed in order to reduce the number of tied mixture Gaussian densities while there is only a small subset of mixture Gaussian densities with larger tying weights. Our preliminary experiments show that the SCSPM incorporated with the pruning techniques can lessen the size of model storage and speed up the system with little degradation in the accuracy compared to the prior continuous model.
Thomas Fang Zheng, Mingxing Xu, Ditang Fang
INTERSPEECH2
2000 Input Chinese sentences using digits
abstract
Chinese character input is always a key issue in a variety of Chinese based applications especially when only a small number keypad is available. Though many kinds of Chinese character encoding schemes are proposed according to Chinese character characteristics, such as the shape, they are not straightforward and will take users a long time to learn. An easy way is to input via Chinese pinyins. In this paper, we establish the mapping between digit string and pinyin as well as the mapping between the pinyin string and the word, referred to as the Syllable-Digit search Tree (SDT) and the Word-Syllable search Tree (WST) respectively. By using these two search trees as well as the word N-gram language model and the syllable-synchronous network search (SSNS) algorithm, any digit string can be easily converted into Chinese word sequence or sentence. Without users ’ selecting from candidates, the character error rate (CER) of digit-to-character (D/C) conversion is 6.6 % across a test text consisting 22,083 characters. 1.
Thomas Fang Zheng, Jian Wu 0034, Wenhu Wu
INTERSPEECH1
2000 Integrating the energy information into MFCC
abstract
The Mel-Frequency Cepstrum Coefficients (MFCC) is a widely used set of feature used in automatic speech recognition systems introduced in 1980 by Davis and Mermelstein [2]. In this traditional implementation, the 0th coefficient is excluded for the reason it is somewhat unreliable. In this paper, we analyze this term and find that it can be regarded as the generalized frequency band energy (FBE) and is hence useful, resulting in the FBE-MFCC. We also propose a better analysis, called the auto-regressive analysis, on the frame energy, which performs better than its 1st and/or 2nd order differential derivatives. Experiments show that, the FBE-MFCC and the frame energy with their corresponding auto-regressive analysis coefficients form the better combination reducing the syllable error rate (SER) by 10.0 % across a giant speech database, compared to the traditional MFCC with its corresponding auto-regressive analysis coefficients. 1.
Thomas Fang Zheng
INTERSPEECH1
2000 Improving the Syllable-Synchronous Network Search Algorithm for Word Decoding in Continuous Chinese Speech Recognition
Thomas Fang Zheng, Jian Wu 0034, Zhanjiang Song
J. Comput. Sci. Technol.1
1999 An new method used in HMM for modeling frame correlation
abstract
We present a novel method to incorporate temporal correlation into a speech recognition system based on conventional hidden Markov model (HMM). In this new model the probability of the current observation not only depends on the current state but also depends on the previous state and the previous observation. The joint conditional probability density (PD) is approximated by a non-linear estimation method. As a result, we can still use the mixture Gaussian density to represent the joint conditional PD for the principle of any PD can be approximated by the mixture Gaussian density. The HMM incorporated temporal correlation by the non-linear estimation method, which we called FC HMM does not need any additional parameters and it only brings a little additional computing quantity. The results of the experiment show that the top 1 recognition rate of FC HMM has been raised by 6 percent compared to the conventional HMM method.
Thomas Fang Zheng, Jian Wu 0034, Wenhu Wu
ICASSP2
1999 A syllable-synchronous network search algorithm for word decoding in Chinese speech recognition
abstract
The Chinese language is syllabic in nature with frequent homonym phenomena and severe word boundary uncertainty problem. This makes the Chinese continuous speech recognition (CSR) slightly difficult. In order to solve these problems, a Chinese syllable-synchronous network search (SSNS) algorithm is proposed. Together with the vocabulary word search tree and the N-gram based language model, the syllable-synchronous network search algorithm gives a good solution to the Chinese syllable-to-word conversion. In addition, this algorithm is a good method for the accent Chinese speech recognition. The experimental results have showed that the SSNS algorithm can achieve a good overall continuous Chinese speech recognition system performance.
Thomas Fang Zheng
ICASSP1
1999 An effective scoring method for speaking skill evaluation system
abstract
The Speaking Skill Evaluation (SSE) technologies are derived from speech recognition technologies and are used for language learning and instructing. In this paper, an effective automatic pronunciation scoring method for SSE systems is proposed. The Center-Distance Continuous Probability Model (CDCPM) is incorporated to model the speech. The Merging-Based Syllable Detection Automaton (MBSDA) and the Non-Linear Partition (NLP) method are used to perform the time alignment. And the Critical Area Percentage (CAP) based scoring method is used to score the learner’s pronunciations or reject invalid utterances. Subjective assessments show that this method is concise, fast, and effective. The SSE system based on it has achieved a satisfying performance.
Zhanjiang Song, Thomas Fang Zheng, Mingxing Xu, Wenhu Wu
EUROSPEECH2
1999 A fast and effective state decoding algorithm
abstract
\n Contains fulltext :\n 75022.pdf (author's version ) (Open Access)\n
Mingxing Xu, Thomas Fang Zheng, Wenhu Wu
EUROSPEECH2
1999 Easytalk: a large-vocabulary speaker-independent Chinese dictation machine
Thomas Fang Zheng, Zhanjiang Song, Mingxing Xu, Jian Wu 0034, Yinfei Huang, Wenhu Wu
EUROSPEECH1
1999 HarkMan - A vocabulary-independent keyword spotter for spontaneous Chinese speech
Thomas Fang Zheng, Mingxing Xu, Xiaolong Mou, Jian Wu 0034, Wenhu Wu, Ditang Fang
J. Comput. Sci. Technol.1
1998 Non-linear probability estimation method used in HMM for modeling frame correlation
abstract
In this paper we present a novel method to incorporate temporal correlation into a speech recognition system based on HMM. An obvious way to incorporate temporal correlation is to condition the probability of the current observation on the current state as well as on the previous observation and the previous state. But use this method directly must lead to unreliable parameter estimates for the number of parameters to be estimated may increase too excessively to limited train data. In this paper, we approximate the joint conditional PD by non-linear estimation method. The HMM incorporated temporal correlation by non-linear estimation method, which we called it FC HMM does not need any additional parameters and it only brings a little additional computing quantity. The results in the experiment show that the top 1 recognition rate of FC HMM has been raised by 6 percent compared to the traditional HMM method. 1.
Thomas Fang Zheng, Jian Wu 0034, Wenhu Wu
ICSLP2
1998 The distance measure for line spectrum pairs applied to speech recognition
abstract
The Line Spectrum Pair (LSP) based on the principle of linear predictive coding (LPC) plays a very important role in the speech synthesis; it has many interesting properties. Several famous speech compression / decompression algorithms, including the famous code excited linear predictive coding (CELP), are based on the LSP analysis, where the information loss or predicting errors are often very small due to the LSP’s characteristics. Unfortunately till now there is not a satisfying kind of distance measure available for LSP so that this kind of features can be used for speech recognition applications. In this paper, the principle of LSP analysis is studied at first, and then several distance measures for LSP are proposed which can describe very well the difference between two groups of different LSP parameters. Experimental results are also given to show the efficiency of the proposed distance measures. 1.
Thomas Fang Zheng, Zhanjiang Song, Wenjian Yu, Fengzhou Zheng, Wenhu Wu
ICSLP1
1998 Center-distance continuous probability models and the distance measure
Thomas Fang Zheng, Wenhu Wu, Ditang Fang
J. Comput. Sci. Technol.1
1997 A log-index weighted cepstral distance measure for speech recognition
Thomas Fang Zheng, Wenhu Wu, Ditang Fang
J. Comput. Sci. Technol.1