EDBT 2026 Demo / reviewers in the wild / expert
Chengyuan Ma
dblp:62/6934
· DBLP profile ↗
30ranked-venue papers
11as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 90% Data mining · 10% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › indexing
differentiable search index |
0.7 | 1 | 2023 | PersonalTM: Transformer Memory for Personalized Retrieval · SIGIR 2023 |
Information retrieval
personalized search |
0.7 | 1 | 2023 | PersonalTM: Transformer Memory for Personalized Retrieval · SIGIR 2023 |
Data mining › predictive modeling
supervised learning |
0.1 | 1 | 2011 | A Regularized Maximum Figure-of-Merit (rMFoM) Approach to Supervised and Semi-Supervised Learning · IEEE Trans. Speech Audio Process. 2011 |
Information retrieval › document retrieval
spoken document retrieval |
0.1 | 1 | 2005 | Vocabulary-Independent Indexing of Spontaneous Speech · IEEE Trans. Speech Audio Process. 2005 |
Data mining
semi-supervised learning |
0.0 | 1 | 2011 | A Regularized Maximum Figure-of-Merit (rMFoM) Approach to Supervised and Semi-Supervised Learning · IEEE Trans. Speech Audio Process. 2011 |
Methods — techniques the papers use, named apart from their topics
hierarchical loss · 0.7adapter architecture · 0.7tikhonov regularization · 0.1m-gram phoneme language model · 0.1inverted index · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Modal Synchronized Dataset for Benchmarking ADAS Responses to Traffic Control Devices
Shixiao Liang, Chengyuan Ma, Handong Yao, Qianwen Li, Xiaopeng Li 0020 |
IV | 5 |
| 2026 | Real-Time Traffic Crash Detection Platform Using Sparse Telematics Data from Connected Vehicle
Shixiao Liang, Chengyuan Ma, Keke Long, Xiaopeng Li 0020 |
IV | 2 |
| 2026 | Prescribed-Time Adaptive Neural Control for Piezoelectric Micropositioning Stages With Event-Triggered Communication
Heyu Hu, Chengyuan Ma, Shengjun Wen, Changan Jiang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Privacy-Preserving and Collusion-Resistant Data Query Scheme for Vehicular PlatoonsabstractData queries play a crucial role in the vehicular platoon, enabling vehicles to obtain traffic information about surrounding road conditions and personalized entertainment information services. However, data query requests from vehicles may expose the vehicle owner’s personal attributes and habits. Although several schemes can address these issues, they are incapable of countering collusion attacks between the coordinating vehicle and roadside units (RSUs). To solve this problem, in this paper we propose a privacy-preserving and collusion-resistant data query scheme, named PCDQ. Specifically, PCDQ uses the Paillier encryption and Chinese Remainder Theorem to protect the query privacy of vehicle owners, allowing the RSU to recover individual data query requests without associating them with the original vehicles. Next, the parameter update mechanism in PCDQ prevents the coordinating vehicle from obtaining the corresponding mapping information between vehicles and query parameters, thereby resisting collusion attacks between the coordinating vehicle and RSUs. In addition, identity-based signcryption is used to ensure secure parameter distribution among vehicles, and the batch verification enables efficient authentication of query requests. Detailed security proofs and analysis demonstrate that PCDQ satisfies multiple security properties, including resistance to collusion attacks and replay attacks, unlinkability, confidentiality, and authentication and data integrity. Experimental results show that, compared to existing solutions, PCDQ performs better in terms of computation overhead, communication overhead, and network performance. Chengyuan Ma, Tianjiao Ni, Liangchen Hu, Kaizhong Zuo, Fulong Chen 0002, Yonglong Luo |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2025 | AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech SynthesisabstractWith the advancement of speech synthesis technology, users have higher expectations for the naturalness and expressiveness of synthesized speech. But previous research ignores the importance of prompt selection. This study proposes a text-to-speech (TTS) framework based on Retrieval-Augmented Generation (RAG) technology, which can dynamically adjust the speech style according to the text content to achieve more natural and vivid communication effects. We have constructed a speech style knowledge database containing high-quality speech samples in various contexts and developed a style matching scheme. This scheme uses embeddings, extracted by Llama, PER-LLM-Embedder, and Moka, to match with samples in the knowledge database, selecting the most appropriate speech style for synthesis. Furthermore, our empirical research validates the effectiveness of the proposed method. Our demo can be viewed at: https://thuhcsi.github.io/icme2025-AutoStyle-TTS Chengyuan Ma, Wei Chen 0071, Zhiyong Wu 0001 |
ICME | 2 |
| 2025 | ESTM: An Enhanced Dual-Branch Spectral-Temporal Mamba for Anomalous Sound Detection
Chengyuan Ma, Hongyue Guo, Wenming Yang |
IEEE Signal Process. Lett. | 1 |
| 2024 | Testing Cellular Vehicle-to-Everything Communication Performance and Feasibility in Automated Vehicles *abstractMany studies have demonstrated the eco-driving capabilities of connected and automated vehicles (CAVs) to significantly enhance mobility systems. The majority of these studies have been conducted using simulations, which fail to capture the effects of practical uncertainties encountered in vehicle-to-anything (V2X) communications. In this paper, we investigated the performance of current cellular V2X (C-V2X) communications through systematic testing and provided a quantitative analysis of key performance indices (e.g., inter-packet gap and packet error rate) across various test scenarios. As one use case to demonstrate the benefits of C-V2X communication on the road, we tested the feasibility of eco-driving for a SAE level 3 (L3) automated vehicle (AV) communicating with a connected urban corridor capable of transmitting traffic light information (i.e., signal phase and timing). To achieve this, we implemented the eco-speed planning algorithm at a high-level in the AV control software system and ensured its interactions with other existing low-level control algorithms, as well as the C-V2X onboard unit. Finally, we experimentally demonstrated eco-driving of the L3 CAV on a scaled-down corridor with two signal-controlled intersections, revealing the AV’s ability to maintain smoother trajectories and avoid unnecessary stops compared to human-driven vehicles. Zhaohui Liang, Xiaopeng Li 0020, Dominik Karbowski, Chengyuan Ma, Aymeric Rousseau |
IV | 5 |
| 2023 | Clicker: Attention-Based Cross-Lingual Commonsense Knowledge TransferabstractRecent advances in cross-lingual commonsense reasoning (CSR) are facilitated by the development of multilingual pre-trained models (mPTMs). While mPTMs show the potential to encode commonsense knowledge for different languages, transferring commonsense knowledge learned in large-scale English corpus to other languages is challenging. To address this problem, we propose the attention-based Cross-LIngual Commonsense Knowledge transfER (CLICKER) framework, which minimizes the performance gaps between English and non-English languages in commonsense question-answering tasks. CLICKER effectively improves commonsense reasoning for non-English languages by differentiating non-commonsense knowledge from commonsense knowledge. Experimental results on public benchmarks demonstrate that CLICKER achieves remarkable improvements in the cross-lingual CSR task for languages other than English. Ruolin Su, Zhongkai Sun, Sixing Lu, Chengyuan Ma, Chenlei Guo |
ICASSP | 4 |
| 2023 | An Avatar Robot Overlaid with the 3D Human Model of a Remote OperatorabstractAlthough telepresence assistive robots have made significant progress, they still lack the sense of realism and physical presence of the remote operator. This results in a lack of trust and adoption of such robots. In this paper, we introduce an Avatar Robot System which is a mixed real/virtual robotic system that physically interacts with a person in proximity of the robot. The robot structure is overlaid with the 3D model of the remote caregiver and visualized through Augmented Reality (AR). In this way, the person receives haptic feedback as the robot touches him/her. We further present an Optimal Non-Iterative Alignment solver that solves for the optimally aligned pose of 3D Human model to the robot (shoulder to the wrist non-iteratively). The proposed alignment solver is stateless, achieves optimal alignment and faster than the baseline solvers (demonstrated in our evaluations). We also propose an evaluation framework that quantifies the alignment quality of the solvers through multifaceted metrics. We show that our solver can consistently produce poses with similar or superior alignments as IK-based baselines without their potential drawbacks. Ravi Tejwani, Chengyuan Ma, Paolo Bonato, H. Harry Asada |
IROS | 2 |
| 2023 | PersonalTM: Transformer Memory for Personalized RetrievalabstractThe Transformer Memory as a Differentiable Search Index (DSI) has been proposed as a new information retrieval paradigm, which aims to address the limitations of dual-encoder retrieval framework based on the similarity score. The DSI framework outperforms strong baselines by directly generating relevant document identifiers from queries without relying on an explicit index. The memorization power of DSI framework makes it suitable for personalized retrieval tasks. Therefore, we propose a Personal Transformer Memory (PersonalTM) architecture for personalized text retrieval. PersonalTM incorporates user-specific profiles and contextual user click behaviors, and introduces hierarchical loss in the decoding process to align with the hierarchical assignment of document identifier. Additionally, PersonalTM also employs an adapter architecture to improve the scalability for index updates and reduce computation costs, compared to the vanilla DSI. Experiments show that PersonalTM outperforms the DSI baseline, BM25, fine-tuned dual-encoder, and other personalized models in terms of precision at top 1st and 10th positions and Mean Reciprocal Rank (MRR). Specifically, PersonalTM improves p@1 by 58%, 49%, and 12% compared to BM25, Dual-encoder, and DSI, respectively. Ruixue Lian, Sixing Lu, Clint Solomon Mathialagan, Gustavo Aguilar, Pragaash Ponnusamy, Jialong Han, Chengyuan Ma, Chenlei Guo |
SIGIR | 7 |
| 2022 | Incremental User Embedding Modeling for Personalized Text ClassificationabstractIndividual user profiles and interaction histories play a significant role in providing customized experiences in real-world applications such as chatbots, social media, retail, and education. Adaptive user representation learning by utilizing user personalized information has be-come increasingly challenging due to ever-growing his-tory data. In this work, we propose an incremental user embedding modeling approach, in which embeddings of user’s recent interaction histories are dynamically integrated into the accumulated history vectors via a trans-former encoder. This modeling paradigm allows us to create generalized user representations in a consecutive manner and also alleviate the challenges of data management. We demonstrate the effectiveness of this approach by applying it to a personalized multi-class classification task based on the Reddit dataset, and achieve 9% and 30% relative improvement on prediction accuracy over a baseline system for two experiment settings through appropriate comment history encoding and task modeling. Ruixue Lian, Che-Wei Huang, Qilong Gu, Chengyuan Ma, Chenlei Guo |
ICASSP | 5 |
| 2018 | Combining Acoustic Embeddings and Decoding Features for End-of-Utterance Detection in Real-Time Far-Field Speech Recognition SystemsabstractWe present an end-of-utterance detector for real-time automatic speech recognition in far-field scenarios. The proposed system consists of three components: a long short-term memory (LSTM) neural network trained on acoustic features, an LSTM trained on l-best recognition hypotheses of the automatic speech recognition (ASR) decoder, and a feedforward deep neural network (DNN) combining embeddings derived from both LSTMs with pause duration features from the ASR decoder. At inference time, lower and upper latency (pause duration) bounds act as safeguards. Within the latency bounds, the utterance end-point is triggered as soon as the DNN posterior reaches a tuned threshold. Our experimental evaluation is carried out on real recordings of natural human interactions with voice-controlled far-field devices. We show that the acoustic embeddings are the single most powerful feature and particularly suitable for cross-lingual applications. We furthermore show the benefit of ASR decoder features, especially as a low cost alternative to ASR hypothesis em-beddings. Roland Maas, Ariya Rastrow, Chengyuan Ma, Guitang Lan, Kyle Goehner, Gautam Tiwari, Shaun Joseph, Björn Hoffmeister |
ICASSP | 3 |
| 2018 | LSTM-Based Whisper DetectionabstractThis article presents a whisper speech detector in the far-field domain. The proposed system consists of a long-short term memory (LSTM) neural network trained on log-filterbank energy (LFBE) acoustic features. This model is trained and evaluated on recordings of human interactions with voice-controlled, far-field devices in whisper and normal phonation modes. We compare multiple inference approaches for utterance-level classification by examining trajectories of the LSTM posteriors. In addition, we engineer a set of features based on the signal characteristics inherent to whisper speech, and evaluate their effectiveness in further separating whisper from normal speech. A benchmarking of these features using multilayer perceptrons (MLP) and LSTMs suggests that the proposed features, in combination with LFBE features, can help us further improve our classifiers. We prove that, with enough data, the LSTM model is indeed as capable of learning whisper characteristics from LFBE features alone compared to a simpler MLP model that uses both LFBE and features engineered for separating whisper and normal speech. In addition, we prove that the LSTM classifiers accuracy can be further improved with the incorporation of the proposed engineered features. Zeynab Raeesy, Kellen Gillespie, Chengyuan Ma, Thomas Drugman, Jiacheng Gu, Roland Maas, Ariya Rastrow, Björn Hoffmeister |
SLT | 3 |
| 2014 | Deep neural network trained with speaker representation for speaker normalizationabstractA method for speaker normalization in deep neural network (DNN) based discriminative feature estimation for automatic speech recognition (ASR) is presented. This method is applied in the context of a DNN configured for auto-encoder based low dimensional bottleneck (AE-BN) feature extraction where the derived features are used as input to a continuous Gaussian density hidden Markov model (HMM/GMM) based ASR decoder. While AE-BN features are known to provide significant reduction in ASR word error rate (WER) with respect to more conventional spectral magnitude based features, there is no general agreement on how these networks can reduce the impact of speaker variability by incorporating prior knowledge of the speaker. An approach is presented in this paper where spectrum based DNN inputs are augmented with speaker inputs that are derived from separate regression based speaker transformations. It is shown the proposed method could reduce the WER by 3% relative to the best speaker adapted AE-BN CDHMM system. Aanchan Mohan, Richard C. Rose, Chengyuan Ma |
ICASSP | 4 |
| 2011 | A Regularized Maximum Figure-of-Merit (rMFoM) Approach to Supervised and Semi-Supervised LearningabstractWe propose a regularized extension to supervised maximum figure-of-merit learning to improve its generalization capability and successfully extend it to semi-supervised learning. The proposed method can be used to approximate any objective function consisting of the commonly used performance metrics. We first derive detailed learning algorithms for supervised learning problems and then extend it to more general semi-supervised scenarios, where only a small part of the training data is labeled. The effectiveness of the proposed approach is justified by several text categorization experiments on different datasets. The novelty of this paper lies in several aspects: 1) Tikhonov regularization is used to alleviate potential overfitting of the maximum figure-of-merit criteria; 2) the regularized maximum figure-of-merit algorithm is successfully extended to semi-supervised learning tasks; 3) the proposed approach has good scalability to large-scale applications. Chengyuan Ma |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | A comparative study on system combination schemes for LVCSRabstractWe present a comparative study on combination schemes for large vocabulary continuous speech recognition by incorporating long-span class posterior probability features into conventional short-time cepstral features. System combination can improve the overall speech recognition performance when multiple systems exhibit different error patterns and multiple knowledge sources encode complementary information. A variety of combination approaches are investigated in this paper, e.g., feature concatenation single stream system, model combination multi-stream system, lattice rescoring and ROVER. These techniques work at different levels of a LVCSR system and have different computational cost. We compared their performance and analyzed their advantages and disadvantages on large vocabulary English broadcast news transcription tasks. Experimental results showed that model combination with independent tree consistently outperforms ROVER, feature concatenation and lattice rescoring. In addition, the phoneme posterior probability features do provide complementary information to short-time cepstral features. Chengyuan Ma, Hong-Kwang Jeff Kuo, Hagen Soltau, Upendra V. Chaudhari, Lidia Mangu |
ICASSP | 1 |
| 2009 | A detection-based approach to broadcast news video story segmentationabstractA detection-based paradigm decomposes a complex system into small pieces, solves each subproblem one by one, and combines the collected evidence to obtain a final solution. In this study of video story segmentation, a set of key events are first detected from heterogeneous multimedia signal sources, including a large scale concept ontology for images, text generated from automatic speech recognition systems, features extracted from audio track, and high-level video transcriptions. Then a discriminative evidence fusion scheme is investigated. We use the maximum figure-of-merit learning approach to directly optimize the performance metrics used in system evaluation, such as precision, recall, and F1 measure. Some experimental evaluations conducted on the TRECVID 2003 dataset demonstrate the effectiveness of the proposed detection-based paradigm. The proposed framework facilitates flexible combination and extensions of event detector design and evidence fusion to enable other related video applications. Chengyuan Ma, Byungki Byun, Ilseo Kim |
ICASSP | 1 |
| 2008 | Unsupervised anchor shot detection using multi-modal spectral clusteringabstractThis paper presents a novel unsupervised method for anchor shot detection using spectral clustering with multi-modal features. Unlike previous unsupervised studies where the acoustic trajectory features can not be combined with visual features directly, only a pairwise distance matrix from each attribute is needed instead of individual samples so that diverse information from heterogeneous features can be integrated in a unified manner. Experimental evaluation on a subset of the TRECVID 2004 dataset showed that an appropriate incorporation of the acoustic information with visual information will improve the F1 score from 0.68 for visual information only system to 0.87 in our unsupervised anchor shot detection system. Also a comparison study on the same dataset with a supervised system showed that the performance of our unsupervised system approach that of the supervised system. Chengyuan Ma |
ICASSP | 1 |
| 2008 | An experimental study on discriminative concept classifier combination for TRECVID high-level feature extractionabstractIn this paper, we present an experimental study on using high-dimensional image features to perform discriminative classifier combination for TRECVID concept detection. We combine a multi-class classifier with binary-class classifiers. After training a multi-class classifier, we train binary-class classifiers by decomposing a multi-class problem into several binary-class classification problems, and fuse them together using a discriminative classifier combination approach. This idea leverages on each classifier's properties; multi-class classifiers emphasize on segmenting a decision space optimally in terms of some overall performance criteria whereas binary classifiers focus on detecting corresponding positive samples locally. Testing on the TRECVID2005 development set with 39 LSCOM-Lite concepts by adding an additional set of 39 pairs of binary concept classifiers, the mean average precision was improved by 34.1% over our baseline system with only 39 multi-class concept classifiers. When compared with state-of-the-art systems our proposed method is quite competitive especially for concepts with a relatively small number of positive samples. Byungki Byun, Chengyuan Ma |
ICIP | 2 |
| 2008 | An efficient gradient computation approach to discriminative fusion optimization in semantic concept detectionabstractIn this paper, we propose an efficient gradient computation approach for discriminative fusion optimization in TRECVID high-level feature extraction. Numerical approximation was exploited in gradient calculation and model parameter update. The gradient of the performance measure was approximated by a sum of instance point-wise gradient instead of instance pair-wise gradient used in maximum figure-of-merit learning such that performance metrics like average precision can be optimized directly and efficiently on large training set. Experiments on the TRECVID 2005 high-level feature extraction test set showed that the proposed algorithm can improve the mean average precision from 0.254 of a state-of-the-art baseline system to 0.285. Chengyuan Ma |
ICPR | 1 |
| 2007 | Finding Speaker Identities with a Conditional Maximum Entropy ModelabstractIn this paper, we address the task of identifying the speakers by name in audio content. Identification of speakers by name helps to improve the readability of the transcript and also provides additional meta-data which can help in finding the audio content of interest. We present a conditional maximum entropy (maxent) framework for this problem which yields superior performance and lends itself well to incorporating different types of information. We take advantage of this property of maxent to explore new features for this task. We show that supplementing standard lexical triggers with information such as speaker gender and position of speaker name mentions afford us large gains in performance. At 95% precision, we increase the recall to 67% from the trigger baseline of 38%. Chengyuan Ma, Patrick Nguyen, Milind Mahajan |
ICASSP (4) | 1 |
| 2007 | Detection-based ASR in the automatic speech attribute transcription projectabstractWe present methods of detector design in the Automatic Speech Attribute Transcription project. This paper details the results of a student-led, cross-site collaboration between Georgia Institute of Technology, The Ohio State University and Rutgers University. The work reported in this paper describes and evaluates the detection-based ASR paradigm and discusses phonetic attribute classes, methods of detecting framewise phonetic attributes and methods of combining attribute detectors for ASR. We use Multi-Layer Perceptrons, Hidden Markov Models and Support Vector Machines to compute confidence scores for several prescribed sets of phonetic attribute classes. We use Conditional Random Fields (CRFs) and knowledge-based rescoring of phone lattices to combine framewise detection scores for continuous phone recognition on the TIMIT database. With CRFs, we achieve a phone accuracy of 70.63%, outperforming the baseline and enhanced HMM systems, by incorporating all of the attribute detectors discussed in the paper Ilana Bromberg, Jinyu Li 0001, Chengyuan Ma, Brett Matthews, Antonio Moreno-Daniel, Jeremy Morris, Sabato Marco Siniscalchi, Yu Tsao 0001, Yu Wang 0001 |
INTERSPEECH | 5 |
| 2007 | A study on word detector design and knowledge-based pruning and rescoringabstractDetection of speech attributes, phones and words is a key component of a detection-based automatic speech recognition framework in the automatic speech attribute transcription project. This paper presents a two-stage approach, keywordfiller network method followed by knowledge-based pruning and rescoring, for detection of any given word in continuous speech. Different from conventional keyword spotting systems, both content words and function words are considered in this study. To reduce the high miss, a modified grammar network for word detection is proposed. Then knowledge sources from landmark detection, attributes detection and other spectral cues were combined together to remove the unlikely putative segments from the hypothesized word candidates. This study has been evaluated on the WSJ0 corpus under matched and mismatched acoustic conditions. When comparing with the conventional keyword spotting system, we found the proposed word detector greatly improves the detection performance. The figure-of-merits for content and function words were improved from 48.8% to 61.5%, and 22.3% to 33.1% respectively. Index Terms: word detection, knowledge-based Chengyuan Ma |
INTERSPEECH | 1 |
| 2006 | A study on detection based automatic speech recognition
Chengyuan Ma, Yu Tsao 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 1 |
| 2005 | Vocabulary-Independent Indexing of Spontaneous SpeechabstractWe present a system for vocabulary-independent indexing of spontaneous speech, i.e., neither do we know the vocabulary of a speech recording nor can we predict which query terms for which a user is going to search. The technique can be applied to information retrieval, information extraction, and data mining. Our specific target is search in recorded conversations in the office/information-worker scenario-teleconferences, meetings, presentations, and voice mails. The focus of this paper is on how to index phonetic lattices. We will show that an index should provide expected term frequencies (ETFs) of query terms. Since, at indexing time, it is unknown which phoneme sequences constitute valid query terms, we will introduce an approximation of ETFs of a query's phoneme sequence by M-gram phoneme language models, which are estimated on lattices and organized in an inverted index-like structure for fast access. We will discuss ranking, estimation, and integration of phoneme/word hybrid approaches. Compared with an unindexed baseline without approximation, our approximation leads only to a 3.4% relative loss of search accuracy on the Linguistic Data Consortium (LDC) voicemail task. We also propose a two-stage method for locating individual keyword occurrences using the above method as a fast match. A 20-times speedup is achieved over unindexed search at under a 2-point accuracy loss. Last, we will briefly introduce a prototype applet based on the above techniques. Kaijiang Chen, Chengyuan Ma, Frank Seide |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | Vocabulary-independent search in spontaneous speechabstractFor efficient organization of speech recordings - meetings, interviews, voice mails, lectures - the ability to search for spoken keywords is an essential capability. Today, most spoken-document retrieval systems use large-vocabulary recognition. For the above scenarios, such systems suffer from both the unpredictable vocabulary/domain and generally high word-error rates (WER). We present a vocabulary-independent system to index and to search rapidly spontaneous speech. A speech recognizer generates lattices of phonetic word fragments, against which keywords are matched phonetically. We first show the need to use recognition alternatives (lattices) in a high-WER context, on a word-based baseline. Then we introduce our new method of phonetic word-fragment lattice generation, which uses longer-span language knowledge than a phoneme recognizer. Last we introduce heuristics to compact the lattices to feasible sizes that can be searched efficiently. On the LDC voice mail corpus, we show that vocabulary/domain-independent phonetic search is as accurate as a vocabulary/domain-dependent word-lattice based baseline system for in-vocabulary keywords (FOMs of 74-75%), but nearly maintains this accuracy also for out-of-vocabulary keywords. Frank Seide, Chengyuan Ma, Eric Chang |
ICASSP (1) | 3 |
| 2003 | Comparison of discriminative training methods for speaker verificationabstractThe maximum likelihood estimation (MLE) and Bayesian maximum a-posteriori (MAP) adaptation methods for Gaussian mixture models (GMM) have proven to be effective and efficient for speaker verification, even though each speaker model is trained using only his own training utterances. Discriminative criteria aim at increasing discriminability by using out-of-class data. In this paper, we consider the speaker verification task using three discriminative training methods to compare performance. Comparisons are discussed for the maximum mutual information (MMI), minimum classification error (MCE) and figure of merit (FOM) criteria. Experiments on the 1996 NIST speaker recognition evaluation data set show that FOM training method outperforms the other two methods for speaker verification in terms of system performance. Meanwhile, logistic regression is investigated and successfully employed as a discriminative score-normalization technique. Chengyuan Ma, Eric Chang |
ICASSP (1) | 1 |
| 2003 | Learning to boost GMM based speaker verificationabstractThe Gaussian mixture models (GMM) has proved to be an effective probabilistic model for speaker verification, and has been widely used in most of state-of-the-art systems. In this paper, we introduce a new method for the task: that using AdaBoost learning based on the GMM. The motivation is the following: While a GMM linearly combines a number of Gaussian models according to a set of mixing weights, we believe that there exists a better means of combining individual Gaussian mixture models. The proposed AdaBoost-GMM method is non-parametric in which a selected set of weak classifiers, each constructed based on a single Gaussian model, is optimally combined to form a strong classifier, the optimality being in the sense of maximum margin. Experiments show that the boosted GMM classifier yields 10.81% relative reduction in equal error rate for the same handsets and 11.24% for different handsets, a significant improvement over the baseline adapted GMM system. Stan Z. Li, Dong Zhang 0001, Chengyuan Ma, Harry Shum, Eric Chang |
INTERSPEECH | 3 |
| 2003 | An improved model-based speaker segmentation system
Frank Seide, Chengyuan Ma, Eric Chang |
INTERSPEECH | 3 |
| 2001 | Sparse image coding with clustering property and its application to face recognition
Jun Sun 0004, Qing Zhuo, Chengyuan Ma |
Pattern Recognit. | 3 |