EDBT 2026 Demo / reviewers in the wild / expert
Fuchun Peng
dblp:60/4377
· DBLP profile ↗
46ranked-venue papers
12as first author
10since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 15 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards measuring fairness in speech recognition: Fair-Speech datasetabstractThe current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups.This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers.Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands.We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models. Irina-Elena Veliche, Zhuangqun Huang, Vineeth Ayyat Kochaniyan, Fuchun Peng, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 4 |
| 2023 | Group Personalized Federated LearningabstractFederated learning (FL) can help promote data privacy by training a shared model in a de-centralized manner on the physical devices of clients. In the presence of heterogeneous distributions of local data, personalized FL strategy is introduced to mitigate the potential client drift. In this paper, we present the group personalization approach for applications of FL in which there exist inherent partitions over clients that are significantly distinct. In our approach, the global FL model is fine-tuned through another FL training process over each homogeneous group of clients, after which each group-specific FL model is further adapted and personalized per client. The proposed method can be well interpreted from a Bayesian hierarchical modeling perspective. With experiments on two real-world datasets for language modeling task, we demonstrate this approach can achieve superior personalization performance than other FL counterparts. Zhe Liu 0011, Yue Hui, Fuchun Peng |
ICASSP | 3 |
| 2023 | Mitigating Unintended Memorization in Language Models Via Alternating TeachingabstractRecent research has shown that language models have a tendency to memorize rare or unique sequences in the training corpora which can thus leak sensitive attributes of user data. We employ a teacher-student framework and propose a novel approach called alternating teaching to mitigate unintended memorization in sequential modeling. In our method, multiple teachers are trained on disjoint training sets whose privacy one wishes to protect, and teachers’ predictions supervise the training of a student model in an alternating manner at each time step. Experiments on LibriSpeech datasets show that the proposed method achieves superior privacy-preserving results than other counterparts. In comparison with no prevention for unintended memorization, the accuracy loss is small when training records are sufficient. Zhe Liu 0011, Fuchun Peng |
ICASSP | 3 |
| 2023 | Modeling Dependent Structure for Utterances in ASR Evaluation
Zhe Liu 0011, Fuchun Peng |
INTERSPEECH | 2 |
| 2022 | Neural-FST Class Language Model for End-to-End Speech RecognitionabstractWe propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transducers (FSTs) in a mathematically consistent framework. Our method utilizes a background NNLM which models generic background text together with a collection of domain-specific entities modeled as individual FSTs. Each output token is generated by a mixture of these components; the mixture weights are estimated with a separately trained neural decider. We show that NFCLM significantly outperforms NNLM by 15.8% relative in terms of Word Error Rate. NFCLM achieves similar performance as traditional NNLM and FST shallow fusion while being less prone to overbiasing and 12 times more compact, making it more suitable for on-device usage. Antoine Bruguier, Rohit Prabhavalkar, Dangna Li, Zhe Liu 0011, Eun Chang, Fuchun Peng, Ozlem Kalinli, Michael L. Seltzer |
ICASSP | 8 |
| 2022 | Model-Based Approach for Measuring the Fairness in ASRabstractThe issue of fairness arises when the automatic speech recognition (ASR) systems do not perform equally well for all subgroups of the population. In any fairness measurement studies for ASR, the open questions of how to control the confounding factors, how to handle unobserved heterogeneity across speakers, and how to trace the source of any word error rate (WER) gap among different subgroups are especially important - if not appropriately accounted for, incorrect conclusions will be drawn. In this paper, we introduce mixed-effects Poisson regression to better measure and interpret any WER difference among subgroups of interest. Particularly, the presented method can effectively address the three problems raised above and is very flexible to use in practical disparity analyses. We demonstrate the validity of proposed model-based approach on both synthetic and real-world speech data. Zhe Liu 0011, Irina-Elena Veliche, Fuchun Peng |
ICASSP | 3 |
| 2021 | On Lattice-Free Boosted MMI Training of HMM and CTC-Based Full-Context ASR ModelsabstractHybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implemented in different frameworks. In this paper, by decoupling the concepts of modeling units and label topologies and building proper numerator/denominator graphs accordingly, we establish a generalized framework for hybrid acoustic modeling (AM). In this framework, we show that LF-MMI is a powerful training criterion applicable to both limited-context and full-context models, for wordpiece/mono-char/bi-char/chenone units, with both HMM/CTC topologies. From this framework, we propose three novel training schemes: chenone(ch)/wordpiece(wp)-CTC-bMMI, and wordpiece(wp)-HMM-bMMI with different advantages in training performance, decoding efficiency and decoding time-stamp accuracy. The advantages of different training schemes are evaluated comprehensively on Librispeech, and wp-CTC-bMMI and ch-CTC-bMMI are evaluated on two real world ASR tasks to show their effectiveness. Besides, we also show bi-char(bc) HMM-MMI models can serve as better alignment models than traditional non-neural GMM-HMMs. Xiaohui Zhang 0007, Vimal Manohar, Frank Zhang 0001, Yangyang Shi, Nayan Singhal, Julian Chan, Fuchun Peng, Yatharth Saraf, Mike Seltzer |
ASRU | 8 |
| 2021 | Analyzing the Forgetting Problem in Pretrain-Finetuning of Open-domain Dialogue Response ModelsabstractTianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, Fuchun Peng. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Tianxing He, Kyunghyun Cho, Myle Ott, Bing Liu 0024, James R. Glass, Fuchun Peng |
EACL | 7 |
| 2021 | Federated Marginal Personalization for ASR RescoringabstractWe introduce federated marginal personalization (FMP), a novel method for continuously updating personalized neural network language models (NNLMs) on private devices using federated learning (FL). Instead of fine-tuning the parameters of NNLMs on personal data, FMP regularly estimates global and personalized marginal distributions of words, and adjusts the probabilities from NNLMs by an adaptation factor that is specific to each word. Our presented approach can overcome the limitations of federated fine-tuning and efficiently learn personalized NNLMs on devices. We study the application of FMP on second-pass ASR rescoring tasks. Experiments on two speech evaluation datasets show modest word error rate (WER) reductions. We also demonstrate that FMP could offer reasonable privacy with only a small decrease in speech recognition accuracy. Zhe Liu 0011, Fuchun Peng |
ICASSP | 2 |
| 2021 | Benchmarking LF-MMI, CTC And RNN-T Criteria For Streaming ASRabstractIn this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on three popular training criteria: LF-MMI, CTC and RNN-T. In transcribing social media videos of 7 languages with training data 3K - 14K hours, we conduct large-scale controlled experimentation across each criterion using identical datasets and encoder model architecture. We find that RNN-T has consistent wins in ASR accuracy, while CTC models excel at inference efficiency. Moreover, we selectively examine various modeling strategies for different training criteria, including modeling units, encoder architectures, pre-training, etc. Given such large-scale real-world streaming ASR application, to our best knowledge, we present the first comprehensive benchmark on these three widely used training criteria across a great many languages. Xiaohui Zhang 0007, Frank Zhang 0001, Chunxi Liu, Kjell Schubert, Julian Chan, Pradyot Prakash, Ching-Feng Yeh, Fuchun Peng, Yatharth Saraf, Geoffrey Zweig |
SLT | 9 |
| 2020 | An Empirical Study of Transformer-Based Neural Language Model AdaptationabstractWe explore two adaptation approaches of deep Transformer based neural language models (LMs) for automatic speech recognition. The first approach is a pretrain-finetune framework, where we first pretrain a Transformer LM on a large-scale text corpus from scratch and then adapt it to relatively small target domains via finetuning. The second approach is a mixer of dynamically weighted models that are separately trained on source and target domains, aiming to improve simple linear interpolation with dynamic weighting. We compare the two approaches with three baselines - without adaptation, merging data, and simple interpolation - on Switchboard (SWBD) and Wall Street Journal (WSJ). Experiments show that the mixer model generally performs better than baselines and finetuning. Compared with no adaptation, finetuning and the mixer approach obtain up to relative 11.5% and 14.1% WER reductions on SWBD, respectively. The mixer model also outperforms linear interpolation and merging data. On WSJ, the mixer approach achieves a new state-of-the-art WER result. Ke Li 0018, Zhe Liu 0011, Tianxing He, Hongzhao Huang, Fuchun Peng, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 5 |
| 2020 | Training ASR Models By Generation of Contextual InformationabstractSupervised ASR models have reached unprecedented levels of accuracy, thanks in part to ever-increasing amounts of labelled training data. However, in many applications and locales, only moderate amounts of data are available, which has led to a surge in semi- and weakly-supervised learning research. In this paper, we conduct a large-scale study evaluating the effectiveness of weakly-supervised learning for speech recognition by using loosely related contextual information as a surrogate for ground-truth labels. For weakly supervised training, we use 50k hours of public English social media videos along with their respective titles and post text to train an encoder-decoder transformer model. Our best encoder-decoder models achieve an average of 20.8% WER reduction over a 1000 hours supervised baseline, and an average of 13.4% WER reduction when using only the weakly supervised encoder for CTC fine-tuning. Our results show that our setup for weak supervision improved both the encoder acoustic representations as well as the decoder language generation abilities. Kritika Singh, Dmytro Okhonko, Yongqiang Wang 0005, Frank Zhang 0001, Ross B. Girshick, Sergey Edunov, Fuchun Peng, Yatharth Saraf, Geoffrey Zweig, Abdel-rahman Mohamed |
ICASSP | 8 |
| 2020 | Statistical Testing on ASR Performance via Blockwise BootstrapabstractA common question being raised in automatic speech recognition (ASR) evaluations is how reliable is an observed word error rate (WER) improvement comparing two ASR systems, where statistical hypothesis testing and confidence interval (CI) can be utilized to tell whether this improvement is real or only due to random chance.The bootstrap resampling method has been popular for such significance analysis which is intuitive and easy to use.However, this method fails in dealing with dependent data, which is prevalent in speech world -for example, ASR performance on utterances from the same speaker could be correlated.In this paper we present blockwise bootstrap approach -by dividing evaluation utterances into nonoverlapping blocks, this method resamples these blocks instead of original data.We show that the resulting variance estimator of absolute WER difference between two ASR systems is consistent under mild conditions.We also demonstrate the validity of blockwise bootstrap method on both synthetic and real-world speech data. Zhe Liu 0011, Fuchun Peng |
INTERSPEECH | 2 |
| 2016 | Learning Personalized Pronunciations for Contact Name Recognition
Antoine Bruguier, Fuchun Peng, Françoise Beaufays |
INTERSPEECH | 2 |
| 2015 | Fix it where it fails: Pronunciation learning by mining error corrections from speech logsabstractThe pronunciation dictionary, or lexicon, is an essential component in an automatic speech recognition (ASR) system in that incorrect pronunciations cause systematic misrecognitions. It typically consists of a list of word-pronunciation pairs written by linguists, and a grapheme-to-phoneme (G2P) engine to generate pronunciations for words not in the list. The hand-generated list can never keep pace with the growing vocabulary of a live speech recognition system, and the G2P is usually of limited accuracy. This is especially true for proper names whose pronunciations may be influenced by various historical or foreign-origin factors. In this paper, we propose a language-independent approach to detect misrecognitions and their corrections from voice search logs. We learn previously unknown pronunciations from this data, and demonstrate that they significantly improve the quality of a production-quality speech recognition system. Zhenzhen Kou, Daisy Stanton, Fuchun Peng, Françoise Beaufays, Trevor Strohman |
ICASSP | 3 |
| 2015 | Automatic pronunciation verification for speech recognitionabstractPronunciations for words are a critical component in an automated speech recognition system (ASR) as mis-recognitions may be caused by missing or inaccurate pronunciations. The need for high quality pronunciations has recently motivated data-driven techniques to generate them [1]. We propose a data-driven and language-independent framework for verification of such pronunciations to further improve the lexicon quality in ASR. New candidate pronunciations are verified by re-recognizing historical audio logs and examining the associated recognition costs. We build an additional pronunciation quality feature from word and pronunciation frequencies in logs. A machine learned classifier trained on these features achieves nearly 90% accuracy in labeling good vs bad pronunciations across all languages we tested. New pronunciations verified as good may be added to a dictionary, while bad pronunciations may be discarded or sent to experts for further evaluation. We simultaneously verify 5,000 to 30,000 new pronunciations within a few hours and show improvements in the ASR performance as a result of including pronunciations verified by this system. Kanishka Rao, Fuchun Peng, Françoise Beaufays |
ICASSP | 2 |
| 2015 | Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networksabstractGrapheme-to-phoneme (G2P) models are key components in speech recognition and text-to-speech systems as they describe how words are pronounced. We propose a G2P model based on a Long Short-Term Memory (LSTM) recurrent neural network (RNN). In contrast to traditional joint-sequence based G2P approaches, LSTMs have the flexibility of taking into consideration the full context of graphemes and transform the problem from a series of grapheme-to-phoneme conversions to a word-to-pronunciation conversion. Training joint-sequence based G2P require explicit grapheme-to-phoneme alignments which are not straightforward since graphemes and phonemes don't correspond one-to-one. The LSTM based approach forgoes the need for such explicit alignments. We experiment with unidirectional LSTM (ULSTM) with different kinds of output delays and deep bidirectional LSTM (DBLSTM) with a connectionist temporal classification (CTC) layer. The DBLSTM-CTC model achieves a word error rate (WER) of 25.8% on the public CMU dataset for US English. Combining the DBLSTM-CTC model with a joint n-gram model results in a WER of 21.3%, which is a 9% relative improvement compared to the previous best WER of 23.4% from a hybrid system. Kanishka Rao, Fuchun Peng, Hasim Sak, Françoise Beaufays |
ICASSP | 2 |
| 2014 | Pronunciation learning for named-entities through crowd-sourcingabstractObtaining good pronunciations for named-entities poses a challenge for automated speech recognition because named-entities are diverse in nature and origin, and new entities come up every day. In this paper, we investigate the feasibility of learning named-entity pronunciations using crowd-sourcing. By collecting audio samples from non-linguistic-expert speak-ers with Mechanical Turk and learning from them, we can quickly derive pronunciations that are more accurate in speech recognition tests than manual pronunciations generated by lin-guistic experts. Compared to traditional approaches of generat-ing pronunciations, this new approach proves to be cheap, fast, and quite accurate. 1. Attapol Rutherford, Fuchun Peng, Françoise Beaufays |
INTERSPEECH | 2 |
| 2013 | Search results based N-best hypothesis rescoring with maximum entropy classificationabstractWe propose a simple yet effective method for improving speech recognition by reranking the N-best speech recognition hypotheses using search results. We model N-best reranking as a binary classification problem and select the hypothesis with the highest classification confidence. We use query-specific features extracted from the search results to encode domain knowledge and use it with a maximum entropy classifier to rescore the N-best list. We show that rescoring even only the top 2 hypotheses, we can obtain a significant 3% absolute sentence accuracy (SACC) improvement over a strong baseline on production traffic from an entertainment domain. Fuchun Peng, Scott Roy, Ben Shahshahani, Françoise Beaufays |
ASRU | 1 |
| 2010 | Personalize web search results with user's locationabstractWe build a probabilistic model to identify implicit local intent queries, and leverage user's physical location to improve Web search results for these queries. Evaluation on commercial search engine shows significant improvement on search relevance and user experience. Yumao Lu, Fuchun Peng, Benoît Dumoulin |
SIGIR | 2 |
| 2009 | Context sensitive synonym discovery for web search queriesabstractWe propose a simple yet effective approach to context sensitive synonym discovery for Web search queries based on co-click analysis; i.e., analyzing queries leading to clicking same documents. In addition to deriving word based synonyms, we also derive concept based synonyms with the help of query segmentation. Evaluation results show that this approach dramatically outperforms the thesaurus based synonym replacement method in keeping search intent, from accuracy of 40% to above 80%. Fuchun Peng, Huihsin Tseng, Yumao Lu, Benoît Dumoulin |
CIKM | 2 |
| 2009 | Improving Web Search Relevance with Semantic Features
Yumao Lu, Fuchun Peng, Gilad Mishne, Benoît Dumoulin |
EMNLP | 2 |
| 2009 | Improving search relevance for implicitly temporal queriesabstractNo abstract available. Donald Metzler, Rosie Jones, Fuchun Peng, Ruiqiang Zhang |
SIGIR | 3 |
| 2008 | Analyzing web text association to disambiguate abbreviation in queriesabstractWe introduce a statistical model for abbreviation disambiguation in Web search, based on analysis of Web data resources, including anchor text, click log and query log. By combining evidence from multiple sources, we are able to accurately disambiguate the abbreviation in queries. Experiments on real Web search queries show promising results. Fuchun Peng, Benoît Dumoulin |
SIGIR | 2 |
| 2008 | Unsupervised query segmentation using generative language models and wikipediaabstractIn this paper, we propose a novel unsupervised approach to query segmentation, an important task in Web search. We use a generative query model to recover a query's underlying concepts that compose its original segmented form. The model's parameters are estimated using an expectation-maximization (EM) algorithm, optimizing the minimum description length objective function on a partial corpus that is specific to the query. To augment this unsupervised learning, we incorporate evidence from Wikipedia. Fuchun Peng |
WWW | 2 |
| 2007 | Context sensitive stemming for web searchabstractTraditionally, stemming has been applied to Information Retrieval tasks by transforming words in documents to the their root form before indexing, and applying a similar transformation to query terms. Although it increases recall, this naive strategy does not work well for Web Search since it lowers precision and requires a significant amount of additional computation. Fuchun Peng, Nawaaz Ahmed, Yumao Lu |
SIGIR | 1 |
| 2006 | Coupling feature selection and machine learning methods for navigational query identificationabstractIt is important yet hard to identify navigational queries in Web search due to a lack of sufficient information in Web queries, which are typically very short. In this paper we study several machine learning methods, including naive Bayes model, maximum entropy model, support vector machine (SVM), and stochastic gradient boosting tree (SGBT), for navigational query identification in Web search. To boost the performance of these machine techniques, we exploit several feature selection methods and propose coupling feature selection with classification approaches to achieve the best performance. Different from most prior work that uses a small number of features, in this paper, we study the problem of identifying navigational queries with thousands of available features, extracted from major commercial search engine results, Web search user click data, query log, and the whole Web's relational content. A multi-level feature extraction system is constructed.Our results on real search data show that 1) Among all the features we tested, user click distribution features are the most important set of features for identifying navigational queries. 2) In order to achieve good performance, machine learning approaches have to be coupled with good feature selection methods. We find that gradient boosting tree, coupled with linear SVM feature selection is most effective. 3) With carefully coupled feature selection and classification approaches, navigational queries can be accurately identified with 88.1% F1 score, which is 33% error rate reduction compared to the best uncoupled system, and 40% error rate reduction compared to a well tuned system without feature selection. Yumao Lu, Fuchun Peng, Nawaaz Ahmed |
CIKM | 2 |
| 2006 | Information extraction from research papers using conditional random fields
Fuchun Peng, Andrew McCallum |
Inf. Process. Manag. | 1 |
| 2005 | Combining Statistical Language Models via the Latent Maximum Entropy Principle
Dale Schuurmans, Fuchun Peng, Yunxin Zhao |
Mach. Learn. | 3 |
| 2004 | Event threading within news topicsabstractWith the overwhelming volume of online news available today, there is an increasing need for automatic techniques to analyze and present news to the user in a meaningful and efficient manner. Previous research focused only on organizing news stories by their topics into a flat hierarchy. We believe viewing a news topic as a flat collection of stories is too restrictive and inefficient for a user to understand the topic quickly. Ramesh Nallapati, Ao Feng, Fuchun Peng, James Allan 0001 |
CIKM | 3 |
| 2004 | Chinese Segmentation and New Word Detection using Conditional Random Fields
Fuchun Peng, Fangfang Feng, Andrew McCallum |
COLING | 1 |
| 2004 | Accurate Information Extraction from Research Papers using Conditional Random Fields
Fuchun Peng, Andrew McCallum |
HLT-NAACL | 1 |
| 2004 | An Integrated, Conditional Model of Information Extraction and Coreference with Appli
Ben Wellner, Andrew McCallum, Fuchun Peng, Michael Hay |
UAI | 3 |
| 2004 | Augmenting Naive Bayes Classifiers with Statistical Language Models
Fuchun Peng, Dale Schuurmans |
Inf. Retr. | 1 |
| 2004 | Dynamic Web log session identification with statistical language modelsabstractAbstract We present a novel session identification method based on statistical language modeling. Unlike standard timeout methods, which use fixed time thresholds for session identification, we use an information theoretic approach that yields more robust results for identifying session boundaries. We evaluate our new approach by learning interesting association rules from the segmented session files. We then compare the performance of our approach to three standard session identification methods—the standard timeout method, the reference length method, and the maximal forward reference method—and find that our statistical language modeling approach generally yields superior results. However, as with every method, the performance of our technique varies with changing parameter settings. Therefore, we also analyze the influence of the two key factors in our language‐modeling–based approach: the choice of smoothing technique and the language model order. We find that all standard smoothing techniques, save one, perform well, and that performance is robust to language model order. Jimmy Huang 0001, Fuchun Peng, Aijun An, Dale Schuurmans |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2004 | Learning mixture models with the regularized latent maximum entropy principleabstractThis paper presents a new approach to estimating mixture models based on a recent inference principle we have proposed: the latent maximum entropy principle (LME). LME is different from Jaynes' maximum entropy principle, standard maximum likelihood, and maximum aposteriori probability estimation. We demonstrate the LME principle by deriving new algorithms for mixture model estimation, and show how robust new variants of the expectation maximization (EM) algorithm can be developed. We show that a regularized version of LME (RLME), is effective at estimating mixture models. It generally yields better results than plain LME, which in turn is often better than maximum likelihood and maximum a posterior estimation, particularly when inferring latent variable models from small amounts of data. Dale Schuurmans, Fuchun Peng, Yunxin Zhao |
IEEE Trans. Neural Networks | 3 |
| 2003 | Language Independent Authorship Attribution with Character Level N-Grams
Fuchun Peng, Dale Schuurmans, Vlado Keselj |
EACL | 1 |
| 2003 | Combining Naive Bayes and n-Gram Language Models for Text Classification
Fuchun Peng, Dale Schuurmans |
ECIR | 1 |
| 2003 | Semantic n-gram language modeling with the latent maximum entropy principleabstractWe describe a unified probabilistic framework for statistical language modeling-the latent maximum entropy principle-which can effectively incorporate various aspects of natural language, such as local word interaction, syntactic structure and semantic document information. Unlike previous work on maximum entropy methods for language modeling, which only allow explicit features to be modeled, our framework also allows relationships over hidden features to be captured, resulting in a more expressive language model. We describe efficient algorithms for marginalization, inference and normalization in our extended models. We then present experimental results for our approach on the Wall Street Journal corpus. Dale Schuurmans, Fuchun Peng, Yunxin Zhao |
ICASSP (1) | 3 |
| 2003 | Learning Mixture Models with the Latent Maximum Entropy Principle
Dale Schuurmans, Fuchun Peng, Yunxin Zhao |
ICML | 3 |
| 2003 | Language and Task Independent Text Categorization with Simple Language Models
Fuchun Peng, Dale Schuurmans |
HLT-NAACL | 1 |
| 2003 | Boltzmann Machine Learning with the Latent Maximum Entropy Principle
Dale Schuurmans, Fuchun Peng, Yunxin Zhao |
UAI | 3 |
| 2003 | Applying Machine Learning to Text Segmentation for Information Retrieval
Jimmy Huang 0001, Fuchun Peng, Dale Schuurmans, Nick Cercone, Stephen E. Robertson |
Inf. Retr. | 2 |
| 2002 | Investigating the Relationship between Word Segmentation Performance and Retrieval Performance in Chinese IR
Fuchun Peng, Jimmy Huang 0001, Dale Schuurmans, Nick Cercone |
COLING | 1 |
| 2002 | Using self-supervised word segmentation in Chinese information retrievalabstractWe propose a self-supervised word-segmentation technique for Chinese information retrieval. This method combines the advantages of traditional dictionary based approaches with character based approaches, while overcoming many of their shortcomings. Experiments on TREC data show comparable performance to both the dictionary based and the character based approaches. However, our method is language independent and unsupervised, which provides a promising avenue for constructing accurate multilingual information retrieval systems that are flexible and adaptive. Fuchun Peng, Jimmy Huang 0001, Dale Schuurmans, Nick Cercone, Stephen E. Robertson |
SIGIR | 1 |
| 2001 | Self-Supervised Chinese Word Segmentation
Fuchun Peng, Dale Schuurmans |
IDA | 1 |