Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hsiao-Wuen Hon

dblp:50/4528 · DBLP profile ↗
← Back
49ranked-venue papers
11as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-authorArtificial intelligence and machine learning · 20 · 3 first-authorDatabases, data management, data science and information retrieval · 12Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 64% Representation and self-supervised learning · 26% Question answering and dialogue systems · 4%
Databases, data mining, and information retrieval
8 papers
Information retrieval · 76% Web and social media mining · 18% Recommender systems · 5%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 62% Audio and music processing · 19% Image and video processing · 19%

Topics — the 27 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
pre-trained language model
0.822020
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training · ICML 2020
Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019
Machine learning › Representation and self-supervised learning › pre-training
unified pre-training
0.822020
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training · ICML 2020
Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019
Natural language and speech › Language models and text generation
masked language modeling
0.412020
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training · ICML 2020
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.412019
Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019
Natural language and speech › Language models and text generation
text generation
0.412019
Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019
Information retrieval
cross-language information retrieval
0.222010
Exploiting query logs for cross-lingual query suggestions · ACM Trans. Inf. Syst. 2010
Cross-lingual query suggestion using query logs of different languages · SIGIR 2007
Information retrieval
query suggestion
0.222010
Exploiting query logs for cross-lingual query suggestions · ACM Trans. Inf. Syst. 2010
Cross-lingual query suggestion using query logs of different languages · SIGIR 2007
Information retrieval › question answering
community question answering
0.212013
Question Difficulty Estimation in Community Question Answering Services · EMNLP 2013
Information retrieval
question answering
0.212013
Question Difficulty Estimation in Community Question Answering Services · EMNLP 2013
Web and social media mining
user identity linkage
0.212013
What's in a name?: an unsupervised approach to link users across communities · WSDM 2013
Natural language and speech › Question answering and dialogue systems › answer generation
generative question answering
0.112019
Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019
Recommender systems › content recommendation
question recommendation
0.112008
Recommending questions using the mdl-based tree cut model · WWW 2008
Information retrieval
ranking
0.112008
Recommending questions using the mdl-based tree cut model · WWW 2008
Information retrieval › query understanding
query ambiguity
0.112007
Identifying ambiguous queries in web search · WWW 2007
Information retrieval › query understanding
query analysis
0.112007
Identifying ambiguous queries in web search · WWW 2007
Information retrieval › cross-language information retrieval
query translation
0.112007
Cross-lingual query suggestion using query logs of different languages · SIGIR 2007
Information retrieval
query understanding
0.112007
Identifying ambiguous queries in web search · WWW 2007
Web and social media mining › web mining
web page understanding
0.112007
Webpage understanding: an integrated approach · KDD 2007
Information retrieval › ranking
learning to rank
0.112006
Adapting ranking SVM to document retrieval · SIGIR 2006
Information retrieval
query log analysis
0.122010
Exploiting query logs for cross-lingual query suggestions · ACM Trans. Inf. Syst. 2010
Identifying ambiguous queries in web search · WWW 2007
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
distributed speech recognition
0.012002
Distributed speech processing in miPad's multimodal user interface · IEEE Trans. Speech Audio Process. 2002
Natural language and speech › Speech recognition and synthesis
noise robustness
0.012002
Distributed speech processing in miPad's multimodal user interface · IEEE Trans. Speech Audio Process. 2002
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition
0.012002
Distributed speech processing in miPad's multimodal user interface · IEEE Trans. Speech Audio Process. 2002
Audio and music processing
acoustic signal processing
0.012009
Mobile media search: has media search finally found its perfect platform? part II · ACM Multimedia 2009
Interaction techniques and input
mobile interaction
0.012009
Mobile media search: has media search finally found its perfect platform? part II · ACM Multimedia 2009
Computational finance and economics › online advertising
search advertising
0.012007
Cross-lingual query suggestion using query logs of different languages · SIGIR 2007
Information retrieval
document retrieval
0.012006
Adapting ranking SVM to document retrieval · SIGIR 2006

Methods — techniques the papers use, named apart from their topics

self-attention · 0.4partially autoregressive modeling · 0.4autoencoding · 0.4transformer · 0.4sequence-to-sequence pre-training · 0.4self-attention masking · 0.4discriminative model · 0.3unsupervised learning · 0.2pairwise comparison · 0.2pagerank · 0.2n-gram probability · 0.2competition-based model · 0.2discriminative probabilistic model · 0.1pseudo-relevance feedback · 0.1tree cut model · 0.1minimum description length · 0.1word co-occurrence statistics · 0.1speech feature enhancement · 0.1
YearPublicationVenuePosition
2020 UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
abstract
We propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with masked tokens, we rely on conventional masks to learn inter-relations between corrupted tokens and context via autoencoding, and pseudo masks to learn intra-relations between masked spans via partially autoregressive modeling. With well-designed position embeddings and self-attention masks, the context encodings are reused to avoid redundant computation. Moreover, conventional masks used for autoencoding provide global masking information, so that all the position embeddings are accessible in partially autoregressive language modeling. In addition, the two tasks pre-train a unified language model as a bidirectional encoder and a sequence-to-sequence decoder, respectively. Our experiments show that the unified language models pre-trained using PMLM achieve new state-of-the-art results on a wide range of language understanding and generation tasks across several widely used benchmarks. The code and pre-trained models are available at https://github.com/microsoft/unilm.
Hangbo Bao, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Nan Yang 0002, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon
ICML11
2019 A Brief History of Intelligence
abstract
Intelligence is the deciding factor of how human beings become the most dominant life forms on earth. Throughout history, human beings have developed tools and technologies which help civilizations evolve and grow. Computers, and by extension, artificial intelligence (AI), has played important roles in that continuum of technologies. Recently artificial intelligence has garnered much interest and discussion. As artificial intelligence are tools that can enhance human capability, a sound understanding of what the technology can and cannot do is also necessary to ensure their appropriate use. While developing artificial intelligence, we also found out the definition and understanding of our own human intelligence continue evolving. The debates of the race between human and artificial intelligence have been ever growing. In this talk, I will describe the history of both artificial intelligence and human intelligence (HI). From the great insights of the such historical perspectives, I would like to illustrate how AI and HI will co-evolve with each other and project the future of AI and HI.
Hsiao-Wuen Hon
ICMI1
2019 Unified Language Model Pre-training for Natural Language Understanding and Generation
abstract
This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at https://github.com/microsoft/unilm.
Li Dong 0004, Nan Yang 0002, Wenhui Wang 0003, Furu Wei, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon
NeurIPS9
2013 Question Difficulty Estimation in Community Question Answering Services
abstract
In this paper, we address the problem of estimating question difficulty in community question answering services.We propose a competition-based model for estimating question difficulty by leveraging pairwise comparisons between questions and users.Our experimental results show that our model significantly outperforms a PageRank-based approach.Most importantly, our analysis shows that the text of question descriptions reflects the question difficulty.This implies the possibility of predicting question difficulty from the text of question descriptions.
Jing Liu 0022, Quan Wang 0002, Chin-Yew Lin, Hsiao-Wuen Hon
EMNLP4
2013 What's in a name?: an unsupervised approach to link users across communities
abstract
In this paper, we consider the problem of linking users across multiple online communities. Specifically, we focus on the alias-disambiguation step of this user linking task, which is meant to differentiate users with the same usernames. We start quantitatively analyzing the importance of the alias-disambiguation step by conducting a survey on 153 volunteers and an experimental analysis on a large dataset of About.me (75,472 users). The analysis shows that the alias-disambiguation solution can address a major part of the user linking problem in terms of the coverage of true pairwise decisions (46.8%). To the best of our knowledge, this is the first study on human behaviors with regards to the usages of online usernames. We then cast the alias-disambiguation step as a pairwise classification problem and propose a novel unsupervised approach. The key idea of our approach is to automatically label training instances based on two observations: (a) rare usernames are likely owned by a single natural person, e.g. pennystar88 as a positive instance; (b) common usernames are likely owned by different natural persons, e.g. tank as a negative instance. We propose using the n-gram probabilities of usernames to estimate the rareness or commonness of usernames. Moreover, these two observations are verified by using the dataset of Yahoo! Answers. The empirical evaluations on 53 forums verify: (a) the effectiveness of the classifiers with the automatically generated training data and (b) that the rareness and commonness of usernames can help user linking. We also analyze the cases where the classifiers fail.
Jing Liu 0022, Fan Zhang 0092, Xinying Song, Young-In Song, Chin-Yew Lin, Hsiao-Wuen Hon
WSDM6
2011 Select-the-Best-Ones: A new way to judge relative relevance
Ruihua Song, Qingwei Guo, Ruochi Zhang, Guomao Xin, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon
Inf. Process. Manag.7
2010 Automatic extraction of web data records containing user-generated content
abstract
In this paper, we are concerned with the problem of automatically extracting web data records that contain user-generated content (UGC). In previous work, web data records are usually assumed to be well-formed with a limited amount of UGC, and thus can be extracted by testing repetitive structure similarity. However, when a web data record includes a large portion of free-format UGC, the similarity test between records may fail, which in turn results in lower performance. In our work, we find that certain domain constraints (e.g., post-date) can be used to design better similarity measures capable of circumventing the influence of UGC. In addition, we also use anchor points provided by the domain constraints to improve the extraction process, which ends in an algorithm called MiBAT (Mining data records Based on Anchor Trees). We conduct extensive experiments on a dataset consisting of forum thread pages which are collected from 307 sites that cover 219 different forum software packages. Our approach achieves a precision of 98.9% and a recall of 97.3% with respect to post record extraction. On page level, it perfectly handles 91.7% of pages without extracting any wrong posts or missing any golden posts. We also apply our approach to comment extraction and achieve good results as well.
Xinying Song, Jing Liu 0022, Yunbo Cao, Chin-Yew Lin, Hsiao-Wuen Hon
CIKM5
2010 Learning Query Ambiguity Models by Using Search Logs
Ruihua Song, Zhicheng Dou, Hsiao-Wuen Hon, Yong Yu 0001
J. Comput. Sci. Technol.3
2010 Exploiting query logs for cross-lingual query suggestions
abstract
Query suggestion aims to suggest relevant queries for a given query, which helps users better specify their information needs. Previous work on query suggestion has been limited to the same language. In this article, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to the scenarios of cross-language information retrieval (CLIR) and other related cross-lingual applications. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, and so on, are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly outperforms a baseline system that uses dictionary-based query translation. Besides, we evaluate CLQS with French-English and Chinese-English CLIR tasks on TREC-6 and NTCIR-4 collections, respectively. The CLIR experiments using typical retrieval models demonstrate that the CLQS-based approach has significantly higher effectiveness than several traditional query translation methods. We find that when combined with pseudo-relevance feedback, the effectiveness of CLIR using CLQS is enhanced for different pairs of languages.
Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon
ACM Trans. Inf. Syst.6
2009 Mobile media search
abstract
This panel paper presents motivations for discussing mobile media search and contains statements from the panelists who are industry research leaders in this field.
Berna Erol, Jordan Cohen, Minoru Etoh, Hsiao-Wuen Hon, Jiebo Luo 0001, Johan Schalkwyk
ICASSP4
2009 Mobile media search: has media search finally found its perfect platform? part II
abstract
Recently, many exciting media search applications have been introduced to take advantage of smart phones' audiovisual capture capabilities and their being always on and connected. These applications address a real pain point for most mobile users and allow them to search with minimal text entry, if any. Is the mobile platform an ideal fit for media search? Are audio and visual signal processing technologies sufficiently accurate to support most mobile search applications? What are the killer applications of mobile media search? Earlier in 2009 at ICASSP, a panel on this topic stirred up great interest and enthusiasm while leaving many questions untouched due to the limited time.
Berna Erol, Jiebo Luo 0001, Shih-Fu Chang, Minoru Etoh, Hsiao-Wuen Hon, Qian Lin 0001, Vidya Setlur
ACM Multimedia5
2009 Identification of ambiguous queries in web search
Ruihua Song, Zhenxiao Luo, Jian-Yun Nie, Yong Yu 0001, Hsiao-Wuen Hon
Inf. Process. Manag.5
2008 Viewing Term Proximity from a Different Perspective
Ruihua Song, Michael J. Taylor 0001, Ji-Rong Wen, Hsiao-Wuen Hon, Yong Yu 0001
ECIR4
2008 Recommending questions using the mdl-based tree cut model
abstract
The paper is concerned with the problem of question recommendation. Specifically, given a question as query, we are to retrieve and rank other questions according to their likelihood of being good recommendations of the queried question. A good recommendation provides alternative aspects around users' interest. We tackle the problem of question recommendation in two steps: first represent questions as graphs of topic terms, and then rank recommendations on the basis of the graphs. We formalize both steps as the tree-cutting problems and then employ the MDL (Minimum Description Length) for selecting the best cuts. Experiments have been conducted with the real questions posted at Yahoo! Answers. The questions are about two domains, 'travel' and 'computers & internet'. Experimental results indicate that the use of the MDL-based tree cut model can significantly outperform the baseline methods of word-based VSM or phrase-based VSM. The results also show that the use of the MDL-based tree cut model is essential to our approach.
Yunbo Cao, Huizhong Duan, Chin-Yew Lin, Yong Yu 0001, Hsiao-Wuen Hon
WWW5
2007 Webpage understanding: an integrated approach
abstract
Recent work has shown the effectiveness of leveraging layout and tag-tree structure for segmenting webpages and labeling HTML elements. However, how to effectively segment and label the text contents inside HTML elements is still an open problem. Since many text contents on a webpage are often text fragments and not strictly grammatical, traditional natural language processing techniques, that typically expect grammatical sentences, are no longer directly applicable. In this paper, we examine how to use layout and tag-tree structure in a principled way to help understand text contents on webpages. We propose to segment and label the page structure and the text content of a webpage in a joint discriminative probabilistic model. In this model, semantic labels of page structure can be leveraged to help text content understanding, and semantic labels ofthe text phrases can be used in page structure understanding tasks such as data record detection. Thus, integration of both page structure and text content understanding leads to an integrated solution of webpage understanding. Experimental results on research homepage extraction show the feasibility and promise of our approach.
Jun Zhu 0001, Bo Zhang 0010, Zaiqing Nie, Ji-Rong Wen, Hsiao-Wuen Hon
KDD5
2007 Cross-lingual query suggestion using query logs of different languages
abstract
Query suggestion aims to suggest relevant queries for a given query, which help users better specify their information needs. Previously, the suggested terms are mostly in the same language of the input query. In this paper, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to scenarios of cross-language information retrieval (CLIR) and cross-lingual keyword bidding for search engine advertisement. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, etc. are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly out performs a baseline system based on dictionary-based query translation. Besides, the resulting CLQS is tested with French to English CLIR tasks on TREC collections. The results demonstrate higher effectiveness than the traditional query translation methods.
Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon
SIGIR7
2007 Identifying ambiguous queries in web search
abstract
It is widely believed that some queries submitted to search engines are by nature ambiguous (e.g., java, apple). However, few studies have investigated the questions of "how many queries are ambiguous?" and "how can we automatically identify an ambiguous query?" This paper deals with these issues. First, we construct the taxonomy of query ambiguity, and ask human annotators to manually classify queries based upon it. From manually labeled results, we find that query ambiguity is to some extent predictable. We then use a supervised learning approach to automatically classify queries as being ambiguous or not. Experimental results show that we can correctly identify 87% of labeled queries. Finally, we estimate that about 16% of queries in a real search log are ambiguous.
Ruihua Song, Zhenxiao Luo, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon
WWW5
2006 Multimedia search - Microsoft's perspective
abstract
Summary form only given. The explosive growth of multimedia content on the Web has brought many new business opportunities to search engines. In this talk, the author describes our vision of multimedia search and discuss the monetization opportunities and technical challenges in front of us. He argues the importance of combining text, metadata, natural language, content-based analysis, and visualization technologies to provide a total solution that enables seamless user's experience and integration of information sharing, organization, and searching. Some of our techniques are transforming today's conventional text based search engines to include multimedia content thus delivering more intelligent search results to users. In addition to these works, he also presents AdCenter, the online advertising platform from Microsoft, to showcase some of the business opportunities for multimedia search
Hsiao-Wuen Hon
MMM1
2006 Adapting ranking SVM to document retrieval
abstract
The paper is concerned with applying learning to rank to document retrieval. Ranking SVM is a typical method of learning to rank. We point out that there are two factors one must consider when applying Ranking SVM, in general a "learning to rank" method, to document retrieval. First, correctly ranking documents on the top of the result list is crucial for an Information Retrieval system. One must conduct training in a way that such ranked results are accurate. Second, the number of relevant documents can vary from query to query. One must avoid training a model biased toward queries with a large number of relevant documents. Previously, when existing methods that include Ranking SVM were applied to document retrieval, none of the two factors was taken into consideration. We show it is possible to make modifications in conventional Ranking SVM, so it can be better used for document retrieval. Specifically, we modify the "Hinge Loss" function in Ranking SVM to deal with the problems described above. We employ two methods to conduct optimization on the loss function: gradient descent and quadratic programming. Experimental results show that our method, referred to as Ranking SVM for IR, can outperform the conventional Ranking SVM and other existing methods for document retrieval on two datasets.
Yunbo Cao, Jun Xu 0001, Tie-Yan Liu, Hang Li 0001, Yalou Huang, Hsiao-Wuen Hon
SIGIR6
2004 Towards Next Generation Web Information Retrieval
Wei-Ying Ma, HongJiang Zhang, Hsiao-Wuen Hon
WISE3
2002 Distributed speech processing in miPad's multimodal user interface
abstract
This paper describes the main components of MiPad (multimodal interactive PAD) and especially its distributed speech processing aspects. MiPad is a wireless mobile PDA prototype that enables users to accomplish many common tasks using a multimodal spoken language interface and wireless-data technologies. It fully integrates continuous speech recognition and spoken language understanding, and provides a novel solution for data entry in PDAs or smart phones, often done by pecking with tiny styluses or typing on minuscule keyboards. Our user study indicates that the throughput of MiPad is significantly superior to that of the existing pen-based PDA interface. Acoustic modeling and noise robustness in distributed speech recognition are key components in MiPad's design and implementation. In a typical scenario, the user speaks to the device at a distance so that he or she can see the screen. The built-in microphone thus picks up a lot of background noise, which requires MiPad be noise robust. For complex tasks, such as dictating e-mails, resource limitations demand the use of a client-server (peer-to-peer) architecture, where the PDA performs primitive feature extraction, feature quantization, and error protection, while the transmitted features to the server are subject to further speech feature enhancement, speech decoding and understanding before a dialog is carried out and actions rendered. Noise robustness can be achieved at the client, at the server or both. Various speech processing aspects of this type of distributed computation as related to MiPad's potential deployment are presented. Previous user interface study results are also described. Finally, we point out future research directions as related to several key MiPad functionalities.
Li Deng 0001, Kuansan Wang, Alex Acero, Hsiao-Wuen Hon, Jasha Droppo, Constantinos Boulis, Ye-Yi Wang, Derek Jacoby, Milind Mahajan, Ciprian Chelba, Xuedong Huang 0001
IEEE Trans. Speech Audio Process.4
2001 MiPad: a multimodal interaction prototype
abstract
Dr. Who is a Microsoft research project aiming at creating a speech-centric multimodal interaction framework, which serves as the foundation for the NET natural user interface. MiPad is the application prototype that demonstrates compelling user advantages for wireless personal digital assistant (PDA) devices, MiPad fully integrates continuous speech recognition (CSR) and spoken language understanding (SLU) to enable users to accomplish many common tasks using a multimodal interface and wireless technologies. It tries to solve the problem of pecking with tiny styluses or typing on minuscule keyboards in today's PDAs. Unlike a cellular phone, MiPad avoids speech-only interaction. It incorporates a built-in microphone that activates whenever a field is selected. As a user taps the screen or uses a built in roller to navigate, the tapping action narrows the number of possible instructions for spoken word understanding. MiPad currently runs on a Windows CE Pocket PC with a Windows 2000 machine where speech recognition is performed. The Dr Who CSR engine uses a unified CFG and n-gram language model. The Dr Who SLU engine is based on a robust chart parser and a plan-based dialog manager. The paper discusses MiPad's design, implementation work in progress, and preliminary user study in comparison to the existing pen-based PDA interface.
Xuedong Huang 0001, Alex Acero, Ciprian Chelba, Li Deng 0001, Jasha Droppo, Doug Duchene, Joshua Goodman 0001, Hsiao-Wuen Hon, Derek Jacoby, Ricky Loynd, Milind Mahajan, Peter Mau, Scott Meredith, Salman Mughal, Salvado Neto, Mike Plumpe, Kuansan Steury, Gina Venolia, Kuansan Wang, Ye-Yi Wang
ICASSP8
2000 Unified frame and segment based models for automatic speech recognition
abstract
In this paper, we propose an analytically tractable framework that integrates the frame and segment based acoustic modeling techniques. We combine the two approaches by jointly modeling their respective hidden Markov processes. Since the joint process is based on the same mathematical framework, conventional search and training techniques, such as Viterbi and EM algorithms, can be directly applied. It also allows the score from either model to contribute to the training and decoding of the other, reaching a jointly optimal decision. We conducted two series of experiments to verify our hypotheses. In the phone-pair classification experiments, our segment models show a 24% error reduction over state-of-the-art HMM-based system. The superior quality of segment models contributes to an 8.2% reduction in word error rates for the unified system on the WSJ dictation task.
Hsiao-Wuen Hon, Kuansan Wang
ICASSP1
2000 Unifying HMM and phone-pair segment models
Hsiao-Wuen Hon, Shankar Kumar, Kuansan Wang
INTERSPEECH1
2000 Mipad: a next generation PDA prototype
abstract
MiPad is one of the application prototypes in a project codenamed Dr Who. As a wireless Personal Digital Assistant (PDA), MiPad fully integrates continuous speech recognition (CSR) and spoken language understanding (SLU) to enable users to accomplish many common tasks using a multimodal interface and wireless technologies. It tries to solve the problem of pecking with tiny styluses or typing on minuscule keyboards in today’s PDAs or smart phones. It also avoids the problem of being a cellular telephone that depends on speech-only interaction. MiPad incorporates a built-in microphone that activates whenever a field is selected. As a user taps the screen or uses a built-in roller to navigate, the tapping action narrows the number of possible instructions for spoken language processing. MiPad currently runs on a Windows CE Pocket PC with a Windows 2000 Server where speech recognition is performed. The Dr Who CSR engine has a 64k word vocabulary with a unified context-free grammar and n-gram language model. The Dr Who SLU engine is based on a robust chart parser and a plan-based dialog manager. This paper discusses MiPad’s design, implementation work in progress, and preliminary user study in comparison to the existing pen-based PDA interface. 1.
Xuedong Huang 0001, Alex Acero, Ciprian Chelba, Li Deng 0001, Doug Duchene, Joshua Goodman 0001, Hsiao-Wuen Hon, Derek Jacoby, Ricky Loynd, Milind Mahajan, Peter Mau, Scott Meredith, Salman Mughal, Salvado Neto, Mike Plumpe, Kuansan Wang, Ye-Yi Wang
INTERSPEECH7
1998 Automatic generation of synthesis units for trainable text-to-speech systems
abstract
The Whistler text-to-speech engine was designed so that we can automatically construct the model parameters from training data. This paper describes in detail the design issues of constructing the synthesis unit inventory automatically from speech databases. The automatic process includes (1) determining the scaleable synthesis unit which can reflect spectral variations of different allophones; (2) segmenting the recording sentences into phonetic segments; (3) select good instances for each synthesis unit to generate best synthesis sentence during the run time. These processes are all derived through the use of probabilistic learning methods which are aimed at the same optimization criteria. Through this automatic unit generation, Whistler can automatically produce synthetic speech that sounds very natural and resembles the acoustic characteristics of the original speaker.
Hsiao-Wuen Hon, Alex Acero, Xuedong Huang 0001, Jingsong Liu, Mike Plumpe
ICASSP1
1998 Word-based acoustic confidence measures for large-vocabulary speech recognition
abstract
Word level confidence measures are of use in many areas of speech recognition. Comparing the hypothesized word score to the score of a ‘filler’ model has been the most popular confidence measure because it is highly efficient, and does not require a large amount of training data. This paper explores an extension of this technique which also compares the hypothesized word score to the scores of words that are commonly confused for it, while maintaining efficiency and the low demand for training data. The proposed method gives a 39% relative false accept rate reduction over the ‘filler’model baseline, at a false reject rate of 5%.
Asela Gunawardana, Hsiao-Wuen Hon
ICSLP2
1998 Japanese large-vocabulary continuous speech recognition system based on microsoft whisper
abstract
Input of Asian ideographic characters has traditionally been one of the biggest impediments for information processing in Asia. Speech is arguably the most effective and efficient input method for Asian non-spelling characters. This paper presents a Japanese large-vocabulary continuous speech recognition system based on Microsoft Whisper technology. We focus on the aspects of the system that are language specific and demonstrate the adaptability of the Whisper system to new languages. In this paper, we demonstrate that our pronunciation/part-of-speech distinguished morpheme based language models and Whisper based Japanese senonic acoustic models are able to yield state-of-the-art Japanese LVCSR recognition performance. The speaker-independent character and Kana error rates on the JNAS database are 10% and 5% respectively.
Hsiao-Wuen Hon, Yun-Cheng Ju, Keiko Otani
ICSLP1
1998 HMM-based smoothing for concatenative speech synthesis
abstract
This paper will focus on our recent efforts to further improve the acoustic quality of the Whistler Text-to-Speech engine. We have developed an advanced smoothing system that a small pilot study indicates significantly improves quality. We represent speech as being composed of a number of frames, where each frame can be synthesized from a parameter vector. Each frame is represented by a state in an HMM, where the output distribution of each state is a Gaussian random vector consisting of x and Dx. The set of vectors that maximizes the HMM probability is the representation of the smoothed speech output. This technique follows our traditional goal of developing methods whose parameters are automatically learned from data with minimal human intervention. The general framework is demonstrated to be robust by maintaining improved quality with a significant reduction in data. 1. INTRODUCTION In contrast to most Text-To-Speech (TTS) systems (including both formant and concatena...
Mike Plumpe, Alex Acero, Hsiao-Wuen Hon, Xuedong Huang 0001
ICSLP3
1997 Recent improvements on Microsoft's trainable text-to-speech system-Whistler
abstract
The Whistler text-to-speech engine was designed so that we can automatically construct the model parameters from training data. This paper focuses on the improvements on prosody and acoustic modeling, which are all derived through the use of probabilistic learning methods. Whistler can produce synthetic speech that sounds very natural and resembles the acoustic and prosodic characteristics of the original speaker. The underlying technologies used in Whistler can significantly facilitate the process of creating generic TTS systems for a new language, a new voice, or a new speech style. Whisper TTS engine supports Microsoft Speech API and requires less than 3 MB of working memory.
Xuedong Huang 0001, Alex Acero, Hsiao-Wuen Hon, Yun-Cheng Ju, Jingsong Liu, Scott Meredith, Mike Plumpe
ICASSP3
1997 Improvements on a trainable letter-to-sound converter
Hsiao-Wuen Hon, Xuedong Huang 0001
EUROSPEECH2
1996 Whistler: a trainable text-to-speech system
Xuedong Huang 0001, Alex Acero, J. Adcock, Hsiao-Wuen Hon, John Goldsmith, Jingsong Liu, Mike Plumpe
ICSLP4
1995 Tangerine: a large vocabulary Mandarin dictation system
abstract
The text input for non-alphabetic languages, such as Chinese, has been a decades-long problem. Chinese dictation using large vocabulary speech recognition provides a convenient mode of text entry. In contrast to a character based dictation system, a word-based Mandarin dictation system has been designed (based on Apple's PlainTalk speech recognition technology for efficient entry of Chinese characters into a computer. New features and improvements to the dictation system are presented. The new features and improvements have produced an overall reduction in recognition error of 50-80%. The vocabulary has also been increased from 5000 words to over 11000 words.
Hsiao-Wuen Hon, Gareth Loudon, S. Yogananthan, Baosheng Yuan
ICASSP2
1994 Towards large vocabulary Mandarin Chinese speech recognition
abstract
Although commercial dictation products are beginning to emerge for English, the existence of a convenient keyboard has prevented pervasive use of dictation. On the other hand, for non alphabetic languages like Chinese, there is no convenient input method. Therefore, dictation may already be a more appealing input method, for Chinese. In this paper, we demonstrate that our sub-syllable HMM recognizer and tone classifier are able to yield state-of-the-art Mandarin Chinese syllable and tone recognition performance (95.7% for syllables and 98.9% for tones). By combining the HMM syllable recognizer and tone classifier, the tonal syllable result (94%) appears adequate for a syllable base dictation machine. Finally, to alleviate the homophone problem of syllable dictation, we developed a high-performance 5,000-word recognition system with 93% accuracy for the correct answer and 99% accuracy for the top 3 candidates.>
Hsiao-Wuen Hon, Baosheng Yuan, Yen-Lu Chow, Shankar Narayan, Kai-Fu Lee
ICASSP (1)1
1993 The SPHINX-II speech recognition system: an overview
Xuedong Huang 0001, Fil Alleva, Hsiao-Wuen Hon, Mei-Yuh Hwang, Kai-Fu Lee, Ronald Rosenfeld
Comput. Speech Lang.3
1993 A comparative study of discrete, semicontinuous, and continuous hidden Markov models
Xuedong Huang 0001, Hsiao-Wuen Hon, Mei-Yuh Hwang, Kai-Fu Lee
Comput. Speech Lang.2
1992 Vocabulary learning and environment normalization in vocabulary-independent speech recognition
abstract
The authors discuss adaptation issues of vocabulary-independent (VI) systems. Just as with speaker-adaptation in a speaker-independent system, two vocabulary learning algorithms are implemented in order to tailor the VI subword models to the target vocabulary. The first algorithm generates vocabulary-adapted clustering decision trees by focusing on relevant allophones during tree generation and reduces the VI error rate by 9%. The second algorithm, vocabulary-bias training, gives the relevant allophones more prominence by assigning more weight to them during Baum-Welch training of the generalized allophonic models and reduces the VI error rate by 15%. Finally, in order to overcome the degradation caused by the different acoustic environments used for VI training and testing, codebook-dependent cepstral normalization (CDCN) and interpolated SNR-dependent cepstral normalization (ISDCN) originally designed for microphone adaptation are incorporated into the VI system, and both reduce the degradation of VI cross-environment recognition by 50%.>
Hsiao-Wuen Hon, Kai-Fu Lee
ICASSP1
1991 CMU robust vocabulary-independent speech recognition system
abstract
Efforts to improve the performance of CMU's robust vocabulary-independent (VI) speech recognition systems on the DARPA speaker-independent resource management task are discussed. The improvements are evaluated on 320 sentences randomly selected from the DARPA June 88, February 89, and October 89 test sets. The first improvement involves more detailed acoustic modeling. The authors incorporated more dynamic features computed from the LPC cepstra and reduced error by 15% over the baseline system. The second improvement comes from a larger database. With more training data, the third improvement comes from a more detailed subword modeling. The authors incorporated the word boundary context into their VI subword modeling and it resulted in a 30% error reduction. Decision-tree allophone clustering was used to find more suitable models for the subword units not covered in the training set and further reduced error by 17%.>
Hsiao-Wuen Hon, Kai-Fu Lee
ICASSP1
1991 Improved acoustic modeling with the SPHINX speech recognition system
abstract
The authors report recent efforts to further improve the performance of the SPHINX system for speaker-independent continuous speech recognition. They adhere to the basic architecture of the SPHINX system and use the DARPA resource management task and training corpus. The improvements are evaluated on the 600 sentences that comprise the DARPA February and October 1989 test sets. Several techniques that substantially reduced SPHINX's error rate are presented. These techniques include dynamic features, semicontinuous hidden Markov models, speaker clustering, and the shared distribution modeling. The error rate of the baseline system was reduced by 45%.>
Xuedong Huang 0001, Kai-Fu Lee, Hsiao-Wuen Hon, Mei-Yuh Hwang
ICASSP3
1990 On vocabulary-independent speech modeling
abstract
The use of vocabulary-independent (VI) models to improve the usability of speech recognizers is described. Initial results using generalized triphones as VI models show that with more training data and more detailed modeling, the error rate of VI models can be reduced substantially. For example, the error rates for VI models with 5000, 10000, and 15000 training sentences, are 23.9%, 15.2%, and 13.3%, respectively. Moreover, if task-specific training data are available, one can interpolate them with VI models. This task adaptation can reduce the error rate by 18% over task-specifying models.>
Hsiao-Wuen Hon, Kai-Fu Lee
ICASSP1
1990 On semi-continuous hidden Markov modeling
abstract
The semicontinuous hidden Markov model is used in a 1000-word speaker-independent continuous speech recognition system and compared with the continuous mixture model and the discrete model. When the acoustic parameter is not well modeled by the continuous probability density, it is observed that the model assumption problems may cause the recognition accuracy of the semicontinuous model to be inferior to the discrete model. A simple method based on the semicontinuous model is investigated, to re-estimate the vector quantization codebook without continuous probability density function assumptions. Preliminary experiments show that such reestimation methods are as effective as the semicontinuous model, especially when the continuous probability density function assumption is inappropriate.>
Xuedong Huang 0001, Kai-Fu Lee, Hsiao-Wuen Hon
ICASSP3
1990 Allophone clustering for continuous speech recognition
abstract
Two methods are presented for subword clustering. The first method is an agglomerative clustering algorithm. This method is completely data-driven and finds clusters without any external guidance. The second method uses decision trees for clustering. This method uses an expert-generated list of questions about contexts and recursively selects the most appropriate question to split the allophones. Preliminary results showed that when the training set has a good coverage of the allophonic variations in the test set, both method are capable of high-performance recognition. However, under vocabulary-independent conditions, the method using tree-based allophones outperformed agglomerative clustering because of its superior generalization capability.>
Kai-Fu Lee, Satoru Hayamizu, Hsiao-Wuen Hon, Cecil Huang, Jonathan Swartz, Robert Weide
ICASSP3
1990 Description of acoustic variations by tree-based phone modeling
Satoru Hayamizu, Kai-Fu Lee, Hsiao-Wuen Hon
ICSLP3
1990 Speech recognition using hidden Markov models: A CMU perspective
Kai-Fu Lee, Hsiao-Wuen Hon, Mei-Yuh Hwang, Xuedong Huang 0001
Speech Commun.2
1989 The SPHINX speech recognition system
abstract
A description is given of SPHINX an accurate large-vocabulary speaker-independent continuous speech recognition system. The authors have made several recent enhancements, including generalized triphone models, word duration modeling, function-phrase modeling, between-word coarticulation modeling, and corrective training. On the 997-word resource management task, SPHINX attained a word accuracy of 96% with a grammar (perplexity 60), and 82% without grammar (perplexity 997).>
Kai-Fu Lee, Hsiao-Wuen Hon, Mei-Yuh Hwang, Sanjoy Mahajan, Raj Reddy
ICASSP2
1989 Towards speech recognition without vocabulary-specific training
abstract
With the emergence of high-performance speaker-independent systems, a great barrier to man-machine interface has been overcome. This work describes our next step to improve the usability of speech recognizers—the use of vocabulary-independent (VI) models. If successful, VI models are trained once and for all. They will completely eliminate task-specific training, and will enable rapid configuration of speech recognizers for new vocabularies. Our initial results using generalized triphones as VI models show that with more training data and more detailed modeling, the error rate of VI models can be reduced substantially. For example, the error rates for VI models with 5,000, 10,000 and 15,000 training sentences are 23.9%, 15.2% and 13.3% respectively. Moreover, if task-specific training data were available, we can interpolate them with VI models. Our prelimenary results show that this interpolation can lead to an 18% error rate reduction over task-specific models.
Hsiao-Wuen Hon, Kai-Fu Lee, Robert Weide
EUROSPEECH1
1989 Large-vocabulary speaker-independent continuous speech recognition with semi-continuous hidden Markov models
abstract
A semi-continuous hidden Markov model based on the multiple vector quantization codebooks is used here for large-vocabulary speaker-independent continuous speech recognition In the techniques employed here, the semi-continuous output probability density function for each codebook is represented by a combination of the corresponding discrete output probabilities of the hidden Markov model and the continuous Gaussian density functions of each individual codebook. Parameters of vector quantization codebook and hidden Markov model are mutually optimized to achieve an optimal model codebook combination under a unified probabilistic framework Another advantages of this approach is the enhanced robustness of the semi-continuous output probability by the combination of multiple codewords and multiple codebooks For a 1000-word speaker-independent continuous speech recognition using a word-pair grammar the recognition error rate of the semi-continuous hidden Markov model was reduced by more than 29% and 41% in comparison to the discrete and continuous mixture hidden Markov model respectively
Xuedong Huang 0001, Hsiao-Wuen Hon, Kai-Fu Lee
EUROSPEECH2
1989 Modeling between-word coarticulation in continuous speech recognition
Mei-Yuh Hwang, Hsiao-Wuen Hon, Kai-Fu Lee
EUROSPEECH2
1988 Large-vocabulary speaker-independent continuous speech recognition using HMM
abstract
SPHINX, the first large-vocabulary speaker-independent continuous-speech recognizer is described. SPHINX is a hidden-Markov-model (HMM)-based recognizer using multiple codebooks of various LPC-derived features. Two types of HMMs are used in SPHINX: context-independent phone models and function-word-dependent phone models. On a 997-word task using a bigram grammar, SPHINX achieved a word accuracy of 93%. This demonstrates the feasibility of speaker-independent continuous-speech recognition, and the appropriateness of hidden Markov models for such a task.>
Kai-Fu Lee, Hsiao-Wuen Hon
ICASSP2