EDBT 2026 Demo / reviewers in the wild / expert
Hsiao-Wuen Hon
dblp:50/4528
· DBLP profile ↗
49ranked-venue papers
11as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-authorArtificial intelligence and machine learning · 20 · 3 first-authorDatabases, data management, data science and information retrieval · 12Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 64% Representation and self-supervised learning · 26% Question answering and dialogue systems · 4% | |
| Databases, data mining, and information retrieval
8 papers |
Information retrieval · 76% Web and social media mining · 18% Recommender systems · 5% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 62% Audio and music processing · 19% Image and video processing · 19% |
Topics — the 27 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
pre-trained language model |
0.8 | 2 | 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training · ICML 2020 Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019 |
Machine learning › Representation and self-supervised learning › pre-training
unified pre-training |
0.8 | 2 | 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training · ICML 2020 Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019 |
Natural language and speech › Language models and text generation
masked language modeling |
0.4 | 1 | 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training · ICML 2020 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.4 | 1 | 2019 | Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2019 | Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019 |
Information retrieval
cross-language information retrieval |
0.2 | 2 | 2010 | Exploiting query logs for cross-lingual query suggestions · ACM Trans. Inf. Syst. 2010 Cross-lingual query suggestion using query logs of different languages · SIGIR 2007 |
Information retrieval
query suggestion |
0.2 | 2 | 2010 | Exploiting query logs for cross-lingual query suggestions · ACM Trans. Inf. Syst. 2010 Cross-lingual query suggestion using query logs of different languages · SIGIR 2007 |
Information retrieval › question answering
community question answering |
0.2 | 1 | 2013 | Question Difficulty Estimation in Community Question Answering Services · EMNLP 2013 |
Information retrieval
question answering |
0.2 | 1 | 2013 | Question Difficulty Estimation in Community Question Answering Services · EMNLP 2013 |
Web and social media mining
user identity linkage |
0.2 | 1 | 2013 | What's in a name?: an unsupervised approach to link users across communities · WSDM 2013 |
Natural language and speech › Question answering and dialogue systems › answer generation
generative question answering |
0.1 | 1 | 2019 | Unified Language Model Pre-training for Natural Language Understanding and Generation · NeurIPS 2019 |
Recommender systems › content recommendation
question recommendation |
0.1 | 1 | 2008 | Recommending questions using the mdl-based tree cut model · WWW 2008 |
Information retrieval
ranking |
0.1 | 1 | 2008 | Recommending questions using the mdl-based tree cut model · WWW 2008 |
Information retrieval › query understanding
query ambiguity |
0.1 | 1 | 2007 | Identifying ambiguous queries in web search · WWW 2007 |
Information retrieval › query understanding
query analysis |
0.1 | 1 | 2007 | Identifying ambiguous queries in web search · WWW 2007 |
Information retrieval › cross-language information retrieval
query translation |
0.1 | 1 | 2007 | Cross-lingual query suggestion using query logs of different languages · SIGIR 2007 |
Information retrieval
query understanding |
0.1 | 1 | 2007 | Identifying ambiguous queries in web search · WWW 2007 |
Web and social media mining › web mining
web page understanding |
0.1 | 1 | 2007 | Webpage understanding: an integrated approach · KDD 2007 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2006 | Adapting ranking SVM to document retrieval · SIGIR 2006 |
Information retrieval
query log analysis |
0.1 | 2 | 2010 | Exploiting query logs for cross-lingual query suggestions · ACM Trans. Inf. Syst. 2010 Identifying ambiguous queries in web search · WWW 2007 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
distributed speech recognition |
0.0 | 1 | 2002 | Distributed speech processing in miPad's multimodal user interface · IEEE Trans. Speech Audio Process. 2002 |
Natural language and speech › Speech recognition and synthesis
noise robustness |
0.0 | 1 | 2002 | Distributed speech processing in miPad's multimodal user interface · IEEE Trans. Speech Audio Process. 2002 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition |
0.0 | 1 | 2002 | Distributed speech processing in miPad's multimodal user interface · IEEE Trans. Speech Audio Process. 2002 |
Audio and music processing
acoustic signal processing |
0.0 | 1 | 2009 | Mobile media search: has media search finally found its perfect platform? part II · ACM Multimedia 2009 |
Interaction techniques and input
mobile interaction |
0.0 | 1 | 2009 | Mobile media search: has media search finally found its perfect platform? part II · ACM Multimedia 2009 |
Computational finance and economics › online advertising
search advertising |
0.0 | 1 | 2007 | Cross-lingual query suggestion using query logs of different languages · SIGIR 2007 |
Information retrieval
document retrieval |
0.0 | 1 | 2006 | Adapting ranking SVM to document retrieval · SIGIR 2006 |
Methods — techniques the papers use, named apart from their topics
self-attention · 0.4partially autoregressive modeling · 0.4autoencoding · 0.4transformer · 0.4sequence-to-sequence pre-training · 0.4self-attention masking · 0.4discriminative model · 0.3unsupervised learning · 0.2pairwise comparison · 0.2pagerank · 0.2n-gram probability · 0.2competition-based model · 0.2discriminative probabilistic model · 0.1pseudo-relevance feedback · 0.1tree cut model · 0.1minimum description length · 0.1word co-occurrence statistics · 0.1speech feature enhancement · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingabstractWe propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with masked tokens, we rely on conventional masks to learn inter-relations between corrupted tokens and context via autoencoding, and pseudo masks to learn intra-relations between masked spans via partially autoregressive modeling. With well-designed position embeddings and self-attention masks, the context encodings are reused to avoid redundant computation. Moreover, conventional masks used for autoencoding provide global masking information, so that all the position embeddings are accessible in partially autoregressive language modeling. In addition, the two tasks pre-train a unified language model as a bidirectional encoder and a sequence-to-sequence decoder, respectively. Our experiments show that the unified language models pre-trained using PMLM achieve new state-of-the-art results on a wide range of language understanding and generation tasks across several widely used benchmarks. The code and pre-trained models are available at https://github.com/microsoft/unilm. Hangbo Bao, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Nan Yang 0002, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
ICML | 11 |
| 2019 | A Brief History of IntelligenceabstractIntelligence is the deciding factor of how human beings become the most dominant life forms on earth. Throughout history, human beings have developed tools and technologies which help civilizations evolve and grow. Computers, and by extension, artificial intelligence (AI), has played important roles in that continuum of technologies. Recently artificial intelligence has garnered much interest and discussion. As artificial intelligence are tools that can enhance human capability, a sound understanding of what the technology can and cannot do is also necessary to ensure their appropriate use. While developing artificial intelligence, we also found out the definition and understanding of our own human intelligence continue evolving. The debates of the race between human and artificial intelligence have been ever growing. In this talk, I will describe the history of both artificial intelligence and human intelligence (HI). From the great insights of the such historical perspectives, I would like to illustrate how AI and HI will co-evolve with each other and project the future of AI and HI. Hsiao-Wuen Hon |
ICMI | 1 |
| 2019 | Unified Language Model Pre-training for Natural Language Understanding and GenerationabstractThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at https://github.com/microsoft/unilm. Li Dong 0004, Nan Yang 0002, Wenhui Wang 0003, Furu Wei, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
NeurIPS | 9 |
| 2013 | Question Difficulty Estimation in Community Question Answering ServicesabstractIn this paper, we address the problem of estimating question difficulty in community question answering services.We propose a competition-based model for estimating question difficulty by leveraging pairwise comparisons between questions and users.Our experimental results show that our model significantly outperforms a PageRank-based approach.Most importantly, our analysis shows that the text of question descriptions reflects the question difficulty.This implies the possibility of predicting question difficulty from the text of question descriptions. Jing Liu 0022, Quan Wang 0002, Chin-Yew Lin, Hsiao-Wuen Hon |
EMNLP | 4 |
| 2013 | What's in a name?: an unsupervised approach to link users across communitiesabstractIn this paper, we consider the problem of linking users across multiple online communities. Specifically, we focus on the alias-disambiguation step of this user linking task, which is meant to differentiate users with the same usernames. We start quantitatively analyzing the importance of the alias-disambiguation step by conducting a survey on 153 volunteers and an experimental analysis on a large dataset of About.me (75,472 users). The analysis shows that the alias-disambiguation solution can address a major part of the user linking problem in terms of the coverage of true pairwise decisions (46.8%). To the best of our knowledge, this is the first study on human behaviors with regards to the usages of online usernames. We then cast the alias-disambiguation step as a pairwise classification problem and propose a novel unsupervised approach. The key idea of our approach is to automatically label training instances based on two observations: (a) rare usernames are likely owned by a single natural person, e.g. pennystar88 as a positive instance; (b) common usernames are likely owned by different natural persons, e.g. tank as a negative instance. We propose using the n-gram probabilities of usernames to estimate the rareness or commonness of usernames. Moreover, these two observations are verified by using the dataset of Yahoo! Answers. The empirical evaluations on 53 forums verify: (a) the effectiveness of the classifiers with the automatically generated training data and (b) that the rareness and commonness of usernames can help user linking. We also analyze the cases where the classifiers fail. Jing Liu 0022, Fan Zhang 0092, Xinying Song, Young-In Song, Chin-Yew Lin, Hsiao-Wuen Hon |
WSDM | 6 |
| 2011 | Select-the-Best-Ones: A new way to judge relative relevance
Ruihua Song, Qingwei Guo, Ruochi Zhang, Guomao Xin, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon |
Inf. Process. Manag. | 7 |
| 2010 | Automatic extraction of web data records containing user-generated contentabstractIn this paper, we are concerned with the problem of automatically extracting web data records that contain user-generated content (UGC). In previous work, web data records are usually assumed to be well-formed with a limited amount of UGC, and thus can be extracted by testing repetitive structure similarity. However, when a web data record includes a large portion of free-format UGC, the similarity test between records may fail, which in turn results in lower performance. In our work, we find that certain domain constraints (e.g., post-date) can be used to design better similarity measures capable of circumventing the influence of UGC. In addition, we also use anchor points provided by the domain constraints to improve the extraction process, which ends in an algorithm called MiBAT (Mining data records Based on Anchor Trees). We conduct extensive experiments on a dataset consisting of forum thread pages which are collected from 307 sites that cover 219 different forum software packages. Our approach achieves a precision of 98.9% and a recall of 97.3% with respect to post record extraction. On page level, it perfectly handles 91.7% of pages without extracting any wrong posts or missing any golden posts. We also apply our approach to comment extraction and achieve good results as well. Xinying Song, Jing Liu 0022, Yunbo Cao, Chin-Yew Lin, Hsiao-Wuen Hon |
CIKM | 5 |
| 2010 | Learning Query Ambiguity Models by Using Search Logs
Ruihua Song, Zhicheng Dou, Hsiao-Wuen Hon, Yong Yu 0001 |
J. Comput. Sci. Technol. | 3 |
| 2010 | Exploiting query logs for cross-lingual query suggestionsabstractQuery suggestion aims to suggest relevant queries for a given query, which helps users better specify their information needs. Previous work on query suggestion has been limited to the same language. In this article, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to the scenarios of cross-language information retrieval (CLIR) and other related cross-lingual applications. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, and so on, are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly outperforms a baseline system that uses dictionary-based query translation. Besides, we evaluate CLQS with French-English and Chinese-English CLIR tasks on TREC-6 and NTCIR-4 collections, respectively. The CLIR experiments using typical retrieval models demonstrate that the CLQS-based approach has significantly higher effectiveness than several traditional query translation methods. We find that when combined with pseudo-relevance feedback, the effectiveness of CLIR using CLQS is enhanced for different pairs of languages. Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon |
ACM Trans. Inf. Syst. | 6 |
| 2009 | Mobile media searchabstractThis panel paper presents motivations for discussing mobile media search and contains statements from the panelists who are industry research leaders in this field. Berna Erol, Jordan Cohen, Minoru Etoh, Hsiao-Wuen Hon, Jiebo Luo 0001, Johan Schalkwyk |
ICASSP | 4 |
| 2009 | Mobile media search: has media search finally found its perfect platform? part IIabstractRecently, many exciting media search applications have been introduced to take advantage of smart phones' audiovisual capture capabilities and their being always on and connected. These applications address a real pain point for most mobile users and allow them to search with minimal text entry, if any. Is the mobile platform an ideal fit for media search? Are audio and visual signal processing technologies sufficiently accurate to support most mobile search applications? What are the killer applications of mobile media search? Earlier in 2009 at ICASSP, a panel on this topic stirred up great interest and enthusiasm while leaving many questions untouched due to the limited time. Berna Erol, Jiebo Luo 0001, Shih-Fu Chang, Minoru Etoh, Hsiao-Wuen Hon, Qian Lin 0001, Vidya Setlur |
ACM Multimedia | 5 |
| 2009 | Identification of ambiguous queries in web search
Ruihua Song, Zhenxiao Luo, Jian-Yun Nie, Yong Yu 0001, Hsiao-Wuen Hon |
Inf. Process. Manag. | 5 |
| 2008 | Viewing Term Proximity from a Different Perspective
Ruihua Song, Michael J. Taylor 0001, Ji-Rong Wen, Hsiao-Wuen Hon, Yong Yu 0001 |
ECIR | 4 |
| 2008 | Recommending questions using the mdl-based tree cut modelabstractThe paper is concerned with the problem of question recommendation. Specifically, given a question as query, we are to retrieve and rank other questions according to their likelihood of being good recommendations of the queried question. A good recommendation provides alternative aspects around users' interest. We tackle the problem of question recommendation in two steps: first represent questions as graphs of topic terms, and then rank recommendations on the basis of the graphs. We formalize both steps as the tree-cutting problems and then employ the MDL (Minimum Description Length) for selecting the best cuts. Experiments have been conducted with the real questions posted at Yahoo! Answers. The questions are about two domains, 'travel' and 'computers & internet'. Experimental results indicate that the use of the MDL-based tree cut model can significantly outperform the baseline methods of word-based VSM or phrase-based VSM. The results also show that the use of the MDL-based tree cut model is essential to our approach. Yunbo Cao, Huizhong Duan, Chin-Yew Lin, Yong Yu 0001, Hsiao-Wuen Hon |
WWW | 5 |
| 2007 | Webpage understanding: an integrated approachabstractRecent work has shown the effectiveness of leveraging layout and tag-tree structure for segmenting webpages and labeling HTML elements. However, how to effectively segment and label the text contents inside HTML elements is still an open problem. Since many text contents on a webpage are often text fragments and not strictly grammatical, traditional natural language processing techniques, that typically expect grammatical sentences, are no longer directly applicable. In this paper, we examine how to use layout and tag-tree structure in a principled way to help understand text contents on webpages. We propose to segment and label the page structure and the text content of a webpage in a joint discriminative probabilistic model. In this model, semantic labels of page structure can be leveraged to help text content understanding, and semantic labels ofthe text phrases can be used in page structure understanding tasks such as data record detection. Thus, integration of both page structure and text content understanding leads to an integrated solution of webpage understanding. Experimental results on research homepage extraction show the feasibility and promise of our approach. Jun Zhu 0001, Bo Zhang 0010, Zaiqing Nie, Ji-Rong Wen, Hsiao-Wuen Hon |
KDD | 5 |
| 2007 | Cross-lingual query suggestion using query logs of different languagesabstractQuery suggestion aims to suggest relevant queries for a given query, which help users better specify their information needs. Previously, the suggested terms are mostly in the same language of the input query. In this paper, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to scenarios of cross-language information retrieval (CLIR) and cross-lingual keyword bidding for search engine advertisement. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, etc. are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly out performs a baseline system based on dictionary-based query translation. Besides, the resulting CLQS is tested with French to English CLIR tasks on TREC collections. The results demonstrate higher effectiveness than the traditional query translation methods. Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon |
SIGIR | 7 |
| 2007 | Identifying ambiguous queries in web searchabstractIt is widely believed that some queries submitted to search engines are by nature ambiguous (e.g., java, apple). However, few studies have investigated the questions of "how many queries are ambiguous?" and "how can we automatically identify an ambiguous query?" This paper deals with these issues. First, we construct the taxonomy of query ambiguity, and ask human annotators to manually classify queries based upon it. From manually labeled results, we find that query ambiguity is to some extent predictable. We then use a supervised learning approach to automatically classify queries as being ambiguous or not. Experimental results show that we can correctly identify 87% of labeled queries. Finally, we estimate that about 16% of queries in a real search log are ambiguous. Ruihua Song, Zhenxiao Luo, Ji-Rong Wen, Yong Yu 0001, Hsiao-Wuen Hon |
WWW | 5 |
| 2006 | Multimedia search - Microsoft's perspectiveabstractSummary form only given. The explosive growth of multimedia content on the Web has brought many new business opportunities to search engines. In this talk, the author describes our vision of multimedia search and discuss the monetization opportunities and technical challenges in front of us. He argues the importance of combining text, metadata, natural language, content-based analysis, and visualization technologies to provide a total solution that enables seamless user's experience and integration of information sharing, organization, and searching. Some of our techniques are transforming today's conventional text based search engines to include multimedia content thus delivering more intelligent search results to users. In addition to these works, he also presents AdCenter, the online advertising platform from Microsoft, to showcase some of the business opportunities for multimedia search Hsiao-Wuen Hon |
MMM | 1 |
| 2006 | Adapting ranking SVM to document retrievalabstractThe paper is concerned with applying learning to rank to document retrieval. Ranking SVM is a typical method of learning to rank. We point out that there are two factors one must consider when applying Ranking SVM, in general a "learning to rank" method, to document retrieval. First, correctly ranking documents on the top of the result list is crucial for an Information Retrieval system. One must conduct training in a way that such ranked results are accurate. Second, the number of relevant documents can vary from query to query. One must avoid training a model biased toward queries with a large number of relevant documents. Previously, when existing methods that include Ranking SVM were applied to document retrieval, none of the two factors was taken into consideration. We show it is possible to make modifications in conventional Ranking SVM, so it can be better used for document retrieval. Specifically, we modify the "Hinge Loss" function in Ranking SVM to deal with the problems described above. We employ two methods to conduct optimization on the loss function: gradient descent and quadratic programming. Experimental results show that our method, referred to as Ranking SVM for IR, can outperform the conventional Ranking SVM and other existing methods for document retrieval on two datasets. Yunbo Cao, Jun Xu 0001, Tie-Yan Liu, Hang Li 0001, Yalou Huang, Hsiao-Wuen Hon |
SIGIR | 6 |
| 2004 | Towards Next Generation Web Information Retrieval
Wei-Ying Ma, HongJiang Zhang, Hsiao-Wuen Hon |
WISE | 3 |
| 2002 | Distributed speech processing in miPad's multimodal user interfaceabstractThis paper describes the main components of MiPad (multimodal interactive PAD) and especially its distributed speech processing aspects. MiPad is a wireless mobile PDA prototype that enables users to accomplish many common tasks using a multimodal spoken language interface and wireless-data technologies. It fully integrates continuous speech recognition and spoken language understanding, and provides a novel solution for data entry in PDAs or smart phones, often done by pecking with tiny styluses or typing on minuscule keyboards. Our user study indicates that the throughput of MiPad is significantly superior to that of the existing pen-based PDA interface. Acoustic modeling and noise robustness in distributed speech recognition are key components in MiPad's design and implementation. In a typical scenario, the user speaks to the device at a distance so that he or she can see the screen. The built-in microphone thus picks up a lot of background noise, which requires MiPad be noise robust. For complex tasks, such as dictating e-mails, resource limitations demand the use of a client-server (peer-to-peer) architecture, where the PDA performs primitive feature extraction, feature quantization, and error protection, while the transmitted features to the server are subject to further speech feature enhancement, speech decoding and understanding before a dialog is carried out and actions rendered. Noise robustness can be achieved at the client, at the server or both. Various speech processing aspects of this type of distributed computation as related to MiPad's potential deployment are presented. Previous user interface study results are also described. Finally, we point out future research directions as related to several key MiPad functionalities. Li Deng 0001, Kuansan Wang, Alex Acero, Hsiao-Wuen Hon, Jasha Droppo, Constantinos Boulis, Ye-Yi Wang, Derek Jacoby, Milind Mahajan, Ciprian Chelba, Xuedong Huang 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2001 | MiPad: a multimodal interaction prototypeabstractDr. Who is a Microsoft research project aiming at creating a speech-centric multimodal interaction framework, which serves as the foundation for the NET natural user interface. MiPad is the application prototype that demonstrates compelling user advantages for wireless personal digital assistant (PDA) devices, MiPad fully integrates continuous speech recognition (CSR) and spoken language understanding (SLU) to enable users to accomplish many common tasks using a multimodal interface and wireless technologies. It tries to solve the problem of pecking with tiny styluses or typing on minuscule keyboards in today's PDAs. Unlike a cellular phone, MiPad avoids speech-only interaction. It incorporates a built-in microphone that activates whenever a field is selected. As a user taps the screen or uses a built in roller to navigate, the tapping action narrows the number of possible instructions for spoken word understanding. MiPad currently runs on a Windows CE Pocket PC with a Windows 2000 machine where speech recognition is performed. The Dr Who CSR engine uses a unified CFG and n-gram language model. The Dr Who SLU engine is based on a robust chart parser and a plan-based dialog manager. The paper discusses MiPad's design, implementation work in progress, and preliminary user study in comparison to the existing pen-based PDA interface. Xuedong Huang 0001, Alex Acero, Ciprian Chelba, Li Deng 0001, Jasha Droppo, Doug Duchene, Joshua Goodman 0001, Hsiao-Wuen Hon, Derek Jacoby, Ricky Loynd, Milind Mahajan, Peter Mau, Scott Meredith, Salman Mughal, Salvado Neto, Mike Plumpe, Kuansan Steury, Gina Venolia, Kuansan Wang, Ye-Yi Wang |
ICASSP | 8 |
| 2000 | Unified frame and segment based models for automatic speech recognitionabstractIn this paper, we propose an analytically tractable framework that integrates the frame and segment based acoustic modeling techniques. We combine the two approaches by jointly modeling their respective hidden Markov processes. Since the joint process is based on the same mathematical framework, conventional search and training techniques, such as Viterbi and EM algorithms, can be directly applied. It also allows the score from either model to contribute to the training and decoding of the other, reaching a jointly optimal decision. We conducted two series of experiments to verify our hypotheses. In the phone-pair classification experiments, our segment models show a 24% error reduction over state-of-the-art HMM-based system. The superior quality of segment models contributes to an 8.2% reduction in word error rates for the unified system on the WSJ dictation task. Hsiao-Wuen Hon, Kuansan Wang |
ICASSP | 1 |
| 2000 | Unifying HMM and phone-pair segment models
Hsiao-Wuen Hon, Shankar Kumar, Kuansan Wang |
INTERSPEECH | 1 |
| 2000 | Mipad: a next generation PDA prototypeabstractMiPad is one of the application prototypes in a project codenamed Dr Who. As a wireless Personal Digital Assistant (PDA), MiPad fully integrates continuous speech recognition (CSR) and spoken language understanding (SLU) to enable users to accomplish many common tasks using a multimodal interface and wireless technologies. It tries to solve the problem of pecking with tiny styluses or typing on minuscule keyboards in today’s PDAs or smart phones. It also avoids the problem of being a cellular telephone that depends on speech-only interaction. MiPad incorporates a built-in microphone that activates whenever a field is selected. As a user taps the screen or uses a built-in roller to navigate, the tapping action narrows the number of possible instructions for spoken language processing. MiPad currently runs on a Windows CE Pocket PC with a Windows 2000 Server where speech recognition is performed. The Dr Who CSR engine has a 64k word vocabulary with a unified context-free grammar and n-gram language model. The Dr Who SLU engine is based on a robust chart parser and a plan-based dialog manager. This paper discusses MiPad’s design, implementation work in progress, and preliminary user study in comparison to the existing pen-based PDA interface. 1. Xuedong Huang 0001, Alex Acero, Ciprian Chelba, Li Deng 0001, Doug Duchene, Joshua Goodman 0001, Hsiao-Wuen Hon, Derek Jacoby, Ricky Loynd, Milind Mahajan, Peter Mau, Scott Meredith, Salman Mughal, Salvado Neto, Mike Plumpe, Kuansan Wang, Ye-Yi Wang |
INTERSPEECH | 7 |
| 1998 | Automatic generation of synthesis units for trainable text-to-speech systemsabstractThe Whistler text-to-speech engine was designed so that we can automatically construct the model parameters from training data. This paper describes in detail the design issues of constructing the synthesis unit inventory automatically from speech databases. The automatic process includes (1) determining the scaleable synthesis unit which can reflect spectral variations of different allophones; (2) segmenting the recording sentences into phonetic segments; (3) select good instances for each synthesis unit to generate best synthesis sentence during the run time. These processes are all derived through the use of probabilistic learning methods which are aimed at the same optimization criteria. Through this automatic unit generation, Whistler can automatically produce synthetic speech that sounds very natural and resembles the acoustic characteristics of the original speaker. Hsiao-Wuen Hon, Alex Acero, Xuedong Huang 0001, Jingsong Liu, Mike Plumpe |
ICASSP | 1 |
| 1998 | Word-based acoustic confidence measures for large-vocabulary speech recognitionabstractWord level confidence measures are of use in many areas of speech recognition. Comparing the hypothesized word score to the score of a ‘filler’ model has been the most popular confidence measure because it is highly efficient, and does not require a large amount of training data. This paper explores an extension of this technique which also compares the hypothesized word score to the scores of words that are commonly confused for it, while maintaining efficiency and the low demand for training data. The proposed method gives a 39% relative false accept rate reduction over the ‘filler’model baseline, at a false reject rate of 5%. Asela Gunawardana, Hsiao-Wuen Hon |
ICSLP | 2 |
| 1998 | Japanese large-vocabulary continuous speech recognition system based on microsoft whisperabstractInput of Asian ideographic characters has traditionally been one of the biggest impediments for information processing in Asia. Speech is arguably the most effective and efficient input method for Asian non-spelling characters. This paper presents a Japanese large-vocabulary continuous speech recognition system based on Microsoft Whisper technology. We focus on the aspects of the system that are language specific and demonstrate the adaptability of the Whisper system to new languages. In this paper, we demonstrate that our pronunciation/part-of-speech distinguished morpheme based language models and Whisper based Japanese senonic acoustic models are able to yield state-of-the-art Japanese LVCSR recognition performance. The speaker-independent character and Kana error rates on the JNAS database are 10% and 5% respectively. Hsiao-Wuen Hon, Yun-Cheng Ju, Keiko Otani |
ICSLP | 1 |
| 1998 | HMM-based smoothing for concatenative speech synthesisabstractThis paper will focus on our recent efforts to further improve the acoustic quality of the Whistler Text-to-Speech engine. We have developed an advanced smoothing system that a small pilot study indicates significantly improves quality. We represent speech as being composed of a number of frames, where each frame can be synthesized from a parameter vector. Each frame is represented by a state in an HMM, where the output distribution of each state is a Gaussian random vector consisting of x and Dx. The set of vectors that maximizes the HMM probability is the representation of the smoothed speech output. This technique follows our traditional goal of developing methods whose parameters are automatically learned from data with minimal human intervention. The general framework is demonstrated to be robust by maintaining improved quality with a significant reduction in data. 1. INTRODUCTION In contrast to most Text-To-Speech (TTS) systems (including both formant and concatena... Mike Plumpe, Alex Acero, Hsiao-Wuen Hon, Xuedong Huang 0001 |
ICSLP | 3 |
| 1997 | Recent improvements on Microsoft's trainable text-to-speech system-WhistlerabstractThe Whistler text-to-speech engine was designed so that we can automatically construct the model parameters from training data. This paper focuses on the improvements on prosody and acoustic modeling, which are all derived through the use of probabilistic learning methods. Whistler can produce synthetic speech that sounds very natural and resembles the acoustic and prosodic characteristics of the original speaker. The underlying technologies used in Whistler can significantly facilitate the process of creating generic TTS systems for a new language, a new voice, or a new speech style. Whisper TTS engine supports Microsoft Speech API and requires less than 3 MB of working memory. Xuedong Huang 0001, Alex Acero, Hsiao-Wuen Hon, Yun-Cheng Ju, Jingsong Liu, Scott Meredith, Mike Plumpe |
ICASSP | 3 |
| 1997 | Improvements on a trainable letter-to-sound converter
Hsiao-Wuen Hon, Xuedong Huang 0001 |
EUROSPEECH | 2 |
| 1996 | Whistler: a trainable text-to-speech system
Xuedong Huang 0001, Alex Acero, J. Adcock, Hsiao-Wuen Hon, John Goldsmith, Jingsong Liu, Mike Plumpe |
ICSLP | 4 |
| 1995 | Tangerine: a large vocabulary Mandarin dictation systemabstractThe text input for non-alphabetic languages, such as Chinese, has been a decades-long problem. Chinese dictation using large vocabulary speech recognition provides a convenient mode of text entry. In contrast to a character based dictation system, a word-based Mandarin dictation system has been designed (based on Apple's PlainTalk speech recognition technology for efficient entry of Chinese characters into a computer. New features and improvements to the dictation system are presented. The new features and improvements have produced an overall reduction in recognition error of 50-80%. The vocabulary has also been increased from 5000 words to over 11000 words. Hsiao-Wuen Hon, Gareth Loudon, S. Yogananthan, Baosheng Yuan |
ICASSP | 2 |
| 1994 | Towards large vocabulary Mandarin Chinese speech recognitionabstractAlthough commercial dictation products are beginning to emerge for English, the existence of a convenient keyboard has prevented pervasive use of dictation. On the other hand, for non alphabetic languages like Chinese, there is no convenient input method. Therefore, dictation may already be a more appealing input method, for Chinese. In this paper, we demonstrate that our sub-syllable HMM recognizer and tone classifier are able to yield state-of-the-art Mandarin Chinese syllable and tone recognition performance (95.7% for syllables and 98.9% for tones). By combining the HMM syllable recognizer and tone classifier, the tonal syllable result (94%) appears adequate for a syllable base dictation machine. Finally, to alleviate the homophone problem of syllable dictation, we developed a high-performance 5,000-word recognition system with 93% accuracy for the correct answer and 99% accuracy for the top 3 candidates.> Hsiao-Wuen Hon, Baosheng Yuan, Yen-Lu Chow, Shankar Narayan, Kai-Fu Lee |
ICASSP (1) | 1 |
| 1993 | The SPHINX-II speech recognition system: an overview
Xuedong Huang 0001, Fil Alleva, Hsiao-Wuen Hon, Mei-Yuh Hwang, Kai-Fu Lee, Ronald Rosenfeld |
Comput. Speech Lang. | 3 |
| 1993 | A comparative study of discrete, semicontinuous, and continuous hidden Markov models
Xuedong Huang 0001, Hsiao-Wuen Hon, Mei-Yuh Hwang, Kai-Fu Lee |
Comput. Speech Lang. | 2 |
| 1992 | Vocabulary learning and environment normalization in vocabulary-independent speech recognitionabstractThe authors discuss adaptation issues of vocabulary-independent (VI) systems. Just as with speaker-adaptation in a speaker-independent system, two vocabulary learning algorithms are implemented in order to tailor the VI subword models to the target vocabulary. The first algorithm generates vocabulary-adapted clustering decision trees by focusing on relevant allophones during tree generation and reduces the VI error rate by 9%. The second algorithm, vocabulary-bias training, gives the relevant allophones more prominence by assigning more weight to them during Baum-Welch training of the generalized allophonic models and reduces the VI error rate by 15%. Finally, in order to overcome the degradation caused by the different acoustic environments used for VI training and testing, codebook-dependent cepstral normalization (CDCN) and interpolated SNR-dependent cepstral normalization (ISDCN) originally designed for microphone adaptation are incorporated into the VI system, and both reduce the degradation of VI cross-environment recognition by 50%.> Hsiao-Wuen Hon, Kai-Fu Lee |
ICASSP | 1 |
| 1991 | CMU robust vocabulary-independent speech recognition systemabstractEfforts to improve the performance of CMU's robust vocabulary-independent (VI) speech recognition systems on the DARPA speaker-independent resource management task are discussed. The improvements are evaluated on 320 sentences randomly selected from the DARPA June 88, February 89, and October 89 test sets. The first improvement involves more detailed acoustic modeling. The authors incorporated more dynamic features computed from the LPC cepstra and reduced error by 15% over the baseline system. The second improvement comes from a larger database. With more training data, the third improvement comes from a more detailed subword modeling. The authors incorporated the word boundary context into their VI subword modeling and it resulted in a 30% error reduction. Decision-tree allophone clustering was used to find more suitable models for the subword units not covered in the training set and further reduced error by 17%.> Hsiao-Wuen Hon, Kai-Fu Lee |
ICASSP | 1 |
| 1991 | Improved acoustic modeling with the SPHINX speech recognition systemabstractThe authors report recent efforts to further improve the performance of the SPHINX system for speaker-independent continuous speech recognition. They adhere to the basic architecture of the SPHINX system and use the DARPA resource management task and training corpus. The improvements are evaluated on the 600 sentences that comprise the DARPA February and October 1989 test sets. Several techniques that substantially reduced SPHINX's error rate are presented. These techniques include dynamic features, semicontinuous hidden Markov models, speaker clustering, and the shared distribution modeling. The error rate of the baseline system was reduced by 45%.> Xuedong Huang 0001, Kai-Fu Lee, Hsiao-Wuen Hon, Mei-Yuh Hwang |
ICASSP | 3 |
| 1990 | On vocabulary-independent speech modelingabstractThe use of vocabulary-independent (VI) models to improve the usability of speech recognizers is described. Initial results using generalized triphones as VI models show that with more training data and more detailed modeling, the error rate of VI models can be reduced substantially. For example, the error rates for VI models with 5000, 10000, and 15000 training sentences, are 23.9%, 15.2%, and 13.3%, respectively. Moreover, if task-specific training data are available, one can interpolate them with VI models. This task adaptation can reduce the error rate by 18% over task-specifying models.> Hsiao-Wuen Hon, Kai-Fu Lee |
ICASSP | 1 |
| 1990 | On semi-continuous hidden Markov modelingabstractThe semicontinuous hidden Markov model is used in a 1000-word speaker-independent continuous speech recognition system and compared with the continuous mixture model and the discrete model. When the acoustic parameter is not well modeled by the continuous probability density, it is observed that the model assumption problems may cause the recognition accuracy of the semicontinuous model to be inferior to the discrete model. A simple method based on the semicontinuous model is investigated, to re-estimate the vector quantization codebook without continuous probability density function assumptions. Preliminary experiments show that such reestimation methods are as effective as the semicontinuous model, especially when the continuous probability density function assumption is inappropriate.> Xuedong Huang 0001, Kai-Fu Lee, Hsiao-Wuen Hon |
ICASSP | 3 |
| 1990 | Allophone clustering for continuous speech recognitionabstractTwo methods are presented for subword clustering. The first method is an agglomerative clustering algorithm. This method is completely data-driven and finds clusters without any external guidance. The second method uses decision trees for clustering. This method uses an expert-generated list of questions about contexts and recursively selects the most appropriate question to split the allophones. Preliminary results showed that when the training set has a good coverage of the allophonic variations in the test set, both method are capable of high-performance recognition. However, under vocabulary-independent conditions, the method using tree-based allophones outperformed agglomerative clustering because of its superior generalization capability.> Kai-Fu Lee, Satoru Hayamizu, Hsiao-Wuen Hon, Cecil Huang, Jonathan Swartz, Robert Weide |
ICASSP | 3 |
| 1990 | Description of acoustic variations by tree-based phone modeling
Satoru Hayamizu, Kai-Fu Lee, Hsiao-Wuen Hon |
ICSLP | 3 |
| 1990 | Speech recognition using hidden Markov models: A CMU perspective
Kai-Fu Lee, Hsiao-Wuen Hon, Mei-Yuh Hwang, Xuedong Huang 0001 |
Speech Commun. | 2 |
| 1989 | The SPHINX speech recognition systemabstractA description is given of SPHINX an accurate large-vocabulary speaker-independent continuous speech recognition system. The authors have made several recent enhancements, including generalized triphone models, word duration modeling, function-phrase modeling, between-word coarticulation modeling, and corrective training. On the 997-word resource management task, SPHINX attained a word accuracy of 96% with a grammar (perplexity 60), and 82% without grammar (perplexity 997).> Kai-Fu Lee, Hsiao-Wuen Hon, Mei-Yuh Hwang, Sanjoy Mahajan, Raj Reddy |
ICASSP | 2 |
| 1989 | Towards speech recognition without vocabulary-specific trainingabstractWith the emergence of high-performance speaker-independent systems, a great barrier to man-machine interface has been overcome. This work describes our next step to improve the usability of speech recognizers—the use of vocabulary-independent (VI) models. If successful, VI models are trained once and for all. They will completely eliminate task-specific training, and will enable rapid configuration of speech recognizers for new vocabularies. Our initial results using generalized triphones as VI models show that with more training data and more detailed modeling, the error rate of VI models can be reduced substantially. For example, the error rates for VI models with 5,000, 10,000 and 15,000 training sentences are 23.9%, 15.2% and 13.3% respectively. Moreover, if task-specific training data were available, we can interpolate them with VI models. Our prelimenary results show that this interpolation can lead to an 18% error rate reduction over task-specific models. Hsiao-Wuen Hon, Kai-Fu Lee, Robert Weide |
EUROSPEECH | 1 |
| 1989 | Large-vocabulary speaker-independent continuous speech recognition with semi-continuous hidden Markov modelsabstractA semi-continuous hidden Markov model based on the multiple vector quantization codebooks is used here for large-vocabulary speaker-independent continuous speech recognition In the techniques employed here, the semi-continuous output probability density function for each codebook is represented by a combination of the corresponding discrete output probabilities of the hidden Markov model and the continuous Gaussian density functions of each individual codebook. Parameters of vector quantization codebook and hidden Markov model are mutually optimized to achieve an optimal model codebook combination under a unified probabilistic framework Another advantages of this approach is the enhanced robustness of the semi-continuous output probability by the combination of multiple codewords and multiple codebooks For a 1000-word speaker-independent continuous speech recognition using a word-pair grammar the recognition error rate of the semi-continuous hidden Markov model was reduced by more than 29% and 41% in comparison to the discrete and continuous mixture hidden Markov model respectively Xuedong Huang 0001, Hsiao-Wuen Hon, Kai-Fu Lee |
EUROSPEECH | 2 |
| 1989 | Modeling between-word coarticulation in continuous speech recognition
Mei-Yuh Hwang, Hsiao-Wuen Hon, Kai-Fu Lee |
EUROSPEECH | 2 |
| 1988 | Large-vocabulary speaker-independent continuous speech recognition using HMMabstractSPHINX, the first large-vocabulary speaker-independent continuous-speech recognizer is described. SPHINX is a hidden-Markov-model (HMM)-based recognizer using multiple codebooks of various LPC-derived features. Two types of HMMs are used in SPHINX: context-independent phone models and function-word-dependent phone models. On a 997-word task using a bigram grammar, SPHINX achieved a word accuracy of 93%. This demonstrates the feasibility of speaker-independent continuous-speech recognition, and the appropriateness of hidden Markov models for such a task.> Kai-Fu Lee, Hsiao-Wuen Hon |
ICASSP | 2 |