VLDB 2026 Research / reviewers in the wild / expert
Guihong Cao
dblp:30/572
· DBLP profile ↗
25ranked-venue papers
8as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Information extraction and text analysis · 27% Question answering and dialogue systems · 19% Language models and text generation · 15% | |
| Databases, data mining, and information retrieval
10 papers |
Information retrieval · 73% Query processing and optimization · 13% Data models and query languages · 13% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.8 | 2 | 2020 | Graph-Based Transformer with Cross-Candidate Verification for Semantic Parsing · AAAI 2020 Weakly Supervised Multi-task Learning for Semantic Parsing · IJCAI 2019 |
Natural language and speech › Question answering and dialogue systems
multimodal question answering |
0.6 | 1 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
commonsense question answering |
0.4 | 1 | 2020 | Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering · AAAI 2020 |
Natural language and speech › Language models and text generation › multilingual language models
cross-lingual pre-training |
0.4 | 1 | 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and Generation · EMNLP (1) 2020 |
Machine learning › Graph learning › graph neural network
graph transformer |
0.4 | 1 | 2020 | Graph-Based Transformer with Cross-Candidate Verification for Semantic Parsing · AAAI 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.4 | 1 | 2020 | Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering · AAAI 2020 |
Machine learning › Deep learning architectures and training
transformer |
0.4 | 1 | 2020 | Graph-Based Transformer with Cross-Candidate Verification for Semantic Parsing · AAAI 2020 |
Machine learning › Learning paradigms › multi-task learning
multi-task learning for semantic parsing |
0.4 | 1 | 2019 | Weakly Supervised Multi-task Learning for Semantic Parsing · IJCAI 2019 |
Natural language and speech › Information extraction and text analysis › semantic parsing
weakly supervised semantic parsing |
0.4 | 1 | 2019 | Weakly Supervised Multi-task Learning for Semantic Parsing · IJCAI 2019 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
text-to-SQL parsing |
0.3 | 1 | 2018 | Semantic Parsing with Syntax- and Table-Aware SQL Generation · ACL (1) 2018 |
Query processing and optimization
semantic parsing |
0.3 | 1 | 2018 | Semantic Parsing with Syntax- and Table-Aware SQL Generation · ACL (1) 2018 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.3 | 1 | 2018 | Semantic Parsing with Syntax- and Table-Aware SQL Generation · ACL (1) 2018 |
Information retrieval
web search |
0.2 | 2 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 Selecting Query Term Alternations for Web Search by Exploiting Query Contexts · ACL 2008 |
Information retrieval › retrieval models
language model |
0.2 | 3 | 2007 | Using query contexts in information retrieval · SIGIR 2007 Integrating word relationships into language models · SIGIR 2005 Dependence language model for information retrieval · SIGIR 2004 |
Computer vision › Vision and language
visual question answering |
0.2 | 1 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 |
Information retrieval
multimodal retrieval |
0.2 | 1 | 2022 | WebQA: Multihop and Multimodal QA · CVPR 2022 |
Information retrieval
retrieval models |
0.2 | 3 | 2006 | Context-Dependent Term Relations for Information Retrieval · EMNLP 2006 Integrating word relationships into language models · SIGIR 2005 Dependence language model for information retrieval · SIGIR 2004 |
Information retrieval › ranking › learning to rank
feature-based ranking |
0.1 | 1 | 2012 | Extracting search-focused key n-grams for relevance ranking in web search · WSDM 2012 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2012 | Extracting search-focused key n-grams for relevance ranking in web search · WSDM 2012 |
Information retrieval › ranking › search ranking
relevance ranking |
0.1 | 1 | 2012 | Extracting search-focused key n-grams for relevance ranking in web search · WSDM 2012 |
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network |
0.1 | 1 | 2020 | Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering · AAAI 2020 |
Machine learning › Graph learning
graph neural network |
0.1 | 1 | 2020 | Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering · AAAI 2020 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and Generation · EMNLP (1) 2020 |
Information retrieval
ranking |
0.1 | 1 | 2020 | Graph-Based Transformer with Cross-Candidate Verification for Semantic Parsing · AAAI 2020 |
Information retrieval
reranking |
0.1 | 1 | 2020 | Graph-Based Transformer with Cross-Candidate Verification for Semantic Parsing · AAAI 2020 |
Information retrieval › retrieval models › language model
term dependency models |
0.1 | 2 | 2005 | Integrating word relationships into language models · SIGIR 2005 Dependence language model for information retrieval · SIGIR 2004 |
Information retrieval › relevance feedback
pseudo-relevance feedback |
0.1 | 1 | 2008 | Selecting good expansion terms for pseudo-relevance feedback · SIGIR 2008 |
Information retrieval › query reformulation
query expansion |
0.1 | 1 | 2008 | Selecting good expansion terms for pseudo-relevance feedback · SIGIR 2008 |
Information retrieval
query processing |
0.1 | 1 | 2008 | Selecting Query Term Alternations for Web Search by Exploiting Query Contexts · ACL 2008 |
Information retrieval › query understanding › query modeling
query context modeling |
0.1 | 1 | 2007 | Using query contexts in information retrieval · SIGIR 2007 |
Methods — techniques the papers use, named apart from their topics
sequence-to-sequence model · 1.5multimodal reasoning · 1.1knowledge aggregation · 1.1graph transformer · 0.9BERT · 0.9graph convolutional network · 0.4graph attention mechanism · 0.4weakly supervised learning · 0.4multi-task learning · 0.4encoder-decoder · 0.4table-aware generation · 0.3syntax-aware decoding · 0.3search log mining · 0.1learning to rank · 0.1query context exploitation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | WebQA: Multihop and Multimodal QAabstractScaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce WEBQA, a challenging new benchmark that proves difficult for large-scale state-of-the-art models which lack language groundable visual representations for novel objects and the ability to reason, yet trivial for humans. WebQA mirrors the way humans use the web: 1) Ask a question, 2) Choose sources to aggregate, and 3) Produce a fluent language response. This is the behavior we should be expecting from IoT devices and digital assistants. Existing work prefers to assume that a model can either reason about knowledge in images or in text. WebQA includes a secondary text-only QA task to ensure improved visual performance does not come at the cost of language understanding. Our challenge for the community is to create unified multimodal reasoning models that answer questions regardless of the source modality, moving us closer to digital assistants that not only query language knowledge, but also the richer visual online world. Yingshan Chang, Guihong Cao, Mridu Narang, Jianfeng Gao 0001, Hisami Suzuki, Yonatan Bisk |
CVPR | 2 |
| 2020 | Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question AnsweringabstractCommonsense question answering aims to answer questions which require background knowledge that is not explicitly expressed in the question. The key challenge is how to obtain evidence from external knowledge and make predictions based on the evidence. Recent studies either learn to generate evidence from human-annotated evidence which is expensive to collect, or extract evidence from either structured or unstructured knowledge bases which fails to take advantages of both sources simultaneously. In this work, we propose to automatically extract evidence from heterogeneous knowledge sources, and answer questions based on the extracted evidence. Specifically, we extract evidence from both structured knowledge base (i.e. ConceptNet) and Wikipedia plain texts. We construct graphs for both sources to obtain the relational structures of evidence. Based on these graphs, we propose a graph-based approach consisting of a graph-based contextual word representation learning module and a graph-based inference module. The first module utilizes graph structural information to re-define the distance between words for learning better contextual word representations. The second module adopts graph convolutional network to encode neighbor information into the representations of nodes, and aggregates evidence with graph attention mechanism for predicting the final answer. Experimental results on CommonsenseQA dataset illustrate that our graph-based approach over both knowledge sources brings improvement over strong baselines. Our approach achieves the state-of-the-art accuracy (75.3%) on the CommonsenseQA dataset. Shangwen Lv, Daya Guo, Jingjing Xu 0001, Duyu Tang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Songlin Hu 0001 |
AAAI | 9 |
| 2020 | Graph-Based Transformer with Cross-Candidate Verification for Semantic ParsingabstractIn this paper, we present a graph-based Transformer for semantic parsing. We separate the semantic parsing task into two steps: 1) Use a sequence-to-sequence model to generate the logical form candidates. 2) Design a graph-based Transformer to rerank the candidates. To handle the structure of logical forms, we incorporate graph information to Transformer, and design a cross-candidate verification mechanism to consider all the candidates in the ranking process. Furthermore, we integrate BERT into our model and jointly train the graph-based Transformer and BERT. We conduct experiments on 3 semantic parsing benchmarks, ATIS, JOBS and Task Oriented semantic Parsing dataset (TOP). Experiments show that our graph-based reranking model achieves results comparable to state-of-the-art models on the ATIS and JOBS datasets. And on the TOP dataset, our model achieves a new state-of-the-art result. Yeyun Gong, Weizhen Qi, Guihong Cao, Jianshu Ji, Xiaola Lin |
AAAI | 4 |
| 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationabstractYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001 |
EMNLP (1) | 10 |
| 2019 | Weakly Supervised Multi-task Learning for Semantic ParsingabstractSemantic parsing is a challenging and important task which aims to convert a natural language sentence to a logical form. Existing neural semantic parsing methods mainly use (Q-L) pairs to train a sequence-to-sequence model. However, the amount of existing Q-L labeled data is limited and hard to obtain. We propose an effective method which substantially utilizes labeling information from other tasks to enhance the training of a semantic parser. We design a multi-task learning model to train question type classification, entity mention detection together with question semantic parsing using a shared encoder. We propose a weakly supervised learning method to enhance our multi-task learning model with paraphrase data, based on the idea that the paraphrased questions should have the same logical form and question type information. Finally, we integrate the weakly supervised multi-task learning method to an encoder-decoder framework. Experiments on a newly constructed dataset and ComplexWebQuestions show that our proposed method outperforms state-of-the-art methods which demonstrates the effectiveness and robustness of our method. Yeyun Gong, Junwei Bao 0001, Jianshu Ji, Guihong Cao, Xiaola Lin, Nan Duan 0001 |
IJCAI | 5 |
| 2018 | Semantic Parsing with Syntax- and Table-Aware SQL GenerationabstractYibo Sun, Duyu Tang, Nan Duan, Jianshu Ji, Guihong Cao, Xiaocheng Feng, Bing Qin, Ting Liu, Ming Zhou. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Duyu Tang, Nan Duan 0001, Jianshu Ji, Guihong Cao, Bing Qin 0001, Ting Liu 0001, Ming Zhou 0001 |
ACL (1) | 5 |
| 2012 | Extracting search-focused key n-grams for relevance ranking in web searchabstractIn web search, relevance ranking of popular pages is relatively easy, because of the inclusion of strong signals such as anchor text and search log data. In contrast, with less popular pages, relevance ranking becomes very challenging due to a lack of information. In this paper the former is referred to as head pages, and the latter tail pages. We address the challenge by learning a model that can extract search-focused key n-grams from web pages, and using the key n-grams for searches of the pages, particularly, the tail pages. To the best of our knowledge, this problem has not been previously studied. Our approach has four characteristics. First, key n-grams are search-focused in the sense that they are defined as those which can compose "good queries" for searching the page. Second, key n-grams are learned in a relative sense using learning to rank techniques. Third, key n-grams are learned using search log data, such that the characteristics of key n-grams in the search log data, particularly in the heads; can be applied to the other data, particularly to the tails. Fourth, the extracted key n-grams are used as features of the relevance ranking model also trained with learning to rank techniques. Experiments validate the effectiveness of the proposed approach with large-scale web search datasets. The results show that our approach can significantly improve relevance ranking performance on both heads and tails; and particularly tails, compared with baseline approaches. Characteristics of our approach have also been fully investigated through comprehensive experiments. Keping Bi, Yunhua Hu, Hang Li 0001, Guihong Cao |
WSDM | 5 |
| 2008 | Selecting Query Term Alternations for Web Search by Exploiting Query Contexts
Guihong Cao, Stephen E. Robertson, Jian-Yun Nie |
ACL | 1 |
| 2008 | Relating dependent indexes using dempster-shafer theoryabstractTraditional information retrieval (IR) approaches assume that the indexing terms are independent, which is not true in reality. Although some previous studies have tried to consider term relationships, strong simplifications had to be made at the very basic indexing step, namely, dependent terms are assigned independent counts or probabilities. Lixin Shi, Jian-Yun Nie, Guihong Cao |
CIKM | 3 |
| 2008 | Selecting good expansion terms for pseudo-relevance feedbackabstractPseudo-relevance feedback assumes that most frequent terms in the pseudo-feedback documents are useful for the retrieval. In this study, we re-examine this assumption and show that it does not hold in reality - many expansion terms identified in traditional approaches are indeed unrelated to the query and harmful to the retrieval. We also show that good expansion terms cannot be distinguished from bad ones merely on their distributions in the feedback documents and in the whole collection. We then propose to integrate a term classification process to predict the usefulness of expansion terms. Multiple additional features can be integrated in this process. Our experiments on three TREC collections show that retrieval effectiveness can be much improved when term classification is used. In addition, we also demonstrate that good terms should be identified directly according to their possible impact on the retrieval effectiveness, i.e. using supervised learning, instead of unsupervised learning. Guihong Cao, Jian-Yun Nie, Jianfeng Gao 0001, Stephen E. Robertson |
SIGIR | 1 |
| 2007 | Extending query translation to cross-language query expansion with markov chain modelsabstractDictionary-based approaches to query translation have been widely used in Cross-Language Information Retrieval (CLIR) experiments. However, translation has been not only limited by the coverage of the dictionary, but also affected by translation ambiguities. In this paper we propose a novel method of query translation that combines other types of term relation to complement the dictionary-based translation. This allows extending the literal query translation to related words, which produce a beneficial effect of query expansion in CLIR. In this paper, we model query translation by Markov Chains (MC), where query translation is viewed as a process of expanding query terms to their semantically similar terms in a different language. In MC, terms and their relationships are modeled as a directed graph, and query translation is performed as a random walk in the graph, which propagates probabilities to related terms. This framework allows us to incorporating different types of term relation, either between two languages or within the source or target languages. In addition, the iterative training process of MC allows us to attribute higher probabilities to the target terms more related to the original query, thus offers a solution to the translation ambiguity problem. We evaluated our method on three CLIR benchmark collections, and obtained significant improvements over traditional dictionary-based approaches. Guihong Cao, Jianfeng Gao 0001, Jian-Yun Nie, Jing Bai 0005 |
CIKM | 1 |
| 2007 | A system to mine large-scale bilingual dictionaries from monolingual web pages
Guihong Cao, Jianfeng Gao 0001, Jian-Yun Nie |
MTSummit | 1 |
| 2007 | Using query contexts in information retrievalabstractUser query is an element that specifies an information need, but it is not the only one. Studies in literature have found many contextual factors that strongly influence the interpretation of a query. Recent studies have tried to consider the user's interests by creating a user profile. However, a single profile for a user may not be sufficient for a variety of queries of the user. In this study, we propose to use query-specific contexts instead of user-centric ones, including context around query and context within query. The former specifies the environment of a query such as the domain of interest, while the latter refers to context words within the query, which is particularly useful for the selection of relevant term relations. In this paper, both types of context are integrated in an IR model based on language modeling. Our experiments on several TREC collections show that each of the context factors brings significant improvements in retrieval effectiveness. Jing Bai 0005, Jian-Yun Nie, Guihong Cao, Hugues Bouchard |
SIGIR | 3 |
| 2006 | Constructing better document and query models with markov chainsabstractDocument and query expansions have been used separately in previous studies to enhance the representation of documents and queries. In this paper, we propose a general method that integrates both of them. Expansion is carried out using multi-stage Markov chains. Our experiments show that this method significantly outperforms the existing approaches. Guihong Cao, Jian-Yun Nie, Jing Bai 0005 |
CIKM | 1 |
| 2006 | Context-Dependent Term Relations for Information Retrieval
Jing Bai 0005, Jian-Yun Nie, Guihong Cao |
EMNLP | 3 |
| 2006 | An Information-Theoretic Approach to Automatic Evaluation of Summaries
Chin-Yew Lin, Guihong Cao, Jianfeng Gao 0001, Jian-Yun Nie |
HLT-NAACL | 2 |
| 2006 | Inferential language models for information retrievalabstractLanguage modeling (LM) has been widely used in IR in recent years. An important operation in LM is smoothing of the document language model. However, the current smoothing techniques merely redistribute a portion of term probability according to their frequency of occurrences only in the whole document collection. No relationships between terms are considered and no inference is involved. In this article, we propose several inferential language models capable of inference using term relationships. The inference operation is carried out through a semantic smoothing either on the document model or query model, resulting in document or query expansion. The proposed models implement some of the logical inference capabilities proposed in the previous studies on logical models, but with necessary simplifications in order to make them tractable. They are a good compromise between inference power and efficiency. The models have been tested on several TREC collections, both in English and Chinese. It is shown that the integration of term relationships into the language modeling framework can consistently improve the retrieval effectiveness compared with the traditional language models. This study shows that language modeling is a suitable framework to implement basic inference operations in IR effectively. Jian-Yun Nie, Guihong Cao, Jing Bai 0005 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2005 | Query expansion using term relationships in language models for information retrievalabstractLanguage Modeling (LM) has been successfully applied to Information Retrieval (IR). However, most of the existing LM approaches only rely on term occurrences in documents, queries and document collections. In traditional unigram based models, terms (or words) are usually considered to be independent. In some recent studies, dependence models have been proposed to incorporate term relationships into LM, so that links can be created between words in the same sentence, and term relationships (e.g. synonymy) can be used to expand the document model. In this study, we further extend this family of dependence models in the following two ways: (1) Term relationships are used to expand query model instead of document model, so that query expansion process can be naturally implemented; (2) We exploit more sophisticated inferential relationships extracted with Information Flow (IF). Information flow relationships are not simply pairwise term relationships as those used in previous studies, but are between a set of terms and another term. They allow for context-dependent query expansion. Our experiments conducted on TREC collections show that we can obtain large and significant improvements with our approach. This study shows that LM is an appropriate framework to implement effective query expansion. Jing Bai 0005, Dawei Song 0001, Peter Bruza, Jian-Yun Nie, Guihong Cao |
CIKM | 5 |
| 2005 | Integrating word relationships into language modelsabstractIn this paper, we propose a novel dependency language modeling approach for information retrieval. The approach extends the existing language modeling approach by relaxing the independence assumption. Our goal is to build a language model in which various word relationships can be integrated. In this work, we integrate two types of relationship extracted from WordNet and co-occurrence relationships respectively. The integrated model has been tested on several TREC collections. The results show that our model achieves substantial and significant improvements with respect to the models without these relationships. These results clearly show the benefit of integrating word relationships into language models for IR. Guihong Cao, Jian-Yun Nie, Jing Bai 0005 |
SIGIR | 1 |
| 2005 | Integrating Compound Terms in Bayesian Text ClassificationabstractText classification usually assumed a word-based document representation. In this paper, we propose a new approach to integrate compound terms in Bayesian text classification. Compound terms are used as complementary features to single words. An acute problem is to consider their dependence with the component words. In this paper, we propose to use smoothing techniques to combine both compound term and word representations. Experiments have been conducted on two corpora. Our results show that this approach can slightly but steadily improve the classification performance on both test corpora. Jing Bai 0005, Jian-Yun Nie, Guihong Cao |
Web Intelligence | 3 |
| 2004 | Applying Machine Learning to Chinese Temporal Relation ResolutionabstractTemporal relation resolution involves extraction of temporal information explicitly or implicitly embedded in a language. This information is often inferred from a variety of interactive grammatical and lexical cues, especially in Chinese. For this purpose, inter-clause relations (temporal or otherwise) in a multiple-clause sentence play an important role. In this paper, a computational model based on machine learning and heterogeneous collaborative bootstrapping is proposed for analyzing temporal relations in a Chinese multiple-clause sentence. The model makes use of the fact that events are represented in different temporal structures. It takes into account the effects of linguistic features such as tense/aspect, temporal connectives, and discourse structures. A set of experiments has been conducted to investigate how linguistic features could affect temporal relation resolution. Wenjie Li 0002, Kam-Fai Wong, Guihong Cao, Chunfa Yuan |
ACL | 3 |
| 2004 | Fuzzy K-Means Clustering on a High Dimensional Semantic Space
Guihong Cao, Dawei Song 0001, Peter Bruza |
APWeb | 1 |
| 2004 | Combining Linguistic Features with Weighted Bayesian Classifier for Temporal Reference Processing
Guihong Cao, Wenjie Li 0002, Kam-Fai Wong, Chunfa Yuan |
COLING | 1 |
| 2004 | Dependence language model for information retrievalabstractThis paper presents a new dependence language modeling approach to information retrieval. The approach extends the basic language modeling approach based on unigram by relaxing the independence assumption. We integrate the linkage of a query as a hidden variable, which expresses the term dependencies within the query as an acyclic, planar, undirected graph. We then assume that a query is generated from a document in two stages: the linkage is generated first, and then each term is generated in turn depending on other related terms according to the linkage. We also present a smoothing method for model parameter estimation and an approach to learning the linkage of a sentence in an unsupervised manner. The new approach is compared to the classical probabilistic retrieval model and the previously proposed language models with and without taking into account term dependencies. Results show that our model achieves substantial and significant improvements on TREC collections. Jianfeng Gao 0001, Jian-Yun Nie, Guangyuan Wu, Guihong Cao |
SIGIR | 4 |
| 2002 | Exploring Asymmetric Clustering for Statistical Language ModelingabstractThe n-gram model is a stochastic model, which predicts the next word (predicted word) given the previous words (conditional words) in a word sequence. The cluster n-gram model is a variant of the n-gram model in which similar words are classified in the same cluster. It has been demonstrated that using different clusters for predicted and conditional words leads to cluster models that are superior to classical cluster models which use the same clusters for both words. This is the basis of the asymmetric cluster model (ACM) discussed in our study. In this paper, we first present a formal definition of the ACM. We then describe in detail the methodology of constructing the ACM. The effectiveness of the ACM is evaluated on a realistic application, namely Japanese Kana-Kanji conversion. Experimental results show substantial improvements of the ACM in comparison with classical cluster models and word n-gram models at the same model size. Our analysis shows that the high-performance of the ACM lies in the asymmetry of the model. Jianfeng Gao 0001, Joshua Goodman 0001, Guihong Cao, Hang Li 0001 |
ACL | 3 |