Hui Fang 0001

dblp:03/2511-1 · DBLP profile ↗
← Back
55ranked-venue papers in the field
7as first author
10since 2021 · last 2025
0009-0003-1904-787XORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 52 (7 first)Database Systems & Data Management · 2Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2025 FAIR-QR: Enhancing Fairness-Aware Information Retrieval Through Query Refinement
Fumian Chen, Hui Fang 0001
ECIR (4)2
2025 CAFE: Context-Aware Applicability-Weighted Fairness Evaluation
Fumian Chen, Hui Fang 0001
NLDB (1)2
2025 CaLuX: A Catalyst and Lubricant Properties Extraction System for Domain Experts
Catalina Riano, Fumian Chen, Hui Fang 0001
NLDB (2)3
2025 CoachGPT: A Scaffolding-based Academic Writing Assistant
abstract
Academic writing skills are crucial for students' success but can feel overwhelming without proper guidance and practice, particularly when writing in a second language. Traditionally, students ask instructors or search dictionaries, which are not universally accessible. Early writing assistants emerged as rule-based systems that focused on detecting misspellings, subject-verb disagreements, and basic punctuation errors but are inaccurate and lack contextual understanding. Machine learning-based assistants demonstrate a strong ability for language understanding but are expensive to train. Large language models (LLMs) have shown remarkable capabilities in generating responses in natural languages based on given prompts, but they have a fundamental limitation in education: they generate essays without teaching, which can have detrimental effects on learning when misused. To address this limitation, we develop CoachGPT, which leverages LLMs to assist academic writing for those with limited educational resources and those who prefer self-paced learning. CoachGPT is an AI agent-based web application that (1) takes instructions from experienced educators (2) converts instructions into sub-tasks, and (3) provides real-time feedback and suggestions using large language models. This unique scaffolding structure makes CoachGPT unique among existing writing assistants. Compared with existing writing assistants, CoachGPT provides a more immersed writing experience with personalized messages. Our user studies prove the usefulness of CoachGPT and the potential of large language models for academic writing.
Fumian Chen, Sotheara Veng, Joshua Wilson, Xiaoming Li 0010, Hui Fang 0001
SIGIR5
2025 ReGeS: Reciprocal Retrieval-Generation Synergy for Conversational Recommender Systems
Dayu Yang, Hui Fang 0001
WISE (2)2
2024 Toward Automatic Group Membership Annotation for Group Fairness Evaluation
Fumian Chen, Dayu Yang, Hui Fang 0001
NLDB (1)3
2024 Behavior Alignment: A New Perspective of Evaluating LLM-based Conversational Recommendation Systems
abstract
Large Language Models (LLMs) have demonstrated great potential in Conversational Recommender Systems (CRS). However, the application of LLMs to CRS has exposed a notable discrepancy in behavior between LLM-based CRS and human recommenders: LLMs often appear inflexible and passive, frequently rushing to complete the recommendation task without sufficient inquiry. This behavior discrepancy can lead to decreased accuracy in recommendations and lower user satisfaction. Despite its importance, existing studies in CRS lack a study about how to measure such behavior discrepancy. To fill this gap, we propose Behavior Alignment, a new evaluation metric to measure how well the recommendation strategies made by a LLM-based CRS are consistent with human recommenders'. Our experiment results show that the new metric is better aligned with human preferences and can better differentiate how systems perform than existing evaluation metrics. As Behavior Alignment requires explicit and costly human annotations on the recommendation strategies, we also propose a classification-based method to implicitly measure the Behavior Alignment based on the responses. The evaluation results confirm the robustness of the method.
Dayu Yang, Fumian Chen, Hui Fang 0001
SIGIR3
2023 Less is More: A Prototypical Framework for Efficient Few-Shot Named Entity Recognition
Yue Zhang 0043, Hui Fang 0001
NLDB2
2022 Axiomatically Regularized Pre-training for Ad hoc Search
abstract
Recently, pre-training methods tailored for IR tasks have achieved great success. However, as the mechanisms behind the performance improvement remain under-investigated, the interpretability and robustness of these pre-trained models still need to be improved. Axiomatic IR aims to identify a set of desirable properties expressed mathematically as formal constraints to guide the design of ranking models. Existing studies have already shown that considering certain axioms may help improve the effectiveness and interpretability of IR models. However, there still lack efforts of incorporating these IR axioms into pre-training methodologies. To shed light on this research question, we propose a novel pre-training method with \underlineA xiomatic \underlineRe gularization for ad hoc \underlineS earch (ARES). In the ARES framework, a number of existing IR axioms are re-organized to generate training samples to be fitted in the pre-training process. These training samples then guide neural rankers to learn the desirable ranking properties. Compared to existing pre-training approaches, ARES is more intuitive and explainable. Experimental results on multiple publicly available benchmark datasets have shown the effectiveness of ARES in both full-resource and low-resource (e.g., zero-shot and few-shot) settings. An intuitive case study also indicates that ARES has learned useful knowledge that existing pre-trained models (e.g., BERT and PROP) fail to possess. This work provides insights into improving the interpretability of pre-trained models and the guidance of incorporating IR axioms or human heuristics into pre-training methods.
Jia Chen 0003, Yiqun Liu 0001, Jiaxin Mao, Hui Fang 0001, Shenghao Yang 0004, Xiaohui Xie, Min Zhang 0006, Shaoping Ma
SIGIR5
2021 Predicting Question Responses to Improve the Performance of Retrieval-Based Chatbot
Disen Wang, Hui Fang 0001
ECIR (2)2
2020 An Adaptive Response Matching Network for Ranking Multi-turn Chatbot Responses
Disen Wang, Hui Fang 0001
NLDB2
2020 Axiomatic thinking for information retrieval: introduction to special issue
Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai
Inf. Retr. J.2
2018 Silent Day Detection on Microblog Data
Kuang Lu, Hui Fang 0001
NLDB2
2018 Are we on the Right Track?: An Examination of Information Retrieval Methodologies
abstract
The unpredictability of user behavior and the need for effectiveness make it difficult to define a suitable research methodology for Information Retrieval (IR). In order to tackle this challenge, we categorize existing IR methodologies along two dimensions: (1) empirical vs. theoretical, and (2) top-down vs. bottom-up. The strengths and drawbacks of the resulting categories are characterized according to 6 desirable aspects. The analysis suggests that different methodologies are complementary and therefore, equally necessary. The categorization of the 167 full papers published in the last SIGIR (2016 and 2017) and ICTIR (2017) conferences suggest that most of existing work is empirical bottom-up, suggesting lack of some desirable aspects. With the hope of improving IR research practice, we propose a general methodology for IR that integrates the strengths of existing research methods.
Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai
SIGIR2
2017 Axiomatic Thinking for Information Retrieval: And Related Tasks
abstract
This is the first workshop on the emerging interdisciplinary research area of applying axiomatic thinking to information retrieval (IR) and related tasks. The workshop aims to help foster collaboration of researchers working on different perspectives of axiomatic thinking and encourage discussion and research on general methodological issues related to applying axiomatic thinking to IR and related tasks.
Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai
SIGIR2
2017 The Lucene for Information Access and Retrieval Research (LIARR) Workshop at SIGIR 2017
abstract
As an empirical discipline, information access and retrieval research requires substantial software infrastructure to index and search large collections. This workshop is motivated by the desire to better align information retrieval research with the practice of building search applications from the perspective of open-source information retrieval systems. Our goal is to promote the use of Lucene for information access and retrieval research.
Leif Azzopardi, Matt Crane, Hui Fang 0001, Grant Ingersoll, Jimmy Lin, Yashar Moshfeghi, Harrisen Scells, Guido Zuccon
SIGIR3
2017 Anserini: Enabling the Use of Lucene for Information Retrieval Research
abstract
Software toolkits play an essential role in information retrieval research. Most open-source toolkits developed by academics are designed to facilitate the evaluation of retrieval models over standard test collections. Efforts are generally directed toward better ranking and less attention is usually given to scalability and other operational considerations. On the other hand, Lucene has become the de facto platform in industry for building search applications (outside a small number of companies that deploy custom infrastructure). Compared to academic IR toolkits, Lucene can handle heterogeneous web collections at scale, but lacks systematic support for evaluation over standard test collections. This paper introduces Anserini, a new information retrieval toolkit that aims to provide the best of both worlds, to better align information retrieval practice and research. Anserini provides wrappers and extensions on top of core Lucene libraries that allow researchers to use more intuitive APIs to accomplish common research tasks. Our initial efforts have focused on three functionalities: scalable, multi-threaded inverted indexing to handle modern web-scale collections, streamlined IR evaluation for ad hoc retrieval on standard test collections, and an extensible architecture for multi-stage ranking. Anserini ships with support for many TREC test collections, providing a convenient way to replicate competitive baselines right out of the box. Experiments verify that our system is both efficient and effective, providing a solid foundation to support future research.
Hui Fang 0001, Jimmy Lin
SIGIR2
2016 OPMES: A Similarity Search Engine for Mathematical Content
Hui Fang 0001
ECIR2
2015 A novel methodology for retrieving infographics utilizing structure and message content
Zhuo Li 0004, Sandra Carberry, Hui Fang 0001, Kathleen F. McCoy, Kelly Peterson, Matthew Stagitis
Data Knowl. Eng.3
2015 Latent entity space: a novel retrieval approach for entity-bearing queries
Xitong Liu, Hui Fang 0001
Inf. Retr. J.2
2015 Opinions matter: a general approach to user profile modeling for contextual suggestion
Hongning Wang, Hui Fang 0001, Deng Cai 0001
Inf. Retr. J.3
2014 Document Prioritization for Scalable Query Processing
abstract
Query latency is an important performance measure of any search engines because it directly affects search users' satisfaction. The key challenge is how to efficiently retrieve top-K ranked results for a query. Current search engines process queries in either the conjunctive or disjunctive modes. However, there is still a large performance gap between these two modes since the conjunctive mode is more efficient with lower search accuracy while the disjunctive mode is more effective but requires more time to process the queries.
Hao Wu 0036, Hui Fang 0001
CIKM2
2014 Analytical Performance Modeling for Top-K Query Processing
abstract
Top-K query processing is one of the most important problems in large-scale Information Retrieval systems. Since query processing time varies for different queries, an accurate run-time performance prediction is critical for online query scheduling and load balancing, which could eventually reduce the query waiting time and improve the throughput. Previous studies estimated the query processing time based on the combination of term-level features. Unfortunately, these features were often selected arbitrarily, and the linear combination of these features might not be able to accurately capture the complexity in the query processing.
Hao Wu 0036, Hui Fang 0001
CIKM2
2014 EntEXPO: An Interactive Search System for Entity-Bearing Queries
Xitong Liu, Hui Fang 0001
ECIR3
2014 An Exploration of Tie-Breaking for Microblog Retrieval
Yue Wang 0037, Hao Wu 0036, Hui Fang 0001
ECIR3
2014 Infographics Retrieval: A New Methodology
Zhuo Li 0004, Sandra Carberry, Hui Fang 0001, Kathleen F. McCoy, Kelly Peterson
NLDB3
2014 VIRLab: a web-based virtual lab for learning and studying information retrieval models
abstract
In this paper, we describe VIRLab, a novel web-based virtual laboratory for Information Retrieval (IR). Unlike existing command line based IR toolkits, the VIRLab system provides a more interactive tool that enables easy implementation of retrieval functions with only a few lines of codes, simplified evaluation process over multiple data sets and parameter settings and straightforward result analysis interface through operational search engines and pair-wise comparisons. These features make VIRLab a unique and novel tool that can help teaching IR models, improving the productivity for doing IR model research, as well as promoting controlled experimental study of IR models.
Hui Fang 0001, Hao Wu 0036, ChengXiang Zhai
SIGIR1
2014 Axiomatic analysis and optimization of information retrieval models
abstract
Axiomatic approach provides a systematic way to think about heuristics, identify the weakness of existing methods, and optimize the existing methods accordingly. This tutorial aims to promote axiomatic thinking that can benefit not only the study of IR models but also the methods for many IR applications.
Hui Fang 0001, ChengXiang Zhai
SIGIR1
2014 Wikimantic: Toward effective disambiguation and expansion of queries
Christopher Boston, Hui Fang 0001, Sandra Carberry, Hao Wu 0036, Xitong Liu
Data Knowl. Eng.2
2014 Exploiting entity relationship for query expansion in enterprise search
Xitong Liu, Hui Fang 0001, Min Wang 0001
Inf. Retr.3
2014 Leveraging integrated information to extract query subtopics for search result diversification
Wei Zheng 0007, Hui Fang 0001, Conglei Yao, Min Wang 0001
Inf. Retr.2
2013 Automatic Detection of Ambiguous Terminology for Software Requirements
Yue Wang 0037, Irene Lizeth Manotas Gutiérrez, Kristina Winbladh, Hui Fang 0001
NLDB4
2013 An incremental approach to efficient pseudo-relevance feedback
abstract
Pseudo-relevance feedback is an important strategy to improve search accuracy. It is often implemented as a two-round retrieval process: the first round is to retrieve an initial set of documents relevant to an original query, and the second round is to retrieve final retrieval results using the original query expanded with terms selected from the previously retrieved documents. This two-round retrieval process is clearly time consuming, which could arguably be one of main reasons that hinder the wide adaptation of the pseudo-relevance feedback methods in real-world IR systems.
Hao Wu 0036, Hui Fang 0001
SIGIR2
2013 iPLUG: Personalized List Recommendation in Twitter
Lijiang Chen, Yibing Zhao, Shimin Chen, Hui Fang 0001, Chengkai Li 0001, Min Wang 0001
WISE (2)4
2013 An exploration of ranking models and feedback method for related entity finding
Xitong Liu, Wei Zheng 0007, Hui Fang 0001
Inf. Process. Manag.3
2012 Entity centric query expansion for enterprise search
abstract
Enterprise search is important, and the search quality has a direct impact on the productivity of an enterprise. Many information needs of enterprise search center around entities. Intuitively, information related to the entities mentioned in the query, such as related entities, would be useful to reformulate the query and improve the retrieval performance. However, most existing studies on query expansion are term-centric. In this paper, we propose a novel entity-centric query expansion framework for enterprise search. Specifically, given a query containing entities, we first utilize both unstructured and structured information to find entities that are related to the ones in the query. We then discuss how to adapt existing feedback methods to use the related entities to improve search quality. Experiment results show that the proposed entity-centric query expansion strategy is more effective to improve the search performance than the state-of-the-art pseudo feedback methods on longer, natural language-like queries with entities.
Xitong Liu, Hui Fang 0001, Min Wang 0001
CIKM2
2012 Exploiting concept hierarchy for result diversification
abstract
The goal of result diversification is to maximize the coverage of query subtopics while minimizing the redundancy in the search results. Intuitively, it is more desirable for a diversification system to cover independent subtopics since it would retrieve sets of non-overlapped relevant documents, which leads to less redundancy in the search results. Unfortunately, existing diversification methods assume that query subtopics are independent and ignore their relations in the diversification process. To overcome this limitation, we propose to exploit concept hierarchies to extract query subtopics and infer their relations. We then apply axiomatic approaches to derive a structural diversification method that can leverage the subtopic relations in result diversification. Experimental results over an enterprise collection show that the relations among query subtopics are useful to improve the diversification performance.
Wei Zheng 0007, Hui Fang 0001, Conglei Yao
CIKM2
2012 Relation Based Term Weighting Regularization
Hao Wu 0036, Hui Fang 0001
ECIR2
2012 Wikimantic: Disambiguation for Short Queries
Christopher Boston, Sandra Carberry, Hui Fang 0001
NLDB3
2012 Diversifying Search Results through Pattern-Based Subtopic Modeling
abstract
Traditional information retrieval models do not necessarily provide users with optimal search experience because the top ranked documents may contain excessively redundant information. Therefore, satisfying search results should be not only relevant to the query but also diversified to cover different subtopics of the query. In this paper, the authors propose a novel pattern-based framework to diversify search results, where each pattern is a set of semantically related terms covering the same subtopic. They first apply a maximal frequent pattern mining algorithm to extract the patterns from retrieval results of the query. The authors then propose to model a subtopic with either a single pattern or a group of similar patterns. A profile-based clustering method is adapted to group similar patterns based on their context information. The search results are then diversified using the extracted subtopics. Experimental results show that the proposed pattern-based methods are effective to diversify the search results.
Wei Zheng 0007, Hui Fang 0001, Hong Cheng 0001, Xuanhui Wang
Int. J. Semantic Web Inf. Syst.2
2012 Coverage-based search result diversification
Wei Zheng 0007, Xuanhui Wang, Hui Fang 0001, Hong Cheng 0001
Inf. Retr.3
2011 Finding relevant information of certain types from enterprise data
abstract
Search over enterprise data is essential to every aspect of an enterprise because it helps users fulfill their information needs. Similar to Web search, most queries in enterprise search are keyword queries. However, enterprise search is a unique research problem because, compared with the data in traditional IR applications (e.g., text data), enterprise data includes information stored in different formats. In particular, enterprise data include both unstructured and structured information, and all the data center around a particular enterprise. As a result, the relevant information from these two data sources could be complementary to each other. Intuitively, such integrated data could be exploited to improve the enterprise search quality. Despite its importance, this problem has received little attention so far. In this paper, we demonstrate the feasibility of leveraging the integrated information in enterprise data to improve search quality through a case study, i.e., finding relevant information of certain types from enterprise data. Enterprise search users often look for different types of relevant information other than documents, e.g., the contact information of per- sons working on a product. When formulating a keyword query, search users may specify both content requirements, i.e., what kind of information is relevant, and type requirements, i.e., what type of information is relevant. Thus, the goal is to find information relevant to both requirements specified in the query. Specifically, we formulate the problem as keyword search over structured or semistructured data, and then propose to leverage the complementary unstructured information in the enterprise data to solve the problem. Experiment results over real world enterprise data and simulated data show that the proposed methods can effectively exploit the unstructured information to find relevant information of certain types from structured and semistructured information in enterprise data.
Xitong Liu, Hui Fang 0001, Conglei Yao, Min Wang 0001
CIKM2
2011 Search result diversification for enterprise data
abstract
Search result diversification aims to return a list of diversified relevant documents in order to satisfy different user information needs. Most of the efforts focused on Web Search, and few studies have considered another important search domain, i.e., enterprise search. Unlike Web search, enterprise search deals with both unstructured and structured data. In this paper, we propose to integrate the structured and unstructured data to discover meaningful query subtopics in search result diversification. Experimental results show that integrating structured and unstructured information allows us to discover high quality query, which are effective in diversifying the retrieval results.
Wei Zheng 0007, Hui Fang 0001, Conglei Yao, Min Wang 0001
CIKM2
2011 Diagnostic Evaluation of Information Retrieval Models
abstract
Developing effective retrieval models is a long-standing central challenge in information retrieval research. In order to develop more effective models, it is necessary to understand the deficiencies of the current retrieval models and the relative strengths of each of them. In this article, we propose a general methodology to analytically and experimentally diagnose the weaknesses of a retrieval function, which provides guidance on how to further improve its performance. Our methodology is motivated by the empirical observation that good retrieval performance is closely related to the use of various retrieval heuristics. We connect the weaknesses and strengths of a retrieval function with its implementations of these retrieval heuristics, and propose two strategies to check how well a retrieval function implements the desired retrieval heuristics. The first strategy is to formalize heuristics as constraints, and use constraint analysis to analytically check the implementation of retrieval heuristics. The second strategy is to define a set of relevance-preserving perturbations and perform diagnostic tests to empirically evaluate how well a retrieval function implements retrieval heuristics. Experiments show that both strategies are effective to identify the potential problems in implementations of the retrieval heuristics. The performance of retrieval functions can be improved after we fix these problems.
Hui Fang 0001, Tao Tao 0003, ChengXiang Zhai
ACM Trans. Inf. Syst.1
2010 Query Aspect Based Term Weighting Regularization in Information Retrieval
Wei Zheng 0007, Hui Fang 0001
ECIR2
2010 Reusable test collections through experimental design
abstract
Portable, reusable test collections are a vital part of research and development in information retrieval. Reusability is difficult to assess, however. The standard approach— simulating judgment collection when groups of systems are held out, then evaluating those held-out systems—only works when there is a large set of relevance judgments to draw on during the simulation. As test collections adapt to larger and larger corpora, it becomes less and less likely that there will be sufficient judgments for such simulation experiments. Thus we propose a methodology for information retrieval experimentation that collects evidence for or against the reusability of a test collection while judgments are being made. Using this methodology along with the appropriate statistical analyses, researchers will be able to estimate the reusability of their test collections while building them and implement “course corrections ” if the collection does not seem to be achieving desired levels of reusability. We show the robustness of our design to inherent sources of variance, and provide a description of an actual implementation of the framework for creating a large test collection.
Ben Carterette, Evangelos Kanoulas, Virgil Pavlu, Hui Fang 0001
SIGIR4
2009 An empirical study of gene synonym query expansion in biomedical information retrieval
Yue Lu 0002, Hui Fang 0001, ChengXiang Zhai
Inf. Retr.2
2008 A study of methods for negative relevance feedback
abstract
Negative relevance feedback is a special case of relevance feedback where we do not have any positive example; this often happens when the topic is difficult and the search results are poor. Although in principle any standard relevance feedback technique can be applied to negative relevance feedback, it may not perform well due to the lack of positive examples. In this paper, we conduct a systematic study of methods for negative relevance feedback. We compare a set of representative negative feedback methods, covering vector-space models and language models, as well as several special heuristics for negative feedback. Evaluating negative feedback methods requires a test set with sufficient difficult topics, but there are not many naturally difficult topics in the existing test collections. We use two sampling strategies to adapt a test collection with easy topics to evaluate negative feedback. Experiment results on several TREC collections show that language model based negative feedback methods are generally more effective than those based on vector-space models, and using multiple negative models is an effective heuristic for negative feedback. Our results also show that it is feasible to adapt test collections with easy topics for evaluating negative feedback methods through sampling.
Xuanhui Wang, Hui Fang 0001, ChengXiang Zhai
SIGIR2
2007 Improve retrieval accuracy for difficult queries using negative feedback
abstract
How to improve search accuracy for difficult topics is an under-addressed, yet important research question. In this paper, we consider a scenario when the search results are so poor that none of the top-ranked documents is relevant to a user's query, and propose to exploit negative feedback to improve retrieval accuracy for such difficult queries. Specifically, we propose to learn from a certain number of top-ranked non-relevant documents to rerank the rest unseen documents. We propose several approaches to penalizing the documents that are similar to the known non-relevant documents in the language modeling framework. To evaluate the proposed methods, we adapt standard TREC collections to construct a test collection containing only difficult queries. Experiment results show that the proposed approaches are effective for improving retrieval accuracy of difficult queries.
Xuanhui Wang, Hui Fang 0001, ChengXiang Zhai
CIKM2
2007 Probabilistic Models for Expert Finding
Hui Fang 0001, ChengXiang Zhai
ECIR1
2007 A study of Poisson query generation model for information retrieval
abstract
Many variants of language models have been proposed for information retrieval. Most existing models are based on multinomial distribution and would score documents based on query likelihood computed based on a query generation probabilistic model. In this paper, we propose and study a new family of query generation models based on Poisson distribution. We show that while in their simplest forms, the new family of models and the existing multinomial models are equivalent. However, based on different smoothing methods, the two families of models behave differently. We show that the Poisson model has several advantages, including naturally accommodating per-term smoothing and modeling accurate background more efficiently. We present several variants of the new model corresponding to different smoothing methods, and evaluate them on four representative TREC test collections. The results show that while their basic models perform comparably, the Poisson model can out perform multinomial model with per-term smoothing. The performance can be further improved with two-stage smoothing.
Qiaozhu Mei, Hui Fang 0001, ChengXiang Zhai
SIGIR2
2007 Term feedback for information retrieval with language models
abstract
I n t hi s paper w e s t udy t er m- based f eedback f or i nf or mat i on r etrieval in the language modeling approach. With term feedback auserdirectly judges the relevance of individual terms without interaction with feedback documents, taking full control of the query expansion process. We propose a cluster-based method for selecting terms to present to the user for judgment, as well as effective algorithms for constructing refined query language models from user term feedback. Our algorithms are shown to bring significant improvement in retrieval accuracy over a non-feedback baseline, and achieve comparable performance to relevance feedback. They are helpful even when there are no relevant documents in the top.
Atulya Velivelli, Hui Fang 0001, ChengXiang Zhai
SIGIR3
2006 Semantic term matching in axiomatic approaches to information retrieval
abstract
A common limitation of many retrieval models, including the recently proposed axiomatic approaches, is that retrieval scores are solely based on exact (i.e., syntactic) matching of terms in the queries and documents, without allowing distinct but semantically related terms to match each other and contribute to the retrieval score. In this paper, we show that semantic term matching can be naturally incorporated into the axiomatic retrieval model through defining the primitive weighting function based on a semantic similarity function of terms. We define several desirable retrieval constraints for semantic term matching and use such constraints to extend the axiomatic model to directly support semantic term matching based on the mutual information of terms computed on some document set. We show that such extension can be efficiently implemented as query expansion. Experiment results on several representative data sets show that, with mutual information computed over the documents in either the target collection for retrieval or an external collection such as the Web, our semantic expansion consistently and substantially improves retrieval accuracy over the baseline axiomatic retrieval model. As a pseudo feedback method, our method also outperforms a state-of-the-art language modeling feedback method.
Hui Fang 0001, ChengXiang Zhai
SIGIR1
2005 An exploration of axiomatic approaches to information retrieval
abstract
Existing retrieval models generally do not offer any guarantee for optimal retrieval performance. Indeed, it is even difficult, if not impossible, to predict a model's empirical performance analytically. This limitation is at least partly caused by the way existing retrieval models are developed where relevance is only coarsely modeled at the level of documents and queries as opposed to a finer granularity level of terms. In this paper, we present a new axiomatic approach to developing retrieval models based on direct modeling of relevance with formalized retrieval constraints defined at the level of terms. The basic idea of this axiomatic approach is to search in a space of candidate retrieval functions for one that can satisfy a set of reasonable retrieval constraints. To constrain the search space, we propose to define a retrieval function inductively and decompose a retrieval function into three component functions. Inspired by the analysis of the existing retrieval functions with the inductive definition, we derive several new retrieval functions using the axiomatic retrieval framework. Experiment results show that the derived new retrieval functions are more robust and less sensitive to parameter settings than the existing retrieval functions with comparable optimal performance.
Hui Fang 0001, ChengXiang Zhai
SIGIR1
2004 A formal study of information retrieval heuristics
abstract
Empirical studies of information retrieval methods show that good retrieval performance is closely related to the use of various retrieval heuristics, such as TF-IDF weighting. One basic research question is thus what exactly are these "necessary" heuristics that seem to cause good retrieval performance. In this paper, we present a formal study of retrieval heuristics. We formally define a set of basic desirable constraints that any reasonable retrieval function should satisfy, and check these constraints on a variety of representative retrieval functions. We find that none of these retrieval functions satisfies all the constraints unconditionally. Empirical results show that when a constraint is not satisfied, it often indicates non-optimality of the method, and when a constraint is satisfied only for a certain range of parameter values, its performance tends to be poor when the parameter is out of the range. In general, we find that the empirical performance of a retrieval formula is tightly related to how well it satisfies these constraints. Thus the proposed constraints provide a good explanation of many empirical observations and make it possible to evaluate any existing or new retrieval formula analytically.
Hui Fang 0001, Tao Tao 0003, ChengXiang Zhai
SIGIR1