Daqing He

dblp:41/134 · DBLP profile ↗
← Back
46ranked-venue papers in the field
8as first author
7since 2021 · last 2026
0000-0002-4645-8696ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 41 (7 first)Other / Interdisciplinary · 4 (1 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
abstract
While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to securely verify document provenance through dynamic GNN-based credibility scoring, effectively preventing stealthy knowledge corruption attacks while preserving essential domain knowledge integrity. Through extensive evaluations and formal analysis, we demonstrate that SecureCollaRAG maintains robustness against attackers under non-IID data distributions.
Daqing He, Zijian Zhang 0001, Ye Liu 0012, Jiamou Liu, Zhirui Zeng, Zhan Qin, Xin Li 0033, Hongwei Yao, Jincheng An, Yi Li 0008, Xiulei Liu, Liehuang Zhu
WWW2
2025 Reason-to-Rank: Distilling Direct and Comparative Reasoning from Large Language Models for Document Reranking
abstract
Reranking documents in information retrieval often relies on black-box models that improve effectiveness but lack explainability. We introduce Reason-to-Rank (R2R), a novel framework that separates direct relevance reasoning from comparison reasoning to provide both direct and comparitive explanations. We first prompt a large language model to produce comprehensive rationales and a ranking order; then we distill both the ranking decisions and textual explanations into a smaller, open-source student model. Our approach not only improves retrieval performance, as demonstrated in MSMARCO, BEIR, and BRIGHT, but also provides interpretable justifications for why one document outranks another. We report NDCG@5 (and NDCG@10) for direct comparisons with prior work, and show that the distilled student model achieves competitive results while significantly reducing computational overhead. By unifying direct and comparative reasoning in a single pipeline, R2R bridges the gap between transparency and effectiveness in modern reranking systems.
Yuelyu Ji, Zhuochun Li, Daqing He
SIGIR4
2025 Guest editorial of the IPM special issue on information science in human-centered AI
Dan Wu 0003, Daqing He, Preben Hansen, Shaobo Liang
Inf. Process. Manag.2
2023 Mapping dementia caregivers' comments on social media with evidence-based care strategies for memory loss and confusion
abstract
Dementia caregivers widely turn to social media for much needed information and support. Prior research on caregivers’ online information exchange has focused on the original questions or posts, without including peer comments in response to those original questions, which does not provide a complete picture of online information exchanges between caregivers and their peers. This paper provides a preliminary analysis of a subset of data from a larger project and suggests how peer comments might match with evidence-based care strategies for memory loss and confusion. A total of 954 peer comments on 114 Reddit posts for dementia memory loss and confusion were collected and mapped with 5 evidence-based strategies generated from 17 care strategies from 3 credible websites. Our results report how well peer comments on Reddit can map to existing evidence-based strategies, and provide preliminary evidence supporting the necessity of providing tailored information based on the patient’s characteristics, stages of disease, and the progression level of memory loss and confusion.
Ning Zou, Yuelyu Ji, Bo Xie 0001, Daqing He, Zhimeng Luo
CHIIR4
2023 Promoting data use through understanding user behaviors: A model for human open government data interaction
abstract
Abstract Recent dramatic increases in the ability to generate, collect, and use datasets have inspired numerous academic and policy discussions regarding the emerging field of human data interaction (HDI). Given the challenges in interacting with open government data (OGD) and the existing research gap in this field, our study intends to explore HDI in the OGD domain and investigate ways HDI can further contribute to OGD promotion. Building upon two existing behavioral models, we proposed an initial conceptual model for OGD interaction, then using this model, conducted two studies to empirically examine users' behaviors when interacting with OGD. Ultimately, we refined this model for OGD interaction and invited three experts to validate it to enhance its understandability, comprehensiveness, and reasonableness. This comprehensive model for human OGD interaction will contribute to the theoretical work of the HDI field as well as the practical design of OGD platforms and data literacy education.
Fanghui Xiao, Yu Chi 0001, Daqing He
J. Assoc. Inf. Sci. Technol.3
2022 HELPeR: An Interactive Recommender System for Ovarian Cancer Patients and Caregivers
abstract
Recommending online resources to patients with ovarian cancer and their caregivers is a challenging task. On one hand, the recommended items must be relevant, recent, and reliable. On the other hand, they need to match the user’s levels of disease-specific health literacy. In this demonstration, we describe the overall architecture and key components of HELPeR, a knowledge-adaptive interactive recommender system for ovarian cancer patients and their caregivers.
Behnam Rahdari, Peter Brusilovsky, Daqing He, Khushboo Thaker, Zhimeng Luo, Young Ji Lee
RecSys3
2022 Together they shall not fade away: Opportunities and challenges of self-tracking for dementia care
Ning Zou, Yu Chi 0001, Daqing He, Bo Xie 0001
Inf. Process. Manag.3
2020 Laypeople's source selection in online health information-seeking process
abstract
Abstract For laypeople, searching online health information resources can be challenging due to topic complexity and the large number of online sources with differing quality. The goal of this article is to examine, among all the available online sources, which online sources laypeople select to address their health‐related information needs, and whether or how much the severity of a health condition influences their selection. Twenty‐four participants were recruited individually, and each was asked (using a retrieval system called HIS) to search for information regarding a severe health condition and a mild health condition, respectively. The selected online health information sources were automatically captured by the HIS system and classified at both the website and webpage levels. Participants' selection behavior patterns were then plotted across the whole information‐seeking process. Our results demonstrate that laypeople's source selection fluctuates during the health information‐seeking process, and also varies by the severity of health conditions. This study reveals laypeople's real usage of different types of online health information sources, and engenders implications to the design of search engines, as well as the development of health literacy programs.
Yu Chi 0001, Daqing He, Wei Jeng
J. Assoc. Inf. Sci. Technol.2
2020 Global health crises are also information crises: A call to action
abstract
Abstract In this opinion paper, we argue that global health crises are also information crises. Using as an example the coronavirus disease 2019 (COVID‐19) epidemic, we (a) examine challenges associated with what we term “global information crises”; (b) recommend changes needed for the field of information science to play a leading role in such crises; and (c) propose actionable items for short‐ and long‐term research, education, and practice in information science.
Bo Xie 0001, Daqing He, Tim Mercer, Youfa Wang, Dan Wu 0003, Kenneth R. Fleischmann, Yan Zhang 0005, Linda H. Yoder, Keri K. Stephens, Michael Mackert, Min Kyung Lee
J. Assoc. Inf. Sci. Technol.2
2019 Challenges and Supports for Accessing Open Government Datasets: Data Guide for Better Open Data Access and Uses
abstract
The importance of open government data is often associated with increased public trust, civic engagement, and accountable administrations. While there is a myriad of benefits, the existing literature suggests that many open government datasets lack accessibility and usability for diverse users. This study seeks to explore what contextual information users require when they access these datasets. Using mixed methods, we aim to discover the challenges of accessing data, and the necessary contextual information needed by the users to overcome these challenges. As the outcome of this study, we propose a framework called "Data Guides", which is composed of the identified important contextual information. In future work, we will test the effectiveness of the Data Guide in aiding users' accessing and understanding open government data.
Fanghui Xiao, Daqing He, Yu Chi 0001, Wei Jeng, Christinger Tomer
CHIIR2
2018 What Sources to Rely on: : Laypeople's Source Selection in Online Health Information Seeking
abstract
In this study, we examined what sources laypeople would select (i.e., visit and adopt) to resolve their health-related information needs, and how different health conditions affect the selection. Twenty-four college students participated in this user study, where they were asked to search for two separate health issues respectively: multiple sclerosis and weight loss. The search logs were collected and analyzed afterwards. We classify the online information sources on both website level and webpage level, and a webpage classification scheme based on genre is proposed. Results suggest that users» selection of sources depends on different types of health issues in terms of urgency and complexity. Health-specific webpage is a popular source and highly adopted for both tasks, but it is particularly helpful for urgent and complex health conditions. Search engines could facilitate users to navigate among scattered health information and support concerns regarding common health issues.
Yu Chi 0001, Daqing He, Shuguang Han
CHIIR2
2018 Concept Enhanced Content Representation for Linking Educational Resources
abstract
The education sector has been undergoing a welcoming change in recent years with the introduction of a wide variety of digital content openly available to students. Due to the volume of this new digital content, it is very difficult for learners to find the needed information at the right time. Digital textbooks, as well-curated domain knowledge sources, could provide a conceptual and physical platform that unites disparate educational resources as one entity. Educational resource linkage, with state-of-the-art techniques, are based on term-level and topic-level representations. However, term-level representations suffer from the term-mismatch problem and often, topics are too broad for linking to other educational resources. To address these challenges, we propose to link educational resources through concept-level representation. The proposed model generates concept embeddings by utilizing domain-specific educational content and external knowledge graph resources to achieve robust and effective concept-level representations. We conducted evaluations of the proposed models on multiple contents linking tasks, and the results demonstrate that concept-level representations perform better than the state-of-the-art representations in helping students to find more learning resources easily. This could increase both students' learning and satisfaction.
Khushboo Thaker, Peter Brusilovsky, Daqing He
WI3
2018 Mobile Information Retrieval. Fabio Crestani, Stefano Mizzaro, and Ivan Scagnetto. Cham, Switzerland: Springer, 2017. 110 pp. $54.99 (softcover). (ISBN 9783319607764)
Daqing He
J. Assoc. Inf. Sci. Technol.1
2017 Understanding Ephemeral State of Relevance
abstract
Despite its dynamic nature, relevance is often measured in a context-independent manner in information retrieval practice. We look into this discrepancy. We propose a contextual relevance/usefulness measurement called ephemeral state of relevance (ESR), which is defined as the amount of useful information a user acquired from a clicked result as assessed just after examining the result during an interactive search session. We collect ESR and context-independent usefulness judgments through a laboratory user study and compare the two. We examine factors related to both judgments and examine their differences.
Jiepu Jiang, Daqing He, Diane Kelly 0001, James Allan 0001
CHIIR2
2017 Semi-Supervised Techniques for Mining Learning Outcomes and Prerequisites
abstract
Educational content of today no longer only resides in textbooks and classrooms; more and more learning material is found in a free, accessible form on the Internet. Our long-standing vision is to transform this web of educational content into an adaptive, web-scale "textbook", that can guide its readers to most relevant "pages" according to their learning goal and current knowledge. In this paper, we address one core, long-standing problem towards this goal: identifying outcome and prerequisite concepts within a piece of educational content (e.g., a tutorial). Specifically, we propose a novel approach that leverages textbooks as a source of distant supervision, but learns a model that can generalize to arbitrary documents (such as those on the web). As such, our model can take advantage of any existing textbook, without requiring expert annotation. At the task of predicting outcome and prerequisite concepts, we demonstrate improvements over a number of baselines on six textbooks, especially in the regime of little to no ground-truth labels available. Finally, we demonstrate the utility of a model learned using our approach at the task of identifying prerequisite documents for adaptive content recommendation --- an important step towards our vision of the "web as a textbook".
Igor Labutov, Yun Huang 0002, Peter Brusilovsky, Daqing He
KDD4
2017 Comparing In Situ and Multidimensional Relevance Judgments
abstract
To address concerns of TREC-style relevance judgments, we explore two improvements. The first one seeks to make relevance judgments contextual, collecting in situ feedback of users in an interactive search session and embracing usefulness as the primary judgment criterion. The second one collects multidimensional assessments to complement relevance or usefulness judgments, with four distinct alternative aspects examined in this paper - novelty, understandability, reliability, and effort.
Jiepu Jiang, Daqing He, James Allan 0001
SIGIR2
2017 Information exchange on an academic social networking site: A multidiscipline comparison on researchgate Q&A
abstract
The increasing popularity of academic social networking sites (ASNSs) requires studies on the usage of ASNSs among scholars and evaluations of the effectiveness of these ASNSs. However, it is unclear whether current ASNSs have fulfilled their design goal, as scholars' actual online interactions on these platforms remain unexplored. To fill the gap, this article presents a study based on data collected from ResearchGate. Adopting a mixed‐method design by conducting qualitative content analysis and statistical analysis on 1,128 posts collected from ResearchGate Q&A, we examine how scholars exchange information and resources, and how their practices vary across three distinct disciplines: library and information services, history of art, and astrophysics. Our results show that the effect of a questioner's intention (i.e., seeking information or discussion) is greater than disciplinary factors in some circumstances. Across the three disciplines, responses to questions provide various resources, including experts' contact details, citations, links to Wikipedia, images, and so on. We further discuss several implications of the understanding of scholarly information exchange and the design of better academic social networking interfaces, which should stimulate scholarly interactions by minimizing confusion, improving the clarity of questions, and promoting scholarly content management.
Wei Jeng, Spencer DesAutels, Daqing He, Lei Li 0037
J. Assoc. Inf. Sci. Technol.3
2016 Contextual Support for Collaborative Information Retrieval
abstract
Recent research shows that Collaborative Information Retrieval (CIR), in which two or more users collaborate on the same search task, has become increasingly popular. The presence of both search and collaboration behaviors makes CIR a complex search format, which further drives a critical need to understand CIR's search context. The contextual support for CIR should consider search contexts derived from both team members' search histories (including users' own search histories and partners' search histories) and their explicit collaboration (e.g., chatting). As it stands, existing studies on contextual search support only focus on Individual Information Retrieval (IIR) and only utilize individuals' own search histories. In this paper, we examine the unique search contexts (e.g., partners' search histories and team collaboration histories) in CIR. Based on a user study data collection with 54 participants, we find that compared to the use of individuals' own search histories, CIR contextual support is more effective when utilizing partners' search histories and teams' collaboration behaviors. More interestingly, though the explicit communication information (i.e., chat content) often involves massive noisy information, involving such noise does not affect the ranking of relevant documents since it also does not appear in relevant documents.
Shuguang Han, Daqing He, Zhen Yue, Jiepu Jiang
CHIIR2
2016 Knowledge-Based Content Linking for Online Textbooks
abstract
Although the volume of online educational resources has dramatically increased in recent years, many of these resources are isolated and distributed in diverse websites and databases. This hinders the discovery and overall usage of online educational resources. By using linking between related subsections of online textbooks as a testbed, this paper explores multiple knowledge-based content linking algorithms for connecting online educational resources. We focus on examining semantic-based methods for identifying important knowledge components in textbooks and their usefulness in linking book subsections. To overcome the data sparsity in representing textbook content, we evaluated the utility of external corpuses, such as more textbooks or other online educational resources in the same domain. Our results show that semantic modeling can be integrated with a term-based approach for additional performance improvement, and that using extra textbooks significantly benefits semantic modeling. Similar results are obtained when we applied the same approach to other domains.
Shuguang Han, Yun Huang 0002, Daqing He, Peter Brusilovsky
WI4
2016 Finding cultural heritage images through a Dual-Perspective Navigation Framework
Peter Brusilovsky, Daqing He
Inf. Process. Manag.3
2015 User participation in an academic social networking service: A survey of open group users on Mendeley
abstract
Although there are a number of social networking services that specifically target scholars, little has been published about the actual practices and the usage of these so‐called academic social networking services (ASNSs). To fill this gap, we explore the populations of academics who engage in social activities using an ASNS; as an indicator of further engagement, we also determine their various motivations for joining a group in ASNSs. Using groups and their members in Mendeley as the platform for our case study, we obtained 146 participant responses from our online survey about users' common activities, usage habits, and motivations for joining groups. Our results show that (a) participants did not engage with social‐based features as frequently and actively as they engaged with research‐based features, and (b) users who joined more groups seemed to have a stronger motivation to increase their professional visibility and to contribute the research articles that they had read to the group reading list. Our results generate interesting insights into Mendeley's user populations, their activities, and their motivations relative to the social features of Mendeley. We also argue that further design of ASNSs is needed to take greater account of disciplinary differences in scholarly communication and to establish incentive mechanisms for encouraging user participation.
Wei Jeng, Daqing He, Jiepu Jiang
J. Assoc. Inf. Sci. Technol.2
2015 The impact of image descriptions on user tagging behavior: A study of the nature and functionality of crowdsourced tags
abstract
Crowdsourcing has emerged as a way to harvest social wisdom from thousands of volunteers to perform a series of tasks online. However, little research has been devoted to exploring the impact of various factors such as the content of a resource or crowdsourcing interface design on user tagging behavior. Although images' titles and descriptions are frequently available in image digital libraries, it is not clear whether they should be displayed to crowdworkers engaged in tagging. This paper focuses on offering insight to the curators of digital image libraries who face this dilemma by examining (i) how descriptions influence the user in his/her tagging behavior and (ii) how this relates to the (a) nature of the tags, (b) the emergent folksonomy, and (c) the findability of the images in the tagging system. We compared two different methods for collecting image tags from Amazon's Mechanical Turk's crowdworkers—with and without image descriptions. Several properties of generated tags were examined from different perspectives: diversity, specificity, reusability, quality, similarity, descriptiveness, and so on. In addition, the study was carried out to examine the impact of image description on supporting users' information seeking with a tag cloud interface. The results showed that the properties of tags are affected by the crowdsourcing approach. Tags from the “with description” condition are more diverse and more specific than tags from the “without description” condition, while the latter has a higher tag reuse rate. A user study also revealed that different tag sets provided different support for search. Tags produced “with description” shortened the path to the target results, whereas tags produced without description increased user success in the search task.
Christoph Trattner, Peter Brusilovsky, Daqing He
J. Assoc. Inf. Sci. Technol.4
2015 Understanding and Supporting Cross-Device Web Search for Exploratory Tasks with Mobile Touch Interactions
abstract
Mobile devices enable people to look for information at the moment when their information needs are triggered. While experiencing complex information needs that require multiple search sessions, users may utilize desktop computers to fulfill information needs started on mobile devices. Under the context of mobile-to-desktop web search, this article analyzes users’ behavioral patterns and compares them to the patterns in desktop-to-desktop web search. Then, we examine several approaches of using Mobile Touch Interactions (MTIs) to infer relevant content so that such content can be used for supporting subsequent search queries on desktop computers. The experimental data used in this article was collected through a user study involving 24 participants and six properly designed cross-device web search tasks. Our experimental results show that (1) users’ mobile-to-desktop search behaviors do significantly differ from desktop-to-desktop search behaviors in terms of information exploration, sense-making and repeated behaviors. (2) MTIs can be employed to predict the relevance of click-through documents, but applying document-level relevant content based on the predicted relevance does not improve search performance. (3) MTIs can also be used to identify the relevant text chunks at a fine-grained subdocument level. Such relevant information can achieve better search performance than the document-level relevant content. In addition, such subdocument relevant information can be combined with document-level relevance to further improve the search performance. However, the effectiveness of these methods relies on the sufficiency of click-through documents. (4) MTIs can also be obtained from the Search Engine Results Pages (SERPs). The subdocument feedbacks inferred from this set of MTIs even outperform the MTI-based subdocument feedback from the click-through documents.
Shuguang Han, Zhen Yue, Daqing He
ACM Trans. Inf. Syst.3
2014 Searching, browsing, and clicking in a search session: changes in user behavior by task and over time
abstract
There are many existing studies of user behavior in simple tasks (e.g., navigational and informational search) within a short duration of 1--2 queries. However, we know relatively little about user behavior, especially browsing and clicking behavior, for longer search session solving complex search tasks. In this paper, we characterize and compare user behavior in relatively long search sessions (10 minutes; about 5 queries) for search tasks of four different types. The tasks differ in two dimensions: (1) the user is locating facts or is pursuing intellectual understanding of a topic; (2) the user has a specific task goal or has an ill-defined and undeveloped goal. We analyze how search behavior as well as browsing and clicking patterns change during a search session in these different tasks. Our results indicate that user behavior in the four types of tasks differ in various aspects, including search activeness, browsing style, clicking strategy, and query reformulation. As a search session progresses, we note that users shift their interests to focus less on the top results but more on results ranked at lower positions in browsing. We also found that results eventually become less and less attractive for the users. The reasons vary and include downgraded search performance of query, decreased novelty of search results, and decaying persistence of users in browsing. Our study highlights the lack of long session support in existing search engines and suggests different strategies of supporting longer sessions according to different task types.
Jiepu Jiang, Daqing He, James Allan 0001
SIGIR2
2013 Supporting exploratory people search: a study of factor transparency and user control
abstract
People search is an active research topic in recent years. Related works includes expert finding, collaborator recommendation, link prediction and social matching. However, the diverse objectives and exploratory nature of those tasks make it difficult to develop a flexible method for people search that works for every task. In this project, we developed PeopleExplorer, an interactive people search system to support exploratory search tasks when looking for people. In the system, users could specify their task objectives by selecting and adjusting key criteria. Three criteria were considered: the content relevance, the candidate authoritativeness and the social similarity between the user and the candidates. This project represents a first attempt to add transparency to exploratory people search, and to give users full control over the search process. The system was evaluated through an experiment with 24 participants undertaking four different tasks. The results show that with comparable time and effort, users of our system performed significantly better in their people search tasks than those using the baseline system. Users of our system also exhibited many unique behaviors in query reformulation and candidate selection. We found that users' general perceptions about three criteria varied during different tasks, which confirms our assumptions regarding modeling task difference and user variance in people search systems.
Shuguang Han, Daqing He, Jiepu Jiang, Zhen Yue
CIKM2
2013 How do users respond to voice input errors?: lexical and phonetic query reformulation in voice search
abstract
Voice search offers users with a new search experience: instead of typing, users can vocalize their search queries. However, due to voice input errors (such as speech recognition errors and improper system interruptions), users need to frequently reformulate queries to handle the incorrectly recognized queries. We conducted user experiments with native English speakers on their query reformulation behaviors in voice search and found that users often reformulate queries with both lexical and phonetic changes to previous queries. In this paper, we first characterize and analyze typical voice input errors in voice search and users' corresponding reformulation strategies. Then, we evaluate the impacts of typical voice input errors on users' search progress and the effectiveness of different reformulation strategies on handling these errors. This study provides a clearer picture on how to further improve current voice search systems.
Jiepu Jiang, Wei Jeng, Daqing He
SIGIR3
2012 Contextual evaluation of query reformulations in a search session by user simulation
abstract
We propose a method to dynamically estimate the utility of documents in a search session by modeling the users' browsing behaviors and novelty. The method can be applied to evaluate query reformulations in a search session.
Jiepu Jiang, Daqing He, Shuguang Han, Zhen Yue, Chaoqun Ni
CIKM2
2012 Where do the query terms come from?: an analysis of query reformulation in collaborative web search
abstract
This paper presents a user study aiming to investigate the query reformulation in collaborative Web search. 7 pairs of participants were recruited and each pair worked as a team on two collaborative exploratory Web search tasks. Through the log analysis, we compared possible sources for participants to draw query terms from. The results show that both search and collaborative actions are possible resources for new query terms. Traditional resources for query expansion such as previous search histories and relevant documents are still important resources for new query terms. The content in chat and workspace generated by participants themselves seems more likely to be the resource for new query terms than that of their partners. Task types also affect the influences on query reformulations. For the academic task, previously saved relevance documents are the most important resources for new query terms while chat histories are the most important resources for the leisure task.
Zhen Yue, Jiepu Jiang, Shuguang Han, Daqing He
CIKM4
2012 Finding readings for scientists from social websites
abstract
Current search systems are designed to find relevant articles, especially topically relevant ones, but the notion of relevance largely depends on search tasks. We study the specific task that scientists are searching for worth-reading articles beneficial for their research. Our study finds: users' perception of relevance and preference of reading are only moderately correlated; current systems can effectively find readings that are highly relevant to the topic, but 36% of the worth-reading articles are only marginally relevant or even non-relevant. Our system can effectively find those worth-reading but marginally relevant or non-relevant articles by taking advantages of scientists' recommendations in social websites.
Jiepu Jiang, Zhen Yue, Shuguang Han, Daqing He
SIGIR4
2011 Enhancing query translation with relevance feedback in translingual information retrieval
Daqing He, Dan Wu 0003
Inf. Process. Manag.1
2010 CiteData: a new multi-faceted dataset for evaluating personalized search performance
abstract
Personalized search systems have evolved to utilize heterogeneous features including document hyperlinks, category labels in various taxonomies and social tags in addition to free-text of the documents. Consequently, classifiers, PageRank algorithms and Collaborative Filtering methods are often used as intermediate steps in such personalized retrieval systems. Thorough comparative evaluation of such complex systems has been difficult due to the lack of appropriate publicly available datasets that provide such diverse feature sets. To remedy the situation, we have created CiteData, a new dataset for benchmark evaluations of personalized search performance, that will be made publicly accessible. CiteData is a collection of academic articles extracted from CiteULike and CiteSeer repositories, with rich feature sets such as authors, author-affiliations, topic labels, social tags and citation information. We further supplement it with personalized queries and relevance judgments which were obtained from volunteer users. This paper starts with a discussion of the design criteria and characteristics of the CiteData dataset in comparison with current benchmark datasets, followed by a set of task-oriented empirical evaluations of popular algorithms in statistical classification, collaborative filtering and link analysis as intermediate steps for personalized search. Our results show significant performance improvement of personalized approaches, over that of unpersonalized approaches. We also observe that a meta personalized search engine that leverages information from multiple sources of features performs better than algorithms that use only one of the constituent source of features.
Abhay Harpale, Yiming Yang 0002, Siddharth Gopal, Daqing He, Zhen Yue
CIKM4
2010 Semantic annotation based exploratory search for information analysts
Jae-wook Ahn, Peter Brusilovsky, Jonathan Grady, Daqing He, Radu Florian
Inf. Process. Manag.4
2009 Pseudo relevance feedback using semantic clustering in relevance language model
abstract
Pseudo relevance feedback has demonstrated to be in general an effective technique for improving retrieval effectiveness, but the noise in the top retrieved documents still can cause topic drift problem that affects the performance of certain topics. By viewing a document as an interaction of a set of independent hidden topics, we propose a novel semantic clustering technique using independent component analysis. Then within the language modeling framework, we apply the obtained semantic topic clusters into the query sampling process so that the sampling depends on the activated topics rather than on the individual document language model. Therefore, we obtain a semantic cluster based relevance language model, which uses pseudo relevance feedback technique without requiring any relevance training information. We applied the model on five TREC data sets. The experiments show that our model can significantly improve retrieval performance over traditional language models including relevance-based and clustering-based retrieval language models. The main contribution of the improvements comes from the estimation of the relevance model on the semantic clusters that are closely related to the query.
Qiang Pu, Daqing He
CIKM2
2008 Translation enhancement: a new relevance feedback method for cross-language information retrieval
abstract
As an effective technique for improving retrieval effectiveness, relevance feedback (RF) has been widely studied in both monolingual and cross-language information retrieval (CLIR) settings. The studies of RF in CLIR have been focused on query expansion (QE), in which queries are reformulated before and/or after they are translated. However, RF in CLIR actually not only can help select better query terms, but also can enhance query translation by adjusting translation probabilities and even resolve some out-of-vocabulary terms. In this paper, we propose a novel RF method called translation enhancement (TE), which uses the extracted translation relationships from relevant documents to revise the translation probabilities of query terms and to identify extra translation alternatives if available so that the translated queries are more tuned to the current search. We studied TE using pseudo relevance feedback (PRF) and interactive relevance feedback (IRF). Our results show that TE can significantly improve CLIR with both types of RF methods, and that the improvement is comparable to that of QE. More importantly, the effects of TE and QE are complementary. Their integration can produce further improvement, and makes CLIR more robust for a variety of queries.
Daqing He, Dan Wu 0003
CIKM1
2008 Ice-tea: an interactive cross-language search engine with translation enhancement
abstract
No abstract available.
Dan Wu 0003, Daqing He
SIGIR2
2008 Personalized web exploration with task models
abstract
Personalized Web search has emerged as one of the hottest topics for both the Web industry and academic researchers. However, the majority of studies on personalized search focused on a rather simple type of search, which leaves an important research topic - the personalization in exploratory searches - as an under-studied area. In this paper, we present a study of personalization in task-based information exploration using a system called TaskSieve. TaskSieve is a Web search system that utilizes a relevance feedback based profile, called a "task model", for personalization. Its innovations include flexible and user controlled integration of queries and task models, task-infused text snippet generation, and on-screen visualization of task models. Through an empirical study using human subjects conducting task-based exploration searches, we demonstrate that TaskSieve pushes significantly more relevant documents to the top of search result lists as compared to a traditional search system. TaskSieve helps users select significantly more accurate information for their tasks, allows the users to do so with higher productivity, and is viewed more favorably by subjects under several usability related characteristics.
Jae-wook Ahn, Peter Brusilovsky, Daqing He, Jonathan Grady
WWW3
2008 An evaluation of adaptive filtering in the context of realistic task-based information exploration
Daqing He, Peter Brusilovsky, Jae-wook Ahn, Jonathan Grady, Rosta Farzan, Yefei Peng, Yiming Yang 0002, Monica Rogati
Inf. Process. Manag.1
2008 User-assisted query translation for interactive cross-language information retrieval
Douglas W. Oard, Daqing He, Jianqiang Wang 0002
Inf. Process. Manag.2
2007 How Up-to-date should it be? the Value of Instant Profiling and Adaptation in Information Filtering
abstract
In profile-based or content-based adaptive systems, one of the open research questions is how frequently the user's profile and the list of recommended items should be updated. Different systems tend to choose one of the two extremes. Some systems do it once per session (thus called between-session update strategy), whereas some others update whenever there is feedback (called instant update strategy). This paper presents our attempt to assess the value of keeping the list of recommended items up-to-date in the context of task-based information exploration. We conducted controlled studies involving human users performing realistic tasks using two systems that have the same adaptive filtering engine but with the above two different update strategies. Our results show that the between-session strategy helped to find better quality information, and received better subjects' responses about its usefulness and usability. However, it prolonged the selection of useful passages, whereas the instant update strategy helped subjects to obtain almost all of their selected passages (>98%) within the first 5 minutes. Based on the results, we hypothesize that the best strategy for updating might be a hybrid between the two update strategies, where both adaptability and stability can be achieved.
Daqing He, Peter Brusilovsky, Jonathan Grady, Jae-wook Ahn
Web Intelligence1
2007 Open user profiles for adaptive news systems: help or harm?
abstract
Over the last five years, a range of projects have focused on progressively more elaborated techniques for adaptive news delivery. However, the adaptation process in these systems has become more complicated and thus less transparent to the users. In this paper, we concentrate on the application of open user models in adding transparency and controllability to adaptive news systems. We present a personalized news system, YourNews, which allows users to view and edit their interest profiles, and report a user study on the system. Our results confirm that users prefer transparency and control in their systems, and generate more trust to such systems. However, similar to previous studies, our study demonstrate that this ability to edit user profiles may also harm the system.s performance and has to be used with caution.
Jae-wook Ahn, Peter Brusilovsky, Jonathan Grady, Daqing He, Sue Yeon Syn
WWW4
2006 Direct comparison of commercial and academic retrieval system: an initial study
abstract
No abstract available.
Yefei Peng, Daqing He
CIKM2
2006 Comparing two blind relevance feedback techniques
abstract
No abstract available.
Daqing He, Yefei Peng
SIGIR1
2006 DiLight: an ontology-based information access system for e-learning environments
abstract
No abstract available.
Ming Mao, Yefei Peng, Daqing He
SIGIR3
2006 Geographic Named Entity Disambiguation with Automatic Profile Generation
abstract
Knowledge rich approach of processing documents has been viewed as a method to improve over simple bag-of-word representation. Extracting location information from documents and link them to some ontology such as world gazetteer through a disambiguation process becomes an interesting and important topic. Lacking of training data is a problem in disambiguation method. In this paper we described a method to automatically extract training data from large collection of documents based on local context disambiguation, and then sense profiles are generated automatically for disambiguation use. Another topic of this paper is to describe a linear combination method to combine different types of evidences of disambiguation. We explored three different evidences including location sense context in training documents, local neighbor context, and the popularity of individual location sense. Our results show that combining the three evidences generates reasonable results
Yefei Peng, Daqing He, Ming Mao
Web Intelligence2
2003 User-assisted query translation for interactive CLIR
abstract
No abstract available.
Daqing He, Jianqiang Wang 0002, Douglas W. Oard, Michael Nossal
SIGIR1
2002 Combining evidence for automatic Web session identification
Daqing He, Ayse Göker, David J. Harper
Inf. Process. Manag.1