VLDB 2026 Research / reviewers in the wild / expert
Daqing He
dblp:41/134
· DBLP profile ↗
46ranked-venue papers in the field
8as first author
7since 2021 · last 2026
0000-0002-4645-8696ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 41 (7 first)Other / Interdisciplinary · 4 (1 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG FrameworkabstractWhile retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents provided by RAG system to manipulate LLM outputs. To counter this threat, we propose SecureCollaRAG, a Byzantine-tolerant collaborative RAG framework leveraging Multi-source Knowledge Validation Mechanism. Our approach enables agent system to securely verify document provenance through dynamic GNN-based credibility scoring, effectively preventing stealthy knowledge corruption attacks while preserving essential domain knowledge integrity. Through extensive evaluations and formal analysis, we demonstrate that SecureCollaRAG maintains robustness against attackers under non-IID data distributions. Daqing He, Zijian Zhang 0001, Ye Liu 0012, Jiamou Liu, Zhirui Zeng, Zhan Qin, Xin Li 0033, Hongwei Yao, Jincheng An, Yi Li 0008, Xiulei Liu, Liehuang Zhu |
WWW | 2 |
| 2025 | Reason-to-Rank: Distilling Direct and Comparative Reasoning from Large Language Models for Document RerankingabstractReranking documents in information retrieval often relies on black-box models that improve effectiveness but lack explainability. We introduce Reason-to-Rank (R2R), a novel framework that separates direct relevance reasoning from comparison reasoning to provide both direct and comparitive explanations. We first prompt a large language model to produce comprehensive rationales and a ranking order; then we distill both the ranking decisions and textual explanations into a smaller, open-source student model. Our approach not only improves retrieval performance, as demonstrated in MSMARCO, BEIR, and BRIGHT, but also provides interpretable justifications for why one document outranks another. We report NDCG@5 (and NDCG@10) for direct comparisons with prior work, and show that the distilled student model achieves competitive results while significantly reducing computational overhead. By unifying direct and comparative reasoning in a single pipeline, R2R bridges the gap between transparency and effectiveness in modern reranking systems. Yuelyu Ji, Zhuochun Li, Daqing He |
SIGIR | 4 |
| 2025 | Guest editorial of the IPM special issue on information science in human-centered AI
Dan Wu 0003, Daqing He, Preben Hansen, Shaobo Liang |
Inf. Process. Manag. | 2 |
| 2023 | Mapping dementia caregivers' comments on social media with evidence-based care strategies for memory loss and confusionabstractDementia caregivers widely turn to social media for much needed information and support. Prior research on caregivers’ online information exchange has focused on the original questions or posts, without including peer comments in response to those original questions, which does not provide a complete picture of online information exchanges between caregivers and their peers. This paper provides a preliminary analysis of a subset of data from a larger project and suggests how peer comments might match with evidence-based care strategies for memory loss and confusion. A total of 954 peer comments on 114 Reddit posts for dementia memory loss and confusion were collected and mapped with 5 evidence-based strategies generated from 17 care strategies from 3 credible websites. Our results report how well peer comments on Reddit can map to existing evidence-based strategies, and provide preliminary evidence supporting the necessity of providing tailored information based on the patient’s characteristics, stages of disease, and the progression level of memory loss and confusion. Ning Zou, Yuelyu Ji, Bo Xie 0001, Daqing He, Zhimeng Luo |
CHIIR | 4 |
| 2023 | Promoting data use through understanding user behaviors: A model for human open government data interactionabstractAbstract Recent dramatic increases in the ability to generate, collect, and use datasets have inspired numerous academic and policy discussions regarding the emerging field of human data interaction (HDI). Given the challenges in interacting with open government data (OGD) and the existing research gap in this field, our study intends to explore HDI in the OGD domain and investigate ways HDI can further contribute to OGD promotion. Building upon two existing behavioral models, we proposed an initial conceptual model for OGD interaction, then using this model, conducted two studies to empirically examine users' behaviors when interacting with OGD. Ultimately, we refined this model for OGD interaction and invited three experts to validate it to enhance its understandability, comprehensiveness, and reasonableness. This comprehensive model for human OGD interaction will contribute to the theoretical work of the HDI field as well as the practical design of OGD platforms and data literacy education. Fanghui Xiao, Yu Chi 0001, Daqing He |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2022 | HELPeR: An Interactive Recommender System for Ovarian Cancer Patients and CaregiversabstractRecommending online resources to patients with ovarian cancer and their caregivers is a challenging task. On one hand, the recommended items must be relevant, recent, and reliable. On the other hand, they need to match the user’s levels of disease-specific health literacy. In this demonstration, we describe the overall architecture and key components of HELPeR, a knowledge-adaptive interactive recommender system for ovarian cancer patients and their caregivers. Behnam Rahdari, Peter Brusilovsky, Daqing He, Khushboo Thaker, Zhimeng Luo, Young Ji Lee |
RecSys | 3 |
| 2022 | Together they shall not fade away: Opportunities and challenges of self-tracking for dementia care
Ning Zou, Yu Chi 0001, Daqing He, Bo Xie 0001 |
Inf. Process. Manag. | 3 |
| 2020 | Laypeople's source selection in online health information-seeking processabstractAbstract For laypeople, searching online health information resources can be challenging due to topic complexity and the large number of online sources with differing quality. The goal of this article is to examine, among all the available online sources, which online sources laypeople select to address their health‐related information needs, and whether or how much the severity of a health condition influences their selection. Twenty‐four participants were recruited individually, and each was asked (using a retrieval system called HIS) to search for information regarding a severe health condition and a mild health condition, respectively. The selected online health information sources were automatically captured by the HIS system and classified at both the website and webpage levels. Participants' selection behavior patterns were then plotted across the whole information‐seeking process. Our results demonstrate that laypeople's source selection fluctuates during the health information‐seeking process, and also varies by the severity of health conditions. This study reveals laypeople's real usage of different types of online health information sources, and engenders implications to the design of search engines, as well as the development of health literacy programs. Yu Chi 0001, Daqing He, Wei Jeng |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2020 | Global health crises are also information crises: A call to actionabstractAbstract In this opinion paper, we argue that global health crises are also information crises. Using as an example the coronavirus disease 2019 (COVID‐19) epidemic, we (a) examine challenges associated with what we term “global information crises”; (b) recommend changes needed for the field of information science to play a leading role in such crises; and (c) propose actionable items for short‐ and long‐term research, education, and practice in information science. Bo Xie 0001, Daqing He, Tim Mercer, Youfa Wang, Dan Wu 0003, Kenneth R. Fleischmann, Yan Zhang 0005, Linda H. Yoder, Keri K. Stephens, Michael Mackert, Min Kyung Lee |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2019 | Challenges and Supports for Accessing Open Government Datasets: Data Guide for Better Open Data Access and UsesabstractThe importance of open government data is often associated with increased public trust, civic engagement, and accountable administrations. While there is a myriad of benefits, the existing literature suggests that many open government datasets lack accessibility and usability for diverse users. This study seeks to explore what contextual information users require when they access these datasets. Using mixed methods, we aim to discover the challenges of accessing data, and the necessary contextual information needed by the users to overcome these challenges. As the outcome of this study, we propose a framework called "Data Guides", which is composed of the identified important contextual information. In future work, we will test the effectiveness of the Data Guide in aiding users' accessing and understanding open government data. Fanghui Xiao, Daqing He, Yu Chi 0001, Wei Jeng, Christinger Tomer |
CHIIR | 2 |
| 2018 | What Sources to Rely on: : Laypeople's Source Selection in Online Health Information SeekingabstractIn this study, we examined what sources laypeople would select (i.e., visit and adopt) to resolve their health-related information needs, and how different health conditions affect the selection. Twenty-four college students participated in this user study, where they were asked to search for two separate health issues respectively: multiple sclerosis and weight loss. The search logs were collected and analyzed afterwards. We classify the online information sources on both website level and webpage level, and a webpage classification scheme based on genre is proposed. Results suggest that users» selection of sources depends on different types of health issues in terms of urgency and complexity. Health-specific webpage is a popular source and highly adopted for both tasks, but it is particularly helpful for urgent and complex health conditions. Search engines could facilitate users to navigate among scattered health information and support concerns regarding common health issues. Yu Chi 0001, Daqing He, Shuguang Han |
CHIIR | 2 |
| 2018 | Concept Enhanced Content Representation for Linking Educational ResourcesabstractThe education sector has been undergoing a welcoming change in recent years with the introduction of a wide variety of digital content openly available to students. Due to the volume of this new digital content, it is very difficult for learners to find the needed information at the right time. Digital textbooks, as well-curated domain knowledge sources, could provide a conceptual and physical platform that unites disparate educational resources as one entity. Educational resource linkage, with state-of-the-art techniques, are based on term-level and topic-level representations. However, term-level representations suffer from the term-mismatch problem and often, topics are too broad for linking to other educational resources. To address these challenges, we propose to link educational resources through concept-level representation. The proposed model generates concept embeddings by utilizing domain-specific educational content and external knowledge graph resources to achieve robust and effective concept-level representations. We conducted evaluations of the proposed models on multiple contents linking tasks, and the results demonstrate that concept-level representations perform better than the state-of-the-art representations in helping students to find more learning resources easily. This could increase both students' learning and satisfaction. Khushboo Thaker, Peter Brusilovsky, Daqing He |
WI | 3 |
| 2018 | Mobile Information Retrieval. Fabio Crestani, Stefano Mizzaro, and Ivan Scagnetto. Cham, Switzerland: Springer, 2017. 110 pp. $54.99 (softcover). (ISBN 9783319607764)
Daqing He |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | Understanding Ephemeral State of RelevanceabstractDespite its dynamic nature, relevance is often measured in a context-independent manner in information retrieval practice. We look into this discrepancy. We propose a contextual relevance/usefulness measurement called ephemeral state of relevance (ESR), which is defined as the amount of useful information a user acquired from a clicked result as assessed just after examining the result during an interactive search session. We collect ESR and context-independent usefulness judgments through a laboratory user study and compare the two. We examine factors related to both judgments and examine their differences. Jiepu Jiang, Daqing He, Diane Kelly 0001, James Allan 0001 |
CHIIR | 2 |
| 2017 | Semi-Supervised Techniques for Mining Learning Outcomes and PrerequisitesabstractEducational content of today no longer only resides in textbooks and classrooms; more and more learning material is found in a free, accessible form on the Internet. Our long-standing vision is to transform this web of educational content into an adaptive, web-scale "textbook", that can guide its readers to most relevant "pages" according to their learning goal and current knowledge. In this paper, we address one core, long-standing problem towards this goal: identifying outcome and prerequisite concepts within a piece of educational content (e.g., a tutorial). Specifically, we propose a novel approach that leverages textbooks as a source of distant supervision, but learns a model that can generalize to arbitrary documents (such as those on the web). As such, our model can take advantage of any existing textbook, without requiring expert annotation. At the task of predicting outcome and prerequisite concepts, we demonstrate improvements over a number of baselines on six textbooks, especially in the regime of little to no ground-truth labels available. Finally, we demonstrate the utility of a model learned using our approach at the task of identifying prerequisite documents for adaptive content recommendation --- an important step towards our vision of the "web as a textbook". Igor Labutov, Yun Huang 0002, Peter Brusilovsky, Daqing He |
KDD | 4 |
| 2017 | Comparing In Situ and Multidimensional Relevance JudgmentsabstractTo address concerns of TREC-style relevance judgments, we explore two improvements. The first one seeks to make relevance judgments contextual, collecting in situ feedback of users in an interactive search session and embracing usefulness as the primary judgment criterion. The second one collects multidimensional assessments to complement relevance or usefulness judgments, with four distinct alternative aspects examined in this paper - novelty, understandability, reliability, and effort. Jiepu Jiang, Daqing He, James Allan 0001 |
SIGIR | 2 |
| 2017 | Information exchange on an academic social networking site: A multidiscipline comparison on researchgate Q&AabstractThe increasing popularity of academic social networking sites (ASNSs) requires studies on the usage of ASNSs among scholars and evaluations of the effectiveness of these ASNSs. However, it is unclear whether current ASNSs have fulfilled their design goal, as scholars' actual online interactions on these platforms remain unexplored. To fill the gap, this article presents a study based on data collected from ResearchGate. Adopting a mixed‐method design by conducting qualitative content analysis and statistical analysis on 1,128 posts collected from ResearchGate Q&A, we examine how scholars exchange information and resources, and how their practices vary across three distinct disciplines: library and information services, history of art, and astrophysics. Our results show that the effect of a questioner's intention (i.e., seeking information or discussion) is greater than disciplinary factors in some circumstances. Across the three disciplines, responses to questions provide various resources, including experts' contact details, citations, links to Wikipedia, images, and so on. We further discuss several implications of the understanding of scholarly information exchange and the design of better academic social networking interfaces, which should stimulate scholarly interactions by minimizing confusion, improving the clarity of questions, and promoting scholarly content management. Wei Jeng, Spencer DesAutels, Daqing He, Lei Li 0037 |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | Contextual Support for Collaborative Information RetrievalabstractRecent research shows that Collaborative Information Retrieval (CIR), in which two or more users collaborate on the same search task, has become increasingly popular. The presence of both search and collaboration behaviors makes CIR a complex search format, which further drives a critical need to understand CIR's search context. The contextual support for CIR should consider search contexts derived from both team members' search histories (including users' own search histories and partners' search histories) and their explicit collaboration (e.g., chatting). As it stands, existing studies on contextual search support only focus on Individual Information Retrieval (IIR) and only utilize individuals' own search histories. In this paper, we examine the unique search contexts (e.g., partners' search histories and team collaboration histories) in CIR. Based on a user study data collection with 54 participants, we find that compared to the use of individuals' own search histories, CIR contextual support is more effective when utilizing partners' search histories and teams' collaboration behaviors. More interestingly, though the explicit communication information (i.e., chat content) often involves massive noisy information, involving such noise does not affect the ranking of relevant documents since it also does not appear in relevant documents. Shuguang Han, Daqing He, Zhen Yue, Jiepu Jiang |
CHIIR | 2 |
| 2016 | Knowledge-Based Content Linking for Online TextbooksabstractAlthough the volume of online educational resources has dramatically increased in recent years, many of these resources are isolated and distributed in diverse websites and databases. This hinders the discovery and overall usage of online educational resources. By using linking between related subsections of online textbooks as a testbed, this paper explores multiple knowledge-based content linking algorithms for connecting online educational resources. We focus on examining semantic-based methods for identifying important knowledge components in textbooks and their usefulness in linking book subsections. To overcome the data sparsity in representing textbook content, we evaluated the utility of external corpuses, such as more textbooks or other online educational resources in the same domain. Our results show that semantic modeling can be integrated with a term-based approach for additional performance improvement, and that using extra textbooks significantly benefits semantic modeling. Similar results are obtained when we applied the same approach to other domains. Shuguang Han, Yun Huang 0002, Daqing He, Peter Brusilovsky |
WI | 4 |
| 2016 | Finding cultural heritage images through a Dual-Perspective Navigation Framework
Peter Brusilovsky, Daqing He |
Inf. Process. Manag. | 3 |
| 2015 | User participation in an academic social networking service: A survey of open group users on MendeleyabstractAlthough there are a number of social networking services that specifically target scholars, little has been published about the actual practices and the usage of these so‐called academic social networking services (ASNSs). To fill this gap, we explore the populations of academics who engage in social activities using an ASNS; as an indicator of further engagement, we also determine their various motivations for joining a group in ASNSs. Using groups and their members in Mendeley as the platform for our case study, we obtained 146 participant responses from our online survey about users' common activities, usage habits, and motivations for joining groups. Our results show that (a) participants did not engage with social‐based features as frequently and actively as they engaged with research‐based features, and (b) users who joined more groups seemed to have a stronger motivation to increase their professional visibility and to contribute the research articles that they had read to the group reading list. Our results generate interesting insights into Mendeley's user populations, their activities, and their motivations relative to the social features of Mendeley. We also argue that further design of ASNSs is needed to take greater account of disciplinary differences in scholarly communication and to establish incentive mechanisms for encouraging user participation. Wei Jeng, Daqing He, Jiepu Jiang |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2015 | The impact of image descriptions on user tagging behavior: A study of the nature and functionality of crowdsourced tagsabstractCrowdsourcing has emerged as a way to harvest social wisdom from thousands of volunteers to perform a series of tasks online. However, little research has been devoted to exploring the impact of various factors such as the content of a resource or crowdsourcing interface design on user tagging behavior. Although images' titles and descriptions are frequently available in image digital libraries, it is not clear whether they should be displayed to crowdworkers engaged in tagging. This paper focuses on offering insight to the curators of digital image libraries who face this dilemma by examining (i) how descriptions influence the user in his/her tagging behavior and (ii) how this relates to the (a) nature of the tags, (b) the emergent folksonomy, and (c) the findability of the images in the tagging system. We compared two different methods for collecting image tags from Amazon's Mechanical Turk's crowdworkers—with and without image descriptions. Several properties of generated tags were examined from different perspectives: diversity, specificity, reusability, quality, similarity, descriptiveness, and so on. In addition, the study was carried out to examine the impact of image description on supporting users' information seeking with a tag cloud interface. The results showed that the properties of tags are affected by the crowdsourcing approach. Tags from the “with description” condition are more diverse and more specific than tags from the “without description” condition, while the latter has a higher tag reuse rate. A user study also revealed that different tag sets provided different support for search. Tags produced “with description” shortened the path to the target results, whereas tags produced without description increased user success in the search task. Christoph Trattner, Peter Brusilovsky, Daqing He |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2015 | Understanding and Supporting Cross-Device Web Search for Exploratory Tasks with Mobile Touch InteractionsabstractMobile devices enable people to look for information at the moment when their information needs are triggered. While experiencing complex information needs that require multiple search sessions, users may utilize desktop computers to fulfill information needs started on mobile devices. Under the context of mobile-to-desktop web search, this article analyzes users’ behavioral patterns and compares them to the patterns in desktop-to-desktop web search. Then, we examine several approaches of using Mobile Touch Interactions (MTIs) to infer relevant content so that such content can be used for supporting subsequent search queries on desktop computers. The experimental data used in this article was collected through a user study involving 24 participants and six properly designed cross-device web search tasks. Our experimental results show that (1) users’ mobile-to-desktop search behaviors do significantly differ from desktop-to-desktop search behaviors in terms of information exploration, sense-making and repeated behaviors. (2) MTIs can be employed to predict the relevance of click-through documents, but applying document-level relevant content based on the predicted relevance does not improve search performance. (3) MTIs can also be used to identify the relevant text chunks at a fine-grained subdocument level. Such relevant information can achieve better search performance than the document-level relevant content. In addition, such subdocument relevant information can be combined with document-level relevance to further improve the search performance. However, the effectiveness of these methods relies on the sufficiency of click-through documents. (4) MTIs can also be obtained from the Search Engine Results Pages (SERPs). The subdocument feedbacks inferred from this set of MTIs even outperform the MTI-based subdocument feedback from the click-through documents. Shuguang Han, Zhen Yue, Daqing He |
ACM Trans. Inf. Syst. | 3 |
| 2014 | Searching, browsing, and clicking in a search session: changes in user behavior by task and over timeabstractThere are many existing studies of user behavior in simple tasks (e.g., navigational and informational search) within a short duration of 1--2 queries. However, we know relatively little about user behavior, especially browsing and clicking behavior, for longer search session solving complex search tasks. In this paper, we characterize and compare user behavior in relatively long search sessions (10 minutes; about 5 queries) for search tasks of four different types. The tasks differ in two dimensions: (1) the user is locating facts or is pursuing intellectual understanding of a topic; (2) the user has a specific task goal or has an ill-defined and undeveloped goal. We analyze how search behavior as well as browsing and clicking patterns change during a search session in these different tasks. Our results indicate that user behavior in the four types of tasks differ in various aspects, including search activeness, browsing style, clicking strategy, and query reformulation. As a search session progresses, we note that users shift their interests to focus less on the top results but more on results ranked at lower positions in browsing. We also found that results eventually become less and less attractive for the users. The reasons vary and include downgraded search performance of query, decreased novelty of search results, and decaying persistence of users in browsing. Our study highlights the lack of long session support in existing search engines and suggests different strategies of supporting longer sessions according to different task types. Jiepu Jiang, Daqing He, James Allan 0001 |
SIGIR | 2 |
| 2013 | Supporting exploratory people search: a study of factor transparency and user controlabstractPeople search is an active research topic in recent years. Related works includes expert finding, collaborator recommendation, link prediction and social matching. However, the diverse objectives and exploratory nature of those tasks make it difficult to develop a flexible method for people search that works for every task. In this project, we developed PeopleExplorer, an interactive people search system to support exploratory search tasks when looking for people. In the system, users could specify their task objectives by selecting and adjusting key criteria. Three criteria were considered: the content relevance, the candidate authoritativeness and the social similarity between the user and the candidates. This project represents a first attempt to add transparency to exploratory people search, and to give users full control over the search process. The system was evaluated through an experiment with 24 participants undertaking four different tasks. The results show that with comparable time and effort, users of our system performed significantly better in their people search tasks than those using the baseline system. Users of our system also exhibited many unique behaviors in query reformulation and candidate selection. We found that users' general perceptions about three criteria varied during different tasks, which confirms our assumptions regarding modeling task difference and user variance in people search systems. Shuguang Han, Daqing He, Jiepu Jiang, Zhen Yue |
CIKM | 2 |
| 2013 | How do users respond to voice input errors?: lexical and phonetic query reformulation in voice searchabstractVoice search offers users with a new search experience: instead of typing, users can vocalize their search queries. However, due to voice input errors (such as speech recognition errors and improper system interruptions), users need to frequently reformulate queries to handle the incorrectly recognized queries. We conducted user experiments with native English speakers on their query reformulation behaviors in voice search and found that users often reformulate queries with both lexical and phonetic changes to previous queries. In this paper, we first characterize and analyze typical voice input errors in voice search and users' corresponding reformulation strategies. Then, we evaluate the impacts of typical voice input errors on users' search progress and the effectiveness of different reformulation strategies on handling these errors. This study provides a clearer picture on how to further improve current voice search systems. Jiepu Jiang, Wei Jeng, Daqing He |
SIGIR | 3 |
| 2012 | Contextual evaluation of query reformulations in a search session by user simulationabstractWe propose a method to dynamically estimate the utility of documents in a search session by modeling the users' browsing behaviors and novelty. The method can be applied to evaluate query reformulations in a search session. Jiepu Jiang, Daqing He, Shuguang Han, Zhen Yue, Chaoqun Ni |
CIKM | 2 |
| 2012 | Where do the query terms come from?: an analysis of query reformulation in collaborative web searchabstractThis paper presents a user study aiming to investigate the query reformulation in collaborative Web search. 7 pairs of participants were recruited and each pair worked as a team on two collaborative exploratory Web search tasks. Through the log analysis, we compared possible sources for participants to draw query terms from. The results show that both search and collaborative actions are possible resources for new query terms. Traditional resources for query expansion such as previous search histories and relevant documents are still important resources for new query terms. The content in chat and workspace generated by participants themselves seems more likely to be the resource for new query terms than that of their partners. Task types also affect the influences on query reformulations. For the academic task, previously saved relevance documents are the most important resources for new query terms while chat histories are the most important resources for the leisure task. Zhen Yue, Jiepu Jiang, Shuguang Han, Daqing He |
CIKM | 4 |
| 2012 | Finding readings for scientists from social websitesabstractCurrent search systems are designed to find relevant articles, especially topically relevant ones, but the notion of relevance largely depends on search tasks. We study the specific task that scientists are searching for worth-reading articles beneficial for their research. Our study finds: users' perception of relevance and preference of reading are only moderately correlated; current systems can effectively find readings that are highly relevant to the topic, but 36% of the worth-reading articles are only marginally relevant or even non-relevant. Our system can effectively find those worth-reading but marginally relevant or non-relevant articles by taking advantages of scientists' recommendations in social websites. Jiepu Jiang, Zhen Yue, Shuguang Han, Daqing He |
SIGIR | 4 |
| 2011 | Enhancing query translation with relevance feedback in translingual information retrieval
Daqing He, Dan Wu 0003 |
Inf. Process. Manag. | 1 |
| 2010 | CiteData: a new multi-faceted dataset for evaluating personalized search performanceabstractPersonalized search systems have evolved to utilize heterogeneous features including document hyperlinks, category labels in various taxonomies and social tags in addition to free-text of the documents. Consequently, classifiers, PageRank algorithms and Collaborative Filtering methods are often used as intermediate steps in such personalized retrieval systems. Thorough comparative evaluation of such complex systems has been difficult due to the lack of appropriate publicly available datasets that provide such diverse feature sets. To remedy the situation, we have created CiteData, a new dataset for benchmark evaluations of personalized search performance, that will be made publicly accessible. CiteData is a collection of academic articles extracted from CiteULike and CiteSeer repositories, with rich feature sets such as authors, author-affiliations, topic labels, social tags and citation information. We further supplement it with personalized queries and relevance judgments which were obtained from volunteer users. This paper starts with a discussion of the design criteria and characteristics of the CiteData dataset in comparison with current benchmark datasets, followed by a set of task-oriented empirical evaluations of popular algorithms in statistical classification, collaborative filtering and link analysis as intermediate steps for personalized search. Our results show significant performance improvement of personalized approaches, over that of unpersonalized approaches. We also observe that a meta personalized search engine that leverages information from multiple sources of features performs better than algorithms that use only one of the constituent source of features. Abhay Harpale, Yiming Yang 0002, Siddharth Gopal, Daqing He, Zhen Yue |
CIKM | 4 |
| 2010 | Semantic annotation based exploratory search for information analysts
Jae-wook Ahn, Peter Brusilovsky, Jonathan Grady, Daqing He, Radu Florian |
Inf. Process. Manag. | 4 |
| 2009 | Pseudo relevance feedback using semantic clustering in relevance language modelabstractPseudo relevance feedback has demonstrated to be in general an effective technique for improving retrieval effectiveness, but the noise in the top retrieved documents still can cause topic drift problem that affects the performance of certain topics. By viewing a document as an interaction of a set of independent hidden topics, we propose a novel semantic clustering technique using independent component analysis. Then within the language modeling framework, we apply the obtained semantic topic clusters into the query sampling process so that the sampling depends on the activated topics rather than on the individual document language model. Therefore, we obtain a semantic cluster based relevance language model, which uses pseudo relevance feedback technique without requiring any relevance training information. We applied the model on five TREC data sets. The experiments show that our model can significantly improve retrieval performance over traditional language models including relevance-based and clustering-based retrieval language models. The main contribution of the improvements comes from the estimation of the relevance model on the semantic clusters that are closely related to the query. Qiang Pu, Daqing He |
CIKM | 2 |
| 2008 | Translation enhancement: a new relevance feedback method for cross-language information retrievalabstractAs an effective technique for improving retrieval effectiveness, relevance feedback (RF) has been widely studied in both monolingual and cross-language information retrieval (CLIR) settings. The studies of RF in CLIR have been focused on query expansion (QE), in which queries are reformulated before and/or after they are translated. However, RF in CLIR actually not only can help select better query terms, but also can enhance query translation by adjusting translation probabilities and even resolve some out-of-vocabulary terms. In this paper, we propose a novel RF method called translation enhancement (TE), which uses the extracted translation relationships from relevant documents to revise the translation probabilities of query terms and to identify extra translation alternatives if available so that the translated queries are more tuned to the current search. We studied TE using pseudo relevance feedback (PRF) and interactive relevance feedback (IRF). Our results show that TE can significantly improve CLIR with both types of RF methods, and that the improvement is comparable to that of QE. More importantly, the effects of TE and QE are complementary. Their integration can produce further improvement, and makes CLIR more robust for a variety of queries. Daqing He, Dan Wu 0003 |
CIKM | 1 |
| 2008 | Ice-tea: an interactive cross-language search engine with translation enhancementabstractNo abstract available. Dan Wu 0003, Daqing He |
SIGIR | 2 |
| 2008 | Personalized web exploration with task modelsabstractPersonalized Web search has emerged as one of the hottest topics for both the Web industry and academic researchers. However, the majority of studies on personalized search focused on a rather simple type of search, which leaves an important research topic - the personalization in exploratory searches - as an under-studied area. In this paper, we present a study of personalization in task-based information exploration using a system called TaskSieve. TaskSieve is a Web search system that utilizes a relevance feedback based profile, called a "task model", for personalization. Its innovations include flexible and user controlled integration of queries and task models, task-infused text snippet generation, and on-screen visualization of task models. Through an empirical study using human subjects conducting task-based exploration searches, we demonstrate that TaskSieve pushes significantly more relevant documents to the top of search result lists as compared to a traditional search system. TaskSieve helps users select significantly more accurate information for their tasks, allows the users to do so with higher productivity, and is viewed more favorably by subjects under several usability related characteristics. Jae-wook Ahn, Peter Brusilovsky, Daqing He, Jonathan Grady |
WWW | 3 |
| 2008 | An evaluation of adaptive filtering in the context of realistic task-based information exploration
Daqing He, Peter Brusilovsky, Jae-wook Ahn, Jonathan Grady, Rosta Farzan, Yefei Peng, Yiming Yang 0002, Monica Rogati |
Inf. Process. Manag. | 1 |
| 2008 | User-assisted query translation for interactive cross-language information retrieval
Douglas W. Oard, Daqing He, Jianqiang Wang 0002 |
Inf. Process. Manag. | 2 |
| 2007 | How Up-to-date should it be? the Value of Instant Profiling and Adaptation in Information FilteringabstractIn profile-based or content-based adaptive systems, one of the open research questions is how frequently the user's profile and the list of recommended items should be updated. Different systems tend to choose one of the two extremes. Some systems do it once per session (thus called between-session update strategy), whereas some others update whenever there is feedback (called instant update strategy). This paper presents our attempt to assess the value of keeping the list of recommended items up-to-date in the context of task-based information exploration. We conducted controlled studies involving human users performing realistic tasks using two systems that have the same adaptive filtering engine but with the above two different update strategies. Our results show that the between-session strategy helped to find better quality information, and received better subjects' responses about its usefulness and usability. However, it prolonged the selection of useful passages, whereas the instant update strategy helped subjects to obtain almost all of their selected passages (>98%) within the first 5 minutes. Based on the results, we hypothesize that the best strategy for updating might be a hybrid between the two update strategies, where both adaptability and stability can be achieved. Daqing He, Peter Brusilovsky, Jonathan Grady, Jae-wook Ahn |
Web Intelligence | 1 |
| 2007 | Open user profiles for adaptive news systems: help or harm?abstractOver the last five years, a range of projects have focused on progressively more elaborated techniques for adaptive news delivery. However, the adaptation process in these systems has become more complicated and thus less transparent to the users. In this paper, we concentrate on the application of open user models in adding transparency and controllability to adaptive news systems. We present a personalized news system, YourNews, which allows users to view and edit their interest profiles, and report a user study on the system. Our results confirm that users prefer transparency and control in their systems, and generate more trust to such systems. However, similar to previous studies, our study demonstrate that this ability to edit user profiles may also harm the system.s performance and has to be used with caution. Jae-wook Ahn, Peter Brusilovsky, Jonathan Grady, Daqing He, Sue Yeon Syn |
WWW | 4 |
| 2006 | Direct comparison of commercial and academic retrieval system: an initial studyabstractNo abstract available. Yefei Peng, Daqing He |
CIKM | 2 |
| 2006 | Comparing two blind relevance feedback techniquesabstractNo abstract available. Daqing He, Yefei Peng |
SIGIR | 1 |
| 2006 | DiLight: an ontology-based information access system for e-learning environmentsabstractNo abstract available. Ming Mao, Yefei Peng, Daqing He |
SIGIR | 3 |
| 2006 | Geographic Named Entity Disambiguation with Automatic Profile GenerationabstractKnowledge rich approach of processing documents has been viewed as a method to improve over simple bag-of-word representation. Extracting location information from documents and link them to some ontology such as world gazetteer through a disambiguation process becomes an interesting and important topic. Lacking of training data is a problem in disambiguation method. In this paper we described a method to automatically extract training data from large collection of documents based on local context disambiguation, and then sense profiles are generated automatically for disambiguation use. Another topic of this paper is to describe a linear combination method to combine different types of evidences of disambiguation. We explored three different evidences including location sense context in training documents, local neighbor context, and the popularity of individual location sense. Our results show that combining the three evidences generates reasonable results Yefei Peng, Daqing He, Ming Mao |
Web Intelligence | 2 |
| 2003 | User-assisted query translation for interactive CLIRabstractNo abstract available. Daqing He, Jianqiang Wang 0002, Douglas W. Oard, Michael Nossal |
SIGIR | 1 |
| 2002 | Combining evidence for automatic Web session identification
Daqing He, Ayse Göker, David J. Harper |
Inf. Process. Manag. | 1 |