EDBT 2026 Demo / reviewers in the wild / expert
Grace Hui Yang
dblp:04/999-1 · also Hui Yang 0001
· DBLP profile ↗
47ranked-venue papers in the field
14as first author
14since 2021 · last 2026
0000-0001-6095-8358ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 42 (12 first)Data Mining & Knowledge Discovery · 5 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information AccessabstractAs AI agents become more capable of anticipating intent and taking initiative, the ways humans seek, interpret, and act on information are being quietly reshaped. Yet at the heart of every interaction lies a human—curious, uncertain, and contextually situated, whose goals and boundaries cannot be fully captured by data alone. This workshop centers on the human experience of proactivity and personalization in interactive information access, asking how agents can assist without overriding agency, adapt without imposing assumptions, and anticipate without eroding trust. Building on CHIIR’s tradition of bridging information retrieval and human–computer interaction, the workshop will explore when and how proactivity supports human information behavior – enhancing exploration, sense-making, and learning – and when it risks diminishing transparency or control. Through co-design sessions and participatory discussions, we will interrogate concrete design and evaluation dimensions of proactive systems, including timing of initiative, transparency of intent, user control, and their effects on exploration, sense-making, and trust. Ultimately, this workshop seeks to reimagine proactivity not as automation of the search process, but as a collaborative partnership where agents act as companions in the human pursuit of understanding. All resources related to this workshop are available at https://proactive-chiir.github.io/. Kirandeep Kaur, Madhura Raju, Tanya G. Roosta, Grace Hui Yang, Chirag Shah 0001 |
CHIIR | 5 |
| 2025 | AgentIR: 2nd Workshop on Agent-based Information RetrievalabstractInformation retrieval (IR) systems are essential in modern society, aiding users to efficiently locate relevant information through query expansion, document retrieval, ranking, and re-ranking. User feedback from ranked outputs forms a dynamic interaction loop with IR systems, which can be modeled as either one-time or sequential decision-making problems. Over the past decade, deep reinforcement learning (DRL) has emerged as a promising approach to decision-making, leveraging the high model capacity of deep learning for complex tasks. While significant research has explored the application of DRL to IR tasks, several fundamental challenges remain underexplored, including the underlying information theory in DRL settings, the limitations of reinforcement learning methods for industrial IR applications, and the simulation of DRL-based IR systems. Concurrently, the advent of large language models (LLMs) has introduced new opportunities for optimizing and simulating IR systems. Building on the success of the Agent-based IR Workshop at SIGIR 2024, we propose hosting the second Agent-based IR Workshop at SIGIR 2025. This workshop will continue to provide a platform for researchers and practitioners from academia and industry to present cutting-edge advances in DRL-based and LLM-based IR systems from an agent-based perspective. By building on the foundation laid in the first workshop, the 2025 edition aims to delve deeper into emerging research challenges, foster collaborations, and explore innovative applications. Through engaging discussions and insightful presentations, the workshop seeks to further expand the boundaries of IR research and solidify its role as a premier venue for advancing agent-based IR systems. Pengyue Jia, Qingpeng Cai 0001, Xiangyu Zhao 0001, Ling Pan, Xin Xin 0003, Jin Huang 0010, Weinan Zhang 0001, Li Zhao 0007, Dawei Yin 0001, Grace Hui Yang |
SIGIR | 10 |
| 2025 | Proactive Conversational AI: A Comprehensive Survey of Advancements and OpportunitiesabstractDialogue systems are designed to offer human users social support or functional services through natural language interactions. Traditional conversation research has put significant emphasis on a system’s response-ability, including its capacity to understand dialogue context and generate appropriate responses. However, the key element of proactive behavior—a crucial aspect of intelligent conversations—is often overlooked in these studies. Proactivity empowers conversational agents to lead conversations towards achieving pre-defined targets or fulfilling specific goals on the system side. Proactive dialogue systems are equipped with advanced techniques to handle complex tasks, requiring strategic and motivational interactions, thus representing a significant step towards artificial general intelligence. Motivated by the necessity and challenges of building proactive dialogue systems, we provide a comprehensive review of various prominent problems and advanced designs for implementing proactivity into different types of dialogue systems, including open-domain dialogues, task-oriented dialogues, and information-seeking dialogues. We also discuss real-world challenges that require further research attention to meet application needs in the future, such as proactivity in dialogue systems that are based on large language models, proactivity in hybrid dialogues, evaluation protocols and ethical considerations for proactive dialogue systems. By providing a quick access and overall picture of the proactive dialogue systems domain, we aim to inspire new research directions and stimulate further advancements towards achieving the next level of conversational AI capabilities, paving the way for more dynamic and intelligent interactions within various application domains. Yang Deng 0002, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, Tat-Seng Chua |
ACM Trans. Inf. Syst. | 4 |
| 2025 | Pre-Trained Models for Search and Recommendation: Introduction to the Special Issue - Part 1
Wenjie Wang 0007, Zheng Liu 0011, Fuli Feng, Zhicheng Dou, Qingyao Ai, Grace Hui Yang, Defu Lian, Lu Hou 0002, Aixin Sun, Hamed Zamani, Donald Metzler, Maarten de Rijke |
ACM Trans. Inf. Syst. | 6 |
| 2025 | Pre-Trained Models for Search and Recommendation: Introduction to the Special Issue - Part 2
Wenjie Wang 0007, Zheng Liu 0011, Fuli Feng, Zhicheng Dou, Qingyao Ai, Grace Hui Yang, Defu Lian, Lu Hou 0002, Aixin Sun, Hamed Zamani, Donald Metzler, Maarten de Rijke |
ACM Trans. Inf. Syst. | 6 |
| 2024 | AgentIR: 1st Workshop on Agent-based Information RetrievalabstractInformation retrieval (IR) systems have become an essential component in modern society to help users find useful information, which consists of a series of processes including query expansion, item recall, item ranking and re-ranking, etc. Based on the ranked information list, users can provide their feedbacks. Such an interaction process between users and IR systems can be naturally formulated as a decision-making problem, which can be either one-step or sequential. In the last ten years, deep reinforcement learning (DRL) has become a promising direction for decision-making, since DRL utilizes the high model capacity of deep learning for complex decision-making tasks. On the one hand, there have been emerging research works focusing on leveraging DRL for IR tasks. However, the fundamental information theory under DRL settings, the challenge of RL methods for Industrial IR tasks, or the simulations of DRL-based IR systems, has not been deeply investigated. On the other hand, the emerging LLM provides new opportunities for optimizing and simulating IR systems. To this end, we propose the first Agent-based IR workshop at SIGIR 2024, as a continuation from one of the most successful IR workshops, DRL4IR. It provides a venue for both academia researchers and industry practitioners to present the recent advances of both DRL-based IR systems and LLM-based IR systems from the agent-based IR's perspective, to foster novel research, interesting findings, and new applications. Qingpeng Cai 0001, Xiangyu Zhao 0001, Ling Pan, Xin Xin 0003, Jin Huang 0010, Weinan Zhang 0001, Li Zhao 0007, Dawei Yin 0001, Grace Hui Yang |
SIGIR | 9 |
| 2024 | Towards Human-centered Proactive Conversational AgentsabstractRecent research on proactive conversational agents (PCAs) mainly focuses on improving the system's capabilities in anticipating and planning action sequences to accomplish tasks and achieve goals before users articulate their requests. This perspectives paper highlights the importance of moving towards building human-centered PCAs that emphasize human needs and expectations, and that considers ethical and social implications of these agents, rather than solely focusing on technological capabilities. The distinction between a proactive and a reactive system lies in the proactive system's initiative-taking nature. Without thoughtful design, proactive systems risk being perceived as intrusive by human users. We address the issue by establishing a new taxonomy concerning three key dimensions of human-centered PCAs, namely Intelligence, Adaptivity, and Civility. We discuss potential research opportunities and challenges based on this new taxonomy upon the five stages of PCA system construction. This perspectives paper lays a foundation for the emerging area of conversational information retrieval research and paves the way towards advancing human-centered proactive conversational systems. Yang Deng 0002, Lizi Liao, Zhonghua Zheng, Grace Hui Yang, Tat-Seng Chua |
SIGIR | 4 |
| 2024 | Improving searcher struggle detection via the reversal theoryabstractSearcher struggle is important feedback to Web search engines. Existing Web search struggle detection methods rely on effort-based features to identify the struggling moments. Their underlying assumption is that the more effort a user spends, the more struggling the user may be. However, studies have shown that this simple association might be incorrect. This paper proposes a new feature modulation method for struggle detection and refers to the Reversal Theory in psychology. Reversal Theory points out that instead of having a static personality trait, people constantly switch between opposite psychological states, complicating the relationship between the efforts they spend and the level of frustration they feel. Supported by the theory, our method modulates the effort-based features based on Reversal Theory’s bi-modal arousal model. After modification, the users’ effort level is better aligned with their struggling experience. Evaluations on Pinterest search logs confirm that the proposed method can statistically significantly improve searcher struggle detection methods. Jiyun Luo, Valerie Nayak, Grace Hui Yang |
Discov. Comput. | 4 |
| 2023 | DRL4IR: 4th Workshop on Deep Reinforcement Learning for Information Retrievalabstract\AcIR is one of the most important fields to help users find relevant information. The interaction between IR systems and users can be naturally formulated as a decision-making problem. In the last decade, deep reinforcement learning (DRL) has become a promising direction to utilize the high model capacity of deep learning to improve long-term gains. On the one hand, there have been emerging research works focusing on leveraging DRL for IR tasks while the fundamental information theory under DRL settings, the principle of RL methods for IR tasks, or the experimental evaluation protocols of DRL-based IR systems, has not been deeply investigated. On the other hand, the emerging ChatGPT also provides new insights and challenges for DRL-based IR. Xin Xin 0003, Xiangyu Zhao 0001, Jin Huang 0010, Weinan Zhang 0001, Li Zhao 0007, Dawei Yin 0001, Grace Hui Yang |
CIKM | 7 |
| 2023 | Proactive Conversational Agents in the Post-ChatGPT WorldabstractChatGPT and similar large language model (LLM) based conversational agents have brought shock waves to the research world. Although astonished by their human-like performance, we find they share a significant weakness with many other existing conversational agents in that they all take a passive approach in responding to user queries. This limits their capacity to understand the users and the task better and to offer recommendations based on a broader context than a given conversation. Proactiveness is still missing in these agents, including their ability to initiate a conversation, shift topics, or offer recommendations that take into account a more extensive context. To address this limitation, this tutorial reviews methods for equipping conversational agents with proactive interaction abilities. Lizi Liao, Grace Hui Yang, Chirag Shah 0001 |
SIGIR | 2 |
| 2023 | Proactive Conversational AgentsabstractConversational agents, or commonly known as dialogue systems, have gained escalating popularity in recent years. Their widespread applications support conversational interactions with users and accomplishing various tasks as personal assistants. However, one key weakness in existing conversational agents is that they only learn to passively answer user queries via training on pre-collected and manually-labeled data. Such passiveness makes the interaction modeling and system-building process relatively easier, but it largely hinders the possibility of being human-like hence lowering the user engagement level. In this tutorial, we introduce and discuss methods to equip conversational agents with the ability to interact with end users in a more proactive way. This three-hour tutorial is divided into three parts and includes two interactive exercises. It reviews and presents recent advancements on the topic, focusing on automatically expanding ontology space, actively driving conversation by asking questions or strategically shifting topics, and retrospectively conducting response quality control. Lizi Liao, Grace Hui Yang, Chirag Shah 0001 |
WSDM | 2 |
| 2022 | DRL4IR: 3rd Workshop on Deep Reinforcement Learning for Information RetrievalabstractInformation retrieval (IR) systems have become an essential component in modern society to help users find useful information, which consists of a series of processes including query expansion, item recall, item ranking and re-ranking, etc. Based on the ranked information list, users can provide their feedbacks. Such an interaction process between users and IR systems can be naturally formulated as a decision-making problem, which can be either one-step or sequential. In the last ten years, deep reinforcement learning (DRL) has become a promising direction for decision-making, since DRL utilizes the high model capacity of deep learning for complex decision-making tasks. Recently, there have been emerging research works focusing on leveraging DRL for IR tasks. However, the fundamental information theory under DRL settings, the principle of RL methods for IR tasks, or the experimental evaluation protocols of DRL-based IR systems, has not been deeply investigated. Xiangyu Zhao 0001, Xin Xin 0003, Weinan Zhang 0001, Li Zhao 0007, Dawei Yin 0001, Grace Hui Yang |
SIGIR | 6 |
| 2022 | A Re-classification of Information Seeking Tasks and Their Computational SolutionsabstractThis article presents a re-classification of information seeking (IS) tasks, concepts, and algorithms. The proposed taxonomy provides new dimensions to look into information seeking tasks and methods. The new dimensions include number of search iterations, search goal types, and procedures to reach these goals. Differences along these dimensions for the information seeking tasks call for suitable computational solutions. The article then reviews machine learning solutions that match each new category. The article ends with a review of evaluation campaigns for IS systems. Zhiwen Tang, Grace Hui Yang |
ACM Trans. Inf. Syst. | 2 |
| 2021 | DRL4IR: 2nd Workshop on Deep Reinforcement Learning for Information RetrievalabstractModern information retrieval (IR) consists of a series of processes, including query expansion, candidate item recall, item ranking, item re-ranking, etc. The final ranked item list will be exposed to the user, which will accordingly provide feedback through some expected actions such as browsing and click. Such a whole process can be formulated as a decision-making process where the agent is the IR system while the environment is the specific user. This decision-making process can be one-step or sequential, depending on the scenarios or the ways of problem formulation. Since 2013, Deep reinforcement learning (DRL) has been a fast-developing technique for decision-making tasks. The high capacity of deep learning models is incorporated in the reinforcement learning framework so that the agent may successfully handle complex decision-making. In recent years, there have been a bunch of publications attempting to leverage DRL techniques for different IR tasks such as ad hoc retrieval, learning to rank and interactive recommendation. Nonetheless, the fundamental theory, the principle of RL methods or the recognized experimental protocols of decision-making in IR, has not been well developed, making it challenging to evaluate the correctness of a proposed method or judge whether the reported experimental performance is valid. We propose the second DRL4IR workshop at SIGIR 2021, which provides a venue to gather the academia researchers and industry practitioners to present the recent progress of DRL techniques for IR. More importantly, people in this workshop are expected to discuss more about the fundamental principles of formulating a decision-making IR task, the underlying theory as well as the practical effectiveness of the experiment protocol design, which would foster further research on novel methodologies, innovative experimental findings and new applications of DRL for information retrieval. DRL4IR organized at SIGIR'20 was one of the most popular workshops and attracted over 200 conference attendees. In this year, we will pay more attention to fundamental research topics and recent applications, and expect about 300 participants. Weinan Zhang 0001, Xiangyu Zhao 0001, Li Zhao 0007, Dawei Yin 0001, Grace Hui Yang |
SIGIR | 5 |
| 2020 | Deep Reinforcement Learning for Information Retrieval: Fundamentals and AdvancesabstractInformation retrieval (IR) techniques, such as search, recommendation and online advertising, satisfying users' information needs by suggesting users personalized objects (information or services) at the appropriate time and place, play a crucial role in mitigating the information overload problem. Since the widely use of mobile applications, more and more information retrieval services have provided interactive functionality and products. Thus, learning from interaction becomes a crucial machine learning paradigm for interactive IR, which is based on reinforcement learning. With recent great advances in deep reinforcement learning (DRL), there have been increasing interests in developing DRL based information retrieval techniques, which could continuously update the information retrieval strategies according to users' real-time feedback, and optimize the expected cumulative long-term satisfaction from users. Our workshop aims to provide a venue, which can bring together academia researchers and industry practitioners (i) to discuss the principles, limitations and applications of DRL for information retrieval, and (ii) to foster research on innovative algorithms, novel techniques, and new applications of DRL to information retrieval. Weinan Zhang 0001, Xiangyu Zhao 0001, Li Zhao 0007, Dawei Yin 0001, Grace Hui Yang, Alex Beutel |
SIGIR | 5 |
| 2020 | Balancing Reinforcement Learning Training Experiences in Interactive Information RetrievalabstractInteractive Information Retrieval (IIR) and Reinforcement Learning (RL) share many commonalities, including an agent who learns while interacts, a long-term and complex goal, and an algorithm that explores and adapts. To successfully apply RL methods to IIR, one challenge is to obtain sufficient relevance labels to train the RL agents, which are infamously known as sample inefficient. However, in a text corpus annotated for a given query, it is not the relevant documents but the irrelevant documents that predominate. This would cause very unbalanced training experiences for the agent and prevent it from learning any policy that is effective. Our paper addresses this issue by using domain randomization to synthesize more relevant documents for the training. Our experimental results on the Text REtrieval Conference (TREC) Dynamic Domain (DD) 2017 Track show that the proposed method is able to boost an RL agent's learning effectiveness by 22% in dealing with unseen situations. Zhiwen Tang, Grace Hui Yang |
SIGIR | 3 |
| 2019 | Modeling Long-Range Context for Concurrent Dialogue Acts RecognitionabstractIn dialogues, an utterance is a chain of consecutive sentences produced by one speaker which ranges from a short sentence to a thousand-word post. When studying dialogues at the utterance level, it is not uncommon that an utterance would serve multiple functions. For instance, "Thank you. It works great." expresses both gratitude and positive feedback in the same utterance. Multiple dialogue acts (DA) for one utterance breeds complex dependencies across dialogue turns. Therefore, DA recognition challenges a model's predictive power over long utterances and complex DA context. We term this problem Concurrent Dialogue Acts (CDA) recognition. Previous work on DA recognition either assumes one DA per utterance or fails to realize the sequential nature of dialogues. In this paper, we present an adapted Convolutional Recurrent Neural Network (CRNN) which models the interactions between utterances of long-range context. Our model significantly outperforms existing work on CDA recognition on a tech forum dataset. Siyao Peng, Grace Hui Yang |
CIKM | 3 |
| 2018 | Minority Report by Lemur: Supporting Search Engine with Virtual RealityabstractIn this paper, we introduce a Virtual Reality (VR) search engine interface. Virtual reality has been explored in the game industry and in the multimedia community. When wearing a VR device, a realistic experience is simulated around the user. In the working environment, VR's potential is still understudied. As a first step to enable VR-supported working environment, we present a search engine with a virtual reality interface. In our system, users can read, search and interact with the search engine with novel experiences. They only need to use their hands to interact with digital content, just like what is shown in the "minority report" movie. Andrew Jie Zhou, Grace Hui Yang |
SIGIR | 2 |
| 2018 | Differential Privacy for Information RetrievalabstractThe concern for privacy is real for any research that uses user data. Information Retrieval (IR) is not an exception. Many IR algorithms and applications require the use of users' personal information, contextual information and other sensitive and private information. Grace Hui Yang, Sicong Zhang |
WSDM | 1 |
| 2018 | Session search modeling by partially observable Markov decision process
Grace Hui Yang, Xuchu Dong, Jiyun Luo, Sicong Zhang |
Inf. Retr. J. | 1 |
| 2016 | Generating risk reduction recommendations to decrease vulnerability of public online profilesabstractPreserving online privacy is becoming increasingly challenging due in large part to the continued growth of social media. Those who choose to share their information publicly may not realize what features of their profiles make their public data more identifiable and potentially vulnerable to cross-site record linkage. This paper proposes a risk reduction recommendation method that suggests removal or modification of a small number of attributes to make a profile less unique, thereby reducing the identifiability and vulnerability of the user. Empirical results on data collected from Google+, LinkedIn, and Foursquare show that users' vulnerability in terms of identifiability and data exposure level can be significantly reduced while public profile utility can be maintained using our proposed approach. Janet Zhu, Sicong Zhang, Lisa Singh, Grace Hui Yang, Micah Sherr |
ASONAM | 4 |
| 2016 | Privacy-Preserving IR 2016: Differential Privacy, Search, and Social MediaabstractDue to lack of mature techniques in privacy-preserving information retrieval (IR), concerns about information privacy and security have become serious obstacles that prevent valuable user data to be used in IR research such as studies on query logs, social media, and medical record retrieval. In SIGIR 2014 and SIGIR 2015, we have run the privacy-preserving IR workshops exploring and understanding the privacy and security risks in information retrieval. This year, we continue the efforts of connecting the two disciplines of IR and privacy/security by organizing this workshop. We target on three themes, differential privacy and IR dataset release, privacy in search and browsing, and privacy in social media. The workshop includes panels with researchers from both fields on these three themes, as well as invite industry speakers for real-world challenges. The goals of this workshop include (1) bringing together the two research fields, and (2) yielding fruitful collaborations. Grace Hui Yang, Ian Soboroff, Li Xiong 0001, Charles L. A. Clarke, Simson L. Garfinkel |
SIGIR | 1 |
| 2016 | Anonymizing Query Logs by Differential PrivacyabstractQuery logs are valuable resources for Information Retrieval (IR) research. However, because they are also rich in private and personal information, the huge concern of leaking user privacy prevents query logs from being shared from the search companies to the broad research community. Bothered by the lack of good research data for years, the authors of this paper are motivated to explore ways to generate anonymized query logs that can still be effectively used to support the search task. We introduce a framework to anonymize query logs by differential privacy, the latest development in privacy research. The framework is empirically evaluated against multiple search algorithms on their retrieval utility, measured in standard IR evaluation metrics, using the anonymized logs. The experiments show that our framework is able to achieve a good balance between retrieval utility and privacy. Sicong Zhang, Grace Hui Yang, Lisa Singh |
SIGIR | 2 |
| 2015 | Public Information Exposure Detection: Helping Users Understand Their Web FootprintsabstractTo help users better understand the potential risks associated with publishing data publicly, as well as the quantity and sensitivity of information that can be obtained by combining data from various online sources, we introduce a novel information exposure detection framework that generates and analyzes the web footprints users leave across the social web. Web footprints are the traces of one's online social activities represented by a set of attributes that are known or can be inferred with a high probability by an adversary who has basic information about a user from his/her public profiles. Our framework employs new probabilistic operators, novel pattern-based attribute extraction from text, and a population-based inference engine to generate web footprints. Using a web footprint, the framework then quantifies a user's level of information exposure relative to others with similar traits, as well as with regard to others in the population. Evaluation over public profiles from multiple sites (Google+, LinkeIn, FourSquare, and Twitter) shows that the proposed framework effectively detects and quantifies information exposure using a small amount of initial knowledge. Lisa Singh, Grace Hui Yang, Micah Sherr, Andrew Hian-Cheong, Kevin Tian, Janet Zhu, Sicong Zhang |
ASONAM | 2 |
| 2015 | Designing States, Actions, and Rewards for Using POMDP in Session Search
Jiyun Luo, Sicong Zhang, Xuchu Dong, Grace Hui Yang |
ECIR | 4 |
| 2015 | Detecting the Eureka Effect in Complex Search
Grace Hui Yang, Jiyun Luo, Christopher Wing |
ECIR | 1 |
| 2015 | Privacy-Preserving IR 2015: When Information Retrieval Meets Privacy and SecurityabstractInformation retrieval (IR) and information privacy/security are two fast-growing computer science disciplines. There are many synergies and connections between these two disciplines. However, there have been very limited efforts to connect the two important disciplines. On the other hand, due to lack of mature techniques in privacy-preserving IR, concerns about information privacy and security have become serious obstacles that prevent valuable user data to be used in IR research such as studies on query logs, social media, tweets, and medical record retrieval. We propose this privacy-preserving IR workshop to connect the two disciplines of information retrieval and information privacy and security. We look forward to spurring research that aims to bring together the research fields of IR and privacy/security. Last year, the first privacy-preserving IR workshop focused on mitigating privacy threats in information retrieval by novel algorithms and tools that enable web users to better understand associated privacy risks. Grace Hui Yang, Ian Soboroff |
SIGIR | 1 |
| 2015 | DUMPLING: A Novel Dynamic Search EngineabstractIn this demo paper, we introduce a new search engine that supports Information Retrieval (IR) in a dynamic setting. A dynamic search engine distinguishes itself by handling rich interactions and temporal dependency among the queries in a session or for a task. The proposed search engine is called Dumpling, named after the development team's favorite food. It implements state-of-the-art dynamic search algorithms and provides: (i) a dynamic search toolkit by integrating the Query Change Retrieval Model (QCM) and the Win-win search algorithm; (ii) a user-friendly interface supporting side-by-side comparison of search results given by a state-of-the-art static search algorithm and the proposed dynamic search algorithms; (iii) and APIs for developers to apply the dynamic search algorithms to index and search over custom datasets. Dumpling is developed under the umbrella of a bigger project in the DARPA Memex program to crawl and search the dark web to support law enforcement and national security. Andrew Jie Zhou, Jiyun Luo, Grace Hui Yang |
SIGIR | 3 |
| 2015 | Dynamic Information Retrieval ModelingabstractIn Dynamic Information Retrieval modeling we model dynamic systems which change or adapt over time or a sequence of events using a range of techniques from artificial intelligence and reinforcement learning. Many of the open problems in current IR research can be described as dynamic systems, for instance, session search or computational advertising. State of the art research provides solutions to these problems that are responsive to a changing environment, learn from past interactions and predict future utility. Advances in IR interface, personalization and ad display demand models that can react to users in real time and in an intelligent, contextual way. The objective of this half-day tutorial is to provide a comprehensive and up-to-date introduction to Dynamic Information Retrieval Modeling. We motivate a conceptual model linking static, interactive and dynamic retrieval and use this to define dynamics within the context of IR. We then cover a number of algorithms and techniques from the artificial intelligence (AI) and online learning literature such as Markov Decision Processes (MDP), their partially observable variation (POMDP) and multi-armed bandits. Grace Hui Yang, Marc Sloan, Jun Wang 0012 |
WSDM | 1 |
| 2015 | A term-based methodology for query reformulation understanding
Marc Sloan, Grace Hui Yang, Jun Wang 0012 |
Inf. Retr. J. | 2 |
| 2015 | Browsing Hierarchy Construction by Minimum EvolutionabstractHierarchies serve as browsing tools to access information in document collections. This article explores techniques to derive browsing hierarchies that can be used as an information map for task-based search. It proposes a novel minimum-evolution hierarchy construction framework that directly learns semantic distances from training data and from users to construct hierarchies. The aim is to produce globally optimized hierarchical structures by incorporating user-generated task specifications into the general learning framework. Both an automatic version of the framework and an interactive version are presented. A comparison with state-of-the-art systems and a user study jointly demonstrate that the proposed framework is highly effective. Grace Hui Yang |
ACM Trans. Inf. Syst. | 1 |
| 2015 | The Query Change Model: Modeling Session Search as a Markov Decision ProcessabstractModern information retrieval (IR) systems exhibit user dynamics through interactivity. These dynamic aspects of IR, including changes found in data, users, and systems, are increasingly being utilized in search engines. Session search is one such IR task—document retrieval within a session. During a session, a user constantly modifies queries to find documents that fulfill an information need. Existing IR techniques for assisting the user in this task are limited in their ability to optimize over changes, learn with a minimal computational footprint, and be responsive. This article proposes a novel query change retrieval model (QCM), which uses syntactic editing changes between consecutive queries, as well as the relationship between query changes and previously retrieved documents, to enhance session search. We propose modeling session search as a Markov decision process (MDP). We consider two agents in this MDP: the user agent and the search engine agent. The user agent’s actions are query changes that we observe, and the search engine agent’s actions are term weight adjustments as proposed in this work. We also investigate multiple query aggregation schemes and their effectiveness on session search. Experiments show that our approach is highly effective and outperforms top session search systems in TREC 2011 and TREC 2012. Grace Hui Yang, Dongyi Guan, Sicong Zhang |
ACM Trans. Inf. Syst. | 1 |
| 2014 | Win-win search: dual-agent stochastic game in session searchabstractSession search is a complex search task that involves multiple search iterations triggered by query reformulations. We observe a Markov chain in session search: user's judgment of retrieved documents in the previous search iteration affects user's actions in the next iteration. We thus propose to model session search as a dual-agent stochastic game: the user agent and the search engine agent work together to jointly maximize their long term rewards. The framework, which we term "win-win search", is based on Partially Observable Markov Decision Process. We mathematically model dynamics in session search, including decision states, query changes, clicks, and rewards, as a cooperative game between the user and the search engine. The experiments on TREC 2012 and 2013 Session datasets show a statistically significant improvement over the state-of-the-art interactive search and session search algorithms. Jiyun Luo, Sicong Zhang, Grace Hui Yang |
SIGIR | 3 |
| 2014 | Privacy-preserving IR: when information retrieval meets privacy and securityabstractInformation retrieval (IR) and information privacy/security are two fast-growing computer science disciplines. There are many synergies and connections between these two disciplines. However, there have been very limited efforts to connect the two important disciplines. On the other hand, due to lack of mature techniques in privacy-preserving IR, concerns about information privacy and security have become serious obstacles that prevent valuable user data to be used in IR research such as studies on query logs, social media, tweets, sessions, and medical record retrieval. This privacy-preserving IR workshop aims to spur research that brings together the research fields of IR and privacy/security, and research that mitigates privacy threats in information retrieval by constructing novel algorithms and tools that enable web users to better understand associated privacy risks. Luo Si, Grace Hui Yang |
SIGIR | 2 |
| 2014 | FitYou: integrating health profiles to real-time contextual suggestionabstractObesity and its associated health consequences such as high blood pressure and cardiac disease affect a significant proportion of the world's population. At the same time, the popularity of location-based services (LBS) and recommender systems is continually increasing with improvements in mobile technology. We observe that the health domain lacks a suggestion system that focuses on healthy lifestyle choices. We introduce the mobile application FitYou, which dynamically generates recommendations according to the user's current location and health condition as a real-time LBS. It utilizes preferences determined from user history and health information from a biometric profile. The system was developed upon a top performing contextual suggestion system in both TREC 2012 and 2013 Contextual Suggestion Tracks. Christopher Wing, Grace Hui Yang |
SIGIR | 2 |
| 2014 | Dynamic information retrieval modelingabstractDynamic aspects of Information Retrieval (IR), including changes found in data, users and systems, are increasingly being utilized in search engines and information filtering systems. Existing IR techniques are limited in their ability to optimize over changes, learn with minimal computational footprint and be responsive and adaptive. The objective of this tutorial is to provide a comprehensive and up-to-date introduction to Dynamic Information Retrieval Modeling, the statistical modeling of IR systems that can adapt to change. It will cover techniques ranging from classic relevance feedback to the latest applications of partially observable Markov decision processes (POMDPs) and a handful of useful algorithms and tools for solving IR problems incorporating dynamics. Grace Hui Yang, Marc Sloan, Jun Wang 0012 |
SIGIR | 1 |
| 2014 | A POMDP model for content-free document re-rankingabstractLog-based document re-ranking is a special form of session search. The task re-ranks documents from Search Engine Results Page (SERP) according to the search logs, in which both the search activities from other users and personalized query log for a user are available. The purpose of re-ranking is to provide the user with a new and better ordering of the initial retrieved documents. We test the system on the WSCD 2014 dataset, in which the actual content of the queries and documents are not available due to privacy concerns. The challenge is to perform effective re-ranking purely based on user behaviors, such as clicks and query reformulations rather than document content. In this paper, we propose to model log-based document re-ranking as a Partially Observable Markov Decision Process (POMDP). Experiments on the document re-ranking task show that our approach is effective and outperforms the baseline rankings provided by a commercial search engine. Sicong Zhang, Jiyun Luo, Grace Hui Yang |
SIGIR | 3 |
| 2013 | The water filling model and the cube test: multi-dimensional evaluation for professional searchabstractProfessional search activities such as patent and legal search are often time sensitive and consist of rich information needs with multiple aspects or subtopics. This paper proposes a 3D water filling model to describe this search process, and derives a new evaluation metric, the Cube Test, to encompass the complex nature of professional search. The new metric is compared against state-of-the-art patent search evaluation metrics as well as Web search evaluation metrics over two distinct patent datasets. The experimental results show that the Cube Test metric effectively captures the characteristics and requirements of professional search. Jiyun Luo, Christopher Wing, Grace Hui Yang, Marti A. Hearst |
CIKM | 3 |
| 2013 | Increasing Stability of Result Organization for Session Search
Dongyi Guan, Grace Hui Yang |
ECIR | 2 |
| 2013 | Utilizing query change for session searchabstractSession search is the Information Retrieval (IR) task that performs document retrieval for a search session. During a session, a user constantly modifies queries in order to find relevant documents that fulfill the information need. This paper proposes a novel query change retrieval model (QCM), which utilizes syntactic editing changes between adjacent queries as well as the relationship between query change and previously retrieved documents to enhance session search. We propose to model session search as a Markov Decision Process (MDP). We consider two agents in this MDP: the user agent and the search engine agent. The user agent's actions are query changes that we observe and the search agent's actions are proposed in this paper. Experiments show that our approach is highly effective and outperforms top session search systems in TREC 2011 and 2012. Dongyi Guan, Sicong Zhang, Grace Hui Yang |
SIGIR | 3 |
| 2013 | InfoLand: information lay-of-land for session searchabstractSearch result clustering (SRC) is a post-retrieval process that hierarchically organizes search results. The hierarchical structure offers overview for the search results and displays an "information lay-of-land" that intents to guide the users throughout a search session. However, SRC hierarchies are sensitive to query changes, which are common among queries in the same session. This instability may leave users seemly random overviews throughout the session. We present a new tool called InfoLand that integrates external knowledge from Wikipedia when building SRC hierarchies and increase their stability. Evaluation on TREC 2010-2011 Session Tracks shows that InfoLand produces more stable results organization than a commercial search engine. Jiyun Luo, Dongyi Guan, Grace Hui Yang |
SIGIR | 3 |
| 2013 | Query change as relevance feedback in session searchabstractSession search is the Information Retrieval (IR) task that performs document retrieval for an entire session. During a session, users often change queries to explore and investigate the information needs. In this paper, we propose to use query change as a new form of relevance feedback for better session search. Evaluation conducted over TREC 2012 Session Track shows that query change is a highly effective form of feedback as compared with existing relevance feedback methods. The proposed method outperforms the state-of-the-art relevance feedback methods for the TREC 2012 Session Track by a significant improvement of >25%. Sicong Zhang, Dongyi Guan, Grace Hui Yang |
SIGIR | 3 |
| 2010 | Collecting high quality overlapping labels at low costabstractThis paper studies quality of human labels used to train search engines' rankers. Our specific focus is performance improvements obtained by using overlapping relevance labels, which is by collecting multiple human judgments for each training sample. The paper explores whether, when, and for which samples one should obtain overlapping training labels, as well as how many labels per sample are needed. The proposed selective labeling scheme collects additional labels only for a subset of training samples, specifically for those that are labeled relevant by a judge. Our experiments show that this labeling scheme improves the NDCG of two Web search rankers on several real-world test sets, with a low labeling overhead of around 1.4 labels per sample. This labeling scheme also outperforms several methods of using overlapping labels, such as simple k-overlap, majority vote, the highest labels, etc. Finally, the paper presents a study of how many overlapping labels are needed to get the best improvement in retrieval accuracy. Grace Hui Yang, Anton Mityagin, Krysta M. Svore, Sergey Markov |
SIGIR | 1 |
| 2009 | Feature selection for automatic taxonomy inductionabstractMost existing automatic taxonomy induction systems exploit one or more features to induce a taxonomy; nevertheless there is no systematic study examining which are the best features for the task under various conditions. This paper studies the impact of using different features on taxonomy induction for different types of relations and for terms at different abstraction levels. The evaluation shows that different conditions need different technologies or different combination of the technologies. In particular, co-occurrence and lexico-syntactic patterns are good features for is-a, sibling and part-of relations; contextual, co-occurrence, patterns, and syntactic features work well for concrete terms; co-occurrence works well for abstract terms. Grace Hui Yang, Jamie Callan |
SIGIR | 1 |
| 2006 | Near-duplicate detection by instance-level constrained clusteringabstractFor the task of near-duplicated document detection, both traditional fingerprinting techniques used in database community and bag-of-word comparison approaches used in information retrieval community are not sufficiently accurate. This is due to the fact that the characteristics of near-duplicated documents are different from that of both “almost-identical ” documents in the data cleaning task and “relevant ” documents in the search task. This paper presents an instance-level constrained clustering approach for near-duplicate detection. The framework incorporates information such as document attributes and content structure into the clustering process to form near-duplicate clusters. Gathered from several collections of public comments sent to U.S. government agencies on proposed new regulations, the experimental results demonstrate that our approach outperforms other near-duplicate detection algorithms and as about as effective as human assessors. Grace Hui Yang, Jamie Callan |
SIGIR | 1 |
| 2004 | Effectiveness of web page classification on finding list answersabstractList question answering (QA) offers a unique challenge in effectively and efficiently locating a complete set of distinct answers from huge corpora or the Web. In TREC-12, the median average F1 performance of list QA systems was only 6.9%. This paper exploits the wealth of freely available text and link structures on the Web to seek complete answers to list questions. We employ natural language parsing, web page classification and clustering to find reliable list answers. We also study the effectiveness of web page classification on both the recall and uniqueness of answers for web-based list QA. Grace Hui Yang, Tat-Seng Chua |
SIGIR | 1 |
| 2003 | Structured use of external knowledge for event-based open domain question answeringabstractOne of the major problems in question answering (QA) is that the queries are either too brief or often do not contain most relevant terms in the target corpus. In order to overcome this problem, our earlier work integrates external knowledge extracted from the Web and WordNet to perform Event-based QA on the TREC-11 task. This paper extends our approach to perform event-based QA by uncovering the structure within the external knowledge. The knowledge structure loosely models different facets of QA events, and is used in conjunction with successive constraint relaxation algorithm to achieve effective QA. Our results obtained on TREC-11 QA corpus indicate that the new approach is more effective and able to attain a confidence-weighted score of above 80%. Grace Hui Yang, Tat-Seng Chua, Shuguang Wang, Chun-Keat Koh |
SIGIR | 1 |