Jiepu Jiang

dblp:36/3589 · DBLP profile ↗
← Back
26ranked-venue papers
16as first author
3since 2021 · last 2022
0000-0002-8207-7452ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 26 · 16 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 How Misinformation Density Affects Health Information Search
abstract
Search engine results can include misinformation that is inaccurate, misleading, or even harmful. But people may not recognize or realize false information results when searching online. We suspect that the percentage of misinformation search results (misinformation density) may influence people’s search activities, learning outcomes, and search experience. We conducted a zoom-mediated “lab” user study to examine this matter. The experiment used a between-subjects design. We asked 60 participants to finish two health information search tasks using search engines with High, Medium, or Low misinformation density levels. To create these experimental settings, we trained task-dependent text classifiers to manipulate the number of correct and misinformation results displayed on SERPs. We collected participants’ search activities, responses to pre-task and post-task surveys, and answers to task-related factual questions before and after searching.
Qiurong Song, Jiepu Jiang
WWW2
2021 Evaluating Human-AI Hybrid Conversational Systems with Chatbot Message Suggestions
abstract
AI chatbots can offer suggestions to help humans answer questions by reducing text entry effort and providing relevant knowledge for unfamiliar questions. We study whether chatbot suggestions can help people answer knowledge-demanding questions in a conversation and influence response quality and efficiency. We conducted a large-scale crowdsourcing user study and evaluated 20 hybrid system variants and a human-only baseline. The hybrid systems used four chatbots of varied response quality and differed in the number of suggestions and whether to preset the message box with top suggestions.
Jiepu Jiang
CIKM2
2021 Utility of Missing Concepts in Query-biased Summarization
abstract
Query-biased Summarization (QBS) aims to produce a query-dependent summary of a retrieved document to reduce the human effort for inspecting the full-text content. Typical summarization approaches extract document snippets that overlap with the query and show them to searchers. Such QBS methods show relevant information in a document but do not inform searchers what is missing. Our study focuses on reducing user effort in finding relevant documents by exposing the information in the query that is missing in the retrieved results. We use a classical approach, DSPApprox, to find terms or phrases relevant to a query. Then, we identify which terms or phrases are missing in a document, present them in a search interface, and ask crowd workers to judge document relevance based on snippets and missing information. Experimental results show both benefits and limitations of our method compared with traditional ones that only show relevant snippets.
Sheikh Muhammad Sarwar, Felipe Moraes, Jiepu Jiang, James Allan 0001
SIGIR3
2020 Response Quality in Human-Chatbot Collaborative Systems
abstract
We report the results of a crowdsourcing user study for evaluating the effectiveness of human-chatbot collaborative conversation systems, which aim to extend the ability of a human user to answer another person's requests in a conversation using a chatbot. We examine the quality of responses from two collaborative systems and compare them with human-only and chatbot-only settings. Our two systems both allow users to formulate responses based on a chatbot's top-ranked results as suggestions. But they encourage the synthesis of human and AI outputs to a different extent. Experimental results show that both systems significantly improved the informativeness of messages and reduced user effort compared with a human-only baseline while sacrificing the fluency and humanlikeness of the responses. Compared with a chatbot-only baseline, the collaborative systems provided comparably informative but more fluent and human-like messages.
Jiepu Jiang, Naman Ahuja
SIGIR1
2019 Understanding the Interpretability of Search Result Summaries
abstract
We examine the interpretability of search results in current web search engines through a lab user study. Particularly, we evaluate search result summary as an interpretable technique that informs users why the system retrieves a result and to which extent the result is useful. We collected judgments about 1,252 search results from 40 users in 160 sessions. Experimental results indicate that the interpretability of a search result summary is a salient factor influencing users' click decisions. Users are less likely to click on a result link if they do not understand why it was retrieved (low transparency) or cannot assess if the result would be useful based on the summary (low assessability). Our findings suggest it is crucial to improve the interpretability of search result summaries and develop better techniques to explain retrieved results to search engine users.
Siyu Mi, Jiepu Jiang
SIGIR2
2017 Understanding Ephemeral State of Relevance
abstract
Despite its dynamic nature, relevance is often measured in a context-independent manner in information retrieval practice. We look into this discrepancy. We propose a contextual relevance/usefulness measurement called ephemeral state of relevance (ESR), which is defined as the amount of useful information a user acquired from a clicked result as assessed just after examining the result during an interactive search session. We collect ESR and context-independent usefulness judgments through a laboratory user study and compare the two. We examine factors related to both judgments and examine their differences.
Jiepu Jiang, Daqing He, Diane Kelly 0001, James Allan 0001
CHIIR1
2017 Similarity-based Distant Supervision for Definition Retrieval
abstract
Recognizing definition sentences from free text corpora often requires hand-crafted patterns or explicitly labeled training instances. We present a distant supervision approach addressing this challenge without using explicitly labeled data. We use plausibly good but imperfect definition sentences from Wikipedia as references to annotate sentences in a target corpus based on text similarity measures such as ROUGE. Experimental results show our approach is highly effective, generating noisy but large, useful, and localized training instances. Definition sentence retrieval models trained using the synthesized training examples are more effective than those learned from manual judgments of a few thousand sentences. We also examine different text similarity measures for annotation, including both unsupervised and supervised ones. We show that our method can significantly benefit from supervised text similarity measures learned from either external training data (from the SemEval Semantic Text Similarity task) or local ones (a few hundred judged sentences on the target corpus). Our method offers a cheap, effective, and flexible solution to this task and can benefit a broad range of applications such as web search engines and QA systems.
Jiepu Jiang, James Allan 0001
CIKM1
2017 Adaptive Persistence for Search Effectiveness Measures
abstract
Many search effectiveness evaluation measures penalize the importance of results at lower ranks. This is usually explained as an attempt to model users' persistence when sequentially examining results---lower ranked results are less important because users are less likely persistent enough to read them. The persistence parameters are usually set to cope with the target cohort and tasks. But during a particular evaluation round, the same parameters are applied to evaluate different ranked lists. In contrast, we present work that adapts the persistence factor according to the ranking and relevance of the ranked lists being evaluated. This is to model that rational users change their browsing behavior according to the search result page, e.g., users avoid wasting time (a low persistence level) if the results look apparently off-topic. Experimental results show that this approach better fits observed user behavior and correlates with users' ratings on their search performance.
Jiepu Jiang, James Allan 0001
CIKM1
2017 Comparing In Situ and Multidimensional Relevance Judgments
abstract
To address concerns of TREC-style relevance judgments, we explore two improvements. The first one seeks to make relevance judgments contextual, collecting in situ feedback of users in an interactive search session and embracing usefulness as the primary judgment criterion. The second one collects multidimensional assessments to complement relevance or usefulness judgments, with four distinct alternative aspects examined in this paper - novelty, understandability, reliability, and effort.
Jiepu Jiang, Daqing He, James Allan 0001
SIGIR1
2016 Contextual Support for Collaborative Information Retrieval
abstract
Recent research shows that Collaborative Information Retrieval (CIR), in which two or more users collaborate on the same search task, has become increasingly popular. The presence of both search and collaboration behaviors makes CIR a complex search format, which further drives a critical need to understand CIR's search context. The contextual support for CIR should consider search contexts derived from both team members' search histories (including users' own search histories and partners' search histories) and their explicit collaboration (e.g., chatting). As it stands, existing studies on contextual search support only focus on Individual Information Retrieval (IIR) and only utilize individuals' own search histories. In this paper, we examine the unique search contexts (e.g., partners' search histories and team collaboration histories) in CIR. Based on a user study data collection with 54 participants, we find that compared to the use of individuals' own search histories, CIR contextual support is more effective when utilizing partners' search histories and teams' collaboration behaviors. More interestingly, though the explicit communication information (i.e., chat content) often involves massive noisy information, involving such noise does not affect the ranking of relevant documents since it also does not appear in relevant documents.
Shuguang Han, Daqing He, Zhen Yue, Jiepu Jiang
CHIIR4
2016 Correlation Between System and User Metrics in a Session
abstract
We investigate the correlations between system-oriented evaluation metrics and a few user experience metrics for a search session. The system-oriented metrics include session-based DCG (sDCG), normalized sDCG (nsDCG), estimated session nDCG (esNDCG), and a few variants of these metrics. We also look into statistics (e.g., the mean, maximum, and minimum values) of individual queries' nDCG scores, as well as the first and the last query's nDCG in a session. These system-oriented metrics are compared with users' self-rated search performance and task difficulty for a session. Experimental results show that nsDCG and esNDCG have reasonable but weak correlations with the user metrics, while the worst and the last query's nDCG in a session have comparably strong correlations. This suggests future work may better measure users' search experience in a session by modeling each query in the session differently.
Jiepu Jiang, James Allan 0001
CHIIR1
2016 What Affects Word Changes in Query Reformulation During a Task-based Search Session?
abstract
This paper performs an analysis on the influence of different factors on users' choices of specific word changes in query reformulation during a search session. We study three types of word changes: whether to remove or retain a word in the current query; whether or not to add a brand-new word to the query; whether or not to reuse a word (included in previous queries, but removed in the current query). Three types of factors are examined: session-level factors measuring task and user characteristics; query-level factors related to past user activities in a session; word-level factors for the characteristics of the examined word and its relation to the current query and search results. Statistical analysis suggests that: word-level factors strongly influence all three types of word changes; query-level factors only show a clear influence on retaining or removing a word; task-level factors exhibit limited direct influence on all three types of word changes. Analysis also disclose reasons for different word changes: users remove a word to stop exploring a subtask, or to correct bad performing queries; they look for related, unused words from recently viewed result summaries and add to queries; reusing a word usually indicates reverting from a subtask to the main task or another subtask.
Jiepu Jiang, Chaoqun Ni
CHIIR1
2016 Understanding User Satisfaction with Intelligent Assistants
abstract
Voice-controlled intelligent personal assistants, such as Cortana, Google Now, Siri and Alexa, are increasingly becoming a part of users' daily lives, especially on mobile devices. They introduce a significant change in information access, not only by introducing voice control and touch gestures but also by enabling dialogues where the context is preserved. This raises the need for evaluation of their effectiveness in assisting users with their tasks. However, in order to understand which type of user interactions reflect different degrees of user satisfaction we need explicit judgements. In this paper, we describe a user study that was designed to measure user satisfaction over a range of typical scenarios of use: controlling a device, web search, and structured search dialogue. Using this data, we study how user satisfaction varied with different usage scenarios and what signals can be used for modeling satisfaction in the different scenarios. We find that the notion of satisfaction varies across different scenarios, and show that, in some scenarios (e.g. making a phone call), task completion is very important while for others (e.g. planning a night out), the amount of effort spent is key. We also study how the nature and complexity of the task at hand affects user satisfaction, and find that preserving the conversation context is essential and that overall task-level satisfaction cannot be reduced to query-level satisfaction alone. Finally, we shed light on the relative effectiveness and usefulness of voice-controlled intelligent agents, explaining their increasing popularity and uptake relative to the traditional query-response interaction.
Julia Kiseleva, Kyle Williams 0001, Jiepu Jiang, Ahmed Awadallah 0001, Aidan C. Crook, Imed Zitouni, Tasos Anastasakos
CHIIR3
2016 Adaptive Effort for Search Evaluation Metrics
Jiepu Jiang, James Allan 0001
ECIR1
2016 Reducing Click and Skip Errors in Search Result Ranking
abstract
Search engines provide result summaries to help users quickly identify whether or not it is worthwhile to click on a result and read in detail. However, users may visit non-relevant results and/or skip relevant ones. These actions are usually harmful to the user experience, but few considered this problem in search result ranking. This paper optimizes relevance of results and user click and skip activities at the same time. Comparing two equally relevant results, our approach learns to rank the one that users are more likely to click on at a higher position. Similarly, it demotes non-relevant web pages with high click probabilities. Experimental results show this approach reduces about 10%-20% of the click and skip errors with a trade off of 2.1% decline in [email protected]
Jiepu Jiang, James Allan 0001
WSDM1
2015 Understanding and Predicting Graded Search Satisfaction
abstract
Understanding and estimating satisfaction with search engines is an important aspect of evaluating retrieval performance. Research to date has modeled and predicted search satisfaction on a binary scale, i.e., the searchers are either satisfied or dissatisfied with their search outcome. However, users' search experience is a complex construct and there are different degrees of satisfaction. As such, binary classification of satisfaction may be limiting. To the best of our knowledge, we are the first to study the problem of understanding and predicting graded (multi-level) search satisfaction. We ex-amine sessions mined from search engine logs, where searcher satisfaction was also assessed on multi-point scale by human annotators. Leveraging these search log data, we observe rich and non-monotonous changes in search behavior in sessions with different degrees of satisfaction. The findings suggest that we should predict finer-grained satisfaction levels. To address this issue, we model search satisfaction using features indicating search outcome, search effort, and changes in both outcome and effort during a session. We show that our approach can predict subtle changes in search satisfaction more accurately than state-of-the-art methods, affording greater insight into search satisfaction. The strong performance of our models has implications for search providers seeking to accu-rately measure satisfaction with their services.
Jiepu Jiang, Ahmed Awadallah 0001, Ryen W. White
WSDM1
2015 Automatic Online Evaluation of Intelligent Assistants
abstract
Voice-activated intelligent assistants, such as Siri, Google Now, and Cortana, are prevalent on mobile devices. However, it is challenging to evaluate them due to the varied and evolving number of tasks supported, e.g., voice command, web search, and chat. Since each task may have its own procedure and a unique form of correct answers, it is expensive to evaluate each task individually. This paper is the first attempt to solve this challenge. We develop consistent and automatic approaches that can evaluate different tasks in voice-activated intelligent assistants. We use implicit feedback from users to predict whether users are satisfied with the intelligent assistant as well as its components, i.e., speech recognition and intent classification. Using this approach, we can potentially evaluate and compare different tasks within and across intelligent assistants ac-cording to the predicted user satisfaction rates. Our approach is characterized by an automatic scheme of categorizing user-system interaction into task-independent dialog actions, e.g., the user is commanding, selecting, or confirming an action. We use the action sequence in a session to predict user satisfaction and the quality of speech recognition and intent classification. We also incorporate other features to further improve our approach, including features derived from previous work on web search satisfaction prediction, and those utilizing acoustic characteristics of voice requests. We evaluate our approach using data collected from a user study. Results show our approach can accurately identify satisfactory and unsatisfactory sessions.
Jiepu Jiang, Ahmed Awadallah 0001, Rosie Jones, Umut Ozertem, Imed Zitouni, Ranjitha Gurunath Kulkarni, Omar Zia Khan
WWW1
2015 User participation in an academic social networking service: A survey of open group users on Mendeley
abstract
Although there are a number of social networking services that specifically target scholars, little has been published about the actual practices and the usage of these so‐called academic social networking services (ASNSs). To fill this gap, we explore the populations of academics who engage in social activities using an ASNS; as an indicator of further engagement, we also determine their various motivations for joining a group in ASNSs. Using groups and their members in Mendeley as the platform for our case study, we obtained 146 participant responses from our online survey about users' common activities, usage habits, and motivations for joining groups. Our results show that (a) participants did not engage with social‐based features as frequently and actively as they engaged with research‐based features, and (b) users who joined more groups seemed to have a stronger motivation to increase their professional visibility and to contribute the research articles that they had read to the group reading list. Our results generate interesting insights into Mendeley's user populations, their activities, and their motivations relative to the social features of Mendeley. We also argue that further design of ASNSs is needed to take greater account of disciplinary differences in scholarly communication and to establish incentive mechanisms for encouraging user participation.
Wei Jeng, Daqing He, Jiepu Jiang
J. Assoc. Inf. Sci. Technol.3
2014 Necessary and frequent terms in queries
abstract
Vocabulary mismatch has long been recognized as one of the major issues affecting search effectiveness. Ineffective queries usually fail to incorporate important terms and/or incorrectly include inappropriate keywords. However, in this paper we show another cause of reduced search performance: sometimes users issue reasonable query terms, but systems cannot identify the correct properties of those terms and take advantages of the properties. Specifically, we study two distinct types of terms that exist in all search queries: (1) necessary terms, for which term occurrence alone is indicative of document relevance; and (2) frequent terms, for which the relative term frequency is indicative of document relevance within the set of documents where the term appears. We evaluate these two properties of query terms in a dataset. Results show that only 1/3 of the terms are both necessary and frequent, while another 1/3 only hold one of the properties and the final third do not hold any of the properties. However, existing retrieval models do not clearly distinguish terms with the two properties and consider them differently. We further show the great potential of improving retrieval models by treating terms with distinct properties differently.
Jiepu Jiang, James Allan 0001
SIGIR1
2014 Searching, browsing, and clicking in a search session: changes in user behavior by task and over time
abstract
There are many existing studies of user behavior in simple tasks (e.g., navigational and informational search) within a short duration of 1--2 queries. However, we know relatively little about user behavior, especially browsing and clicking behavior, for longer search session solving complex search tasks. In this paper, we characterize and compare user behavior in relatively long search sessions (10 minutes; about 5 queries) for search tasks of four different types. The tasks differ in two dimensions: (1) the user is locating facts or is pursuing intellectual understanding of a topic; (2) the user has a specific task goal or has an ill-defined and undeveloped goal. We analyze how search behavior as well as browsing and clicking patterns change during a search session in these different tasks. Our results indicate that user behavior in the four types of tasks differ in various aspects, including search activeness, browsing style, clicking strategy, and query reformulation. As a search session progresses, we note that users shift their interests to focus less on the top results but more on results ranked at lower positions in browsing. We also found that results eventually become less and less attractive for the users. The reasons vary and include downgraded search performance of query, decreased novelty of search results, and decaying persistence of users in browsing. Our study highlights the lack of long session support in existing search engines and suggests different strategies of supporting longer sessions according to different task types.
Jiepu Jiang, Daqing He, James Allan 0001
SIGIR1
2013 Supporting exploratory people search: a study of factor transparency and user control
abstract
People search is an active research topic in recent years. Related works includes expert finding, collaborator recommendation, link prediction and social matching. However, the diverse objectives and exploratory nature of those tasks make it difficult to develop a flexible method for people search that works for every task. In this project, we developed PeopleExplorer, an interactive people search system to support exploratory search tasks when looking for people. In the system, users could specify their task objectives by selecting and adjusting key criteria. Three criteria were considered: the content relevance, the candidate authoritativeness and the social similarity between the user and the candidates. This project represents a first attempt to add transparency to exploratory people search, and to give users full control over the search process. The system was evaluated through an experiment with 24 participants undertaking four different tasks. The results show that with comparable time and effort, users of our system performed significantly better in their people search tasks than those using the baseline system. Users of our system also exhibited many unique behaviors in query reformulation and candidate selection. We found that users' general perceptions about three criteria varied during different tasks, which confirms our assumptions regarding modeling task difference and user variance in people search systems.
Shuguang Han, Daqing He, Jiepu Jiang, Zhen Yue
CIKM3
2013 How do users respond to voice input errors?: lexical and phonetic query reformulation in voice search
abstract
Voice search offers users with a new search experience: instead of typing, users can vocalize their search queries. However, due to voice input errors (such as speech recognition errors and improper system interruptions), users need to frequently reformulate queries to handle the incorrectly recognized queries. We conducted user experiments with native English speakers on their query reformulation behaviors in voice search and found that users often reformulate queries with both lexical and phonetic changes to previous queries. In this paper, we first characterize and analyze typical voice input errors in voice search and users' corresponding reformulation strategies. Then, we evaluate the impacts of typical voice input errors on users' search progress and the effectiveness of different reformulation strategies on handling these errors. This study provides a clearer picture on how to further improve current voice search systems.
Jiepu Jiang, Wei Jeng, Daqing He
SIGIR1
2013 Venue-author-coupling: A measure for identifying disciplines through author communities
abstract
Conceptualizations of disciplinarity often focus on the social aspects of disciplines; that is, disciplines are defined by the set of individuals who participate in their activities and communications. However, operationalizations of disciplinarity often demarcate the boundaries of disciplines by standard classification schemes, which may be inflexible to changes in the participation profile of that discipline. To address this limitation, a metric called venue‐author‐coupling ( VAC ) is proposed and illustrated using journals from the Journal Citation Report's ( JCR ) library science and information science category. As JCRs are some of the most frequently used categories in bibliometric analyses, this allows for an examination of the extent to which the journals in JCR categories can be considered as proxies for disciplines. By extending the idea of bibliographic coupling, VAC identifies similarities among journals based on the similarities of their author profiles. The employment of this method using information science and library science journals provides evidence of four distinct subfields, that is, management information systems, specialized information and library science, library science‐focused, and information science‐focused research. The proposed VAC method provides a novel way to examine disciplinarity from the perspective of author communities.
Chaoqun Ni, Cassidy R. Sugimoto, Jiepu Jiang
J. Assoc. Inf. Sci. Technol.3
2012 Contextual evaluation of query reformulations in a search session by user simulation
abstract
We propose a method to dynamically estimate the utility of documents in a search session by modeling the users' browsing behaviors and novelty. The method can be applied to evaluate query reformulations in a search session.
Jiepu Jiang, Daqing He, Shuguang Han, Zhen Yue, Chaoqun Ni
CIKM1
2012 Where do the query terms come from?: an analysis of query reformulation in collaborative web search
abstract
This paper presents a user study aiming to investigate the query reformulation in collaborative Web search. 7 pairs of participants were recruited and each pair worked as a team on two collaborative exploratory Web search tasks. Through the log analysis, we compared possible sources for participants to draw query terms from. The results show that both search and collaborative actions are possible resources for new query terms. Traditional resources for query expansion such as previous search histories and relevant documents are still important resources for new query terms. The content in chat and workspace generated by participants themselves seems more likely to be the resource for new query terms than that of their partners. Task types also affect the influences on query reformulations. For the academic task, previously saved relevance documents are the most important resources for new query terms while chat histories are the most important resources for the leisure task.
Zhen Yue, Jiepu Jiang, Shuguang Han, Daqing He
CIKM2
2012 Finding readings for scientists from social websites
abstract
Current search systems are designed to find relevant articles, especially topically relevant ones, but the notion of relevance largely depends on search tasks. We study the specific task that scientists are searching for worth-reading articles beneficial for their research. Our study finds: users' perception of relevance and preference of reading are only moderately correlated; current systems can effectively find readings that are highly relevant to the topic, but 36% of the worth-reading articles are only marginally relevant or even non-relevant. Our system can effectively find those worth-reading but marginally relevant or non-relevant articles by taking advantages of scientists' recommendations in social websites.
Jiepu Jiang, Zhen Yue, Shuguang Han, Daqing He
SIGIR1