VLDB 2026 Research / reviewers in the wild / expert
Jaime Arguello
dblp:90/4580
· DBLP profile ↗
61ranked-venue papers in the field
22as first author
22since 2021 · last 2026
0000-0002-7645-0556ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 61 (22 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Effects of Thinking Aloud on Participants during Search and SensemakingabstractDuring a think-aloud study, participants verbalize their thoughts as they complete a specified task. Think-aloud comments provide insights into what a participant is doing and experiencing in the moment. Interactive information retrieval studies have used think-aloud protocols to explore different research questions. However, think-aloud protocols can also influence participants. We report on a qualitative analysis of data collected during a think-aloud study. Participants were asked to learn about a complex topic by searching online and taking notes. After the task, participants were asked whether and how thinking aloud influenced their approach to the task. Open-ended responses revealed 21 (positive and negative) ways in which thinking aloud influenced participants. Additionally, during the study, we measured participants’ working memory (WM) capacity. We did not find that thinking aloud had different influences for low- vs. high-WM participants. Bogeum Choi, Jaime Arguello |
CHIIR | 2 |
| 2026 | Workshop on Generative AI and Academic Search (GAI&AS)abstractThe rapid development of artificial intelligence (AI) is reshaping how people seek, access, and use information, with significant implications for researchers, educators, and students. Increasingly, academic search engines, bibliographic databases, and digital libraries are integrating AI features, including generated and synthesized content, conversational interfaces, and intelligent recommendation. These tools promise to support discovery, synthesis, and learning, yet they also raise critical questions about search integrity, fairness, accountability, transparency, and ethics (FATE). In academic contexts, where reliability and credibility are paramount, the design and use of AI-mediated search systems require novel ideas and approaches. Building on previous work in interactive information retrieval (IIR), search as learning, and search user interface design, this workshop invites the CHIIR community to examine opportunities and challenges in developing and using AI-powered academic search systems for research and higher education. Jaime Arguello, Orland Hoeber, Chang Liu 0007, Soo Young Rieh, Luanne Sinnamon |
CHIIR | 2 |
| 2026 | The effects of goal-setting on learning during information seeking with generative AIabstractOur research in this paper lies at the intersection of Generative AI (GenAI) and search-as-learning (SAL). GenAI technologies (e.g., ChatGPT) have revolutionized how people search for and interact with information. However, we do not yet fully understand how people use GenAI systems to learn about complex topics. SAL research has studied how different tools can support learning with traditional document retrieval systems. Our research closely relates to SAL work that has investigated the effects of goal-setting on learning during search. We explore the influence of goal-setting on learning during information-seeking sessions with a GenAI system. We report on a between-subjects crowdsourced study (N = 120) in which participants were asked to learn about a complex topic using a GenAI system. The study had four conditions that varied along two factors (a 2 × 2 design). The first factor involved displaying related web results in addition to the GenAI output. The second factor involved giving participants access to the Subgoal Manager (SM), a tool designed to help people develop subgoals and take notes. We investigated the effects of both factors on: (RQ1) perceptions; (RQ2) behaviors; (RQ3) learning and retention; (RQ4) the types of requests issued to the system; and (RQ5) participants’ motivations for engaging (or not engaging) with the related web results. Results found that participants with access to the SM had higher post-task learning outcomes, did less copy/pasting into their notes, perceived the task as more difficult, and requested more examples and support for differentiating concepts from the GenAI system. Kelsey Urgo, Yuan Li 0035, Jaime Arguello, Robert G. Capra |
CHIIR | 3 |
| 2026 | A Questionnaire for Capturing Perceptions of Self-Regulated Learning Processes during Information SeekingabstractWe present the SRL Perceptions Questionnaire (SPQ), developed to measure perceptions of self-regulated learning (SRL) after information seeking and learning sessions. In a crowd-sourced study (N = 127), participants completed the SPQ after searching to learn about a complex topic. The SPQ asked participants to report their perceptions of particular SRL constructs (e.g., planning, monitoring, strategy use, adapting). A principal component analysis supported a five-factor structure with high reliability (α > =.87). Perceived SRL did not correlate with normalized learning gains, yet pre-task and post-task perceptions showed correlations with several SPQ dimensions. We offer both the SPQ as an instrument for measuring SRL (processes critical to supporting human learning) after information seeking and insights into how perceptions of SRL constructs align with objective learning outcomes, pre-task perceptions, and post-task perceptions while learning during search. Kelsey Urgo, Jaime Arguello, Robert G. Capra |
CHIIR | 2 |
| 2026 | Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated EvaluationabstractTip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingual ToT test collections for Chinese, Japanese, Korean, and English, using an LLM-based query simulation framework. We systematically study how prompt language and source document language affect the fidelity of simulated ToT queries, validating synthetic queries through system rank correlation against real user queries. Our results show that effective ToT simulation requires language-aware design choices: non-English language sources are generally important, while English Wikipedia can be beneficial when non-English sources provide insufficient information for query generation. Based on these findings, we release four ToT test collections with 5,000 queries per language across multiple domains. This work provides the first large-scale multilingual ToT benchmark and offers practical guidance for constructing realistic ToT datasets beyond English. Xuhong He, To Eun Kim, Maik Fröbe, Jaime Arguello, Bhaskar Mitra 0001, Fernando Diaz 0001 |
SIGIR | 4 |
| 2025 | The Effects of Working Memory during a Search and Sensemaking TaskabstractWorking memory (WM) is involved in high-level cognitive tasks such as comprehension, reasoning, and learning.Search and sensemaking (SSM) is no exception-a wide range of (meta)cognitive activities are involved in the process of making sense of a complex topic by gathering information.Prior studies have found that WM can influence search behaviors, perceptions, and outcomes.However, little work has been done to gain insights into how WM might affect the SSM process.We report on a lab study (𝑁 = 44) in which participants were binned into low-and high-WM groups.During the study, participants were asked to learn about a complex and multifaceted topic by gathering information using a web search engine and taking notes.After the search session, participants were asked to produce a summary of everything they learned.The study investigated four research questions.RQ1 and RQ2 investigate the effects of WM on post-task perceptions and search behaviors.RQ3 investigates the effects of WM on the extent to which participants engaged in specific search, sensemaking, and cognitive activities.To address RQ3, the study used a think-aloud protocol.Search sessions (i.e., recorded actions and think-aloud comments) were then analyzed using qualitative techniques.Finally, RQ4 investigates the effects of WM on learning outcomes.Our RQ4 results found that high-WM participants had better learning outcomes.Our RQ2 and RQ3 results point to possible reasons why.Despite differences for RQ2-RQ4, there were no differences in post-task perceptions (RQ1). Bogeum Choi, Jaime Arguello |
CHIIR | 2 |
| 2025 | Search+Chat: Integrating Search and GenAI to Support Users with Learning-oriented Search TasksabstractGenerative AI (GenAI) technologies such as ChatGPT are changing the ways people interact with information.To illustrate, popular search engines (e.g., Google) have started integrating responses from GenAI tools with the traditional search results.In this paper, we explore the integration of GenAI technology with traditional search in the context of a learning-oriented task.We report on a between-subjects study (𝑁 = 40) in which participants completed a complex, learning-oriented search task.Participants were assigned to one of two conditions.In the SearchOnly condition, participants used a traditional web search system to gather information.In the Search+Chat condition, participants used an experimental system that combined a traditional web search component and an interactive GenAI-based chat component (Chat AI).The study investigated seven research questions.RQ1-RQ3 focused on differences between groups: (RQ1) post-task perceptions, (RQ2) search behaviors, and (RQ3) learning outcomes.To measure learning, participants completed a multiple-choice test before the search task, immediately after, and one week later (to measure retention).RQ4-RQ7 delved deeper into participants' behaviors and experiences in the Search+Chat condition: (RQ4) motivations for (and gains from) engaging with the Chat AI; (RQ5) the phases during which participants engaged with the Chat AI; (RQ6) the types of queries issued to each component; and (RQ7) perceptions about the information returned by each component. Yuyu Yang, Kelsey Urgo, Jaime Arguello, Robert G. Capra |
CHIIR | 3 |
| 2025 | Tip of the Tongue Query Elicitation for Simulated EvaluationabstractTip-of-the-tongue (TOT) search occurs when a user struggles to recall a specific identifier, such as a document title. While common, existing search systems often fail to effectively support TOT scenarios. Research on TOT retrieval is further constrained by the challenge of collecting queries, as current approaches rely heavily on community question-answering (CQA) websites, leading to labor-intensive evaluation and domain bias. To overcome these limitations, we introduce two methods for eliciting TOT queries-leveraging large language models (LLMs) and human participants-to facilitate simulated evaluations of TOT retrieval systems. Our LLM-based TOT user simulator generates synthetic TOT queries at scale, achieving high correlations with how CQA-based TOT queries rank TOT retrieval systems when tested in the Movie domain. Additionally, these synthetic queries exhibit high linguistic similarity to CQA-derived queries. For human-elicited queries, we developed an interface that uses visual stimuli to place participants in a TOT state, enabling the collection of natural queries. In the Movie domain, system rank correlation and linguistic similarity analyses confirm that human-elicited queries are both effective and closely resemble CQA-based queries. These approaches reduce reliance on CQA-based data collection while expanding coverage to underrepresented domains, such as Landmark and Person. LLM-elicited queries for the Movie, Landmark, and Person domains have been released as test queries in the TREC 2024 TOT track, with human-elicited queries scheduled for inclusion in the TREC 2025 TOT track. Additionally, we provide source code for synthetic query generation and the human query collection interface, along with curated visual stimuli used for eliciting TOT queries. To Eun Kim, Fernando Diaz 0001, Jaime Arguello, Bhaskar Mitra 0001 |
SIGIR | 4 |
| 2024 | The Effects of Goal-setting on Learning Outcomes and Self-Regulated Learning ProcessesabstractWe present a user study (N = 40) that investigated the role of goal-setting on learning during search. To this end, we developed a tool called the Subgoal Manager (SM). The SM was designed to help searchers break apart a learning-oriented search task into smaller subgoals. The tool enabled participants to add, delete, and modify subgoals; take notes with respect to subgoals; and mark subgoals as completed. During the study, participants completed a single learning-oriented search task and were assigned to one of two subgoal conditions. In the Subgoals condition, participants had access to the SM; were instructed to develop at least three subgoals before the search session; and could add, delete, and modify subgoals during the search session. In the NoSubgoals condition, participants were not instructed to set subgoals and were simply provided with a text editor to take notes. We investigate the effects of the subgoal condition on: (RQ1) learning and retention and (RQ2) the extent to which participants engaged in specific self-regulated learning (SRL) processes during the search session. Our results found two important trends. First, participants in the Subgoals condition had better learning outcomes, especially with respect to retention. Second, based on a qualitative analysis of participants’ search sessions, participants in the Subgoals condition engaged in more self-regulated learning (SRL) processes. Combined, our results suggest that goal-setting improves learning during search by encouraging and supporting greater engagement with SRL processes. Kelsey Urgo, Jaime Arguello |
CHIIR | 2 |
| 2024 | Better Understanding Procedural Search Tasks: Perceptions, Behaviors, and ChallengesabstractPeople often search for information to acquire procedural knowledge–“how to” knowledge about step-by-step procedures, methods, algorithms, techniques, heuristics, and skills. A procedural search task might involve implementing a solution to a problem, evaluating different approaches to a problem, and brainstorming on the types of problems that can be solved with a specific resource. We report on a study ( N =36) that aimed to better understand how people search for procedural knowledge. Much research has investigated how search task characteristics impact people’s perceptions and behaviors. Along these lines, we manipulated procedural search tasks along two orthogonal dimensions: product and goal. The product dimension relates to the main outcome of the task and the goal dimension relates to task’s success criteria. We manipulated tasks across three product categories and two goal categories. The study investigated four research questions. First, we examined the effects of the product and goal on participants’ (RQ1) pre-task perceptions, (RQ2) post-task perceptions, and (RQ3) search behaviors. Second, regardless of the task product and goal, by analyzing participants’ think-aloud comments and screen activities we closely examined how people search for procedural knowledge. Specifically, we report on (RQ4) important relevance criteria, types of information sought, and challenges. Bogeum Choi, Sarah Casteel, Jaime Arguello, Robert G. Capra |
ACM Trans. Inf. Syst. | 3 |
| 2023 | Understanding Procedural Search Tasks "in the Wild"abstractPeople often search online for procedural (i.e., “how-to”) knowledge. A procedural search task might involve a do-it-yourself project, cooking a dish, fixing a problem, or learning a new skill. Prior research has studied procedural search tasks from different perspectives: estimating the frequency of procedural searches online, understanding how people acquire procedural knowledge in specific contexts, and developing tools to support procedural search. Less research has aimed at deeply understanding procedural search tasks “in the wild”. To bridge this gap, we conducted a survey (N = 128) on Amazon Mechanical Turk. Participants were asked to recall a recent procedural task for which they searched online. Participants were asked open-ended questions about the task itself and their unique situation (e.g., constraints and needs). Additionally, participants provided webpages they found useful in their searches and described the characteristics of the page that made it useful. Finally, they provided useful pieces of information from each selected page and explained what they gained from the information. Using an inductive coding approach, we analyzed participants’ responses to gain insights about: (1) procedural task characteristics, (2) goals, (3) constraints, (4) contextual factors, (5) relevance criteria, and (6) gains obtained from useful information. Based on our results, we discuss important implications for future research and system design. Bogeum Choi, Jaime Arguello, Robert G. Capra |
CHIIR | 2 |
| 2023 | Understanding the Cognitive Influences of Interpretability Features on How Users Scrutinize Machine-Predicted CategoriesabstractThe goal of interpretable machine learning (ML) is to design tools and visualizations to help users scrutinize a system’s predictions. Prior studies have mostly employed quantitative methods to investigate the effects of specific tools/visualizations on outcomes related to objective performance—a human’s ability to correctly agree or disagree with the system—and subjective perceptions of the system. Few studies have employed qualitative methods to investigate how and why specific tools/visualizations influence performance, perceptions, and behaviors. We report on a lab study (N = 30) that investigated the influences of two interpretability features: confidence values and sentence highlighting. Participants judged whether medical articles belong to a predicted medical topic and were exposed to two interface conditions—one with and one without interpretability features. We investigate the effects of our interpretability features on participants’ performance and perceptions. Additionally, we report on a qualitative analysis of participants’ responses during an exit interview. Specifically, we report on how our interpretability features impacted different cognitive activities that participants engaged with during the task—reading, learning, and decision making. We also describe ways in which the interpretability features introduced challenges and sometimes led participants to make mistakes. Insights gained from our results point to future directions for interpretable ML research. Jiaming Qu, Jaime Arguello, Yue Wang 0035 |
CHIIR | 2 |
| 2023 | Goal-setting in support of learning during search: An exploration of learning outcomes and searcher perceptions
Kelsey Urgo, Jaime Arguello |
Inf. Process. Manag. | 2 |
| 2023 | The Influences of a Knowledge Representation Tool on Searchers with Varying Cognitive AbilitiesabstractWhile current systems are effective in helping searchers resolve simple information needs (e.g., fact-finding), they provide less support for searchers working on complex information-seeking tasks. Complex search tasks involve a wide range of (meta)cognitive activities, including goal-setting, organizing information, drawing inferences, monitoring progress, and revising mental models and search strategies. We report on a lab study ( N = 32) that investigated the influences of a knowledge representation tool called the OrgBox, developed to support searchers with complex tasks. The OrgBox tool was integrated into a custom-built search system and allowed study participants to drag-and-drop textual passages into the tool, organize passages into logical groupings called “boxes”, and make notes on passages and boxes. The OrgBox was compared to a baseline tool (called the Bookmark) that allowed participants to save textual passages, but not organize them nor make notes. Knowledge representation tools such as the OrgBox may provide special benefits for users with different cognitive profiles. In this article, we explore two cognitive abilities: (1) working memory (WM) capacity and (2) switching (SW) ability. Participants in the study were asked to gather information on a complex subject and produce an outline for a hypothetical research article. We investigate the influences of the tool (OrgBox vs. Bookmark) and the participant’s working memory capacity and switching ability on three types of outcomes: (RQ1) search behaviors, (RQ2) post-task perceptions, and (RQ3) the quality of outlines produces by participants. Bogeum Choi, Jaime Arguello, Robert G. Capra, Austin R. Ward |
ACM Trans. Inf. Syst. | 2 |
| 2022 | Procedural Knowledge Search by Intelligence AnalystsabstractPrior studies have explored the information-seeking practices of specific professional communities, including lawyers, physicians, engineers, recruiters, and government workers. In this research, we investigate the information-seeking practices of intelligence analysts (IAs) employed by a U.S. government agency. Specifically, we focus on the needs, practices, and challenges related to IAs searching for procedural knowledge using an internal system called the Tradecraft Hub (TC Hub). The TC Hub is a searchable repository of procedural knowledge documents written by agency employees. Procedural knowledge (as opposed to factual and conceptual knowledge) includes knowledge about step-by-step procedures, techniques, methods, tools, technologies, and skills, and is inherently task-oriented. We report on a survey study involving 22 IAs who routinely use the TC Hub. Our survey was designed to address four research questions. In RQ1, we investigate the types of work-related objectives that motivate IAs to search the TC Hub. In RQ2, we investigate the types of information IAs seek when they search the TC Hub. In RQ3, we investigate important relevance criteria used by IAs when judging the usefulness of information. Finally, in RQ4, we investigate the challenges faced by IAs when searching the TC Hub. Based on our findings, we discuss implications for improving and extending searchable knowledge base systems such as the TC Hub that exist in many organizations. Bogeum Choi, Sarah Casteel, Robert G. Capra, Jaime Arguello |
CHIIR | 4 |
| 2022 | Learner, Assignment, and Domain: Contextualizing Search for ComprehensionabstractModern search systems are largely designed and optimized for simple navigational or fact-finding tasks, with little support for complex tasks involving comprehension and learning. In response, the search-as-learning research community has undertaken a wide range of research questions focused on understanding how various types of learning outcomes are affected by searcher characteristics, the search task, and the search system. Typically, these views embed learning within a search system. In this paper we take a different view, embedding search within a framework for an end-to-end learning system designed to support learning in a formal educational context. Our central goal is to motivate research questions aligned to advance progress on techniques for active support of comprehension and formal learning. Thus we intentionally set aside goals for informal and surface learning. We argue that to be effective, such a search-centric learning system must model four key components: individual students (searcher factors), the educational domain (topic factors), academic assignments (task factors), and progress toward learning goals (the objective function of the end-to-end system). In modeling these components, our hypothetical system makes inferences about students’ learning histories, knowledge states, comprehension, and the utilities of different types of information resources. We present examples of possible techniques and data sources for each model. We also introduce the novel concept of leveraging school assignments as rich task context. Our intention is not to propose a functional system, but to frame search-as-learning in the context of comprehension and to inspire research questions arising from an end-to-end view of this important research domain. Catherine L. Smith, Kelsey Urgo, Jaime Arguello, Robert G. Capra |
CHIIR | 3 |
| 2022 | Learning assessments in search-as-learning: A survey of prior work and opportunities for future research
Kelsey Urgo, Jaime Arguello |
Inf. Process. Manag. | 2 |
| 2022 | Understanding the "Pathway" Towards a Searcher's Learning ObjectiveabstractSearch systems are often used to support learning-oriented goals. This trend has given rise to the “search-as-learning” movement, which proposes that search systems should be designed to support learning. To this end, an important research question is: How does a searcher’s type of learning objective (LO) influence their trajectory (or pathway ) toward that objective? We report on a lab study (N = 36) in which participants gathered information to meet a specific type of LO. To characterize LOs and pathways , we leveraged Anderson and Krathwohl’s (A&K’s) taxonomy [ 3 ]. A&K’s taxonomy situates LOs at the intersection of two orthogonal dimensions: (1) cognitive process (CP) (remember, understand, apply, analyze, evaluate, and create) and (2) knowledge type (factual, conceptual, procedural, and metacognitive knowledge). Participants completed learning-oriented search tasks that varied along three CPs (apply, evaluate, and create) and three knowledge types (factual, conceptual, and procedural knowledge). A pathway is defined as a sequence of learning instances (e.g., subgoals) that were also each classified into cells from A&K’s taxonomy. Our study used a think-aloud protocol, and pathways were generated through a qualitative analysis of participants’ think-aloud comments and recorded screen activities. We investigate three research questions. First, in RQ1, we study the impact of the LO on pathway characteristics (e.g., pathway length). Second, in RQ2, we study the impact of the LO on the types of A&K cells traversed along the pathway. Third, in RQ3, we study common and uncommon transitions between A&K cells along pathways conditioned on the knowledge type of the objective. We discuss implications of our results for designing search systems to support learning. Kelsey Urgo, Jaime Arguello |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Tip of the Tongue Known-Item Retrieval: A Case Study in Movie IdentificationabstractWhile current information retrieval systems are effective for known-item retrieval where the searcher provides a precise name or identifier for the item being sought, systems tend to be much less effective for cases where the searcher is unable to express a precise name or identifier. We refer to this as tip of the tongue (TOT) known-item retrieval, named after the cognitive state of not being able to retrieve an item from memory. Using movie search as a case study, we explore the characteristics of questions posed by searchers in TOT states in a community question answering website. We analyze how searchers express their information needs during TOT states in the movie domain. Specifically, what information do searchers remember about the item being sought and how do they convey this information? Our results suggest that searchers use a combination of information about: (1) the content of the item sought, (2) the context in which they previously engaged with the item, and (3) previous attempts to find the item using other resources (e.g., search engines). Additionally, searchers convey information by sometimes expressing uncertainty (i.e., hedging), opinions, emotions, and by performing relative (vs. absolute) comparisons with attributes of the item. As a result of our analysis, we believe that searchers in TOT states may require specialized query understanding methods or document representations. Finally, our preliminary retrieval experiments show the impact of each information type presented in information requests on retrieval performance. Jaime Arguello, Adam Ferguson, Emery Fine, Bhaskar Mitra 0001, Hamed Zamani, Fernando Diaz 0001 |
CHIIR | 1 |
| 2021 | OrgBox: A Knowledge Representation Tool to Support Complex Search TasksabstractCurrent search systems are effective in helping users complete simple search tasks (e.g., fact-finding). However, they provide less support for users completing complex search tasks. Complex search tasks involve a diverse set of cognitive and metacognitive activities, such as goal-setting, organizing information, drawing inferences, monitoring progress, and updating mental models. We report on a lab study (N=32) that investigated the uses and influences of a novel knowledge representation tool called the "OrgBox'', developed to support searchers with complex tasks. The OrgBox was integrated into a custom-built search system and allowed participants to save information by drag-and-dropping textual passages into the tool, organize passages into "boxes'', and make notes on passages and boxes. The OrgBox tool was compared to a baseline tool (the "Bookmark'') that allowed participants to save passages, but not organize them nor make notes. We investigate four research questions. In RQ1, we investigate the effects of the knowledge representation tool on participants' post-task perceptions. In RQ2-RQ4, we investigate: (RQ2) how participants used different features of each tool; (RQ3) the perceived benefits and challenges of each tool; and (RQ4) the influences of each tool on the approaches taken by participants to complete the task. To address RQ2-RQ4, we conducted a qualitative analysis of participants' responses during an exit interview. We discuss implications from our results for designing tools to support users with complex search tasks. Bogeum Choi, Jaime Arguello, Robert G. Capra, Austin R. Ward |
CHIIR | 2 |
| 2021 | A Study of Explainability Features to Scrutinize Faceted Filtering ResultsabstractFaceted search systems enable users to filter results by selecting values along different dimensions or facets. Traditionally, facets have corresponded to properties of information items that are part of the document metadata. Recently, faceted search systems have begun to use machine learning to automatically associate documents with facet-values that are more subjective and abstract. Examples include search systems that support topic-based filtering of research articles, concept-based filtering of medical documents, and tag-based filtering of images. While machine learning can be used to infer facet-values when the collection is too large for manual annotation, machine-learned classifiers make mistakes. In such cases, it is desirable to have a scrutable system that explains why a filtered result is relevant to a facet-value. Such explanations are missing from current systems. In this paper, we investigate how explainability features can help users interpret results filtered using machine-learned facets. We consider two explainability features: (1) showing prediction confidence values and (2) highlighting rationale sentences that played an influential role in predicting a facet-value. We report on a crowdsourced study involving 200 participants. Participants were asked to scrutinize movie plot summaries predicted to satisfy multiple genres and indicate their agreement or disagreement with the system. Participants were exposed to four interface conditions. We found that both explainability features had a positive impact on participants' perceptions and performance. While both features helped, the sentence-highlighting feature played a more instrumental role in enabling participants to reject false positive cases. We discuss implications for designing tools to help users scrutinize automatically assigned facet-values. Jiaming Qu, Jaime Arguello, Yue Wang 0035 |
CIKM | 2 |
| 2021 | A Deep Analysis of an Explainable Retrieval Model for Precision Medicine Literature Search
Jiaming Qu, Jaime Arguello, Yue Wang 0035 |
ECIR (1) | 2 |
| 2020 | Sources of Evidence for Interactive Table CompletionabstractAn important question in interactive information retrieval (IIR) is: How can we support searchers with specific types of search tasks? We describe an auxiliary support tool referred to as the "Matrix''. The Matrix tool was designed to support searchers with comparative search tasks, which require comparing items along different dimensions. The Matrix was designed as a grid of rows and columns representing the items and dimensions related to a comparative task. The Matrix was integrated with a custom-built search interface, which allowed users to search for information and drag-and-drop relevant passages directly into cells in the Matrix. We investigate the following general question: Given a partially completed Matrix, can a system automatically populate empty cells in the Matrix with relevant passages? To this end, we conducted two crowdsourced studies in which participants were assigned comparative tasks and asked to use our system (integrated search interface + Matrix) to populate every cell in the Matrix. After gathering this data, we evaluated machine-learned models for ranking passages in response to an empty Matrix cell and partially completed Matrix. We address two research questions: (RQ1) What are useful types of features for this predictive task? and (RQ2) How does performance vary based on the level of Matrix completion? We view our research as a step towards designing support tools that: (1) help users organize information while searching and (2) can autocomplete search tasks by exploiting the task structure and a searcher's partial solution. Jaime Arguello, Robert G. Capra |
CHIIR | 1 |
| 2020 | Wizard of Oz Interface to Study System Initiative for Conversational SearchabstractWe describe a Wizard of Oz (WoZ) system, a WebApp, which we use to study how a conversational search system should take the initiative when engaging with users during collaborative search. This system integrates directly into Slack, a chat-messaging platform where users will collaborate. Through our system, the Wizard plays the role of a conversational search system that can search for information, send relevant web results, and message users. In our research, we study three Wizard conditions: bot\_info, bot\_dialog, and bot\_task, which differ in terms of how the Wizard can intervene in a conversation. The intervention modes follow the mixed-initiative framework by Chu-Carroll and Brown~\citechu1997tracking and provide us a foundation to study system initiative for conversational search. In this paper, we describe our design decisions and technical details on how we implemented the system. Sandeep Avula, Jaime Arguello |
CHIIR | 2 |
| 2020 | A Qualitative Analysis of the Effects of Task Complexity on the Functional Role of InformationabstractAn important question in interactive information retrieval (IIR) is: How do task characteristics influence users' needs? In this paper, we investigate the effects of cognitive task complexity on the types of information considered useful for a task. We characterize information types from two perspectives. From one perspective, we classify task-related information items based on inherent characteristics (referred to as info-types): factual statements, concepts/definitions, opinionated statements, and insights---tips/advice related to the task domain. From a second perspective, we used Byströ m and J\"a rvelin's framework~\citebystrom1995 to define information types based on how the information might be used to complete the task (referred to as functional roles): (1) to help the task doer understand the task requirements (problem information); (2) to help the task doer strategize on how to approach the task (problem-solving information); and (3) to help the task doer learn about the task domain (domain information). Our results suggest that: (1) cognitive task complexity influences the functional roles of information items deemed useful for the task (RQ1); (2) certain info-types are more (or less) likely to play certain functional roles (RQ2); and task complexity influences the variety of functional roles played by info-types (RQ3). Bogeum Choi, Jaime Arguello |
CHIIR | 2 |
| 2020 | Towards Explainable Retrieval Models for Precision Medicine Literature SearchabstractIn professional search tasks such as precision medicine literature search, queries often involve multiple aspects. To assess the relevance of a document, a searcher often painstakingly validates each aspect in the query and follows a task-specific logic to make a relevance decision. In such scenarios, we say the searcher makes a structured relevance judgment, as opposed to the traditional univariate (binary or graded) relevance judgment. Ideally, a search engine can support searcher's workflow and follow the same steps to predict document relevance. This approach may not only yield highly effective retrieval models, but also open up opportunities for the model to explain its decision in the same "lingo" as the searcher. Using structured relevance judgment data from the TREC Precision Medicine track, we propose novel retrieval models that emulate how medical experts make structured relevance judgments. Our experiments demonstrate that these simple, explainable models can outperform complex, black-box learning-to-rank models. Jiaming Qu, Jaime Arguello, Yue Wang 0035 |
SIGIR | 2 |
| 2020 | An empirical study of interest, task complexity, and search behaviour on user engagement
Heather L. O'Brien, Jaime Arguello, Robert G. Capra |
Inf. Process. Manag. | 2 |
| 2020 | The Effects of Task Complexity on the Use of Different Types of Information in a Search Assistance ToolabstractIn interactive information retrieval, an important research question is: How do task characteristics influence users’ needs and behaviors? We report on a laboratory study ( N =32) that investigated the effects of task complexity on the types of information used by participants while searching. Participants completed tasks of four complexity levels and had access to four different types of information provided through a search-assistance tool referred to as the InfoBoxes (IB). The IB tool presented the following types of task-related information ( info-types ) on different tabs: (1) facts, (2) concepts, (3) opinions, and (4) insights. Facts (and opinions) were defined as objective (and subjective) statements relevant to the task. Concepts were defined as important ideas, principles, or entities related to the task. Insights were defined as tips or advice about the task. The study investigated six research questions that considered the effects of task complexity on: (RQ1) participants’ pre-/post-task perceptions about useful info-types; (RQ2) use of different info-types during the task; (RQ3) motivations for engaging with the IB; (RQ4) gains from using it; (RQ5) the search stage participants were in while engaging with the IB; and (RQ6) motivations for sometimes avoiding the IB. Our results suggest that task complexity influenced all six types of outcomes. We discuss implications of our results for designing search assistance tools and systems that favor certain types of content based on task characteristics. Bogeum Choi, Austin R. Ward, Yuan Li 0035, Jaime Arguello, Robert G. Capra |
ACM Trans. Inf. Syst. | 4 |
| 2019 | Embedding Search into a Conversational Platform to Support Collaborative SearchabstractPopular messaging platforms such as Slack have given rise to thousands of applications (or bots) that users can engage with individually or as a group. In this paper, we study the use of searchbots (i.e., bots that perform specific types of searches) during collaborative information-seeking tasks mediated through Slack. We report on a user study in which 27 pairs of participants were exposed to three searchbot conditions (a within-subjects design). In the first condition, participants completed the task by searching independently and coordinating through Slack (no searchbot). In the second condition, participants could only search inside of Slack using the searchbot. In the third condition, participants could both search inside of Slack using the searchbot and outside of Slack using their own independent search interfaces. We investigate four research questions focusing on the influence of the searchbot condition on outcomes associated with: (RQ1) participants' levels of workload, (RQ2) collaborative awareness, (RQ3) experiences interacting with the searchbot, and (RQ4) search behaviors. Our results suggest opportunities and challenges in designing searchbots to support collaborative search. On one hand, access to the searchbot resulted in more collaborative awareness, ease of coordination, and fewer duplicated searches. On the other hand, forcing participants to share the querying environment resulted in fewer overall queries, fewer query refinements by individuals, and greater levels of effort. We discuss the implications of our findings for designing effective searchbots to support collaborative search. Sandeep Avula, Jaime Arguello, Robert G. Capra, Jordan Dodson, Yuhui Huang, Filip Radlinski |
CHIIR | 2 |
| 2019 | The Effects of Working Memory during Search Tasks of Varying ComplexityabstractWe report on a study that evaluated the effects of working memory and task complexity on participants' perceptions, behaviors, and outcomes. Twenty-four participants performed two search tasks of varying complexity and completed a psychometric test to measure working memory ability. Our results found several important trends. First, task complexity had an effect on participants' perceptions about temporal demand and satisfaction with the time spent on the task. Second, participants with higher working memory exerted more search effort (e.g., issued more queries). Third, participants with higher working memory had better outcomes, particularly during more complex tasks. Finally, while participants with lower working memory exerted less effort (engaged in satisficing behaviors) and had weaker outcomes, working memory did not affect participants' post-task perceptions about workload and satisfaction. We discuss implications of our results for developing search tools to support users with varying levels of working memory. Bogeum Choi, Robert G. Capra, Jaime Arguello |
CHIIR | 3 |
| 2019 | Using Trails to Support Users with Tasks of Varying ScopeabstractA search trail is an interactive visualization of how a previous searcher approached a related task. Using search trails to assist users requires understanding aspects of the task, user, and trails. In this paper, we examine two questions. First, what are task characteristics that influence a user's ability to gain benefits from others' trails? Second, what is the impact of a "mismatch" between a current user's task and previous user's task which originated the trail? We report on a study that investigated the influence of two factors on participants' perceptions and behaviors while using search trails to complete tasks. Our first factor, task scope, focused on the scope of the task assigned to the participant (broad to narrow). Our manipulation of this factor involved varying the number of constraints associated with tasks. Our second factor, trail scope, focused on the scope of the task that originated the search trails given to participants. We investigated how task scope and trail scope affected participants' (RQ1) pre-task perceptions, (RQ2) post-task perceptions, and (RQ3) search behaviors. We discuss implications of our results for systems that use search trails to provide assistance. Robert G. Capra, Jaime Arguello |
SIGIR | 2 |
| 2019 | The Effects of Working Memory, Perceptual Speed, and Inhibition in Aggregated SearchabstractPrior work has studied how different characteristics of individual users (e.g., personality traits and cognitive abilities) can impact search behaviors and outcomes. We report on a laboratory study ( N = 32) that investigated the effects of three different cognitive abilities (perceptual speed, working memory, and inhibition) in the context of aggregated search. Aggregated search systems combine results from multiple heterogeneous sources (or verticals ) in a unified presentation. Participants in our study interacted with two different aggregated search interfaces (a within-subjects design) that differed based on the extent to which the layout distinguished between results originating from different verticals. The interleaved interface merged results from different verticals in a fairly unconstrained fashion. Conversely, the blocked interface displayed results from the same vertical as a group, displayed each group of vertical results in the same region on the SERP for every query, and used a border around each group of vertical results to help distinguish among results from different sources. We investigated three research questions (RQ1--RQ3). Specifically, we investigated the effects of the interface condition and each cognitive ability on three types of outcomes: (RQ1) participants’ levels of workload, (RQ2) participants’ levels of user engagement, and (RQ3) participants’ search behaviors. Our results found different main and interaction effects. Perceptual speed and inhibition did not significantly affect participants’ workload and user engagement but significantly affected their search behaviors. Specifically, with the interleaved interface, participants with lower perceptual speed had more difficulty finding relevant results on the SERP, and participants with lower inhibitory attention control searched at a slower pace. Working memory did not have a strong effect on participants’ behaviors but had several significant effects on the levels of workload and user engagement reported by participants. Specifically, participants with lower working memory reported higher levels of workload and lower levels of user engagement. We discuss implications of our results for designing aggregated search interfaces that are well suited for users with different cognitive abilities. Jaime Arguello, Bogeum Choi |
ACM Trans. Inf. Syst. | 1 |
| 2018 | SearchBots: User Engagement with ChatBots during Collaborative SearchabstractPopular messaging platforms such as Slack have given rise to hundreds of chatbots that users can engage with individually or as a group. We present a Wizard of Oz study on the use of searchbots (i.e., chatbots that perform specific types of searches) during collaborative information-seeking tasks. Specifically, we study searchbots that intervene dynamically and compare between two intervention types: (1) the searchbot presents questions to users to gather the information it needs to produce results, and (2) the searchbot monitors the conversation among the collaborators, infers the necessary information, and then displays search results with no additional input from the users. We investigate three research questions: (RQ1) What is the effect of a searchbot (and its intervention type) on participants» collaborative experience' (RQ2) What is the effect of a searchbot»s intervention type on participants» perceptions about the searchbot and level of engagement with the searchbot' and (RQ3) What are participants» impressions of a dynamic searchbot? Our results suggest that dynamic searchbots can enhance users» collaborative experience and that the intervention type does not greatly affect users» perceptions and level of engagement. Participants» impressions of the searchbot suggest unique opportunities and challenges for future work. Sandeep Avula, Gordon Chadwick, Jaime Arguello, Robert G. Capra |
CHIIR | 3 |
| 2018 | Second International Workshop on Conversational Approaches to Information Retrieval (CAIR'18): Workshop at SIGIR 2018abstractThe CAIR'18 workshop will bring together academic and industrial researchers to create a forum for research on conversational approaches to search and recommendation. A specific focus will be on techniques that support complex and multi-turn user-machine dialogues for information access and retrieval, and multi-modal interfaces for interacting with such systems. Jaime Arguello, Filip Radlinski, Hideo Joho, Damiano Spina, Julia Kiseleva |
SIGIR | 1 |
| 2018 | The Effects of Manipulating Task Determinability on Search Behaviors and OutcomesabstractAn important area of IR research involves understanding how task characteristics influence search behaviors and outcomes. Task complexity is one characteristic that has received considerable attention. One view of task complexity is through the lens of a priori determinability -- the level of uncertainty about task outcomes and processes experienced by the searcher. In this work, we manipulated the determinability of comparative tasks. Our task manipulation involved modifying the scope of the task by specifying exact items and/or exact (objective or subjective) dimensions to consider as part of the task. This paper reports on a within-subject study (N=144) where we investigated how our task manipulation influenced participants' perceptions, levels of engagement, search effort, and choice of search strategies. Our results suggest a complex relationship between task scope, determinability, and different outcome measures. Our most open-ended tasks were perceived to have low determinability (high uncertainty), but were the least challenging for participants due to satisficing. Furthermore, narrowing the scope of tasks by specifying items had a different effect than by specifying dimensions. Specifying items increased the task determinability (lower uncertainty) and made the task easier, while specifying dimensions did not increase the task determinability and made the task more challenging. A qualitative analysis of participants' queries suggests that searching for dimensions is more challenging than for items. Finally, we observed subtle differences between objective and subjective dimensions. We discuss implications for the design of IIR studies and tools to support users. Robert G. Capra, Jaime Arguello, Heather L. O'Brien, Yuan Li 0035, Bogeum Choi |
SIGIR | 2 |
| 2018 | Factors Influencing Users' Information Requests: Medium, Target, and Extra-Topical DimensionabstractWe report on a crowdsourced study that investigated how two factors influence the way people formulate information requests. Our first factor, medium , considers whether the request is produced using text or voice. Our second factor, target , considers whether the request is intended for a search engine or a human intermediary (i.e., someone who will search on the user’s behalf). In particular, we study how these two factors influence the way people formulate requests in situations where the information need has a specific type of extra-topical dimension (i.e., a type of constraint that is independent from the information need’s topic). We focus on six extra-topical dimensions: (1) domain knowledge, (2) viewpoint, (3) experiential, (4) venue location, (5) source location, and (6) temporal. The extra-topical dimension was manipulated by giving participants carefully constructed search tasks. We analyzed a large number of information requests produced by study participants, and address three research questions. We study the effects of our two factors (medium and target) on (RQ1) participants’ perceptions about their own information requests, (RQ2) the different characteristics of their information requests (e.g., natural language structure, retrieval performance), and (RQ3) participants’ strategies for requesting information when the search task has a specific type of extra-topical dimension. Our results found that both factors influenced participants’ perceptions about their own information requests, the characteristics of participants’ requests, and the strategies adopted by participants to request information matching the extra-topical dimension. Our results have implications for future research on methods that can harness (rather than ignore) extra-topical query terms to retrieve relevant information. Jaime Arguello, Bogeum Choi, Robert G. Capra |
ACM Trans. Inf. Syst. | 1 |
| 2017 | Time Limits, Information Search and the Use of Search AssistanceabstractIn this paper, we analyze the impact of task time limits as described by 24 participants in an experimental study investigating the use of a search assistance tool, the Search Guide (SG), during four search tasks with varying levels of cognitive complexity. Participants were given a 12 minute task time limit and a warning notification when 3 minutes remained. In post-experiment interviews, participants reported two different impacts of the time limit on use of the SG: ten described using the SG as a way to save time given the time limit while five reported less SG use due to uncertainty whether it would contain useful information. Participants also described skimming and reading pages more shallowly, selecting easier to read search results, and bookmarking pages more freely because of the limited time they had to complete the task. Anita Crescenzi, Robert G. Capra, Jaime Arguello |
CHIIR | 3 |
| 2017 | Using Query Performance Predictors to Reduce Spoken Queries
Jaime Arguello, Sandeep Avula, Fernando Diaz 0001 |
ECIR | 1 |
| 2017 | The Effects of Search Task Determinability on Search Behavior
Robert G. Capra, Jaime Arguello |
ECIR | 2 |
| 2017 | First International Workshop on Conversational Approaches to Information Retrieval (CAIR'17)abstractRecent advances in commercial conversational services that allow naturally spoken and typed interaction, particularly for well-formulated questions and commands, have increased the need for more human-centric interactions in information retrieval. The First International Workshop on Conversational Approaches to Information Retrieval (CAIR`17) brings together academic and industrial researchers to create a forum for research on conversational approaches to search. A specific focus is on techniques that support complex and multi-turn user-machine dialogues for information access and retrieval, and multi-model interfaces for interacting with such systems. We invite submissions addressing all modalities of conversation, including speech-based, text-based, and multimodal interaction. We also welcome studies of human-human interaction (e.g., collaborative search) that can inform the design of conversational search applications, and work on evaluation of conversational approaches. Hideo Joho, Lawrence Cavedon, Jaime Arguello, Milad Shokouhi, Filip Radlinski |
SIGIR | 3 |
| 2016 | Using Query Performance Predictors to Improve Spoken Queries
Jaime Arguello, Sandeep Avula, Fernando Diaz 0001 |
ECIR | 1 |
| 2016 | To Blend or Not to Blend?: Perceptual Speed, Visual Memory and Aggregated SearchabstractWhile aggregated search interfaces that present vertical results to searchers are fairly common in today's search environments, little is known about how searchers' cognitive abilities impact how they use and evaluate these interfaces. This study evaluates the relationship between two cognitive abilities ? perceptual speed and visual memory ? and searchers' behaviors and interface preferences when using two aggregated search interfaces: one that blends vertical results into the search results (blended) and one that does not (non-blended). Cognitive tests were administered to sixteen participants who subsequently performed four search tasks using the two interfaces. Participants' search interactions were logged and after searching, they rated the usability, engagement and effectiveness of each interface, as well as made comparative evaluations. Results showed that participants with low perceptual speed spent significantly more time completing tasks when using the blended interface, while those with high perceptual speed spent roughly equivalent amounts of time completing tasks with the two interfaces. Those with low perceptual speed also rated both interfaces as significantly less usable along many measures, and were less satisfied with their searches. There were also main effects for interface: participants rated the non-blended interface significantly more usable than the blended interface. Lauren Turpin, Diane Kelly 0001, Jaime Arguello |
SIGIR | 3 |
| 2016 | The Effects of Aggregated Search Coherence on Search BehaviorabstractAggregated search is the task of combining results from multiple independent search systems in a single Search Engine Results Page (SERP). Aggregated search coherence refers to the extent to which different sources on the SERP focus on similar senses of an ambiguous or underspecified query. In previous studies, we found that the query senses in a set of vertical results can influence user engagement with the web results (the so-called “spillover” effect). In this work, we investigate five research questions (RQ1--RQ5) that extend our prior work. First, we investigate the extent to which results from different sources focus on different senses of an ambiguous query (RQ1). Second, we investigate how the vertical-to-web spillover effect varies across different verticals (RQ2). Then, we examine whether the level of spillover depends on the vertical position (RQ3) and on whether the vertical results are displayed with a border and different-colored background to distinguish them from the web results (RQ4). Finally, we propose a new method for displaying results from a particular vertical that are more consistent with the query senses in the web results (RQ5). We evaluate this new method based on how it influences users to make more correct decisions with respect to the web results—to engage with the web results when at least one of them is relevant and to avoid engaging with the web results otherwise. Our results show the following trends. In terms of RQ1, our analysis suggests that the top results from the web search engine are more diversified than the top results from our four different verticals considered (images, news, shopping, and video). In terms of RQ2, we found a stronger spillover effect for the images vertical than the news, shopping, and video verticals. In terms of RQ3, we found a stronger level of spillover when the vertical was positioned at the top of the SERP versus to the right side of the web results. In terms of RQ4, we found an interesting additive effect between the vertical’s position and displaying the vertical results enclosed in a border and with a different-colored background—the image vertical had no spillover when presented to the right side of the web results and with a border and background. Finally, in terms of RQ5, we found that our proposed vertical results selection approach can influence users to make more correct predictions about their level of engagement with the web results. Jaime Arguello, Robert G. Capra |
ACM Trans. Inf. Syst. | 1 |
| 2015 | Improving Aggregated Search Coherence
Jaime Arguello |
ECIR | 1 |
| 2015 | Predicting Speech Acts in MOOC Forum Posts
Jaime Arguello, Kyle Shaffer |
ICWSM | 1 |
| 2015 | SIGIR 2015 Workshop on Reproducibility, Inexplicability, and Generalizability of Results (RIGOR)abstractNo abstract available. Jaime Arguello, Fernando Diaz 0001, Jimmy Lin, Andrew Trotman |
SIGIR | 1 |
| 2015 | Differences in the Use of Search Assistance for Tasks of Varying ComplexityabstractIn this paper, we study how users interact with a search assistance tool while completing tasks of varying complexity. We designed a novel tool referred to as the search guide (SG) that displays the search trails (queries issued, results clicked, pages bookmarked) from three previous users who completed the task. We report on a laboratory study with 48 participants that investigates different factors that may influence user interaction with the SG and the effects of the SG on different outcome measures. Participants were asked to find and bookmark pages for four tasks of varying complexity and the SG was made available to half the participants. We collected log data and conducted retrospective stimulated recall interviews to learn about participants' use of the SG. Our results suggest the following trends. First, interaction with the SG was greater for more complex tasks. Second, the a priori determinability of the task (i.e., whether the task was perceived to be well-defined) helped predict whether participants gained a bookmark from the SG. Third, participants who interacted with the SG, but did not gain a bookmark, felt less system support than those who gained a bookmark and those who did not interact. Finally, a qualitative analysis of our interviews suggests differences in motivation and benefits from SG use for different levels of task complexity. Our findings extend prior research on search assistance tools and provide insights for the design of systems to help users with complex search tasks. Robert G. Capra, Jaime Arguello, Anita Crescenzi, Emily Vardell |
SIGIR | 2 |
| 2014 | The Effects of Vertical Rank and Border on Aggregated Search Coherence and Search BehaviorabstractAggregated search is the task of blending results from different search services, or verticals, into a set of web search results. Aggregated search coherence is the extent to which results from different sources focus on similar senses of an ambiguous or underspecified query. Prior work investigated the "spill-over" effect between a set of blended vertical results and the web results. These studies found that users are more likely to interact with the web results when the vertical results are more consistent with the user's intended query-sense. We extend this prior work by investigating three new research questions: (1) Does the spill-over effect generalize across different verticals? (2) Does the vertical rank moderate the level of spill-over? and (3) Does the presence of a border around the vertical results moderate the level of spill-over? We investigate four different verticals (images, news, shopping, and video) and measure spill-over using interaction measures associated with varying levels of engagement with the web results (bookmarks, clicks, scrolls, and mouseovers). Results from a large-scale crowdsourced study suggest that: (1) The spill-over effect generalizes across verticals, but is stronger for some verticals than others, (2) Vertical rank has a stronger moderating effect for verticals with a mid-level of spill-over, and (3) Including a border around the vertical results has a subtle moderating effect for those verticals with a low level of spill-over. Jaime Arguello, Robert G. Capra |
CIKM | 1 |
| 2014 | Predicting Search Task Difficulty
Jaime Arguello |
ECIR | 1 |
| 2013 | Factors affecting aggregated search coherence and search behaviorabstractAggregated search is the task of incorporating results from different search services, or verticals, into the web search results. Aggregated search coherence refers to the extent to which results from different sources focus on similar senses of a given query. Prior research investigated aggregated search coherence between images and web results. A user study showed that users are more likely to interact with the web results when the images are more consistent with the intended query-sense. We build upon this work and address three outstanding research questions about aggregated search coherence: (1) Does the same "spill-over" effect generalize to other verticals besides images? (2) Is the effect stronger when the vertical results include image thumbnails? and (3) What factors influence if and when a spill-over occurs from a user's perspective? We investigate these questions using a large-scale crowdsourcing study and a smaller-scale laboratory study. Results suggest that the spill-over effect occurs for some verticals (images, shopping, video), but not others (news), and that including thumbnails in the vertical results has little effect. Qualitative data from our laboratory study provides insights about participants' actions and thought-processes when faced with (in)coherent results. Jaime Arguello, Robert G. Capra, Wan-Ching Wu |
CIKM | 1 |
| 2013 | Augmenting web search surrogates with imagesabstractWhile images are commonly used in search result presentation for vertical domains such as shopping and news, web search results surrogates remain primarily text-based. In this paper, we present results of two large-scale user studies to examine the effects of augmenting text-based surrogates with images extracted from the underlying webpage. We evaluate effectiveness and efficiency at both the individual surrogate level and at the results page level. Additionally, we investigate the influence of two factors: the goodness of the image in terms of representing the underlying page content, and the diversity of the results on a results page. Our results show that at the individual surrogate level, good images provide only a small benefit in judgment accuracy versus text-only surrogates, with a slight increase in judgment time. At the results page level, surrogates with good images had similar effectiveness and efficiency compared to the text-only condition. However, in situations where the results page items had diverse senses, surrogates with images had higher click precision versus text-only ones. Results of these studies show tradeoffs in the use of images in web search surrogates, and highlight particular situations where they can provide benefits. Robert G. Capra, Jaime Arguello, Falk Scholer |
CIKM | 2 |
| 2012 | The effect of aggregated search coherence on search behaviorabstractAggregated search is the task of blending results from different specialized search services, or verticals, into the web search results. Aggregated search coherence refers to the degree to which results from different systems focus on similar senses of the query. While cross-component coherence has been cited as an important criterion for whole-page evaluation, its effect on search behavior has not been deeply investigated in prior research. In this work, we focus on the coherence between two aggregated search components: images and web results. In particular, we investigate whether the query-senses associated with the blended image results can influence user interaction with the web results. For example, if a user wants web results about "jaguar" the animal, are they more likely to examine the web results if the image results contain pictures of the animal instead of pictures of the car? Based on two large user studies, our results show that the image results can systematically affect user interaction with the web results. If the web results are largely consistent with the search task, then the effect of the image results is small. However, if the web results are only marginally consistent with the search task, such as when they are highly diversified across query-senses, the image results have a significant effect on user interaction with the web results. Our findings have implications on current research in whole-page evaluation, aggregated search, and diversity ranking. Jaime Arguello, Robert G. Capra |
CIKM | 1 |
| 2012 | Task complexity, vertical display and user interaction in aggregated searchabstractAggregated search is the task of blending results from specialized search services or verticals into the Web search results. While many studies have focused on aggregated search techniques, few studies have tried to better understand how users interact with aggregated search results. This study investigates how task complexity and vertical display (the blending of vertical results into the web results) affect the use of vertical content. Twenty-nine subjects completed six search tasks of varying levels of task complexity using two aggregated search interfaces: one that blended vertical results into the web results and one that only provided indirect vertical access. Our results show that more complex tasks required significantly more interaction and that subjects completing these tasks examined more vertical results. While the amount of interaction was the same between interfaces, subjects clicked on more vertical results when these were blended into the web results. Our results also show an interaction between task complexity and vertical display; subjects clicked on more verticals when completing the more complex tasks with the interface that blended vertical results. Subjects' evaluations of the two interfaces were nearly identical, but when analyzed with respect to their interface preferences, we found a positive relationship between system evaluations and individual preferences. Subjects justified their preference using similar rationales and their comments illustrate how the display itself can influence judgments of information quality, especially in cases when the vertical results might not be relevant to the search task. Jaime Arguello, Wan-Ching Wu, Diane Kelly 0001, Ashlee Edwards |
SIGIR | 1 |
| 2011 | Learning to aggregate vertical results into web search resultsabstractAggregated search is the task of integrating results from potentially multiple specialized search services, or verticals, into the Web search results. The task requires predicting not only which verticals to present (the focus of most prior research), but also predicting where in the Web results to present them (i.e., above or below the Web results, or somewhere in between). Learning models to aggregate results from multiple verticals is associated with two major challenges. First, because verticals retrieve different types of results and address different search tasks, results from different verticals are associated with different types of predictive evidence (or features). Second, even when a feature is common across verticals, its predictiveness may be vertical-specific. Therefore, approaches to aggregating vertical results require handling an inconsistent feature representation across verticals, and, potentially, a vertical-specific relationship between features and relevance. We present 3 general approaches that address these challenges in different ways and compare their results across a set of 13 verticals and 1070 queries. We show that the best approaches are those that allow the learning algorithm to learn a vertical-specific relationship between features and relevance. Jaime Arguello, Fernando Diaz 0001, Jamie Callan |
CIKM | 1 |
| 2011 | A Methodology for Evaluating Aggregated Search Results
Jaime Arguello, Fernando Diaz 0001, Jamie Callan, Ben Carterette |
ECIR | 1 |
| 2010 | Vertical selection in the presence of unlabeled verticalsabstractVertical aggregation is the task of incorporating results from specialized search engines or verticals (e.g., images, video, news) into Web search results. Vertical selection is the subtask of deciding, given a query, which verticals, if any, are relevant. State of the art approaches use machine learned models to predict which verticals are relevant to a query. When trained using a large set of labeled data, a machine learned vertical selection model outperforms baselines which require no training data. Unfortunately, whenever a new vertical is introduced, a costly new set of editorial data must be gathered. In this paper, we propose methods for reusing training data from a set of existing (source) verticals to learn a predictive model for a new (target) vertical. We study methods for learning robust, portable, and adaptive cross-vertical models. Experiments show the need to focus on different types of features when maximizing portability (the ability for a single model to make accurate predictions across multiple verticals) than when maximizing adaptability (the ability for a single model to make accurate predictions for a specific vertical). We demonstrate the efficacy of our methods through extensive experimentation for 11 verticals Jaime Arguello, Fernando Diaz 0001, Jean-François Paiement |
SIGIR | 1 |
| 2009 | Classification-based resource selectionabstractIn some retrieval situations, a system must search across multiple collections. This task, referred to as federated search, occurs for example when searching a distributed index or aggregating content for web search. Resource selection refers to the subtask of deciding, given a query, which collections to search. Most existing resource selection methods rely on evidence found in collection content. We present an approach to resource selection that combines multiple sources of evidence to inform the selection decision. We derive evidence from three different sources: collection documents, the topic of the query, and query click-through data. We combine this evidence by treating resource selection as a multiclass machine learning problem. Although machine learned approaches often require large amounts of manually generated training data, we present a method for using automatically generated training data. We make use of and compare against prior resource selection work and evaluate across three experimental testbeds. Jaime Arguello, Jamie Callan, Fernando Diaz 0001 |
CIKM | 1 |
| 2009 | Sources of evidence for vertical selectionabstractWeb search providers often include search services for domain-specific subcollections, called verticals, such as news, images, videos, job postings, company summaries, and artist profiles. We address the problem of vertical selection, predicting relevant verticals (if any) for queries issued to the search engine's main web search page. In contrast to prior query classification and resource selection tasks, vertical selection is associated with unique resources that can inform the classification decision. We focus on three sources of evidence: (1) the query string, from which features are derived independent of external resources, (2) logs of queries previously issued directly to the vertical, and (3) corpora representative of vertical content. We focus on 18 different verticals, which differ in terms of semantics, media type, size, and level of query traffic. We compare our method to prior work in federated search and retrieval effectiveness prediction. An in-depth error analysis reveals unique challenges across different verticals and provides insight into vertical selection for future work. Jaime Arguello, Fernando Diaz 0001, Jamie Callan, Jean-François Crespo |
SIGIR | 1 |
| 2009 | Adaptation of offline vertical selection predictions in the presence of user feedbackabstractWeb search results often integrate content from specialized corpora known as verticals. Given a query, one important aspect of aggregated search is the selection of relevant verticals from a set of candidate verticals. One drawback to previous approaches to vertical selection is that methods have not explicitly modeled user feedback. However, production search systems often record a variety of feedback information. In this paper, we present algorithms for vertical selection which adapt to user feedback. We evaluate algorithms using a novel simulator which models performance of a vertical selector situated in realistic query traffic. Fernando Diaz 0001, Jaime Arguello |
SIGIR | 2 |
| 2008 | Document Representation and Query Expansion Models for Blog Recommendation
Jaime Arguello, Jonathan L. Elsas, Jamie Callan, Jaime G. Carbonell |
ICWSM | 1 |
| 2008 | Retrieval and feedback models for blog feed searchabstractBlog feed search poses different and interesting challenges from traditional ad hoc document retrieval. The units of retrieval, the blogs, are collections of documents, the blog posts. In this work we adapt a state-of-the-art federated search model to the feed retrieval task, showing a significant improvement over algorithms based on the best performing submissions in the TREC 2007 Blog Distillation task[12]. We also show that typical query expansion techniques such as pseudo-relevance feedback using the blog corpus do not provide any significant performance improvement and in many cases dramatically hurt performance. We perform an in-depth analysis of the behavior of pseudo-relevance feedback for this task and develop a novel query expansion technique using the link structure in Wikipedia. This query expansion technique provides significant and consistent performance improvements for this task, yielding a 22% and 14% improvement in MAP over the unexpanded query for our baseline and federated algorithms respectively. Jonathan L. Elsas, Jaime Arguello, Jamie Callan, Jaime G. Carbonell |
SIGIR | 2 |