Eric Horvitz

dblp:h/EricHorvitz · also Eric Joel Horvitz · DBLP profile ↗
← Back
256ranked-venue papers
36as first author
28since 2021 · last 2026
0000-0002-8823-0614ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 126 · 24 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 73 · 10 first-author · 8 since 2021Databases, data management, data science and information retrieval · 57 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 3 first-author · 7 since 2021Computer networks · 5Software engineering, systems software and programming languages · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case Preparation
abstract
Multidisciplinary tumor boards (MTBs) bring specialists together to identify therapies for complex cancer cases, but preparing for them is time-intensive. Clinicians must extract key details from extensive records and evaluate treatment options. While large language models (LLMs) show promise in medicine for basic tasks like summarizing notes, little is known about their role in high-stakes tasks like MTB preparation. We conducted a mixed-methods study with 16 oncologists using two AI systems to prepare patient cases for MTB: an off-the-shelf assistant (Copilot) and a task-specific multi-agent system (Healthcare Agent Orchestrator, HAO). We analyzed oncologist prompts, AI responses, and oncologists’ perception of AI. Participants showed greater willingness to adopt HAO but were often overconfident in AI summaries and skeptical of AI-recommended therapies. Trust calibration strategies, such as source links and agent-trajectories, failed to align trust with system capabilities. We conclude with how AI systems should be built to support clinicians in high-stakes tasks.
Amanda K. Hall, Ruican Rachel Zhong, Selin S. Everett, Alyssa Unell, Matthias Blondeel, Jonathan Carlson, Katie Claveau, Thulasee Jose, Tristan Naumann, David C. Rhew, Naiteek Sangani, Frank Tuan, James Weinstein, Varun Mishra 0001, Elizabeth D. Mynatt, T. Scott Saponas, Leonardo Schettini, J. Samuel Preston, Yu Gu 0017, Naoto Usuyama, Zelalem Gero, Cliff Wong, Noel Codella, Hoifung Poon, Shrey Jain, Matthew P. Lungren, Eric Horvitz
CHI30
2026 Learning protein representations with conformational dynamics
abstract
MOTIVATION: Proteins change shape as they work, and these changing states control whether binding sites are exposed, signals are relayed, and catalysis proceeds. Most protein language models (PLMs) pair a sequence with a single structural snapshot, which can miss state-dependent features central to interaction, localization, and enzyme activity. Studies also indicate that many proteins assume multiple, functionally relevant shapes, motivating approaches that learn from this variability. RESULTS: We present DynamicsPLM, a PLM conditioned on ensembles of computationally generated conformations to derive state-aware representations. DynamicsPLM improves predictive performance across protein-protein interaction, subcellular localization, enzyme classification, and metal-ion binding. On a widely used protein-protein interaction benchmark, it achieves a four-point accuracy gain over the strongest baseline. On a curated test set enriched for proteins with multiple conformational states, the margin increases to eleven points. These findings argue for a shift from static to dynamics-aware modeling, in which conformational variability is treated as informative. By elevating conformational state to a central element of machine learning in protein biology, this work advances modeling toward mechanisms that better reflect how proteins operate in cells and provides a route to actionable hypotheses about when and how binding, signaling, and catalysis occur. AVAILABILITY AND IMPLEMENTATION: Code, model weights, and inference scripts are available at https://github.com/kalifadan/DynamicsPLM (DOI: https://doi.org/10.5281/zenodo.17668302).
Dan Kalifa, Eric Horvitz, Kira Radinsky
Bioinform.2
2026 AI, humanity, and the open world of health care: enduring imperatives for the next century
abstract
This address was delivered by Eric Horvitz, MD, PhD, at the 2026 graduation ceremony of Columbia University School of Nursing on May 19, 2026, where he received the Second Century Award for Excellence in Health Care. The address considers the responsibilities of clinicians in shaping the future of artificial intelligence in medicine. It frames health care as an "open world," where information is incomplete, time is limited, and decisions are made under uncertainty. As AI transforms biomedicine and clinical care, the address emphasizes the importance of clinician engagement in guiding how these technologies are developed and used, and calls for systems that strengthen clinical judgment, support care teams, and advance human health, dignity, connection, and trust.
Eric Horvitz
J. Am. Medical Informatics Assoc.1
2025 Navigating Rifts in Human-LLM Grounding: Study and Benchmark
abstract
Language models excel at following instructions but often struggle with the collaborative aspects of conversation that humans naturally employ.This limitation in groundingthe process by which conversation participants establish mutual understanding-can lead to outcomes ranging from frustrated users to serious consequences in high-stakes scenarios.To systematically study grounding challenges in human-LLM interactions, we analyze logs from three human-assistant datasets: WildChat, MultiWOZ, and Bing Chat.We develop a taxonomy of grounding acts and build models to annotate and forecast grounding behavior.Our findings reveal significant differences in humanhuman and human-LLM grounding: LLMs were three times less likely to initiate clarification and sixteen times less likely to provide follow-up requests than humans.Additionally, we find that early grounding failures predict later interaction breakdowns.Building on these insights, we introduce RIFTS, a benchmark derived from publicly available LLM interaction data containing situations where LLMs fail to initiate grounding.We note that current frontier models perform poorly on RIFTS, highlighting the need to reconsider how we train and prompt LLMs for human interaction.To this end, we develop a preliminary intervention aimed at mitigating grounding failures.* Research performed during an internship at Microsoft. 1 We use grounding to refer to Clark's formulation of the language, gestures, and signaling that participants in a conver-
Omar Shaikh, Hussein Mozannar, Gagan Bansal, Adam Fourney, Eric Horvitz
ACL (1)5
2025 Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty Quantification
abstract
There is increasing interest in ``decision-focused" machine learning methods which train models to account for how their predictions are used in downstream optimization problems. Doing so can often improve performance on subsequent decision problems. However, current methods for uncertainty quantification do not incorporate any information at all about downstream decisions. We develop a framework based on conformal prediction to produce prediction sets that account for a downstream decision loss function, making them more appropriate to inform high-stakes decision-making. Our approach harnesses the strengths of conformal methods—modularity, model-agnosticism, and statistical coverage guarantees—while incorporating downstream decisions and user-specified utility functions. We prove that our methods retain standard coverage guarantees. Empirical evaluation across a range of datasets and utility metrics demonstrates that our methods achieve significantly lower decision loss compared to standard conformal methods. Additionally, we present a real-world use case in healthcare diagnosis, where our method effectively incorporates the hierarchical structure of dermatological diseases. It successfully generates sets with coherent diagnostic meaning, aiding the triage process during dermatology diagnosis and illustrating how our method can ground high-stakes decision-making on external domain knowledge.
Santiago Cortes-Gomez, Carlos Miguel Patiño, Yewon Byun, Steven Z. Wu, Eric Horvitz, Bryan Wilder
ICLR5
2025 Improving Instruction-Following in Language Models through Activation Steering
abstract
The ability to follow instructions is crucial for numerous real-world applications of language models. In pursuit of deeper insights and more powerful capabilities, we derive instruction-specific vector representations from language models and use them to steer models accordingly. These vectors are computed as the difference in activations between inputs with and without instructions, enabling a modular approach to activation steering. We demonstrate how this method can enhance model adherence to constraints such as output format, length, and word inclusion, providing inference-time control over instruction following. Our experiments across four models demonstrate how we can use the activation vectors to guide models to follow constraints even without explicit instructions and to enhance performance when instructions are present. Additionally, we explore the compositionality of activation steering, successfully applying multiple instructions simultaneously. Finally, we demonstrate that steering vectors computed on instruction-tuned models can transfer to improve base models. Our findings demonstrate that activation steering offers a practical and scalable approach for fine-grained control in language generation. Our code and data are available at https://github.com/microsoft/llm-steer-instruct.
Alessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz, Besmira Nushi
ICLR4
2025 Creating General User Models from Computer Use
Omar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz, Joon Sung Park 0001, Diyi Yang, Michael S. Bernstein
UIST4
2025 Beyond the leaderboard: leveraging predictive modeling for protein-ligand insights and discovery
abstract
MOTIVATION: Ligands are biomolecules that bind to specific sites on target proteins, often inducing conformational changes important in the protein's function. Knowledge about ligand interactions with proteins are fundamental to understanding biological mechanisms and advancing drug discovery. Traditional protein language models focus on amino acid sequences and 3D structures, overlooking the structural and functional changes induced by protein-ligand interactions. We investigate the value of integrating ligand-protein binding data in several predictive challenges and leverage findings to frame research directions and questions. RESULTS: We show how the integration of protein-ligand interaction data in protein representation learning can increase predictive power. We evaluate the methodology across diverse biological tasks, demonstrating consistent improvements over state-of-the-art models. We further demonstrate how the study of the specific boosts in predictive capabilities coming with the introduction of the ligand modality can serve to focus attention and provide insights on biological mechanisms. By leveraging large pretrained protein language models and enriching them with interaction-specific features through a tailored learning process, we capture functional and structural nuances of proteins in their biochemical context. AVAILABILITY AND IMPLEMENTATION: The full code and data are freely available at https://github.com/kalifadan/ProtLigand (DOI: https://doi.org/10.5281/zenodo.15808053).
Dan Kalifa, Kira Radinsky, Eric Horvitz
Bioinform.3
2025 Reformulating patient stratification for targeting interventions by accounting for severity of downstream outcomes resulting from disease onset: a case study in sepsis
abstract
OBJECTIVES: To quantify differences between (1) stratifying patients by predicted disease onset risk alone and (2) stratifying by predicted disease onset risk and severity of downstream outcomes. We perform a case study of predicting sepsis. MATERIALS AND METHODS: We performed a retrospective analysis using observational data from Michigan Medicine at the University of Michigan (U-M) between 2016 and 2020 and the Beth Israel Deaconess Medical Center (BIDMC) between 2008 and 2012. We measured the correlation between the estimated sepsis risk and the estimated effect of sepsis on mortality using Spearman's correlation. We compared patients stratified by sepsis risk with patients stratified by sepsis risk and effect of sepsis on mortality. RESULTS: The U-M and BIDMC cohorts included 7282 and 5942 ICU visits; 7.9% and 8.1% developed sepsis, respectively. Among visits with sepsis, 21.9% and 26.3% experienced mortality at U-M and BIDMC. The effect of sepsis on mortality was weakly correlated with sepsis risk (U-M: 0.35 [95% CI: 0.33-0.37], BIDMC: 0.31 [95% CI: 0.28-0.34]). High-risk patients identified by both stratification approaches overlapped by 66.8% and 52.8% at U-M and BIDMC, respectively. Accounting for risk of mortality identified an older population (U-M: age = 66.0 [interquartile range-IQR: 55.0-74.0] vs age = 63.0 [IQR: 51.0-72.0], BIDMC: age = 74.0 [IQR: 61.0-83.0] vs age = 68.0 [IQR: 59.0-78.0]). DISCUSSION: Predictive models that guide selective interventions ignore the effect of disease on downstream outcomes. Reformulating patient stratification to account for the estimated effect of disease on downstream outcomes identifies a different population compared to stratification on disease risk alone. CONCLUSION: Models that predict the risk of disease and ignore the effects of disease on downstream outcomes could be suboptimal for stratification.
Fahad Kamran, Donna Tjandra, Thomas S. Valley, Hallie C. Prescott, Nigam H. Shah, Vincent X. Liu, Eric Horvitz, Jenna Wiens
J. Am. Medical Informatics Assoc.7
2025 Perplexity and proximity: Large language model perplexity complements semantic distance metrics for the detection of incoherent speech
abstract
OBJECTIVE: Semantic coherence in speech is characterized by a logical, connected flow of ideas. A lack of coherence in speech may reflect disorganized thinking, a core feature of psychosis in schizophrenia spectrum disorders (SSDs). Developing tools that could help with automated assessment of semantic coherence in language could facilitate early detection of SSDs and improved monitoring of symptoms, enabling more timely intervention. Large language models (LLMs) have demonstrated strong capabilities on numerous language-centric tasks and have shown promise for analyzing semantic coherence due to the natural fit between their innate measures of language perplexity and the surprising turns that incoherent narrative often takes. This study aims to develop a novel representation and associated measure of semantic coherence using LLM-based perplexity metrics and to compare this measure with traditional vector distance-based coherence metrics. METHOD: We evaluated "bag" and "chain" models based on LLM perplexities as measures of semantic coherence. Regression models were trained using both single and paired combinations of perplexity- and proximity-based features to predict human ratings of semantic coherence using standardized instruments. Performance was evaluated on held-out examples from a training set of speeches from individuals experiencing psychotic symptoms and a test set of clinical interviews with patients diagnosed with SSDs, both with labels from human assessments of disorganized thinking severity. RESULTS: The best performance was achieved using a combination of perplexity and proximity features, yielding a Spearman correlation with human ratings of 0.61 (vs. 0.56 with proximity features alone) on leave-one-out cross-validation in the training set, and 0.54 (vs. 0.52 with proximity features alone) on the test set. CONCLUSION: We developed novel methods for assessing semantic coherence using LLM perplexities and found them complementary to proximity-based methods. Combined, these methods showed improved performance across two datasets, highlighting LLM's potential in enhancing automated diagnosis and monitoring of SSDs.
Weizhe Xu, Serguei V. S. Pakhomov, Patrick Heagerty, Eric Horvitz, Ellen Bradley, Joshua Woolley, Andrew T. Campbell, Alex S. Cohen, Dror Ben-Zeev, Trevor Cohen
J. Biomed. Informatics4
2024 When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
abstract
AI powered code-recommendation systems, such as Copilot and CodeWhisperer, provide code suggestions inside a programmer's environment (e.g., an IDE) with the aim of improving productivity. We pursue mechanisms for leveraging signals about programmers' acceptance and rejection of code suggestions to guide recommendations. We harness data drawn from interactions with GitHub Copilot, a system used by millions of programmers, to develop interventions that can save time for programmers. We introduce a utility-theoretic framework to drive decisions about suggestions to display versus withhold. The approach, conditional suggestion display from human feedback (CDHF), relies on a cascade of models that provide the likelihood that recommended code will be accepted. These likelihoods are used to selectively hide suggestions, reducing both latency and programmer verification time. Using data from 535 programmers, we perform a retrospective evaluation of CDHF and show that we can avoid displaying a significant fraction of suggestions that would have been rejected. We further demonstrate the importance of incorporating the programmer's latent unobserved state in decisions about when to display suggestions through an ablation study. Finally, we showcase how using suggestion acceptance as a reward signal for guiding the display of suggestions can lead to suggestions of reduced quality, indicating an unexpected pitfall.
Hussein Mozannar, Gagan Bansal, Adam Fourney, Eric Horvitz
AAAI4
2024 Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
abstract
Code-recommendation systems, such as Copilot and CodeWhisperer, have the potential to improve programmer productivity by suggesting and auto-completing code. However, to fully realize their potential, we must understand how programmers interact with these systems and identify ways to improve that interaction. To seek insights about human-AI collaboration with code recommendations systems, we studied GitHub Copilot, a code-recommendation system used by millions of programmers daily. We developed CUPS, a taxonomy of common programmer activities when interacting with Copilot. Our study of 21 programmers, who completed coding tasks and retrospectively labeled their sessions with CUPS, showed that CUPS can help us understand how programmers interact with code-recommendation systems, revealing inefficiencies and time costs. Our insights reveal how programmers interact with Copilot and motivate new interface designs and metrics.
Hussein Mozannar, Gagan Bansal, Adam Fourney, Eric Horvitz
CHI4
2024 AI-Powered Reminders for Collaborative Tasks: Experiences and Futures
abstract
Email continues to serve as a central medium for managing collaborations. While unstructured email messaging is lightweight and conducive to coordination, it is easy to overlook commitments and requests for collaborations that are embedded in the text of free-flowing communications. Twenty-one years ago, Bellotti et al. proposed TaskMaster with the goal of redesigning the email interface to have explicit task management capabilities. Recently, AI-based task recognition and reminder services have been introduced in major email systems as one approach to managing asynchronous collaborations. While these services have been provided to millions of people around the world, there is little understanding of how people interact with and benefit from them. We explore knowledge workers' experiences with Microsoft's Viva Daily Briefing Email to better understand how AI-powered reminders can support asynchronous collaborations. Through semi-structured interviews and surveys, we shed light on how AI-powered reminders are incorporated into workflows to support asynchronous collaborations. We identify what knowledge workers prefer AI-powered reminders to remind them about and how they would like to interact with these reminders. Using mixed methods and a self-assessment methodology, we investigate the relationship between information workers' work styles and the perceived value of the Viva Daily Briefing Email to identify users who are more likely to benefit from AI-powered reminders for asynchronous collaborations. We conclude by discussing the experiences and futures of AI-powered reminders for collaborative tasks and asynchronous collaborations.
Katelyn Morrison, Shamsi T. Iqbal, Eric Horvitz
Proc. ACM Hum. Comput. Interact.3
2023 Ideal Abstractions for Decision-Focused Learning
abstract
We present a methodology for formulating simplifying abstractions in machine learning systems by identifying and harnessing the utility structure of decisions. Machine learning tasks commonly involve high-dimensional output spaces (e.g., predictions for every pixel in an image or node in a graph), even though a coarser output would often suffice for downstream decision-making (e.g., regions of an image instead of pixels). Developers often hand-engineer abstractions of the output space, but numerous abstractions are possible and it is unclear how the choice of output space for a model impacts its usefulness in downstream decision-making. We propose a method that configures the output space automatically in order to minimize the loss of decision-relevant information. Taking a geometric perspective, we formulate a step of the algorithm as a projection of the probability simplex, termed fold, that minimizes the total loss of decision-related information in the H-entropy sense. Crucially, learning in the abstracted outcome space requires significantly less data, leading to a net improvement in decision quality. We demonstrate the method in two domains: data acquisition for deep neural network training and a closed-loop wildfire management task.
Michael Poli, Stefano Massaroli, Stefano Ermon, Bryan Wilder, Eric Horvitz
AISTATS5
2022 A Search Engine for Discovery of Scientific Challenges and Directions
abstract
Keeping track of scientific challenges, advances and emerging directions is a fundamental part of research. However, researchers face a flood of papers that hinders discovery of important knowledge. In biomedicine, this directly impacts human lives. To address this problem, we present a novel task of extraction and search of scientific challenges and directions, to facilitate rapid knowledge discovery. We construct and release an expert-annotated corpus of texts sampled from full-length papers, labeled with novel semantic categories that generalize across many types of challenges and directions. We focus on a large corpus of interdisciplinary work relating to the COVID-19 pandemic, ranging from biomedicine to areas such as AI and economics. We apply a model trained on our data to identify challenges and directions across the corpus and build a dedicated search engine. In experiments with 19 researchers and clinicians using our system, we outperform a popular scientific search engine in assisting knowledge discovery. Finally, we show that models trained on our resource generalize to the wider biomedical domain and to AI papers, highlighting its broad utility. We make our data, model and search engine publicly available.
Dan Lahav, Jon Saad-Falcon, Bailey Kuehl, Sophie Johnson, Sravanthi Parasa, Noam Shomron, Polo Chau, Diyi Yang, Eric Horvitz, Daniel S. Weld, Tom Hope
AAAI9
2022 Bursting Scientific Filter Bubbles: Boosting Innovation via Novel Author Discovery
abstract
Isolated silos of scientific research and the growing challenge of information overload limit awareness across the literature and hinder innovation. Algorithmic curation and recommendation, which often prioritize relevance, can further reinforce these informational “filter bubbles.” In response, we describe Bridger, a system for facilitating discovery of scholars and their work. We construct a faceted representation of authors with information gleaned from their papers and inferred author personas, and use it to develop an approach that locates commonalities and contrasts between scientists to balance relevance and novelty. In studies with computer science researchers, this approach helps users discover authors considered useful for generating novel research directions. We also demonstrate an approach for displaying information about authors, boosting the ability to understand the work of new, unfamiliar scholars. Our analysis reveals that Bridger connects authors who have different citation profiles and publish in different venues, raising the prospect of bridging diverse scientific communities.
Jason Portenoy, Marissa Radensky, Jevin D. West, Eric Horvitz, Daniel S. Weld, Tom Hope
CHI4
2022 Continual Learning about Objects in the Wild: An Interactive Approach
abstract
We introduce a mixed-reality, interactive approach for continually learning to recognize an open-ended set of objects in a user’s surrounding environment. The proposed approach leverages the multimodal sensing, interaction, and rendering affordances of a mixed-reality headset, and enables users to label nearby objects via speech, gaze, and gestures. Image views of each labeled object are automatically captured from varying viewpoints over time, as the user goes about their everyday tasks. The labels provided by the user can be propagated forward and backwards in time and paired with the collected views to update an object recognition model, in order to continually adapt it to the user’s specific objects and environment. We review key challenges for the proposed interactive continual learning approach, present details of an end-to-end system implementation, and report on results and lessons learned from an initial, exploratory case study using the system.
Dan Bohus, Sean Andrist, Ashley Feniello, Nick Saw, Eric Horvitz
ICMI5
2022 On the Horizon: Interactive and Compositional Deepfakes
abstract
Over a five-year period, computing methods for generating high-fidelity, fictional depictions of people and events moved from exotic demonstrations by computer science research teams into ongoing use as a tool of disinformation. The methods, referred to with the portmanteau of “deepfakes," have been used to create compelling audiovisual content. Here, I share challenges ahead with malevolent uses of two classes of deepfakes that we can expect to come into practice with costly implications for society: interactive and compositional deepfakes. Interactive deepfakes have the capability to impersonate people with realistic interactive behaviors, taking advantage of advances in multimodal interaction. Compositional deepfakes leverage synthetic content in larger disinformation plans that integrate sets of deepfakes over time with observed, expected, and engineered world events to create persuasive synthetic histories. Synthetic histories can be constructed manually but may one day be guided by adversarial generative explanation (AGE) techniques. In the absence of mitigations, interactive and compositional deepfakes threaten to move us closer to a post-epistemic world, where fact cannot be distinguished from fiction. I shall describe interactive and compositional deepfakes and reflect about cautions and potential mitigations to defend against them.
Eric Horvitz
ICMI1
2022 Gender-sensitive word embeddings for healthcare
abstract
OBJECTIVE: To analyze gender bias in clinical trials, to design an algorithm that mitigates the effects of biases of gender representation on natural-language (NLP) systems trained on text drawn from clinical trials, and to evaluate its performance. MATERIALS AND METHODS: We analyze gender bias in clinical trials described by 16 772 PubMed abstracts (2008-2018). We present a method to augment word embeddings, the core building block of NLP-centric representations, by weighting abstracts by the number of women participants in the trial. We evaluate the resulting gender-sensitive embeddings performance on several clinical prediction tasks: comorbidity classification, hospital length of stay prediction, and intensive care unit (ICU) readmission prediction. RESULTS: For female patients, the gender-sensitive model area under the receiver-operator characteristic (AUROC) is 0.86 versus the baseline of 0.81 for comorbidity classification, mean absolute error 4.59 versus the baseline of 4.66 for length of stay prediction, and AUROC 0.69 versus 0.67 for ICU readmission. All results are statistically significant. DISCUSSION: Women have been underrepresented in clinical trials. Thus, using the broad clinical trials literature as training data for statistical language models could result in biased models, with deficits in knowledge about women. The method presented enables gender-sensitive use of publications as training data for word embeddings. In experiments, the gender-sensitive embeddings show better performance than baseline embeddings for the clinical tasks studied. The results highlight opportunities for recognizing and addressing gender and other representational biases in the clinical trials literature. CONCLUSION: Addressing representational biases in data for training NLP embeddings can lead to better results on downstream tasks for underrepresented populations.
Shunit Agmon, Plia Gillis, Eric Horvitz, Kira Radinsky
J. Am. Medical Informatics Assoc.3
2021 Is the Most Accurate AI the Best Teammate? Optimizing AI for Teamwork
abstract
AI practitioners typically strive to develop the most accurate systems, making an implicit assumption that the AI system will function autonomously. However, in practice, AI systems often are used to provide advice to people in domains ranging from criminal justice and finance to healthcare. In such AI-advised decision making, humans and machines form a team, where the human is responsible for making final decisions. But is the most accurate AI the best teammate? We argue "not necessarily" --- predictable performance may be worth a slight sacrifice in AI accuracy. Instead, we argue that AI systems should be trained in a human-centered manner, directly optimized for team performance. We study this proposal for a specific type of human-AI teaming, where the human overseer chooses to either accept the AI recommendation or solve the task themselves. To optimize the team performance for this setting we maximize the team's expected utility, expressed in terms of the quality of the final decision, cost of verifying, and individual accuracies of people and machines. Our experiments with linear and non-linear models on real-world, high-stakes datasets show that the most accuracy AI may not lead to highest team performance and show the benefit of modeling teamwork during training through improvements in expected team utility across datasets, considering parameters such as human skill and the cost of mistakes. We discuss the shortcoming of current optimization approaches beyond well-studied loss functions such as log-loss, and encourage future work on AI optimization problems motivated by human-AI collaboration.
Gagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz, Daniel S. Weld
AAAI4
2021 Understanding Failures of Deep Networks via Robust Feature Extraction
abstract
Traditional evaluation metrics for learned models that report aggregate scores over a test set are insufficient for surfacing important and informative patterns of failure over features and instances. We introduce and study a method aimed at characterizing and explaining failures by identifying visual attributes whose presence or absence results in poor performance. In distinction to previous work that relies upon crowdsourced labels for visual attributes, we leverage the representation of a separate robust model to extract interpretable features and then harness these features to identify failure modes. We further propose a visualization method aimed at enabling humans to understand the meaning encoded in such features and we test the comprehensibility of the features. An evaluation of the methods on the ImageNet dataset demonstrates that: (i) the proposed workflow is effective for discovering important failure modes, (ii) the visualization techniques help humans to understand the extracted features, and (iii) the extracted insights can assist engineers with error analysis and debugging.
Sahil Singla 0002, Besmira Nushi, Shital Shah, Ece Kamar, Eric Horvitz
CVPR5
2021 Exploiting structured data for learning contagious diseases under incomplete testing
abstract
One of the ways that machine learning algorithms can help control the spread of an infectious disease is by building models that predict who is likely to become infected making them good candidates for preemptive interventions. In this work we ask: can we build reliable infection prediction models when the observed data is collected under limited, and biased testing that prioritizes testing symptomatic individuals? Our analysis suggests that when the infection is highly transmissible, incomplete testing might be sufficient to achieve good out-of-sample prediction error. Guided by this insight, we develop an algorithm that predicts infections, and show that it outperforms baselines on simulated data. We apply our model to data from a large hospital to predict Clostridioides difficile infections; a communicable disease that is characterized by both symptomatically infected and asymptomatic (i.e., untested) carriers. Using a proxy instead of the unobserved untested-infected state, we show that our model outperforms benchmarks in predicting infections.
Maggie Makar, Lauren West, David Hooper, Eric Horvitz, Erica Shenoy, John V. Guttag
ICML4
2021 Domain-Specific Pretraining for Vertical Search: Case Study on Biomedical Literature
abstract
Information overload is a prevalent challenge in many high-value domains. A prominent case in point is the explosion of the biomedical literature on COVID-19, which swelled to hundreds of thousands of papers in a matter of months. In general, biomedical literature expands by two papers every minute, totalling over a million new papers every year. Search in the biomedical realm, and many other vertical domains is challenging due to the scarcity of direct supervision from click logs. Self-supervised learning has emerged as a promising direction to overcome the annotation bottleneck. We propose a general approach for vertical search based on domain-specific pretraining and present a case study for the biomedical domain. Despite being substantially simpler and not using any relevance labels for training or development, our method performs comparably or better than the best systems in the official TREC-COVID evaluation, a COVID-related biomedical search competition. Using distributed computing in modern cloud infrastructure, our system can scale to tens of millions of articles on PubMed and has been deployed as Microsoft Biomedical Search, a new search experience for biomedical literature: https://aka.ms/biomedsearch.
Yu Wang 0009, Jinchao Li, Tristan Naumann, Chenyan Xiong, Hao Cheng 0002, Robert Tinn, Cliff Wong, Naoto Usuyama, Richard Rogahn, Zhihong Shen, Eric Horvitz, Paul N. Bennett, Jianfeng Gao 0001, Hoifung Poon
KDD12
2021 AMP: authentication of media via provenance
abstract
Advances in graphics and machine learning have led to the general availability of easy-to-use tools for modifying and synthesizing media. The proliferation of these tools threatens to cast doubt on the veracity of all media. One approach to thwarting the flow of fake media is to detect modified or synthesized media through machine learning methods. While detection may help in the short term, we believe that it is destined to fail as the quality of fake media generation continues to improve. Soon, neither humans nor algorithms will be able to reliably distinguish fake versus real content. Thus, pipelines for assuring the source and integrity of media will be required---and increasingly relied upon. We present AMP, a system that ensures the authentication of media via certifying provenance. AMP creates one or more publisher-signed manifests for a media instance uploaded by a content provider. These manifests are stored in a database allowing fast lookup from applications such as browsers. For reference, the manifests are also registered and signed by a permissioned ledger, implemented using the Confidential Consortium Framework (CCF). CCF employs both software and hardware techniques to ensure the integrity and transparency of all registered manifests. AMP, through its use of CCF, enables a consortium of media providers to govern the service while making all its operations auditable. The authenticity of the media can be communicated to the user via visual elements in the browser, indicating that an AMP manifest has been successfully located and verified.
Paul England, Henrique S. Malvar, Eric Horvitz, Jack W. Stokes, Cédric Fournet, Rebecca Burke-Aguero, Amaury Chamayou, Sylvan Clebsch, Manuel Costa, John Deutscher, Shabnam Erfani, Matt Gaylor, Andrew Jenks, Kevin Kane, Elissa M. Redmiles, Alex Shamis, Isha Sharma, John C. Simmons, Sam Wenker, Anika Zaman
MMSys3
2021 Extracting a Knowledge Base of Mechanisms from COVID-19 Papers
abstract
Tom Hope, Aida Amini, David Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel Weld, Roy Schwartz, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tom Hope, Aida Amini, Dave Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel S. Weld, Roy Schwartz 0001, Hannaneh Hajishirzi
NAACL-HLT6
2021 Population-Scale Study of Human Needs During the COVID-19 Pandemic: Analysis and Implications
abstract
Most work to date on mitigating the COVID-19 pandemic is focused urgently on biomedicine and epidemiology. Yet, pandemic-related policy decisions cannot be made on health information alone. Decisions need to consider the broader impacts on people and their needs. Quantifying human needs across the population is challenging as it requires high geo-temporal granularity, high coverage across the population, and appropriate adjustment for seasonal and other external effects. Here, we propose a computational methodology, building on Maslow's hierarchy of needs, that can capture a holistic view of relative changes in needs following the pandemic through a difference-in-differences approach that corrects for seasonality and volume variations. We apply this approach to characterize changes in human needs across physiological, socioeconomic, and psychological realms in the US, based on more than 35 billion search interactions spanning over 36,000 ZIP codes over a period of 14 months. The analyses reveal that the expression of basic human needs has increased exponentially while higher-level aspirations declined during the pandemic in comparison to the pre-pandemic period. In exploring the timing and variations in statewide policies, we find that the durations of shelter-in-place mandates have influenced social and emotional needs significantly. We demonstrate that potential barriers to addressing critical needs, such as support for unemployment and domestic violence, can be identified through web search interactions. Our approach and results suggest that population-scale monitoring of shifts in human needs can inform policies and recovery efforts for current and anticipated needs.
Jina Suh, Eric Horvitz, Ryen W. White, Tim Althoff
WSDM2
2021 On biases of attention in scientific discovery
abstract
SUMMARY: How do nuances of scientists' attention influence what they discover? We pursue an understanding of the influences of patterns of attention on discovery with a case study about confirmations of protein-protein interactions over time. We find that modeling and accounting for attention can help us to recognize and interpret biases in large-scale and widely used databases of confirmed interactions and to better understand missing data and unknowns. Additionally, we present an analysis of how awareness of patterns of attention and use of debiasing techniques can foster earlier discoveries. AVAILABILITY AND IMPLEMENTATION: The data is freely available at https://github.com/urielsinger/PPI-unbias.
Uriel Singer, Kira Radinsky, Eric Horvitz
Bioinform.3
2021 Formation of Social Ties Influences Food Choice: A Campus-wide Longitudinal Study
abstract
Nutrition is a key determinant of long-term health, and social influence has long been theorized to be a key determinant of nutrition. It has been difficult to quantify the postulated role of social influence on nutrition using traditional methods such as surveys, due to the typically small scale and short duration of studies. To overcome these limitations, we leverage a novel source of data: logs of 38 million food purchases made over an 8-year period on the Ecole Polytechnique Federale de Lausanne (EPFL) university campus, linked to anonymized individuals via the smartcards used to make on-campus purchases. In a longitudinal observational study, we ask: How is a person's food choice affected by eating with someone else whose own food choice is healthy vs. unhealthy? To estimate causal effects from the passively observed log data, we control confounds in a matched quasi-experimental design: we identify focal users who at first do not have any regular eating partners but then start eating with a fixed partner regularly, and we match focal users into comparison pairs such that paired users are nearly identical with respect to covariates measured before acquiring the partner, where the two focal users' new eating partners diverge in the healthiness of their respective food choice. A difference-in-differences analysis of the paired data yields clear evidence of social influence: focal users acquiring a healthy-eating partner change their habits significantly more toward healthy foods than focal users acquiring an unhealthy-eating partner. We further identify foods whose purchase frequency is impacted significantly by the eating partner's healthiness of food choice. Beyond the main results, the work demonstrates the utility of passively sensed food purchase logs for deriving insights, with the potential of informing the design of public health interventions and food offerings, especially on university campuses.
Kristina Gligoric, Ryen W. White, Emre Kiciman, Eric Horvitz, Arnaud Chiolero, Robert West 0001
Proc. ACM Hum. Comput. Interact.4
2020 Metareasoning in Modular Software Systems: On-the-Fly Configuration Using Reinforcement Learning with Rich Contextual Representations
Aditya Modi 0002, Debadeepta Dey, Alekh Agarwal, Adith Swaminathan, Besmira Nushi, Sean Andrist, Eric Horvitz
AAAI7
2020 Recollection versus Imagination: Exploring Human Memory and Cognition via Neural Language Models
abstract
We investigate the use of NLP as a measure of the cognitive processes involved in storytelling, contrasting imagination and recollection of events. To facilitate this, we collect and release Hippocorpus, a dataset of 7,000 stories about imagined and recalled events. We introduce a measure of narrative flow and use this to examine the narratives for imagined and recalled events. Additionally, we measure the differential recruitment of knowledge attributed to semantic memory versus episodic memory (Tulving, 1972) for imagined and recalled storytelling by comparing the frequency of descriptions of general commonsense events with more specific realis events. Our analyses show that imagined stories have a substantially more linear narrative flow, compared to recalled stories in which adjacent sentences are more disconnected. In addition, while recalled stories rely more on autobiographical events based on episodic memory, imagined stories express more commonsense knowledge based on semantic memory. Finally, our measures reveal the effect of narrativization of memories in stories (e.g., stories about frequently recalled memories flow more linearly; Bartlett, 1932). Our findings highlight the potential of using NLP tools to study the traces of human cognition in language.
Maarten Sap, Eric Horvitz, Yejin Choi 0001, Noah A. Smith, James W. Pennebaker
ACL2
2020 SQuINTing at VQA Models: Introspecting VQA Models With Sub-Questions
abstract
Existing VQA datasets contain questions with varying levels of complexity. While the majority of questions in these datasets require perception for recognizing existence, properties, and spatial relationships of entities, a significant portion of questions pose challenges that correspond to reasoning tasks - tasks that can only be answered through a synthesis of perception and knowledge about the world, logic and / or reasoning. Analyzing performance across this distinction allows us to notice when existing VQA models have consistency issues - they answer the reasoning questions correctly but fail on associated low-level perception questions. For example, in Figure 1, models answer the complex reasoning question “Is the banana ripe enough to eat?” correctly, but fail on the associated perception question “Are the bananas mostly green or yellow?” indicating that the model likely answered the reasoning question correctly but for the wrong reason. We quantify the extent to which this phenomenon occurs by creating a new Reasoning split of the VQA dataset and collecting VQAintrospect, a new dataset1 which currently consists of 200K new perception questions which serve as sub questions corresponding to the set of perceptual tasks needed to effectively answer the complex reasoning questions in the Reasoning split. Our evaluation shows that state-of-the-art VQA models have comparable performance in answering perception and reasoning questions, but suffer from consistency problems. To address this shortcoming, we propose an approach called Sub-Question Importance-aware Network Tuning (SQuINT), which encourages the model to attend to the same parts of the image when answering the reasoning question and the perception sub question. We show that SQuINT improves model consistency by ~7%, also marginally improving performance on the Reasoning questions in VQA, while also displaying better attention maps.
Ramprasaath R. Selvaraju, Purva Tendulkar, Devi Parikh, Eric Horvitz, Marco Túlio Ribeiro, Besmira Nushi, Ece Kamar
CVPR4
2020 Learning to Complement Humans
abstract
A rising vision for AI in the open world centers on the development of systems that can complement humans for perceptual, diagnostic, and reasoning tasks. To date, systems aimed at complementing the skills of people have employed models trained to be as accurate as possible in isolation. We demonstrate how an end-to-end learning strategy can be harnessed to optimize the combined performance of human-machine teams by considering the distinct abilities of people and machines. The goal is to focus machine learning on problem instances that are difficult for humans, while recognizing instances that are difficult for the machine and seeking human input on them. We demonstrate in two real-world domains (scientific discovery and medical diagnosis) that human-machine teams built via these methods outperform the individual performance of machines and people. We then analyze conditions under which this complementarity is strongest, and which training methods amplify it. Taken together, our work provides the first systematic investigation of how machine learning systems can be trained to complement human reasoning.
Bryan Wilder, Eric Horvitz, Ece Kamar
IJCAI2
2020 An Empirical Analysis of Backward Compatibility in Machine Learning Systems
abstract
In many applications of machine learning (ML), updates are performed with the goal of enhancing model performance. However, current practices for updating models rely solely on isolated, aggregate performance analyses, overlooking important dependencies, expectations, and needs in real-world deployments. We consider how updates, intended to improve ML models, can introduce new errors that can significantly affect downstream systems and users. For example, updates in models used in cloud-based classification services, such as image recognition, can cause unexpected erroneous behavior in systems that make calls to the services. Prior work has shown the importance of "backward compatibility" for maintaining human trust. We study challenges with backward compatibility across different ML architectures and datasets, focusing on common settings including data shifts with structured noise and ML employed in inferential pipelines. Our results show that (i) compatibility issues arise even without data shift due to optimization stochasticity, (ii) training on large-scale noisy datasets often results in significant decreases in backward compatibility even when model accuracy increases, and (iii) distributions of incompatible points align with noise bias, motivating the need for compatibility aware de-noising and robustness methods.
Megha Srivastava, Besmira Nushi, Ece Kamar, Shital Shah, Eric Horvitz
KDD5
2020 Characterizing Search-Engine Traffic to Internet Research Agency Web Properties
abstract
The Russia-based Internet Research Agency (IRA) carried out a broad information campaign in the U.S. before and after the 2016 presidential election. The organization created an expansive set of internet properties: web domains, Facebook pages, and Twitter bots, which received traffic via purchased Facebook ads, tweets, and search engines indexing their domains. In this paper, we focus on IRA activities that received exposure through search engines, by joining data from Facebook and Twitter with logs from the Internet Explorer 11 and Edge browsers and the Bing.com search engine.
Alexander Spangher, Gireeja Ranade, Besmira Nushi, Adam Fourney, Eric Horvitz
WWW5
2020 Blind Spot Detection for Safe Sim-to-Real Transfer
abstract
Agents trained in simulation may make errors when performing actions in the real world due to mismatches between training and execution environments. These mistakes can be dangerous and difficult for the agent to discover because the agent is unable to predict them a priori. In this work, we propose the use of oracle feedback to learn a predictive model of these blind spots in order to reduce costly errors in real-world applications. We focus on blind spots in reinforcement learning (RL) that occur due to incomplete state representation: when the agent lacks necessary features to represent the true state of the world, and thus cannot distinguish between numerous states. We formalize the problem of discovering blind spots in RL as a noisy supervised learning problem with class imbalance. Our system learns models for predicting blind spots within unseen regions of the state space by combining techniques for label aggregation, calibration, and supervised learning. These models take into consideration noise emerging from different forms of oracle feedback, including demonstrations and corrections. We evaluate our approach across two domains and demonstrate that it achieves higher predictive performance than baseline methods, and also that the learned model can be used to selectively query an oracle at execution time to prevent errors. We also empirically analyze the biases of various feedback types and how these biases influence the discovery of blind spots. Further, we include analyses of our approach that incorporate relaxed initial optimality assumptions. (Interestingly, relaxing the assumptions of an optimal oracle and an optimal simulator policy helped our models to perform better.) We also propose extensions to our method that are intended to improve performance when using corrections and demonstrations data.
Ramya Ramakrishnan, Ece Kamar, Debadeepta Dey, Eric Horvitz, Julie A. Shah
J. Artif. Intell. Res.4
2020 Predicting severe clinical events by learning about life-saving actions and outcomes using distant supervision
Dae Hyun Lee, Meliha Yetisgen, Lucy Vanderwende, Eric Horvitz
J. Biomed. Informatics4
2020 Configuring Audiences: A Case Study of Email Communication
abstract
When people communicate with each other, their choice of what to say is tied to their perceptions of the audience. For many communication channels, people have some ability to explicitly specify their audience members and the different roles they can play. While existing accounts of communication behavior have largely focused on how people tailor the content of their messages, we focus on the configuring of the audience as a complementary family of decisions in communication. We formulate a general description of audience configuration choices, highlighting key aspects of the audience that people could configure to reflect a range of communicative goals. We then illustrate these ideas via a case study of email usage-a realistic domain where audience configuration choices are particularly fine-grained and explicit in how email senders fill the To and Cc address fields. In a large collection of enterprise emails, we explore how people configure their audiences, finding salient patterns relating a sender's choice of configuration to the types of participants in the email exchange, the content of the message, and the nature of the subsequent interactions. Our formulation and findings show how analyzing audience configurations can enrich and extend existing accounts of communication behavior, and frame research directions on audience configuration decisions in communication and collaboration.
Justine Zhang, James W. Pennebaker, Susan T. Dumais, Eric Horvitz
Proc. ACM Hum. Comput. Interact.4
2019 Reverse-Engineering Satire, or "Paper on Computational Humor Accepted despite Making Serious Advances"
abstract
Humor is an essential human trait. Efforts to understand humor have called out links between humor and the foundations of cognition, as well as the importance of humor in social engagement. As such, it is a promising and important subject of study, with relevance for artificial intelligence and human– computer interaction. Previous computational work on humor has mostly operated at a coarse level of granularity, e.g., predicting whether an entire sentence, paragraph, document, etc., is humorous. As a step toward deep understanding of humor, we seek fine-grained models of attributes that make a given text humorous. Starting from the observation that satirical news headlines tend to resemble serious news headlines, we build and analyze a corpus of satirical headlines paired with nearly identical but serious headlines. The corpus is constructed via Unfun.me, an online game that incentivizes players to make minimal edits to satirical headlines with the goal of making other players believe the results are serious headlines. The edit operations used to successfully remove humor pinpoint the words and concepts that play a key role in making the original, satirical headline funny. Our analysis reveals that the humor tends to reside toward the end of headlines, and primarily in noun phrases, and that most satirical headlines follow a certain logical pattern, which we term false analogy. Overall, this paper deepens our understanding of the syntactic and semantic structure of satirical news headlines and provides insights for building humor-producing systems.
Robert West 0001, Eric Horvitz
AAAI2
2019 Updates in Human-AI Teams: Understanding and Addressing the Performance/Compatibility Tradeoff
abstract
AI systems are being deployed to support human decision making in high-stakes domains such as healthcare and criminal justice. In many cases, the human and AI form a team, in which the human makes decisions after reviewing the AI’s inferences. A successful partnership requires that the human develops insights into the performance of the AI system, including its failures. We study the influence of updates to an AI system in this setting. While updates can increase the AI’s predictive performance, they may also lead to behavioral changes that are at odds with the user’s prior experiences and confidence in the AI’s inferences. We show that updates that increase AI performance may actually hurt team performance. We introduce the notion of the compatibility of an AI update with prior user experience and present methods for studying the role of compatibility in human-AI teams. Empirical results on three high-stakes classification tasks show that current machine learning algorithms do not produce compatible updates. We propose a re-training objective to improve the compatibility of an update by penalizing new errors. The objective offers full leverage of the performance/compatibility tradeoff across different datasets, enabling more compatible yet accurate updates.
Gagan Bansal, Besmira Nushi, Ece Kamar, Daniel S. Weld, Walter S. Lasecki, Eric Horvitz
AAAI6
2019 Traffic Updates: Saying a Lot While Revealing a Little
abstract
Taking speed reports from vehicles is a proven, inexpensive way to infer traffic conditions. However, due to concerns about privacy and bandwidth, not every vehicle occupant may want to transmit data about their location and speed in real time. We show how to drastically reduce the number of transmissions in two ways, both based on a Markov random field for modeling traffic speed and flow. First, we show that a only a small number of vehicles need to report from each location. We give a simple, probabilistic method that lets a group of vehicles decide on which subset will transmit a report, preserving privacy by coordinating without any communication. The second approach computes the potential value of any location’s speed report, emphasizing those reports that will most affect the overall speed inferences, and omitting those that contribute little value. Both methods significantly reduce the amount of communication necessary for accurate speed inferences on a road network.
John Krumm, Eric Horvitz
AAAI2
2019 Separating Wheat from Chaff: Joining Biomedical Knowledge and Patient Data for Repurposing Medications
abstract
We present a system that jointly harnesses large-scale electronic health records data and a concept graph mined from the medical literature to guide drug repurposing—the process of applying known drugs in new ways to treat diseases. Our study is unique in methods and scope, per the scale of the concept graph and the quantity of data. We harness 10 years of nation-wide medical records of more than 1.5 million people and extract medical knowledge from all of PubMed, the world’s largest corpus of online biomedical literature. We employ links on the concept graph to provide causal signals to prioritize candidate influences between medications and target diseases. We show results of the system on studies of drug repurposing for hypertension and diabetes. In both cases, we present drug families identified by the algorithm which were previously unknown. We verify the results via clinical expert opinion and by prospective clinical trials on hypertension.
Galia Nordon, Gideon Koren, Varda Shalev, Eric Horvitz, Kira Radinsky
AAAI4
2019 Overcoming Blind Spots in the Real World: Leveraging Complementary Abilities for Joint Execution
abstract
Simulators are being increasingly used to train agents before deploying them in real-world environments. While training in simulation provides a cost-effective way to learn, poorly modeled aspects of the simulator can lead to costly mistakes, or blind spots. While humans can help guide an agent towards identifying these error regions, humans themselves have blind spots and noise in execution. We study how learning about blind spots of both can be used to manage hand-off decisions when humans and agents jointly act in the real-world in which neither of them are trained or evaluated fully. The formulation assumes that agent blind spots result from representational limitations in the simulation world, which leads the agent to ignore important features that are relevant for acting in the open world. Our approach for blind spot discovery combines experiences collected in simulation with limited human demonstrations. The first step applies imitation learning to demonstration data to identify important features that the human is using but that the agent is missing. The second step uses noisy labels extracted from action mismatches between the agent and the human across simulation and demonstration data to train blind spot models. We show through experiments on two domains that our approach is able to learn a succinct representation that accurately captures blind spot regions and avoids dangerous errors in the real world through transfer of control between the agent and the human.
Ramya Ramakrishnan, Ece Kamar, Besmira Nushi, Debadeepta Dey, Julie A. Shah, Eric Horvitz
AAAI6
2019 Guidelines for Human-AI Interaction
abstract
Advances in artificial intelligence (AI) frame opportunities and challenges for user interface design. Principles for human-AI interaction have been discussed in the human-computer interaction community for over two decades, but more study and innovation are needed in light of advances in AI and the growing uses of AI technologies in human-facing applications. We propose 18 generally applicable design guidelines for human-AI interaction. These guidelines are validated through multiple rounds of evaluation including a user study with 49 design practitioners who tested the guidelines against 20 popular AI-infused products. The results verify the relevance of the guidelines over a spectrum of interaction scenarios and reveal gaps in our knowledge, highlighting opportunities for further research. Based on the evaluations, we believe the set of design guidelines can serve as a resource to practitioners working on the design of applications and features that harness AI technologies, and to researchers interested in the further development of human-AI interaction design principles.
Saleema Amershi, Daniel S. Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi T. Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, Eric Horvitz
CHI13
2019 Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance
abstract
Decisions made by human-AI teams (e.g., AI-advised humans) are increasingly common in high-stakes domains such as healthcare, criminal justice, and finance. Achieving high team performance depends on more than just the accuracy of the AI system: Since the human and the AI may have different expertise, the highest team performance is often reached when they both know how and when to complement one another. We focus on a factor that is crucial to supporting such complementary: the human’s mental model of the AI capabilities, specifically the AI system’s error boundary (i.e. knowing “When does the AI err?”). Awareness of this lets the human decide when to accept or override the AI’s recommendation. We highlight two key properties of an AI’s error boundary, parsimony and stochasticity, and a property of the task, dimensionality. We show experimentally how these properties affect humans’ mental models of AI capabilities and the resulting team performance. We connect our evaluations to related work and propose goals, beyond accuracy, that merit consideration during model selection and optimization to improve overall human-AI team performance.
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S. Lasecki, Daniel S. Weld, Eric Horvitz
HCOMP6
2019 Bias Correction of Learned Generative Models using Likelihood-Free Importance Weighting
abstract
A learned generative model often produces biased statistics relative to the underlying data distribution. A standard technique to correct this bias is importance sampling, where samples from the model are weighted by the likelihood ratio under model and true distributions. When the likelihood ratio is unknown, it can be estimated by training a probabilistic classifier to distinguish samples from the two distributions. We employ this likelihood-free importance weighting method to correct for the bias in generative models. We find that this technique consistently improves standard goodness-of-fit metrics for evaluating the sample quality of state-of-the-art deep generative models, suggesting reduced bias. Finally, we demonstrate its utility on representative applications in a) data augmentation for classification using generative adversarial networks, and b) model-based policy evaluation using off-policy data.
Aditya Grover, Jiaming Song, Ashish Kapoor, Kenneth Tran, Alekh Agarwal, Eric Horvitz, Stefano Ermon
NeurIPS6
2019 Efficient Forward Architecture Search
abstract
We propose a neural architecture search (NAS) algorithm, Petridish, to iteratively add shortcut connections to existing network layers. The added shortcut connections effectively perform gradient boosting on the augmented layers. The proposed algorithm is motivated by the feature selection algorithm forward stage-wise linear regression, since we consider NAS as a generalization of feature selection for regression, where NAS selects shortcuts among layers instead of selecting features. In order to reduce the number of trials of possible connection combinations, we train jointly all possible connections at each stage of growth while leveraging feature selection techniques to choose a subset of them. We experimentally show this process to be an efficient forward architecture search algorithm that can find competitive models using few GPU days in both the search space of repeatable network modules (cell-search) and the space of general networks (macro-search). Petridish is particularly well-suited for warm-starting from existing models crucial for lifelong-learning scenarios.
Hanzhang Hu, John Langford 0001, Rich Caruana, Saurajit Mukherjee, Eric Horvitz, Debadeepta Dey
NeurIPS5
2019 Staying up to Date with Online Content Changes Using Reinforcement Learning for Scheduling
abstract
From traditional Web search engines to virtual assistants and Web accelerators, services that rely on online information need to continually keep track of remote content changes by explicitly requesting content updates from remote sources (e.g., web pages). We propose a novel optimization objective for this setting that has several practically desirable properties, and efficient algorithms for it with optimality guarantees even in the face of mixed content change observability and initially unknown change model parameters. Experiments on 18.5M URLs crawled daily for 14 weeks show significant advantages of this approach over prior art.
Andrey Kolobov, Yuval Peres, Eric Horvitz
NeurIPS4
2019 Optimal Freshness Crawl Under Politeness Constraints
abstract
A Web crawler is an essential part of a search engine that procures information subsequently served by the search engine to its users. As the Web is becoming increasingly more dynamic, in addition to discovering new web pages a crawler needs to keep revisiting those already in the search engine's index, in order to keep the index fresh by picking up the pages' changed content. Determining how often to recrawl pages requires making tradeoffs based on the pages' relative importance and change rates, subject to multiple resource constraints - the limited daily budget of crawl requests on the search engine's end and politeness constraints restricting the rate at which pages can be requested from a given host. In this paper, we introduce PoliteBinaryLambdaCrawl, the first optimal algorithm for freshness crawl scheduling in the presence of politeness constraints as well as non-uniform page importance scores and the crawler's own crawl request limit. We also propose an approximation for it, stating its theoretical optimality conditions and in the process discovering a connection to an approach previously thought of as a mere heuristic for freshness crawl scheduling. We explore the relative performance of PoliteBinaryLambdaCrawl and other methods for handling politeness constraints on a dataset collected by crawling over 18.5M URLs daily over 14 weeks.
Andrey Kolobov, Yuval Peres, Eyal Lubetzky, Eric Horvitz
SIGIR4
2018 Optimizing Interventions via Offline Policy Evaluation: Studies in Citizen Science
abstract
Volunteers who help with online crowdsourcing such as citizen science tasks typically make only a few contributions before exiting. We propose a computational approach for increasing users' engagement in such settings that is based on optimizing policies for displaying motivational messages to users. The approach, which we refer to as Trajectory Corrected Intervention (TCI), reasons about the tradeoff between the long-term influence of engagement messages on participants' contributions and the potential risk of disrupting their current work. We combine model-based reinforcement learning with off-line policy evaluation to generate intervention policies, without relying on a fixed representation of the domain. TCI works iteratively to learn the best representation from a set of random intervention trials and to generate candidate intervention policies. It is able to refine selected policies off-line by exploiting the fact that users can only be interrupted once per session.We implemented TCI in the wild with Galaxy Zoo, one of the largest citizen science platforms on the web. We found that TCI was able to outperform the state-of-the-art intervention policy for this domain, and significantly increased the contributions of thousands of users. This work demonstrates the benefit of combining traditional AI planning with off-line policy methods to generate intelligent intervention strategies.
Avi Segal, Kobi Gal, Ece Kamar, Eric Horvitz, Grant Miller
AAAI4
2018 Learning about Life-Saving Interventions to Predict the Risk of Acute Organ Failures
Dae Hyun Lee, Meliha Yetisgen, Eric Horvitz
AMIA3
2018 On the value of spatiotemporal information: principles and scenarios
abstract
Location data from mobile devices is a sensitive yet valuable commodity for location-based services and advertising. We investigate the intrinsic value of location data in the context of strong privacy, where location information is only available from end users via purchase. We present an algorithm to compute the expected value of location data from a user, without access to the specific coordinates of the location data point. We use decision-theoretic techniques to provide a principled way for a potential buyer to make purchasing decisions about private user location data. We illustrate our approach in two scenarios: the delivery of targeted ads specific to a user's home location and the estimation of traffic speed. In both cases, the methodology leads to quantifiably better purchasing decisions than competing methods.
Heba Aly 0001, John Krumm, Gireeja Ranade, Eric Horvitz
SIGSPATIAL/GIS4
2018 Towards Accountable AI: Hybrid Human-Machine Analyses for Characterizing System Failure
abstract
As machine learning systems move from computer-science laboratories into the open world, their accountability becomes a high priority problem. Accountability requires deep understanding of system behavior and its failures. Current evaluation methods such as single-score error metrics and confusion matrices provide aggregate views of system performance that hide important shortcomings. Understanding details about failures is important for identifying pathways for refinement, communicating the reliability of systems in different settings, and for specifying appropriate human oversight and engagement. Characterization of failures and shortcomings is particularly complex for systems composed of multiple machine learned components. For such systems, existing evaluation methods have limited expressiveness in describing and explaining the relationship among input content, the internal states of system components, and final output quality. We present Pandora, a set of hybrid human-machine methods and tools for describing and explaining system failures. Pandora leverages both human and system-generated observations to summarize conditions of system malfunction with respect to the input content and system architecture. We share results of a case study with a machine learning pipeline for image captioning that show how detailed performance views can be beneficial for analysis and debugging.
Besmira Nushi, Ece Kamar, Eric Horvitz
HCOMP3
2017 Long-Term Trends in the Public Perception of Artificial Intelligence
abstract
Analyses of text corpora over time can reveal trends in beliefs, interest, and sentiment about a topic. We focus on views expressed about artificial intelligence (AI) in the New York Times over a 30-year period. General interest, awareness, and discussion about AI has waxed and waned since the field was founded in 1956. We present a set of measures that captures levels of engagement, measures of pessimism and optimism, the prevalence of specific hopes and concerns, and topics that are linked to discussions about AI over decades. We find that discussion of AI has increased sharply since 2009, and that these discussions have been consistently more optimistic than pessimistic. However, when we examine specific concerns, we find that worries of loss of control of AI, ethical concerns for AI, and the negative impact of AI on work have grown in recent years. We also find that hopes for AI in healthcare and education have increased over time.
Ethan Fast, Eric Horvitz
AAAI2
2017 Risk-Aware Planning: Methods and Case Study for Safer Driving Routes
John Krumm, Eric Horvitz
AAAI2
2017 Identifying Unknown Unknowns in the Open World: Representations and Policies for Guided Exploration
abstract
Predictive models deployed in the real world may assign incorrect labels to instances with high confidence. Such errors or unknown unknowns are rooted in model incompleteness, and typically arise because of the mismatch between training data and the cases encountered at test time. As the models are blind to such errors, input from an oracle is needed to identify these failures. In this paper, we formulate and address the problem of informed discovery of unknown unknowns of any given predictive model where unknown unknowns occur due to systematic biases in the training data.We propose a model-agnostic methodology which uses feedback from an oracle to both identify unknown unknowns and to intelligently guide the discovery. We employ a two-phase approach which first organizes the data into multiple partitions based on the feature similarity of instances and the confidence scores assigned by the predictive model, and then utilizes an explore-exploit strategy for discovering unknown unknowns across these partitions. We demonstrate the efficacy of our framework by varying the underlying causes of unknown unknowns across various applications. To the best of our knowledge, this paper presents the first algorithmic approach to the problem of discovering unknown unknowns of predictive models.
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, Eric Horvitz
AAAI4
2017 Predicting Mortality of Intensive Care Patients via Learning about Hazard
abstract
Patients in intensive care units (ICU) are acutely ill and have the highest mortality rates for hospitalized patients. Predictive models and planning system could forecast and guide interventions to prevent the hazardous deterioration of patients’ physiologies, thereby giving the opportunity of employing machine learning and inference to assist with the care of ICU patients. We report on the construction of a prediction pipeline that estimates the probability of death by inferring rates of hazard over time, based on patients’ physiological measurements. The inferred model provided the contribution of each variable and information about the influence of sets of observations on the overall risks and expected trajectories of patients.
Dae Hyun Lee, Eric Horvitz
AAAI2
2017 On Human Intellect and Machine Failures: Troubleshooting Integrative Machine Learning Systems
abstract
We study the problem of troubleshooting machine learning systems that rely on analytical pipelines of distinct components. Understanding and fixing errors that arise in such integrative systems is difficult as failures can occur at multiple points in the execution workflow. Moreover, errors can propagate, become amplified or be suppressed, making blame assignment difficult. We propose a human-in-the-loop methodology which leverages human intellect for troubleshooting system failures. The approach simulates potential component fixes through human computation tasks and measures the expected improvements in the holistic behavior of the system. The method provides guidance to designers about how they can best improve the system. We demonstrate the effectiveness of the approach on an automated image captioning system that has been pressed into real-world use.
Besmira Nushi, Ece Kamar, Eric Horvitz, Donald Kossmann
AAAI3
2017 Geographic and Temporal Trends in Fake News Consumption During the 2016 US Presidential Election
abstract
We present an analysis of traffic to websites known for publishing fake news in the months preceding the 2016 US presidential election. The study is based on the combined instrumentation data from two popular desktop web browsers: Internet Explorer 11 and Edge. We find that social media was the primary outlet for the circulation of fake news stories and that aggregate voting patterns were strongly correlated with the average daily fraction of users visiting websites serving fake news. This correlation was observed both at the state level and at the county level, and remained stable throughout the main election season. We propose a simple model based on homophily in social networks to explain the linear association. Finally, we highlight examples of different types of fake news stories: while certain stories continue to circulate in the population, others are short-lived and die out in a few days.
Adam Fourney, Miklós Z. Rácz, Gireeja Ranade, Markus Mobius, Eric Horvitz
CIKM5
2017 Filling the Blanks (hint: plural noun) for Mad Libs Humor
abstract
Computerized generation of humor is a notoriously difficult AI problem.We develop an algorithm called Libitum that helps humans generate humor in a Mad Lib R , which is a popular fill-in-the-blank game.The algorithm is based on a machine learned classifier that determines whether a potential fill-in word is funny in the context of the Mad Lib story.We use Amazon Mechanical Turk to create ground truth data and to judge humor for our classifier to mimic, and we make this data freely available.Our testing shows that Libitum successfully aids humans in filling in Mad Libs that are usually judged funnier than those filled in by humans with no computerized help.We go on to analyze why some words are better than others at making a Mad Lib funny.
Nabil Hossain, John Krumm, Lucy Vanderwende, Eric Horvitz, Henry A. Kautz
EMNLP4
2017 Wi-Fly: Widespread Opportunistic Connectivity via Commercial Air Transport
abstract
More than half of the world's population face barriers in accessing the Internet. A recent ITU study estimates that 2.6 billion people cannot afford connectivity and that 3.8 billion do not have access. Recent proposals for providing low-cost connectivity include fielding of drones and long-lasting balloons in the stratosphere. We propose a more economical alternative, which we refer to as Wi-Fly, that leverages existing commercial planes to provide Internet connectivity to remote regions. In Wi-Fly we enable communication between a lightweight Wi-Fi device on commercial planes and ground stations, resulting in connectivity in regions that do not otherwise have low-cost Internet connectivity. Wi-Fly leverages existing ADS-B signals from planes as a control channel to ensure that there is a strong link from the plane to the ground, and that the stations intelligently wake up and associate to the appropriate AP. For our experimentation, we have customized two airplanes to conduct measurements. Through empirical experiments with test flights and simulations, we show that Wi-Fly and its extensions have the potential to provide connectivity to the most remote regions of the world at a significantly lower cost than existing alternatives.
Talal Ahmad, Ranveer Chandra, Ashish Kapoor, Eric Horvitz
HotNets5
2017 Estimating Accuracy from Unlabeled Data: A Probabilistic Logic Approach
abstract
We propose an efficient method to estimate the accuracy of classifiers using only unlabeled data. We consider a setting with multiple classification problems where the target classes may be tied together through logical constraints. For example, a set of classes may be mutually exclusive, meaning that a data instance can belong to at most one of them. The proposed method is based on the intuition that: (i) when classifiers agree, they are more likely to be correct, and (ii) when the classifiers make a prediction that violates the constraints, at least one classifier must be making an error. Experiments on four real-world data sets produce accuracy estimates within a few percent of the true accuracy, using solely unlabeled data. Our models also outperform existing state-of-the-art solutions in both estimating accuracies, and combining multiple classifier outputs. The results emphasize the utility of logical constraints in estimating accuracy, thus validating our intuition.
Emmanouil A. Platanios, Hoifung Poon, Tom M. Mitchell, Eric Horvitz
NIPS4
2017 Harnessing the Web for Population-Scale Physiological Sensing: A Case Study of Sleep and Performance
abstract
Human cognitive performance is critical to productivity, learning, and accident avoidance. Cognitive performance varies throughout each day and is in part driven by intrinsic, near 24-hour circadian rhythms. Prior research on the impact of sleep and circadian rhythms on cognitive performance has typically been restricted to small-scale laboratory-based studies that do not capture the variability of real-world conditions, such as environmental factors, motivation, and sleep patterns in real-world settings. Given these limitations, leading sleep researchers have called for larger in situ monitoring of sleep and performance. We present the largest study to date on the impact of objectively measured real-world sleep on performance enabled through a reframing of everyday interactions with a web search engine as a series of performance tasks. Our analysis includes 3 million nights of sleep and 75 million interaction tasks. We measure cognitive performance through the speed of keystroke and click interactions on a web search engine and correlate them to wearable device-defined sleep measures over time. We demonstrate that real-world performance varies throughout the day and is influenced by both circadian rhythms, chronotype (morning/evening preference), and prior sleep duration and timing. We develop a statistical model that operationalizes a large body of work on sleep and performance and demonstrates that our estimates of circadian rhythms, homeostatic sleep drive, and sleep inertia align with expectations from laboratory-based sleep studies. Further, we quantify the impact of insufficient sleep on real-world performance and show that two consecutive nights with less than six hours of sleep are associated with decreases in performance which last for a period of six days. This work demonstrates the feasibility of using online interactions for large-scale physiological sensing.
Tim Althoff, Eric Horvitz, Ryen W. White, Jamie M. Zeitzer
WWW2
2017 Toward multimodal signal detection of adverse drug reactions
Rave Harpaz, William DuMouchel, Martijn J. Schuemie, Olivier Bodenreider, Carol Friedman, Eric Horvitz, Anna Ripple, Alfred Sorbello, Ryen W. White, Rainer Winnenburg, Nigam H. Shah
J. Biomed. Informatics6
2016 Toward a Learning Science for Complex Crowdsourcing Tasks
abstract
We explore how crowdworkers can be trained to tackle complex crowdsourcing tasks. We are particularly interested in training novice workers to perform well on solving tasks in situations where the space of strategies is large and workers need to discover and try different strategies to be successful. In a first experiment, we perform a comparison of five different training strategies. For complex web search challenges, we show that providing expert examples is an effective form of training, surpassing other forms of training in nearly all measures of interest. However, such training relies on access to domain expertise, which may be expensive or lacking. Therefore, in a second experiment we study the feasibility of training workers in the absence of domain expertise. We show that having workers validate the work of their peer workers can be even more effective than having them review expert examples if we only present solutions filtered by a threshold length. The results suggest that crowdsourced solutions of peer workers may be harnessed in an automated training pipeline.
Shayan Doroudi, Ece Kamar, Emma Brunskill, Eric Horvitz
CHI4
2016 Identifying Dogmatism in Social Media: Signals and Models
abstract
We explore linguistic and behavioral features of dogmatism in social media and construct statistical models that can identify dogmatic comments.Our model is based on a corpus of Reddit posts, collected across a diverse set of conversational topics and annotated via paid crowdsourcing.We operationalize key aspects of dogmatism described by existing psychology theories (such as over-confidence), finding they have predictive power.We also find evidence for new signals of dogmatism, such as the tendency of dogmatic posts to refrain from signaling cognitive processes.When we use our predictive model to analyze millions of other Reddit posts, we find evidence that suggests dogmatism is a deeper personality trait, present for dogmatic users across many different domains, and that users who engage on dogmatic comments tend to show increases in dogmatic posts themselves.
Ethan Fast, Eric Horvitz
EMNLP2
2016 Are You Messing with Me?: Querying about the Sincerity of Interactions in the Open World
abstract
When interacting with robots deployed in the open world, people may often attempt to engage with them in a playful manner or test their competencies. Such engagements are often associated with language and behaviors that fall outside of designed task capabilities and can lead to interaction failures. Detecting when users are driven by play and curiosity can help a robot to understand why some interactions are breaking down, respond more appropriately by conveying its capabilities to its users, and enhance perceptions of its situational awareness and social intelligence. We have been studying the intentions of everyday users in their engagement with a long-lived robot system that provides directions within an office building. We report on a pilot field-study exploring the use of direct queries to elicit the sincerity of user requests, in terms of their actual need for directions. We discuss early results from this initial study and frame research directions and design implications for robots deployed in the wild.
Sean Andrist, Dan Bohus, Eric Horvitz
HRI4
2016 Intervention Strategies for Increasing Engagement in Crowdsourcing: Platform, Predictions, and Experiments
Avi Segal, Kobi Gal, Ece Kamar, Eric Horvitz, Alex Bowyer, Grant Miller
IJCAI4
2016 Detecting Devastating Diseases in Search Logs
abstract
Web search queries can offer a unique population-scale window onto streams of evidence that are useful for detecting the emergence of health conditions. We explore the promise of harnessing behavioral signals in search logs to provide advance warning about the presence of devastating diseases such as pancreatic cancer. Pancreatic cancer is often diagnosed too late to be treated effectively as the cancer has usually metastasized by the time of diagnosis. Symptoms of the early stages of the illness are often subtle and nonspecific. We identify searchers who issue credible, first-person diagnostic queries for pancreatic cancer and we learn models from prior search histories that predict which searchers will later input such queries. We show that we can infer the likelihood of seeing the rise of diagnostic queries months before they appear and characterize the tradeoff between predictivity and false positive rate. The findings highlight the potential of harnessing search logs for the early detection of pancreatic cancer and more generally for harnessing search systems to reduce health risks for individuals.
John Paparrizos, Ryen W. White, Eric Horvitz
KDD3
2016 Analyzing and Predicting Task Reminders
abstract
Automated personal assistants such as Siri, Cortana, and Google Now provide services to help users accomplish tasks, including tools to set reminders. We study how people specify and use reminders. Our study analyzes a sample of six months of logs of user-specified reminders from Cortana (Microsoft's intelligent personal assistant), the first large-scale analysis of such reminders. We focus our analyses on time-based reminders, the most common type of reminder found in the logs. We perform a data-driven analysis to identify common categories of tasks that give rise to these reminders across a large number of users, and we arrange these tasks into a taxonomy. We identify temporal patterns linked to the type of task, time of creation, and terms in the reminder text. Finally, we show that these patterns generalize by addressing a prediction task. Specifically, we show that a reminder's creation time is a strong feature in predicting the notification time, and that including the reminder text further improves prediction accuracy. The results have implications for the design of systems aimed at helping people to complete tasks and to plan future activities.
David Graus, Paul N. Bennett, Ryen W. White, Eric Horvitz
UMAP4
2016 Early identification of adverse drug reactions from search log data
abstract
The timely and accurate identification of adverse drug reactions (ADRs) following drug approval is a persistent and serious public health challenge. Aggregated data drawn from anonymized logs of Web searchers has been shown to be a useful source of evidence for detecting ADRs. However, prior studies have been based on the analysis of established ADRs, the existence of which may already be known publically. Awareness of these ADRs can inject existing knowledge about the known ADRs into online content and online behavior, and thus raise questions about the ability of the behavioral log-based methods to detect new ADRs. In contrast to previous studies, we investigate the use of search logs for the early detection of known ADRs. We use a large set of recently labeled ADRs and negative controls to evaluate the ability of search logs to accurately detect ADRs in advance of their publication. We leverage the Internet Archive to estimate when evidence of an ADR first appeared in the public domain and adjust the index date in a backdated analysis. Our results demonstrate how search logs can be used to detect new ADRs, the central challenge in pharmacovigilance.
Ryen W. White, Sheng Wang 0012, Apurv Pant, Rave Harpaz, Pushpraj Shukla, Walter Sun, William DuMouchel, Eric Horvitz
J. Biomed. Informatics8
2016 Patient Risk Stratification with Time-Varying Parameters: A Multitask Learning Approach
abstract
The proliferation of electronic health records (EHRs) frames opportunities for using machine learning to build models that help healthcare providers improve patient outcomes. However, building useful risk stratification models presents many technical challenges including the large number of factors (both intrinsic and extrinsic) influencing a patient's risk of an adverse outcome and the inherent evolution of that risk over time. We address these challenges in the context of learning a risk stratification model for predicting which patients are at risk of acquiring a Clostridium difficile infection (CDI). We take a novel data-centric approach, leveraging the contents of EHRs from nearly 50,000 hospital admissions. We show how, by adapting techniques from multitask learning, we can learn models for patient risk stratification with unprecedented classification performance. Our model, based on thousands of variables, both time-varying and time-invariant, changes over the course of a patient admission. Applied to a held out set of approximately 25,000 patient admissions, we achieve an area under the receiver operating characteristic curve of 0.81 (95% CI 0.78-0.84). The model has been integrated into the health record system at a large hospital in the US, and can be used to produce daily risk estimates for each inpatient. While more complex than traditional risk stratification methods, the widespread development and use of such data-driven models could ultimately enable cost-effective, targeted prevention strategies that lead to better patient outcomes.
Jenna Wiens, John V. Guttag, Eric Horvitz
J. Mach. Learn. Res.3
2016 Scalable Semisupervised Functional Neurocartography Reveals Canonical Neurons in Behavioral Networks
abstract
Large-scale data collection efforts to map the brain are underway at multiple spatial and temporal scales, but all face fundamental problems posed by high-dimensional data and intersubject variability. Even seemingly simple problems, such as identifying a neuron/brain region across animals/subjects, become exponentially more difficult in high dimensions, such as recognizing dozens of neurons/brain regions simultaneously. We present a framework and tools for functional neurocartography-the large-scale mapping of neural activity during behavioral states. Using a voltage-sensitive dye (VSD), we imaged the multifunctional responses of hundreds of leech neurons during several behaviors to identify and functionally map homologous neurons. We extracted simple features from each of these behaviors and combined them with anatomical features to create a rich medium-dimensional feature space. This enabled us to use machine learning techniques and visualizations to characterize and account for intersubject variability, piece together a canonical atlas of neural activity, and identify two behavioral networks. We identified 39 neurons (18 pairs, 3 unpaired) as part of a canonical swim network and 17 neurons (8 pairs, 1 unpaired) involved in a partially overlapping preparatory network. All neurons in the preparatory network rapidly depolarized at the onsets of each behavior, suggesting that it is part of a dedicated rapid-response network. This network is likely mediated by the S cell, and we referenced VSD recordings to an activity atlas to identify multiple cells of interest simultaneously in real time for further experiments. We targeted and electrophysiologically verified several neurons in the swim network and further showed that the S cell is presynaptic to multiple neurons in the preparatory network. This study illustrates the basic framework to map neural activity in high dimensions with large-scale recordings and how to extract the rich information necessary to perform analyses in light of intersubject variability.
Edward Paxon Frady, Ashish Kapoor, Eric Horvitz, William B. Kristan Jr.
Neural Comput.3
2016 Search and Breast Cancer: On Episodic Shifts of Attention over Life Histories of an Illness
abstract
We seek to understand the evolving needs of people who are faced with a life-changing medical diagnosis based on analyses of queries extracted from an anonymized search query log. Focusing on breast cancer, we manually tag a set of Web searchers as showing patterns of search behavior consistent with someone grappling with the screening, diagnosis, and treatment of breast cancer. We build and apply probabilistic classifiers to detect these searchers from multiple sessions and to identify the timing of diagnosis using temporal and statistical features. We explore the changes in information seeking over time before and after an inferred diagnosis of breast cancer by aligning multiple searchers by the estimated time of diagnosis. We employ the classifier to automatically identify 1,700 candidate searchers with an estimated 90% precision, and we predict the day of diagnosis within 15 days with an 88% accuracy. We show that the geographic and demographic attributes of searchers identified with high probability are strongly correlated with ground truth of reported incidence rates. We then analyze the content of queries over time for inferred cancer patients, using a detailed ontology of cancer-related search terms. The analysis reveals the rich temporal structure of the evolving queries of people likely diagnosed with breast cancer. Finally, we focus on subtypes of illness based on inferred stages of cancer and show clinically relevant dynamics of information seeking based on the dominant stage expressed by searchers.
Michael J. Paul, Ryen W. White, Eric Horvitz
ACM Trans. Web3
2015 Exploring Time-Dependent Concerns about Pregnancy and Childbirth from Search Logs
abstract
We study time-dependent patterns of information seeking about pregnancy, birth, and the first several weeks of caring for newborns via analyses of queries drawn from anonymized search engine logs. We show how we can detect and align web search behavior for a population of searchers with the natural clock of gestational physiology via proxies for ground truth based on searchers' self-report queries (e.g., [I am 30 weeks pregnant and my baby is moving a lot]). Then, we present a methodology for performing additional alignments, that are valuable for learning about the concerns, curiosities, and needs that arise over time with pregnancy and early parenting. Our findings have implications for learning about the temporal dynamics of pregnancy-related interests and concerns, and also for the design of systems that tailor their responses to point estimates of each searcher's current stage in pregnancy.
Adam Fourney, Ryen W. White, Eric Horvitz
CHI3
2015 Eyewitness: identifying local events via space-time signals in twitter feeds
abstract
We present a methodology for automatically extracting and summarizing reports of significant local events from large-scale Twitter feeds. While previous work has relied on an analysis of tweet text to identify local events, we show how to reliably detect events using only time series analysis of geotagged tweet volumes from localized regions. The algorithm sweeps through different spatial and temporal resolutions and finds events as anomalous spikes in the rate of geotagged tweets. We applied the approach to a corpus of over 733 million geotagged tweets. Using a panel of 103 crowdsourced judges who tagged 2400 detected events, we achieved a local event detection precision of 70%. Using these judged events as ground truth, a decision tree classifier was able to raise the detection precision to 93%.
John Krumm, Eric Horvitz
SIGSPATIAL/GIS2
2015 Identifying and Accounting for Task-Dependent Bias in Crowdsourcing
abstract
Models for aggregating contributions by crowd workers have been shown to be challenged by the rise of task-specific biases or errors. Task-dependent errors in assessment may shift the majority opinion of even large numbers of workers to an incorrect answer. We introduce and evaluate probabilistic models that can detect and correct task-dependent bias automatically. First, we show how to build and use probabilistic graphical models for jointly modeling task features, workers' biases, worker contributions and ground truth answers of tasks so that task-dependent bias can be corrected. Second, we show how the approach can perform a type of transfer learning among workers to address the issue of annotation sparsity. We evaluate the models with varying complexity on a large data set collected from a citizen science project and show that the models are effective at correcting the task-dependent worker bias. Finally, we investigate the use of active learning to guide the acquisition of expert assessments to enable automatic detection and correction of worker bias.
Ece Kamar, Ashish Kapoor, Eric Horvitz
HCOMP3
2015 Learning to Hire Teams
abstract
Crowdsourcing and human computation are being employed in sophisticated projects that require the solution of a heterogeneous set of tasks. We explore the challenge of composing or hiring an effective team from an available pool of applicants for performing tasks required for such projects on an ongoing basis. How can one optimally spend budget to learn the expertise of workers as part of recruiting a team? How can one exploit the similarities among tasks as well as underlying social ties or commonalities among the workers for faster learning? We tackle these decision-theoretic challenges by casting them as an instance of online learning for best action selection with side-observations. We present algorithms with PAC bounds on the required budget to hire a near-optimal team with high confidence. We evaluate our methodology on simulated problem instances using crowdsourcing data collected from the Upwork platform.
Adish Singla, Eric Horvitz, Pushmeet Kohli, Andreas Krause 0001
HCOMP2
2015 Connections: 2015 ICMI Sustained Accomplishment Award Lecture
abstract
Our community has long pursued principles and methods for enabling fluid and effortless collaborations between people and computing systems. Forging deep connections between people and machines has come into focus over the last 25 years as a grand challenge at the intersection of artificial intelligence, human-computer interaction, and cognitive psychology. I will review experiences and directions with leveraging advances in perception, learning, and reasoning in pursuit of our shared dreams.
Eric Horvitz
ICMI1
2015 Metareasoning for Planning Under Uncertainty
Christopher H. Lin, Andrey Kolobov, Ece Kamar, Eric Horvitz
IJCAI4
2015 Information Gathering in Networks via Active Exploration
Adish Singla, Eric Horvitz, Pushmeet Kohli, Ryen W. White, Andreas Krause 0001
IJCAI2
2015 A Deep Hybrid Model for Weather Forecasting
abstract
Weather forecasting is a canonical predictive challenge that has depended primarily on model-based methods. We explore new directions with forecasting weather as a data-intensive challenge that involves inferences across space and time. We study specifically the power of making predictions via a hybrid approach that combines discriminatively trained predictive models with a deep neural network that models the joint statistics of a set of weather-related variables. We show how the base model can be enhanced with spatial interpolation that uses learned long-range spatial dependencies. We also derive an efficient learning and inference procedure that allows for large scale optimization of the model parameters. We evaluate the methods with experiments on real-world meteorological data that highlight the promise of the approach.
Aditya Grover, Ashish Kapoor, Eric Horvitz
KDD3
2015 Inside Jokes: Identifying Humorous Cartoon Captions
abstract
Humor is an integral aspect of the human experience. Motivated by the prospect of creating computational models of humor, we study the influence of the language of cartoon captions on the perceived humorousness of the cartoons. Our studies are based on a large corpus of crowdsourced cartoon captions that were submitted to a contest hosted by the New Yorker. Having access to thousands of captions submitted for the same image allows us to analyze the breadth of responses of people to the same visual stimulus.
Dafna Shahaf, Eric Horvitz, Robert Mankoff
KDD2
2015 Incremental Coordination: Attention-Centric Speech Production in a Physically Situated Conversational Agent
abstract
Inspired by studies of human-human conversations, we present methods for incrementally coordinating speech production with listeners' visual foci of attention.We introduce a model that considers the demands and availability of listeners' attention at the onset and throughout the production of system utterances, and that incrementally coordinates speech synthesis with the listener's gaze.We present an implementation and deployment of the model in a physically situated dialog system and discuss lessons learned.
Dan Bohus, Eric Horvitz
SIGDIAL Conference3
2015 Events and Controversies: Influences of a Shocking News Event on Information Seeking
abstract
It has been suggested that online search and retrieval contributes to the intellectual isolation of users within their preexisting ideologies, where people's prior views are strengthened and alternative viewpoints are infrequently encountered. This so-called "filter bubble" phenomenon has been called out as especially detrimental when it comes to dialog among people on controversial, emotionally charged topics, such as the labeling of genetically modified food, the right to bear arms, the death penalty, and online privacy. We seek to identify and study information-seeking behavior and access to alternative versus reinforcing viewpoints following shocking, emotional, and large-scale news events. We choose for a case study to analyze search and browsing on gun control/rights, a strongly polarizing topic for both citizens and leaders of the United States. We study the period of time preceding and following a mass shooting to understand how its occurrence, follow-on discussions, and debate may have been linked to changes in the patterns of searching and browsing. We employ information-theoretic measures to quantify the diversity of Web domains of interest to users and understand the browsing patterns of users. We use these measures to characterize the influence of news events on these web search and browsing patterns.
Danai Koutra, Paul N. Bennett, Eric Horvitz
WWW3
2015 Diagnoses, Decisions, and Outcomes: Web Search as Decision Support for Cancer
abstract
People diagnosed with a serious illness often turn to the Web for their rising information needs, especially when decisions are required. We analyze the search and browsing behavior of searchers who show a surge of interest in prostate cancer. Prostate cancer is the most common serious cancer in men and is a leading cause of cancer-related death. Diagnoses of prostate cancer typically involve reflection and decision making about treatment based on assessments of preferences and outcomes. We annotated timelines of treatment-related queries from nearly 300 searchers with tags indicating different phases of treatment, including decision making, preparation, and recovery. Using this corpus, we present a variety of analyses toward the goal of understanding search and decision making about treatments. We characterize search queries and the content of accessed pages for different treatment phases, model search behavior during the decision-making phase, and create an aggregate alignment of treatment timelines illustrated with a variety of visualizations. The experiments provide insights about how people who are engaged in intensive searches about prostate cancer over an extended period of time pursue and access information from the Web.
Michael J. Paul, Ryen W. White, Eric Horvitz
WWW3
2015 Belief Dynamics and Biases in Web Search
abstract
We investigate how beliefs about the efficacy of medical interventions are influenced by searchers' exposure to information on retrieved Web pages. We present a methodology for measuring participants' beliefs and confidence about the efficacy of treatment before, during, and after search episodes. We consider interventions studied in the Cochrane collection of meta-analyses. We extract related queries from search engine logs and consider the Cochrane assessments as ground truth. We analyze the dynamics of belief over time and show the influence of prior beliefs and confidence at the end of sessions. We present evidence for confirmation bias and for anchoring-and-adjustment during search and retrieval. Then, we build predictive models to estimate postsearch beliefs using sets of features about behavior and content. The findings provide insights about the influence of Web content on the beliefs of people and have implications for the design of search systems.
Ryen W. White, Eric Horvitz
ACM Trans. Inf. Syst.2
2014 Signals in the Silence: Models of Implicit Feedback in a Recommendation System for Crowdsourcing
abstract
We exploit the absence of signals as informative observations in the context of providing task recommendations in crowdsourcing. Workers on crowdsourcing platforms do not provide explicit ratings about tasks. We present methods that enable a system to leverage implicit signals about task preferences. These signals include types of tasks that have been available and have been displayed, and the number of tasks workers select and complete. In contrast to previous work, we present a general model that can represent both positive and negative implicit signals. We introduce algorithms that can learn these models without exceeding the computational complexity of existing approaches. Finally, using data from a high-throughput crowdsourcing platform, we show that reasoning about both positive and negative implicit feedback can improve the quality of task recommendations.
Christopher H. Lin, Ece Kamar, Eric Horvitz
AAAI3
2014 Stochastic Privacy
abstract
Online services such as web search and e-commerce applications typically rely on the collection of data about users, including details of their activities on the web. Such personal data is used to maximize revenues via targeting of advertisements and longer engagements of users, and to enhance the quality of service via personalization of content. To date, service providers have largely followed the approach of either requiring or requesting consent for collecting user data. Users may be willing to share private information in return for incentives, enhanced services, or assurances about the nature and extent of the logged data. We introduce stochastic privacy, an approach to privacy centering on the simple concept of providing people with a guarantee that the probability that their personal data will be shared does not exceed a given bound. Such a probability, which we refer to as the privacy risk, can be given by users as a preference or communicated as a policy by a service provider. Service providers can work to personalize and to optimize revenues in accordance with preferences about privacy risk. We present procedures, proofs, and an overall system for maximizing the quality of services, while respecting bounds on privacy risk. We demonstrate the methodology with a case study and evaluation of the procedures applied to web search personalization. We show how we can achieve near-optimal utility of accessing information with provable guarantees on the probability of sharing data.
Adish Singla, Eric Horvitz, Ece Kamar, Ryen W. White
AAAI2
2014 Characterizing and predicting postpartum depression from shared facebook data
abstract
The birth of a child is a major milestone in the life of parents. We leverage Facebook data shared voluntarily by 165 new mothers as streams of evidence for characterizing their postnatal experiences. We consider multiple measures including activity, social capital, emotion, and linguistic style in participants' Facebook data in pre- and postnatal periods. Our study includes detecting and predicting onset of post-partum depression (PPD). The work complements recent work on detecting and predicting significant postpartum changes in behavior, language, and affect from Twitter data. In contrast to prior studies, we gain access to ground truth on postpartum experiences via self-reports and a common psychometric instrument used to evaluate PPD. We develop a series of statistical models to predict, from data available before childbirth, a mother's likelihood of PPD. We corroborate our quantitative findings through interviews with mothers experiencing PPD. We find that increased social isolation and lowered availability of social capital on Facebook, are the best predictors of PPD in mothers.
Munmun De Choudhury, Scott Counts, Eric Horvitz, Aaron Hoff
CSCW3
2014 Managing Human-Robot Engagement with Forecasts and... um... Hesitations
abstract
We explore methods for managing conversational engagement in open-world, physically situated dialog systems. We investigate a self-supervised methodology for constructing forecasting models that aim to anticipate when participants are about to terminate their interactions with a situated system. We study how these models can be leveraged to guide a disengagement policy that uses linguistic hesitation actions, such as filled and non-filled pauses, when uncertainty about the continuation of engagement arises. The hesitations allow for additional time for sensing and inference, and convey the system's uncertainty. We report results from a study of the proposed approach with a directions-giving robot deployed in the wild.
Dan Bohus, Eric Horvitz
ICMI2
2014 Natural Communication about Uncertainties in Situated Interaction
abstract
Physically situated, multimodal interactive systems must often grapple with uncertainties about properties of the world, people, and their intentions and actions. We present methods for estimating and communicating about different uncertainties in situated interaction, leveraging the affordances of an embodied conversational agent. The approach harnesses a representation that captures both the magnitude and the sources of uncertainty, and a set of policies that select and coordinate the production of nonverbal and verbal behaviors to communicate the system's uncertainties to conversational participants. The methods are designed to enlist participants' help in a natural manner to resolve uncertainties arising during interactions. We report on a preliminary implementation of the proposed methods in a deployed system and illustrate the functionality with a trace from a sample interaction.
Tomislav Pejsa, Dan Bohus, Michael F. Cohen, Chit W. Saw, James Mahoney, Eric Horvitz
ICMI6
2014 Airplanes aloft as a sensor network for wind forecasting
Ashish Kapoor, Zachary Horvitz, Spencer Laube, Eric Horvitz
IPSN4
2014 Data, predictions, and decisions in support of people and society
abstract
Deep societal benefits will spring from advances in data availability and in computational procedures for mining insights and inferences from large data sets. I will describe efforts to harness data for making predictions and guiding decisions, touching on work in transportation, healthcare, online services, and interactive systems. I will start with efforts to learn and field predictive models that forecast flows of traffic in greater city regions. Moving from the ground to the air, I will discuss fusing data from aircraft to make inferences about atmospheric conditions and using these results to enhance air transport. I will then focus on experiences with building and fielding predictive models in clinical medicine. I will show how inferences about outcomes and interventions can provide insights and guide decision making. Moving beyond data captured by hospitals, I will discuss the promise of transforming anonymized behavioral data drawn from web services into large-scale sensor networks for public health, including efforts to identify adverse effects of medications and to understand illness in populations. I will conclude by describing how we can use machine learning to leverage the complementarity of human and machine intellect to solve challenging problems in science and society.
Eric Horvitz
KDD1
2014 Time-critical search
abstract
We study time-critical search, where users have urgent information needs in the context of an acute problem. As examples, users may need to know how to stem a severe bleed, help a baby who is choking on a foreign object, or respond to an epileptic seizure. While time-critical situations and actions have been studied in the realm of decision-support systems, little has been done with time-critical search and retrieval, and little direct support is offered by search systems. Critical challenges with time-critical search include accurately inferring when users have urgent needs and providing relevant information that can be understood and acted upon quickly. We leverage surveys and search log data from a large mobile search provider to (a) characterize the use of search engines for time-critical situations, and (b) develop predictive models to accurately predict urgent information needs, given a query and a diverse set of features spanning topical, temporal, behavioral, and geospatial attributes. The methods and findings highlight opportunities for extending search and retrieval to consider the urgency of queries.
Nina Mishra, Ryen W. White, Samuel Ieong, Eric Horvitz
SIGIR4
2014 Enhancing personalization via search activity attribution
abstract
Online services rely on machine identifiers to tailor services such as personalized search and advertising to individual users. The assumption made is that each identifier comprises the behavior of a single person. However, shared machine usage is common, and in these cases, the activities of multiple users may be generated under a single identifier, creating a potentially noisy signal for applications such as search personalization. We propose enhancing Web search personalization with methods that can disambiguate among different users of a machine, thus connecting the current query with the appropriate search history. Using logs containing both person and machine identifiers, and logs from a popular commercial search engine, we learn models that accurately assign observed search behaviors to each of different users. This information is then used to augment existing personalization methods that are currently based only on machine identifiers. We show that this new capability to infer users can be used to improve the performance of existing personalization methods. The early findings of our research are promising and have implications for search personalization.
Adish Singla, Ryen W. White, Ahmed Awadallah 0001, Eric Horvitz
SIGIR4
2014 From devices to people: attribution of search activity in multi-user settings
abstract
Online services rely on unique identifiers of machines to tailor offerings to their users. An implicit assumption is made that each machine identifier maps to an individual. However, shared ma-chines are common, leading to interwoven search histories and noisy signals for applications such as personalized search and ad-vertising. We present methods for attributing search activity to individual searchers. Using ground truth data for a sample of almost four million U.S. Web searchers-containing both machine identifiers and person identifiers-we show that over half of the machine identifiers comprise the queries of multiple people. We characterize variations in features of topic, time, and other aspects such as the complexity of the information sought per the number of searchers on a machine, and show significant differences in all measures. Based on these insights, we develop models to accurately estimate when multiple people contribute to the logs ascribed to a single machine identifier. We also develop models to cluster search behavior on a machine, allowing us to attribute historical data accurately and automatically assign new search activity to the correct searcher. The findings have implications for the design of applications such as personalized search and advertising that rely heavily on machine identifiers to custom-tailor their services.
Ryen W. White, Ahmed Awadallah 0001, Adish Singla, Eric Horvitz
WWW4
2014 An Interactive Approach to Solving Correspondence Problems
Stefanie Jegelka, Ashish Kapoor, Eric Horvitz
Int. J. Comput. Vis.3
2014 From health search to healthcare: explorations of intention and utilization via query logs and user surveys
abstract
OBJECTIVE: To better understand the relationship between online health-seeking behaviors and in-world healthcare utilization (HU) by studies of online search and access activities before and after queries that pursue medical professionals and facilities. MATERIALS AND METHODS: We analyzed data collected from logs of online searches gathered from consenting users of a browser toolbar from Microsoft (N=9740). We employed a complementary survey (N=489) to seek a deeper understanding of information-gathering, reflection, and action on the pursuit of professional healthcare. RESULTS: We provide insights about HU through the survey, breaking out its findings by different respondent marginalizations as appropriate. Observations made from search logs may be explained by trends observed in our survey responses, even though the user populations differ. DISCUSSION: The results provide insights about how users decide if and when to utilize healthcare resources, and how online health information seeking transitions to in-world HU. The findings from both the survey and the logs reveal behavioral patterns and suggest a strong relationship between search behavior and HU. Although the diversity of our survey respondents is limited and we cannot be certain that users visited medical facilities, we demonstrate that it may be possible to infer HU from long-term search behavior by the apparent influence that health concerns and professional advice have on search activity. CONCLUSIONS: Our findings highlight different phases of online activities around queries pursuing professional healthcare facilities and services. We also show that it may be possible to infer HU from logs without tracking people's physical location, based on the effect of HU on pre- and post-HU search behavior. This allows search providers and others to develop more robust models of interests and preferences by modeling utilization rather than simply the intention to utilize that is expressed in search queries.
Ryen W. White, Eric Horvitz
J. Am. Medical Informatics Assoc.2
2014 A study in transfer learning: leveraging data from multiple hospitals to enhance hospital-specific predictions
abstract
BACKGROUND: Data-driven risk stratification models built using data from a single hospital often have a paucity of training data. However, leveraging data from other hospitals can be challenging owing to institutional differences with patients and with data coding and capture. OBJECTIVE: To investigate three approaches to learning hospital-specific predictions about the risk of hospital-associated infection with Clostridium difficile, and perform a comparative analysis of the value of different ways of using external data to enhance hospital-specific predictions. MATERIALS AND METHODS: We evaluated each approach on 132 853 admissions from three hospitals, varying in size and location. The first approach was a single-task approach, in which only training data from the target hospital (ie, the hospital for which the model was intended) were used. The second used only data from the other two hospitals. The third approach jointly incorporated data from all hospitals while seeking a solution in the target space. RESULTS: The relative performance of the three different approaches was found to be sensitive to the hospital selected as the target. However, incorporating data from all hospitals consistently had the highest performance. DISCUSSION: The results characterize the challenges and opportunities that come with (1) using data or models from collections of hospitals without adapting them to the site at which the model will be used, and (2) using only local data to build models for small institutions or rare events. CONCLUSIONS: We show how external data from other hospitals can be successfully and efficiently incorporated into hospital-specific models.
Jenna Wiens, John V. Guttag, Eric Horvitz
J. Am. Medical Informatics Assoc.3
2014 Geospatial Structure of a Planetary-Scale Social Network
abstract
Little is known about geographic properties of large-scale social networks. In this paper, we examine the geospatial attributes of a planetary-scale social network of 240 million people and 1.3 billion edges. We study the interplay among topological, geographical, and algorithmically generated paths connecting pairs of nodes in a social network. Starting in the realm of cyberspace, we find that topologically shortest paths of average length of 6.6 exist between pairs of nodes in the network and that the average degree of separation among nodes is robust to removal of hub nodes. Moving to the realm of locations and distances in geographic space, we find that topologically shortest paths in the social graph grow with increasing geographic distance between path's endpoint nodes. We discover that shortest topological paths are geographically inefficient, but that geography provides an important cue for local algorithmic policies for navigating between source and target nodes. Local algorithmic strategies for navigating the larger network structure in the absence of global navigation procedures have varying success. At the early stages of the navigation, navigating to a hub node helps, while in the middle stage, geography provides the most important clue. While local algorithms for navigating have trouble reaching the target node, they are successful in reaching nodes that are geographically close to the target. Taken together, our results demonstrate a complex interplay between topological and geographical properties of social networks and explain the success of local strategies for navigating such networks.
Jure Leskovec, Eric Horvitz
IEEE Trans. Comput. Soc. Syst.2
2013 Automated Workflow Synthesis
abstract
By coordinating efforts from humans and machines, human computation systems can solve problems that machines cannot tackle alone. A general challenge is to design efficient human computation algorithms or workflows with which to coordinate the work of the crowd. We introduce a method for automated workflow synthesis aimed at ideally harnessing human efforts by learning about the crowd's performance on tasks and synthesizing an optimal workflow for solving a problem. We present experimental results for human sorting tasks, which demonstrate both the benefit of understanding and optimizing the structure of workflows based on observations. Results also demonstrate the benefits of using value of information to guide experiments for identifying efficient workflows with fewer experiments.
Eric Horvitz, David C. Parkes
AAAI2
2013 Predicting postpartum changes in emotion and behavior via social media
abstract
We consider social media as a promising tool for public health, focusing on the use of Twitter posts to build predictive models about the forthcoming influence of childbirth on the behavior and mood of new mothers. Using Twitter posts, we quantify postpartum changes in 376 mothers along dimensions of social engagement, emotion, social network, and linguistic style. We then construct statistical models from a training set of observations of these measures before and after the reported childbirth, to forecast significant postpartum changes in mothers. The predictive models can classify mothers who will change significantly following childbirth with an accuracy of 71%, using observations about their prenatal behavior, and as accurately as 80-83% when additionally leveraging the initial 2-3 weeks of postnatal data. The study is motivated by the opportunity to use social media to identify mothers at risk of postpartum depression, an underreported health concern among large populations, and to inform the design of low-cost, privacy-sensitive early-warning systems and intervention programs aimed at promoting wellness postpartum.
Munmun De Choudhury, Scott Counts, Eric Horvitz
CHI3
2013 Major life changes and behavioral markers in social media: case of childbirth
abstract
We explore the harnessing of social media as a window on changes around major life events in individuals and larger populations. We specifically examine patterns of activity, emotional, and linguistic correlates for childbirth and postnatal course. After identifying childbirth events on Twitter, we analyze daily posting patterns and language usage before and after birth by new mothers, and make inferences about the status and dynamics of changes in emotions expressed following childbirth. We find that childbirth is associated with some changes for most new mothers, but approximately 15% of new mothers show significant changes in their online activity and emotional expression postpartum. We observe that these mothers can be distinguished by linguistic changes captured by shifts in a relatively small number of words in their social media posts. We introduce a greedy differencing procedure to identify the type of language that characterizes significant changes in these mothers during postpartum. We conclude with a discussion about how such characterizations might be applied to recognizing and understanding health and well-being in women following childbirth.
Munmun De Choudhury, Scott Counts, Eric Horvitz
CSCW3
2013 Volunteering Versus Work for Pay: Incentives and Tradeoffs in Crowdsourcing
abstract
Paid and volunteer crowd work have emerged as a means for harnessing human intelligence for performing diverse tasks. However, little is known about the relative performance of volunteer versus paid crowd work, and how financial incentives influence the quality and efficiency of output. We study the performance of volunteers as well as workers paid with different monetary schemes on a difficult real-world crowdsourcing task. We observe that performance by unpaid and paid workers can be compared in carefully designed tasks, that financial incentives can be used to trade quality for speed, and that the compensation system on Amazon Mechanical Turk creates particular indirect incentives for workers. Our methodology and results have implications for the ideal choice of financial incentives and motivates further study on how monetary incentives influence worker behavior in crowdsourcing.
Andrew Mao, Ece Kamar, Yiling Chen 0001, Eric Horvitz, Megan E. Schwamb, Chris J. Lintott, Arfon M. Smith
HCOMP4
2013 Why Stop Now? Predicting Worker Engagement in Online Crowdsourcing
abstract
We present studies of the attention and time, or engagement, invested by crowd workers on tasks. Consideration of worker engagement is especially important in volunteer settings such as online citizen science. Using data from Galaxy Zoo, a prominent citizen science project, we design and construct statistical models that provide predictions about the forthcoming engagement of volunteers. We characterize the accuracy of predictions with respect to different sets of features that describe user behavior and study the sensitivity of predictions to variations in the amount of data and retraining. We design our model for guiding system actions in real-time settings, and discuss the prospect for harnessing predictive models of engagement to enhance user attention and effort on volunteer tasks.
Andrew Mao, Ece Kamar, Eric Horvitz
HCOMP3
2013 Execution memory for grounding and coordination
Stephanie Rosenthal, Sarjoun Skaff, Manuela M. Veloso, Dan Bohus, Eric Horvitz
HRI5
2013 Predicting Depression via Social Media
Munmun De Choudhury, Michael Gamon, Scott Counts, Eric Horvitz
ICWSM4
2013 Crowdphysics: Planned and Opportunistic Crowdsourcing for Physical Tasks
Adam Sadilek, John Krumm, Eric Horvitz
ICWSM3
2013 Lifelong Learning for Acquiring the Wisdom of the Crowd
Ece Kamar, Ashish Kapoor, Eric Horvitz
IJCAI3
2013 Look versus Leap: Computing Value of Information with High-Dimensional Streaming Evidence
Stephanie Rosenthal, Dan Bohus, Ece Kamar, Eric Horvitz
IJCAI4
2013 Here and there: goals, activities, and predictions about location from geotagged queries
abstract
A significant portion of Web search is performed in mobile settings. We explore the links between users' queries on mobile devices and their locations and movement, with a focus on interpreting queries about addresses. We find that users tend to have a primary location, likely corresponding to home or workplace, and that a user's location relative to this primary location systematically influences the patterns of address searches. We apply our findings to construct a statistical model that can predict with high accuracy whether a user will be soon observed at an address that had been recently retrieved via search. Such an ability to predict that a user will transition to a location can be harnessed for multiple uses including provision of directions and traffic information, the rendering of competitive advertising, and guiding the opportunistic completion of pending tasks that can be accomplished en route to a target location.
Robert West 0001, Ryen W. White, Eric Horvitz
SIGIR3
2013 Workshop on health search and discovery: helping users and advancing medicine
abstract
This workshop brings together researchers and practitioners from industry and academia to discuss search and discovery in the medi-cal domain. The event focuses on ways to make medical and health information more accessible to laypeople (including enhancements to ranking algorithms and search interfaces), and how we can dis-cover new medical facts and phenomena from information sought online, as evidenced in query streams and other sources such as social media. This domain also offers many opportunities for appli-cations that monitor and improve quality of life of those affected by medical conditions, by providing tools to support their health-related information behavior.
Ryen W. White, Elad Yom-Tov, Eric Horvitz, Eugene Agichtein, William R. Hersh
SIGIR3
2013 Pursuing insights about healthcare utilization via geocoded search queries
abstract
Mobile devices provide people with a conduit to the rich infor-mation resources of the Web. With consent, the devices can also provide streams of information about search activity and location that can be used in population studies and real-time assistance. We analyzed geotagged mobile queries in a privacy-sensitive study of potential transitions from health information search to in-world healthcare utilization. We note differences in people's health infor-mation seeking before, during, and after the appearance of evidence that a medical facility has been visited. We find that we can accu-rately estimate statistics about such potential user engagement with healthcare providers. The findings highlight the promise of using geocoded search for sensing and predicting activities in the world.
Shuang-Hong Yang, Ryen W. White, Eric Horvitz
SIGIR3
2013 Pairwise ranking aggregation in a crowdsourced setting
abstract
Inferring rankings over elements of a set of objects, such as documents or images, is a key learning problem for such important applications as Web search and recommender systems. Crowdsourcing services provide an inexpensive and efficient means to acquire preferences over objects via labeling by sets of annotators. We propose a new model to predict a gold-standard ranking that hinges on combining pairwise comparisons via crowdsourcing. In contrast to traditional ranking aggregation methods, the approach learns about and folds into consideration the quality of contributions of each annotator. In addition, we minimize the cost of assessment by introducing a generalization of the traditional active learning scenario to jointly select the annotator and pair to assess while taking into account the annotator quality, the uncertainty over ordering of the pair, and the current model uncertainty. We formalize this as an active learning strategy that incorporates an exploration-exploitation tradeoff and implement it using an efficient online Bayesian updating scheme. Using simulated and real-world data, we demonstrate that the active learning strategy achieves significant reductions in labeling cost while maintaining accuracy.
Paul N. Bennett, Kevyn Collins-Thompson, Eric Horvitz
WSDM4
2013 Mining the web to predict future events
abstract
We describe and evaluate methods for learning to forecast forthcoming events of interest from a corpus containing 22 years of news stories. We consider the examples of identifying significant increases in the likelihood of disease outbreaks, deaths, and riots in advance of the occurrence of these events in the world. We provide details of methods and studies, including the automated extraction and generalization of sequences of events from news corpora and multiple web resources. We evaluate the predictive power of the approach on real-world events withheld from the system.
Kira Radinsky, Eric Horvitz
WSDM2
2013 From cookies to cooks: insights on dietary patterns via analysis of web usage logs
abstract
Nutrition is a key factor in people's overall health. Hence, understanding the nature and dynamics of population-wide dietary preferences over time and space can be valuable in public health. To date, studies have leveraged small samples of participants via food intake logs or treatment data. We propose a complementary source of population data on nutrition obtained via Web logs. Our main contribution is a spatiotemporal analysis of population-wide dietary preferences through the lens of logs gathered by a widely distributed Web-browser add-on, using the access volume of recipes that users seek via search as a proxy for actual food consumption. We discover that variation in dietary preferences as expressed via recipe access has two main periodic components, one yearly and the other weekly, and that there exist characteristic regional differences in terms of diet within the United States. In a second study, we identify users who show evidence of having made an acute decision to lose weight. We characterize the shifts in interests that they express in their search queries and focus on changes in their recipe queries in particular. Last, we correlate nutritional time series obtained from recipe queries with time-aligned data on hospital admissions, aimed at understanding how behavioral data captured in Web logs might be harnessed to identify potential relationships between diet and acute health problems. In this preliminary study, we focus on patterns of sodium identified in recipes over time and patterns of admission for congestive heart failure, a chronic illness that can be exacerbated by increases in sodium intake.
Robert West 0001, Ryen W. White, Eric Horvitz
WWW3
2013 From web search to healthcare utilization: privacy-sensitive studies from mobile data
abstract
OBJECTIVE: We explore relationships between health information seeking activities and engagement with healthcare professionals via a privacy-sensitive analysis of geo-tagged data from mobile devices. MATERIALS AND METHODS: We analyze logs of mobile interaction data stripped of individually identifiable information and location data. The data analyzed consist of time-stamped search queries and distances to medical care centers. We examine search activity that precedes the observation of salient evidence of healthcare utilization (EHU) (ie, data suggesting that the searcher is using healthcare resources), in our case taken as queries occurring at or near medical facilities. RESULTS: We show that the time between symptom searches and observation of salient evidence of seeking healthcare utilization depends on the acuity of symptoms. We construct statistical models that make predictions of forthcoming EHU based on observations about the current search session, prior medical search activities, and prior EHU. The predictive accuracy of the models varies (65%-90%) depending on the features used and the timeframe of the analysis, which we explore via a sensitivity analysis. DISCUSSION: We provide a privacy-sensitive analysis that can be used to generate insights about the pursuit of health information and healthcare. The findings demonstrate how large-scale studies of mobile devices can provide insights on how concerns about symptomatology lead to the pursuit of professional care. CONCLUSION: We present new methods for the analysis of mobile logs and describe a study that provides evidence about how people transition from mobile searches on symptoms and diseases to the pursuit of healthcare in the world.
Ryen W. White, Eric Horvitz
J. Am. Medical Informatics Assoc.2
2013 Web-scale pharmacovigilance: listening to signals from the crowd
abstract
Adverse drug events cause substantial morbidity and mortality and are often discovered after a drug comes to market. We hypothesized that Internet users may provide early clues about adverse drug events via their online information-seeking. We conducted a large-scale study of Web search log data gathered during 2010. We pay particular attention to the specific drug pairing of paroxetine and pravastatin, whose interaction was reported to cause hyperglycemia after the time period of the online logs used in the analysis. We also examine sets of drug pairs known to be associated with hyperglycemia and those not associated with hyperglycemia. We find that anonymized signals on drug interactions can be mined from search logs. Compared to analyses of other sources such as electronic health records (EHR), logs are inexpensive to collect and mine. The results demonstrate that logs of the search activities of populations of computer users can contribute to drug safety surveillance.
Ryen W. White, Nicholas P. Tatonetti, Nigam H. Shah, Russ B. Altman, Eric Horvitz
J. Am. Medical Informatics Assoc.5
2013 Behavioral dynamics on the web: Learning, modeling, and prediction
abstract
The queries people issue to a search engine and the results clicked following a query change over time. For example, after the earthquake in Japan in March 2011, the query japan spiked in popularity and people issuing the query were more likely to click government-related results than they would prior to the earthquake. We explore the modeling and prediction of such temporal patterns in Web search behavior. We develop a temporal modeling framework adapted from physics and signal processing and harness it to predict temporal patterns in search behavior using smoothing, trends, periodicities, and surprises. Using current and past behavioral data, we develop a learning procedure that can be used to construct models of users' Web search activities. We also develop a novel methodology that learns to select the best prediction model from a family of predictive models for a given query or a class of queries. Experimental results indicate that the predictive models significantly outperform baseline models that weight historical evidence the same for all queries. We present two applications where new methods introduced for the temporal modeling of user behavior significantly improve upon the state of the art. Finally, we discuss opportunities for using models of temporal dynamics to enhance other areas of Web search and information retrieval.
Kira Radinsky, Krysta M. Svore, Susan T. Dumais, Milad Shokouhi, Jaime Teevan, Alex Bocharov, Eric Horvitz
ACM Trans. Inf. Syst.7
2013 Captions and biases in diagnostic search
abstract
People frequently turn to the Web with the goal of diagnosing medical symptoms. Studies have shown that diagnostic search can often lead to anxiety about the possibility that symptoms are explained by the presence of rare, serious medical disorders, rather than far more common benign syndromes. We study the influence of the appearance of potentially-alarming content, such as severe illnesses or serious treatment options associated with the queried for symptoms, in captions comprising titles, snippets, and URLs. We explore whether users are drawn to results with potentially-alarming caption content, and if so, the implications of such attraction for the design of search engines. We specifically study the influence of the content of search result captions shown in response to symptom searches on search-result click-through behavior. We show that users are significantly more likely to examine and click on captions containing potentially-alarming medical terminology such as “heart attack” or “medical emergency” independent of result rank position and well-known positional biases in users' search examination behaviors. The findings provide insights about the possible effects of displaying implicit correlates of searchers' goals in search-result captions, such as unexpressed concerns and fears. As an illustration of the potential utility of these results, we developed and evaluated an enhanced click prediction model that incorporates potentially-alarming caption features and show that it significantly outperforms models that ignore caption content. Beyond providing additional understanding of the effects of Web content on medical concerns, the methods and findings have implications for search engine design. As part of our discussion on the implications of this research, we propose procedures for generating more representative captions that may be less likely to cause alarm, as well as methods for learning to more appropriately rank search results from logged search behavior, for examples, by also considering the presence of potentially-alarming content in the captions that motivate observed clicks and down-weighting clicks seemingly driven by searchers' health anxieties.
Ryen W. White, Eric Horvitz
ACM Trans. Web2
2012 Learning to Learn: Algorithmic Inspirations from Human Problem Solving
abstract
We harness the ability of people to perceive and interact with visual patterns in order to enhance the performance of a machine learning method. We show how we can collect evidence about how people optimize the parameters of an ensemble classification system using a tool that provides a visualization of misclassification costs. Then, we use these observations about human attempts to minimize cost in order to extend the performance of a state-of-the-art ensemble classification system. The study highlights opportunities for learning from evidence collected about human problem solving to refine and extend automated learning and inference.
Ashish Kapoor, Bongshin Lee, Desney S. Tan, Eric Horvitz
AAAI4
2012 Performance and Preferences: Interactive Refinement of Machine Learning Procedures
abstract
Problem-solving procedures have been typically aimed at achieving well-defined goals or satisfying straightforward preferences. However, learners and solvers may often generate rich multiattribute results with procedures guided by sets of controls that define different dimensions of quality. We explore methods that enable people to explore and express preferences about the operation of classification models in supervised multiclass learning. We leverage a leave-one-out confusion matrix that provides users with views and real-time controls of a model space. The approach allows people to consider in an interactive manner the global implications of local changes in decision boundaries. We focus on kernel classifiers and show the effectiveness of the methodology on a variety of tasks.
Ashish Kapoor, Bongshin Lee, Desney S. Tan, Eric Horvitz
AAAI4
2012 Direct answers for search queries in the long tail
abstract
Web search engines now offer more than ranked results. Queries on topics like weather, definitions, and movies may return inline results called answers that can resolve a searcher's information need without any additional interaction. Despite the usefulness of answers, they are limited to popular needs because each answer type is manually authored. To extend the reach of answers to thousands of new information needs, we introduce Tail Answers: a large collection of direct answers that are unpopular individually, but together address a large proportion of search traffic. These answers cover long-tail needs such as the average body temperature for a dog, substitutes for molasses, and the keyboard shortcut for a right-click. We introduce a combination of search log mining and paid crowdsourcing techniques to create Tail Answers. A user study with 361 participants suggests that Tail Answers significantly improved users' subjective ratings of search quality and their ability to solve needs without clicking through to a result. Our findings suggest that search engines can be extended to directly respond to a large new class of queries.
Michael S. Bernstein, Jaime Teevan, Susan T. Dumais, Daniel J. Liebling, Eric Horvitz
CHI5
2012 Human computation tasks with global constraints
abstract
An important class of tasks that are underexplored in current human computation systems are complex tasks with global constraints. One example of such a task is itinerary planning, where solutions consist of a sequence of activities that meet requirements specified by the requester. In this paper, we focus on the crowdsourcing of such plans as a case study of constraint-based human computation tasks and introduce a collaborative planning system called Mobi that illustrates a novel crowdware paradigm. Mobi presents a single interface that enables crowd participants to view the current solution context and make appropriate contributions based on current needs. We conduct experiments that explain how Mobi enables a crowd to effectively and collaboratively resolve global constraints, and discuss how the design principles behind Mobi can more generally facilitate a crowd to tackle problems involving global constraints.
Edith Law, Rob Miller 0001, Krzysztof Z. Gajos, David C. Parkes, Eric Horvitz
CHI6
2012 Memory constrained face recognition
abstract
Real-time recognition may be limited by scarce memory and computing resources for performing classification. Although, prior research has addressed the problem of training classifiers with limited data and computation, few efforts have tackled the problem of memory constraints on recognition. We explore methods that can guide the allocation of limited storage resources for classifying streaming data so as to maximize discriminatory power. We focus on computation of the expected value of information with nearest neighbor classifiers for online face recognition. Experiments on real-world datasets show the effectiveness and power of the approach. The methods provide a principled approach to vision under bounded resources, and have immediate application to enhancing recognition capabilities in consumer devices with limited memory.
Ashish Kapoor, Simon Baker, Sumit Basu, Eric Horvitz
CVPR4
2012 Some help on the way: opportunistic routing under uncertainty
abstract
We investigate opportunistic routing, centering on the recommendation of ideal diversions on trips to a primary destination when an unplanned waypoint, such as a rest stop or a refueling station, is desired. In the general case, an automated routing assistant may not know the driver's final destination and may need to consider probabilities over destinations in identifying the ideal waypoint along with the revised route that includes the waypoint. We consider general principles of opportunistic routing and present the results of several studies with a corpus of real-world trips. Then, we describe how we can compute the expected value of asking a user about the primary destination so as to remove uncertainly about the goal and show how this measure can guide an automated system's engagements with users when making recommendations for navigation and analogous settings in ubiquitous computing.
Eric Horvitz, John Krumm
UbiComp1
2012 Metro maps of science
abstract
As the number of scientific publications soars, even the most enthusiastic reader can have trouble staying on top of the evolving literature. It is easy to focus on a narrow aspect of one's field and lose track of the big picture. Information overload is indeed a major challenge for scientists today, and is especially daunting for new investigators attempting to master a discipline and scientists who seek to cross disciplinary borders. In this paper, we propose metrics of influence, coverage and connectivity for scientific literature. We use these metrics to create structured summaries of information, which we call metro maps. Most importantly, metro maps explicitly show the relations between papers in a way which captures developments in the field. Pilot user studies demonstrate that our method helps researchers acquire new knowledge efficiently: map users achieved better precision and recall scores and found more seminal papers while performing fewer searches.
Dafna Shahaf, Carlos Guestrin, Eric Horvitz
KDD3
2012 Patient Risk Stratification for Hospital-Associated C. diff as a Time-Series Classification Task
abstract
A patient's risk for adverse events is affected by temporal processes including the nature and timing of diagnostic and therapeutic activities, and the overall evolution of the patient's pathophysiology over time. Yet many investigators ignore this temporal aspect when modeling patient risk, considering only the patient's current or aggregate state. We explore representing patient risk as a time series. In doing so, patient risk stratification becomes a time-series classification task. The task differs from most applications of time-series analysis, like speech processing, since the time series itself must first be extracted. Thus, we begin by defining and extracting approximate \textit{risk processes}, the evolving approximate daily risk of a patient. Once obtained, we use these signals to explore different approaches to time-series classification with the goal of identifying high-risk patterns. We apply the classification to the specific task of identifying patients at risk of testing positive for hospital acquired colonization with \textit{Clostridium Difficile}. We achieve an area under the receiver operating characteristic curve of 0.79 on a held-out set of several hundred patients. Our two-stage approach to risk stratification outperforms classifiers that consider only a patient's current state (p$<$0.05).
Jenna Wiens, John V. Guttag, Eric Horvitz
NIPS3
2012 Market user interface design
abstract
Despite the pervasiveness of markets in our lives, little is known about the role of user interfaces (UIs) in promoting good decisions in market domains. How does the way we display market information to end users, and the set of choices we offer, influence users' decisions? In this paper, we introduce a new research agenda on "market user interface design." Our goal is to find the optimal market UI, taking into account that users incur cognitive costs and are boundedly rational. Via lab experiments we systematically explore the market UI design space, and we study the automatic optimization of market UIs given a behavioral (quantal response) model of user behavior. Surprisingly, we find that the behaviorally-optimized UI performs worse than the standard UI, suggesting that the quantal response model did not predict user behavior well. Subsequently, we identify important behavioral factors that are missing from the user model, including loss aversion and position effects, which motivates follow-up studies. Furthermore, we find significant differences between individual users in terms of rationality. This suggests future research on personalized UI designs, with interfaces that are tailored towards each individual user's needs, capabilities, and preferences.
Sven Seuken, David C. Parkes, Eric Horvitz, Kamal Jain, Mary Czerwinski, Desney S. Tan
EC3
2012 Studies of the onset and persistence of medical concerns in search logs
abstract
The Web provides a wealth of information about medical symptoms and disorders. Although this content is often valuable to consumers, studies have found that interaction with Web content may heighten anxiety and stimulate healthcare utilization. We present a longitudinal log-based study of medical search and browsing behavior on the Web. We characterize how users focus on particular medical concerns and how concerns persist and influence future behavior, including changes in focus of attention in searching and browsing for health information. We build and evaluate models that predict transitions from searches on symptoms to searches on health conditions, and escalations from symptoms to serious illnesses. We study the influence that the prior onset of concerns may have on future behavior, including sudden shifts back to searching on the concern amidst other searches. Our findings have implications for refining Web search and retrieval to support people pursuing diagnostic information.
Ryen W. White, Eric Horvitz
SIGIR2
2012 Crowdsourcing the acquisition of natural language corpora: Methods and observations
abstract
We study the opportunity for using crowdsourcing methods to acquire language corpora for use in natural language processing systems. Specifically, we empirically investigate three methods for eliciting natural language sentences that correspond to a given semantic form. The methods convey frame semantics to crowd workers by means of sentences, scenarios, and list-based descriptions. We discuss various performance measures of the crowdsourcing process, and analyze the semantic correctness, naturalness, and biases of the collected language. We highlight research challenges and directions in applying these methods to acquire corpora for natural language processing applications.
William Yang Wang, Dan Bohus, Ece Kamar, Eric Horvitz
SLT4
2012 Modeling and predicting behavioral dynamics on the web
abstract
User behavior on the Web changes over time. For example, the queries that people issue to search engines, and the underlying informational goals behind the queries vary over time. In this paper, we examine how to model and predict this temporal user behavior. We develop a temporal modeling framework adapted from physics and signal processing that can be used to predict time-varying user behavior using smoothing and trends. We also explore other dynamics of Web behaviors, such as the detection of periodicities and surprises. We develop a learning procedure that can be used to construct models of users' activities based on features of current and historical behaviors. The results of experiments indicate that by using our framework to predict user behavior, we can achieve significant improvements in prediction compared to baseline models that weight historical evidence the same for all queries. We also develop a novel learning algorithm that explicitly learns when to apply a given prediction model among a set of such models. Our improved temporal modeling of user behavior can be used to enhance query suggestions, crawling policies, and result ranking.
Kira Radinsky, Krysta M. Svore, Susan T. Dumais, Jaime Teevan, Alex Bocharov, Eric Horvitz
WWW6
2012 Trains of thought: generating information maps
abstract
When information is abundant, it becomes increasingly difficult to fit nuggets of knowledge into a single coherent picture. Complex stories spaghetti into branches, side stories, and intertwining narratives. In order to explore these stories, one needs a map to navigate unfamiliar territory. We propose a methodology for creating structured summaries of information, which we call metro maps. Our proposed algorithm generates a concise structured set of documents maximizing coverage of salient pieces of information. Most importantly, metro maps explicitly show the relations among retrieved pieces in a way that captures story development. We first formalize characteristics of good maps and formulate their construction as an optimization problem. Then we provide efficient methods with theoretical guarantees for generating maps. Finally, we integrate user interaction into our framework, allowing users to alter the maps to better reflect their interests. Pilot user studies with a real-world dataset demonstrate that the method is able to produce maps which help users acquire knowledge efficiently.
Dafna Shahaf, Carlos Guestrin, Eric Horvitz
WWW3
2011 Peripheral computing during presentations: perspectives on costs and preferences
abstract
Despite the common use of mobile computing devices to communicate and access information, the effects of peripheral computing tasks on people's attention is not well understood. Studies that have identified consequences of multitasking in diverse domains have largely focused on influences on productivity. We have yet to understand perceptions and preferences regarding the use of computing devices for potentially extraneous tasks in settings such as presentations at seminars and colloquia. We explore costs and attitudes about the use of computing devices by people attending presentations. We find that audience members who use devices believe that they are missing content being presented and are concerned about social costs. Other attendees report being less offended by multitasking around them than the device users may realize.
Shamsi T. Iqbal, Jonathan Grudin, Eric Horvitz
CHI3
2011 Hang on a sec!: effects of proactive mediation of phone conversations while driving
abstract
Conversing on cell phones while driving is a risky, yet commonplace activity. State legislatures in the U.S. have enacted rules that limit hand-held phone conversations while driving but that allow for hands-free conversations. However, studies have demonstrated that the cognitive load of conversation is a significant source of distraction that increases the likelihood of accidents. We explore in a controlled study with a driving simulator the effectiveness of proactive alerting and mediation of communications during phone conversations while driving. We study the use of auditory messages indicating upcoming critical road conditions and placing calls on hold. We found that such actions reduce driving errors and that alerts sharing details about situations were more effective than general alerts. Drivers found such a system valuable in most situations for maintaining driving safety. These results provide evidence that context-sensitive mediation systems could play a valuable role in focusing drivers' attention on the road during phone conversations.
Shamsi T. Iqbal, Eric Horvitz, Yun-Cheng Ju, Ella Mathews
CHI2
2011 Characterizing patient-friendly "micro-explanations"of medical events
abstract
Patients' basic understanding of clinical events has been shown to dramatically improve patient care. We propose that the automatic generation of very short micro-explanations, suitable for real-time delivery in clinical settings, can transform patient care by giving patients greater awareness of key events in their electronic medical record. We present results of a survey study indicating that it may be possible to automatically generate such explanations by extracting individual sentences from consumer-facing Web pages. We further inform future work by characterizing physician and non-physician responses to a variety of Web-extracted explanations of medical lab tests.
Lauren Wilcox, Dan Morris 0001, Desney S. Tan, Justin Gatewood, Eric Horvitz
CHI5
2011 Decisions about turns in multiparty conversation: from perception to action
abstract
We present a decision-theoretic approach for guiding turn taking in a spoken dialog system operating in multiparty settings. The proposed methodology couples inferences about multiparty conversational dynamics with assessed costs of different outcomes, to guide turn-taking decisions. Beyond considering uncertainties about outcomes arising from evidential reasoning about the state of a conversation, we endow the system with awareness and methods for handling uncertainties stemming from computational delays in its own perception and production. We illustrate via sample cases how the proposed approach makes decisions, and we investigate the behaviors of the proposed methods via a retrospective analysis on logs collected in a multiparty interaction study.
Dan Bohus, Eric Horvitz
ICMI2
2011 Bricks, arches, and cathedrals: reflections on paths to deeper human-computer symbioses
abstract
I will share thoughts about achievements to date and opportunities moving forward on harnessing advances in machine intelligence to enable new forms of competent and fluid human-computer collaboration. I will discuss the promise of assembling key building blocks of methods in machine perception, learning, and inference into larger integrative solutions that draw upon a symphony of skills and that operate over extended periods of time. Explorations of such integrative machine intelligence frames research on the coordination of multiple components for sensing and reasoning to create higher-level functionalities and abstractions. I will discuss the promise of these efforts to advance us toward realizing dreams of deeper human-computer symbioses as imagined by such visionaries as Licklider and Engelbart.
Eric Horvitz
IUI1
2011 Multiparty Turn Taking in Situated Dialog: Study, Lessons, and Directions
Dan Bohus, Eric Horvitz
SIGDIAL Conference2
2011 Intentions and attention in exploratory health search
abstract
We study information goals and patterns of attention in explorato-ry search for health information on the Web, reporting results of a large-scale log-based study. We examine search activity associated with the goal of diagnosing illness from symptoms versus more general information-seeking about health and illness. We decom-pose exploratory health search into evidence-based and hypothe-sis-directed information seeking. Evidence-based search centers on the pursuit of details and relevance of signs and symptoms. Hypothesis-directed search includes the pursuit of content on one or more illnesses, including risk factors, treatments, and therapies for illnesses, and on the discrimination among different diseases under the uncertainty that exists in advance of a confirmed diag-nosis. These different goals of exploratory health search are not independent, and transitions can occur between them within or across search sessions. We construct a classifier that identifies medically-related search sessions in log data. Given a set of search sessions flagged as health-related, we show how we can identify different intentions persisting as foci of attention within those sessions. Finally, we discuss how insights about foci dynamics can help us better understand exploratory health search behavior and better support health search on the Web.
Marc-Allen Cartright, Ryen W. White, Eric Horvitz
SIGIR3
2011 The effects of choice in routing relevance judgments
abstract
The emergence of human computation systems, including Mechanical Turk and games with a purpose, has made it feasible to distribute relevance judgment tasks to workers over the Web. Most human computation systems assign tasks to individuals randomly, and such assignments may match workers with tasks that they may be unqualified or unmotivated to perform. We compare two groups of workers, those given a choice of queries to judge versus those who are not, in terms of their self-rated competence and their actual performance. Results show that when given a choice of task, workers choose ones for which they have greater expertise, interests, confidence, and understanding.
Edith Law, Paul N. Bennett, Eric Horvitz
SIGIR3
2010 Generalized Task Markets for Human and Machine Computation
abstract
We discuss challenges and opportunities for developing generalized task markets where human and machine intelligence are enlisted to solve problems, based on a consideration of the competencies, availabilities, and pricing of different problem-solving resources. The approach couples human computation with machine learning and planning, and is aimed at optimizing the flow of subtasks to people and to computational problem solvers. We illustrate key ideas in the context of Lingua Mechanica, a project focused on harnessing human and machine translation skills to perform translation among languages. We present infrastructure and methods for enlisting and guiding human and machine computation for language translation, including details about the hardness of generating plans for assigning tasks to solvers. Finally, we discuss studies performed with machine and human solvers, focusing on components of a Lingua Mechanica prototype.
Dafna Shahaf, Eric Horvitz
AAAI2
2010 What's your idea?: a case study of a grassroots innovation pipeline within a large software company
abstract
Establishing a grassroots innovation pipeline has come to the fore as strategy for nurturing innovation within large organizations. A key element of such pipelines is the use of an idea management system that enables and encourages community ideation on defined business problems. The value of these systems can be highly sensitive to design choices, as different designs may influence participation. We report the results of a case study examining the use of one particular idea management system and pipeline. We analyzed the content, interaction, and participation from three creativity challenges organized via the pipeline and conducted interviews with users to uncover motivations for participating and perceptions of the outcomes. Additional interviews were conducted with senior managers to learn about the objectives, successes, and unique nature of the pipeline. From the results, we formulate recommendations for improving the design of idea management systems and execution of the pipelines within organizations.
Brian P. Bailey, Eric Horvitz
CHI2
2010 Cars, calls, and cognition: investigating driving and divided attention
abstract
Conversing on cell phones while driving an automobile is a common practice. We examine the interference of the cognitive load of conversational dialog with driving tasks, with the goal of identifying better and worse times for conversations during driving. We present results from a controlled study involving 18 users using a driving simulator. The driving complexity and conversation type were manipulated in the study, and performance was measured for factors related to both the primary driving task and secondary conversation task. Results showed significant interactions between the primary and secondary tasks, where certain combinations of complexity and conversations were found especially detrimental to driving. We present the studies and analyses and relate the findings to prior work on multiple resource models of cognition. We discuss how the results can frame thinking about policies and technologies aimed at enhancing driving safety.
Shamsi T. Iqbal, Yun-Cheng Ju, Eric Horvitz
CHI3
2010 Interactive optimization for steering machine classification
abstract
Interest has been growing within HCI on the use of machine learning and reasoning in applications to classify such hidden states as user intentions, based on observations. HCI researchers with these interests typically have little expertise in machine learning and often employ toolkits as relatively fixed "black boxes" for generating statistical classifiers. However, attempts to tailor the performance of classifiers to specific application requirements may require a more sophisticated understanding and custom-tailoring of methods. We present ManiMatrix, a system that provides controls and visualizations that enable system builders to refine the behavior of classification systems in an intuitive manner. With ManiMatrix, users directly refine parameters of a confusion matrix via an interactive cycle of re-classification and visualization. We present the core methods and evaluate the effectiveness of the approach in a user study. Results show that users are able to quickly and effectively modify decision boundaries of classifiers to tai-lor the behavior of classifiers to problems at hand.
Ashish Kapoor, Bongshin Lee, Desney S. Tan, Eric Horvitz
CHI4
2010 Notifications and awareness: a field study of alert usage and preferences
abstract
Desktop notifications are designed to provide awareness of information while a user is attending to a primary task. Unfortunately the awareness can come with the price of disruption to the focal task. We review results of a field study on the use and perceived value of email notifications in the workplace. We recorded users' interactions with software applications for two weeks and studied how notifications or their forced absence influenced users' quest for awareness of new email arrival, as well as the impact of notifications on their overall task focus. Results showed that users view notifications as a mechanism to provide passive awareness rather than a trigger to switch tasks. Turing off notifications cause some users to self interrupt more to explicitly monitor email arrival, while others appear to be able to better focus on their tasks. Users acknowledge notifications as disruptive, yet opt for them because of their perceived value in providing awareness.
Shamsi T. Iqbal, Eric Horvitz
CSCW2
2010 Predicting escalations of medical queries based on web page structure and content
abstract
Logs of users' searches on Web health topics can exhibit signs of escalation of medical concerns, where initial queries about common symptoms are followed by queries about serious, rare illnesses. We present an effort to predict such escalations based on the structure and content of pages encountered during medical search sessions. We construct and then characterize the performance of classifiers that predict whether an escalation will occur after the access of a page. Our findings have implications for ranking algorithms and the design of search interfaces.
Ryen W. White, Eric Horvitz
SIGIR2
2010 A Utility-Theoretic Approach to Privacy in Online Services
abstract
Online offerings such as web search, news portals, and e-commerce applications face the challenge of providing high-quality service to a large, heterogeneous user base. Recent efforts have highlighted the potential to improve performance by introducing methods to personalize services based on special knowledge about users and their context. For example, a user's demographics, location, and past search and browsing may be useful in enhancing the results offered in response to web search queries. However, reasonable concerns about privacy by both users, providers, and government agencies acting on behalf of citizens, may limit access by services to such information. We introduce and explore an economics of privacy in personalization, where people can opt to share personal information, in a standing or on-demand manner, in return for expected enhancements in the quality of an online service. We focus on the example of web search and formulate realistic objective functions for search efficacy and privacy. We demonstrate how we can find a provably near-optimal optimization of the utility-privacy tradeoff in an efficient manner. We evaluate our methodology on data drawn from a log of the search activity of volunteer participants. We separately assess users’ preferences about privacy and utility via a large-scale survey, aimed at eliciting preferences about peoples’ willingness to trade the sharing of personal data in returns for gains in search efficiency. We show that a significant level of personalization can be achieved using a relatively small amount of information about users.
Andreas Krause 0001, Eric Horvitz
J. Artif. Intell. Res.2
2010 Personalization via friendsourcing
abstract
When information is known only to friends in a social network, traditional crowdsourcing mechanisms struggle to motivate a large enough user population and to ensure accuracy of the collected information. We thus introduce friendsourcing, a form of crowdsourcing aimed at collecting accurate information available only to a small, socially-connected group of individuals. Our approach to friendsourcing is to design socially enjoyable interactions that produce the desired information as a side effect. We focus our analysis around Collabio, a novel social tagging game that we developed to encourage friends to tag one another within an online social network. Collabio encourages friends, family, and colleagues to generate useful information about each other. We describe the design space of incentives in social tagging games and evaluate our choices by a combination of usage log analysis and survey data. Data acquired via Collabio is typically accurate and augments tags that could have been found on Facebook or the Web. To complete the arc from data collection to application, we produce a trio of prototype applications to demonstrate how Collabio tags could be utilized: an aggregate tag cloud visualization, a personalized RSS feed, and a question and answer system. The social data powering these applications enables them to address needs previously difficult to support, such as question answering for topics comprehensible only to a few of a user's friends.
Michael S. Bernstein, Desney S. Tan, Greg Smith, Mary Czerwinski, Eric Horvitz
ACM Trans. Comput. Hum. Interact.5
2010 Potential for personalization
abstract
Current Web search tools do a good job of retrieving documents that satisfy the most common intentions associated with a query, but do not do a very good job of discerning different individuals' unique search goals. We explore the variation in what different people consider relevant to the same query by mining three data sources: (1) explicit relevance judgments, (2) clicks on search results (a behavior-based implicit measure of relevance), and (3) the similarity of desktop content to search results (a content-based implicit measure of relevance). We find that people's explicit judgments for the same queries differ greatly. As a result, there is a large gap between how well search engines could perform if they were to tailor results to the individual, and how well they currently perform by returning results designed to satisfy everyone. We call this gap the potential for personalization . The two implicit indicators we studied provide complementary value for approximating this variation in result relevance among people. We discuss several uses of our findings, including a personalized search system that takes advantage of the implicit measures by ranking personally relevant results more highly and improving click-through rates.
Jaime Teevan, Susan T. Dumais, Eric Horvitz
ACM Trans. Comput. Hum. Interact.3
2009 Experiences with Web Search on Medical Concerns and Self Diagnosis
Ryen W. White, Eric Horvitz
AMIA2
2009 Dialog in the open world: platform and applications
abstract
We review key challenges of developing spoken dialog systems that can engage in interactions with one or multiple participants in relatively unconstrained environments. We outline a set of core competencies for open-world dialog, and describe three prototype systems. The systems are built on a common underlying conversational framework which integrates an array of predictive models and component technologies, including speech recognition, head and pose tracking, probabilistic models for scene analysis, multiparty engagement and turn taking, and inferences about user goals and activities. We discuss the current models and showcase their function by means of a sample recorded interaction, and we review results from an observational study of open-world, multiparty dialog in the wild.
Dan Bohus, Eric Horvitz
ICMI2
2009 Collaboration and Shared Plans in the Open World: Studies of Ridesharing
Ece Kamar, Eric Horvitz
IJCAI2
2009 Investigations of Continual Computation
Dafna Shahaf, Eric Horvitz
IJCAI2
2009 Breaking Boundaries Between Induction Time and Diagnosis Time Active Information Acquisition
abstract
There has been a clear distinction between induction or training time and diagnosis time active information acquisition. While active learning during induction focuses on acquiring data that promises to provide the best classification model, the goal at diagnosis time focuses completely on next features to observe about the test case at hand in order to make better predictions about the case. We introduce a model and inferential methods that breaks this distinction. The methods can be used to extend case libraries under a budget but, more fundamentally, provide a framework for guiding agents to collect data under scarce resources, focused by diagnostic challenges. This extension to active learning leads to a new class of policies for real-time diagnosis, where recommended information-gathering sequences include actions that simultaneously seek new data for the case at hand and for cases in the training set.
Ashish Kapoor, Eric Horvitz
NIPS2
2009 Models for Multiparty Engagement in Open-World Dialog
Dan Bohus, Eric Horvitz
SIGDIAL Conference2
2009 Learning to Predict Engagement with a Spoken Dialog System in Open-World Settings
Dan Bohus, Eric Horvitz
SIGDIAL Conference2
2009 Collabio: a game for annotating people within social networks
abstract
We present Collabio, a social tagging game within an online social network that encourages friends to tag one another. Collabio's approach of incentivizing members of the social network to generate information about each other produces personalizing information about its users. We report usage log analysis, survey data, and a rating exercise demonstrating that Collabio tags are accurate and augment information that could have been scraped online.
Michael S. Bernstein, Desney S. Tan, Greg Smith, Mary Czerwinski, Eric Horvitz
UIST5
2009 Cyberchondria: Studies of the escalation of medical concerns in Web search
abstract
The World Wide Web provides an abundant source of medical information. This information can assist people who are not healthcare professionals to better understand health and illness, and to provide them with feasible explanations for symptoms. However, the Web has the potential to increase the anxieties of people who have little or no medical training, especially when Web search is employed as a diagnostic procedure. We use the term cyberchondria to refer to the unfounded escalation of concerns about common symptomatology, based on the review of search results and literature on the Web. We performed a large-scale, longitudinal, log-based study of how people search for medical information online, supported by a survey of 515 individuals' health-related search experiences. We focused on the extent to which common, likely innocuous symptoms can escalate into the review of content on serious, rare conditions that are linked to the common symptoms. Our results show that Web search engines have the potential to escalate medical concerns. We show that escalation is associated with the amount and distribution of medical content viewed by users, the presence of escalatory terminology in pages visited, and a user's predisposition to escalate versus to seek more reasonable explanations for ailments. We also demonstrate the persistence of postsession anxiety following escalations and the effect that such anxieties can have on interrupting user's activities across multiple sessions. Our findings underscore the potential costs and challenges of cyberchondria and suggest actionable design implications that hold opportunity for improving the search and navigation experience for people turning to the Web to interpret common symptoms.
Ryen W. White, Eric Horvitz
ACM Trans. Inf. Syst.2
2008 A Utility-Theoretic Approach to Privacy and Personalization
Andreas Krause 0001, Eric Horvitz
AAAI2
2008 Experience sampling for building predictive user models: a comparative study
abstract
Experience sampling has been employed for decades to collect assessments of subjects' intentions, needs, and affective states. In recent years, investigators have employed automated experience sampling to collect data to build predictive user models. To date, most procedures have relied on random sampling or simple heuristics. We perform a comparative analysis of several automated strategies for guiding experience sampling, spanning a spectrum of sophistication, from a random sampling procedure to increasingly sophisticated active learning. The more sophisticated methods take a decision-theoretic approach, centering on the computation of the expected value of information of a probe, weighing the cost of the short-term disruptiveness of probes with their benefits in enhancing the long-term performance of predictive models. We test the different approaches in a field study, focused on the task of learning predictive models of the cost of interruption.
Ashish Kapoor, Eric Horvitz
CHI2
2008 Understanding the relationship between searchers' queries and information goals
abstract
We describe results from Web search log studies aimed at elucidating user behaviors associated with queries and destination URLs that appear with different frequencies. We note the diversity of information goals that searchers have and the differing ways that goals are specified. We examine rare and common information goals that are specified using rare or common queries. We identify several significant differences in user behavior depending on the rarity of the query and the destination URL. We find that searchers are more likely to be successful when the frequencies of the query and destination URL are similar. We also establish that the behavioral differences observed for queries and goals of varying rarity persist even after accounting for potential confounding variables, including query length, search engine ranking, session duration, and task difficulty. Finally, using an information-theoretic measure of search difficulty, we show that the benefits obtained by search and navigation actions depend on the frequency of the information goal.
Doug Downey, Susan T. Dumais, Daniel J. Liebling, Eric Horvitz
CIKM4
2008 Complementary computing for visual tasks: Meshing computer vision with human visual processing
abstract
We explore the opportunity to harness electroencephalograph (EEG) signals generated during human visual processing to enhance computer vision systems. We review the challenging task of categorizing objects, such as faces, in images and then describe methods that can be used to combine the complementary competencies of human and machine computation to achieve improved recognition performance. We present the results of several experiments where brain signals, recorded from people examining images, are used to enhance the performance of vision systems on categorization tasks. We find that significant gains in classification accuracy can be achieved with the human-aided vision systems.
Ashish Kapoor, Desney S. Tan, Pradeep Shenoy, Eric Horvitz
FG4
2008 Toward Community Sensing
abstract
A great opportunity exists to fuse information from populations of privately-held sensors to create useful sensing applications. For example, GPS devices, embedded in cellphones and automobiles, might one day be employed as distributed networks of velocity sensors for traffic monitoring and routing. Unfortunately, privacy and resource considerations limit access to such data streams. We describe principles of community sensing that offer mechanisms for sharing data from privately held sensors. The methods take into account the likely availability of sensors, the context-sensitive value of sensor information, based on models of phenomena and demand, and sensor owners' preferences about privacy and resource usage. We present efficient and well-characterized approximations of optimal sensing policies. We provide details on key principles of community sensing and highlight their use within a case study for road traffic monitoring.
Andreas Krause 0001, Eric Horvitz, Aman Kansal, Feng Zhao 0001
IPSN2
2008 Planetary-scale views on a large instant-messaging network
abstract
We present a study of anonymized data capturing a month of high-level communication activities within the whole of the Microsoft Messenger instant-messaging system. We examine characteristics and patterns that emerge from the collective dynamics of large numbers of people, rather than the actions and characteristics of individuals. The dataset contains summary properties of 30 billion conversations among 240 million people. From the data, we construct a communication graph with 180 million nodes and 1.3 billion undirected edges, creating the largest social network constructed and analyzed to date. We report on multiple aspects of the dataset and synthesized graph. We find that the graph is well-connected and robust to node removal. We investigate on a planetary-scale the oft-cited report that people are separated by "six degrees of separation" and find that the average path length among Messenger users is 6.6. We find that people tend to communicate more with each other when they have similar age, language, and location, and that cross-gender conversations are both more frequent and of longer duration than conversations with the same gender.
Jure Leskovec, Eric Horvitz
WWW2
2007 Disruption and recovery of computing tasks: field study, analysis, and directions
abstract
We report on a field study of the multitasking behavior of computer users focused on the suspension and resumption of tasks. Data was collected with a tool that logged users' interactions with software applications and their associated windows, as well as incoming instant messaging and email alerts. We describe methods, summarize results, and discuss design guidelines suggested by the findings.
Shamsi T. Iqbal, Eric Horvitz
CHI2
2007 Towards mixed-initiative access control
abstract
The difficult task of providing access to shared objects is, typically, carried out individually by access authorizers. We motivate and explain here the idea of using distributed collaborative environments to perform this activity. In these environments, the initiative in distributing access rights to shared objects can be taken by information guardians, information consumers, and tools that act as agents of the guardians and consumers. Information consumers are responsible for sending access requests to information guardians; their agents (partially or completely) automate this task for them. Information guardians are responsible for authorizing accesses; their agents automate this task for them. The agents interact with collaborative and communication tools, which must be extended to support the new access-control paradigm.
Prasun Dewan, Jonathan Grudin, Eric Horvitz
CollaborateCom3
2007 Models of Searching and Browsing: Languages, Studies, and Application
Doug Downey, Susan T. Dumais, Eric Horvitz
IJCAI3
2007 Selective Supervision: Guiding Supervised Learning with Decision-Theoretic Active Learning
Ashish Kapoor, Eric Horvitz, Sumit Basu
IJCAI2
2007 S3: Storable, Shareable Search
Meredith Ringel Morris, Eric Horvitz
INTERACT (1)2
2007 Heads and tails: studies of web search with common and rare queries
abstract
A large fraction of queries submitted to Web search enginesoccur very infrequently. We describe search log studiesaimed at elucidating behaviors associated with rare andcommon queries. We present several analyses and discussresearch directions.
Doug Downey, Susan T. Dumais, Eric Horvitz
SIGIR3
2007 Characterizing the value of personalizing search
abstract
We investigate the diverse goals that people have when they issue the same query to a search engine, and the ability of current search engines to address such diversity. We quantify the potential value of personalizing search results based on this analysis. Great variance was found in the results that different individuals rated as relevant for the same query -- even when the same information goal was expressed. Our analysis suggests that while search engines do a good job of ranking results to maximize global happiness, they do not do a very good job for specific individuals.
Jaime Teevan, Susan T. Dumais, Eric Horvitz
SIGIR3
2007 On Discarding, Caching, and Recalling Samples in Active Learning
Ashish Kapoor, Eric Horvitz
UAI2
2007 SearchTogether: an interface for collaborative web search
abstract
Studies of search habits reveal that people engage in many search tasks involving collaboration with others, such as travel planning, organizing social events, or working on a homework assignment. However, current Web search tools are designed for a single user, working alone. We introduce SearchTogether, a prototype that enables groups of remote users to synchronously or asynchronously collaborate when searching the Web. We describe an example usage scenario, and discuss the ways SearchTogether facilitates collaboration by supporting awareness, division of labor, and persistence. We then discuss the findings of our evaluation of SearchTogether, analyzing which aspects of its design enabled successful collaboration among study participants.
Meredith Ringel Morris, Eric Horvitz
UIST2
2007 Web projections: learning from contextual subgraphs of the web
abstract
Graphical relationships among Web pages have been exploited inmethods for ranking search results. To date, specific graphicalproperties have been used in these analyses. We introduce a WebProjection methodology that generalizes prior efforts of graphicalrelationships of the web in several ways. With the approach, wecreate subgraphs by projecting sets of pages and domains onto thelarger web graph, and then use machine learning to constructpredictive models that consider graphical properties as evidence. Wedescribe the method and then present experiments that illustrate theconstruction of predictive models of search result quality and userquery reformulation.
Jure Leskovec, Susan T. Dumais, Eric Horvitz
WWW3
2007 Complementary computing: policies for transferring callers from dialog systems to human receptionists
Eric Horvitz, Tim Paek
User Model. User Adapt. Interact.1
2006 Trip Router with Individualized Preferences (TRIP): Incorporating Personalization into Route Planning
Julia Letchner, John Krumm, Eric Horvitz
AAAI3
2006 Predestination: Inferring Destinations from Partial Trajectories
John Krumm, Eric Horvitz
UbiComp2
2006 Considering Cost Asymmetry in Learning Classifiers
abstract
Receiver Operating Characteristic (ROC) curves are a standard way to display the performance of a set of binary classifiers for all feasible ratios of the costs associated with false positives and false negatives. For linear classifiers, the set of classifiers is typically obtained by training once, holding constant the estimated slope and then varying the intercept to obtain a parameterized set of classifiers whose performances can be plotted in the ROC plane. We consider the alternative of varying the asymmetry of the cost function used for training. We show that the ROC curve obtained by varying both the intercept and the asymmetry, and hence the slope, always outperforms the ROC curve obtained by varying only the intercept. In addition, we present a path-following algorithm for the support vector machine (SVM) that can compute efficiently the entire ROC curve, and that has the same computational complexity as training a single classifier. Finally, we provide a theoretical analysis of the relationship between the asymmetric cost model assumed when training a classifier and the cost model assumed in applying the classifier. In particular, we show that the mismatch between the step function used for testing and its convex upper bounds, usually used for training, leads to a provable and quantifiable difference around extreme asymmetries.
Francis R. Bach, David Heckerman, Eric Horvitz
J. Mach. Learn. Res.3
2005 Personalizing search via automated analysis of interests and activities
abstract
We formulate and study search algorithms that consider a user's prior interactions with a wide variety of content to personalize that user's current Web search. Rather than relying on the unrealistic assumption that people will precisely specify their intent when searching, we pursue techniques that leverage implicit information about the user's interests. This information is used to re-rank Web search results within a relevance feedback framework. We explore rich models of user interests, built from both search-related information, such as previously issued queries and previously visited Web pages, and other information about the user such as documents and email the user has read and created. Our research suggests that rich representations of the user and the corpus are important for personalization, but that it is possible to approximate these representations and provide efficient client-side algorithms for personalizing search. We show that such personalization algorithms can significantly improve on current Web search.
Jaime Teevan, Susan T. Dumais, Eric Horvitz
SIGIR3
2005 Prediction, Expectation, and Surprise: Methods, Designs, and Study of a Deployed Traffic Forecasting Service
Eric Horvitz, Johnson Apacible, Raman Sarin, Lin Liao
UAI1
2005 Computing location from ambient FM radio signals [commercial radio station signals]
abstract
We present a method for computing the location of a device down to a radius of several miles within a greater metropolitan area by analyzing the signal strengths observed from commercial FM radio stations. The use of ambient commercial radio signals allows for wide coverage, both indoor and outdoor reception, client-side computing for privacy, and the feasibility of employing inexpensive, low-power measurement hardware. Our technique is based on a model for computing the likelihood of locations using both received signal strengths and information from a simulated signal strength map. Using simulated signal strengths relieves the burden of manually measuring signal strength as a function of location. We account for the inevitable measurement variations among devices by comparing rankings of radio stations by signal strength. Our experiments show we can measure location down to a median error of about 8 kilometers (5 miles) in the greater Seattle area by listening to seven different radio stations.
Adel Youssef, John Krumm, Ed Miller, Gerry Cermak, Eric Horvitz
WCNC5
2005 Selective perception policies for guiding sensing and computation in multimodal systems: A comparative analysis
Nuria Oliver, Eric Horvitz
Comput. Vis. Image Underst.2
2005 The Combination of Text Classifiers Using Reliability Indicators
Paul N. Bennett, Susan T. Dumais, Eric Horvitz
Inf. Retr.3
2005 Foreground and background interaction with sensor-enhanced mobile devices
abstract
Building on Buxton's foreground/background model, we discuss the importance of explicitly considering both foreground interaction and background interaction, as well as transitions between foreground and background, in the design and implementation of sensing techniques for sensor-enhanced mobile devices. Our view is that the foreground concerns deliberate user activity where the user is attending to the device, while the background is the realm of inattention or split attention, using naturally occurring user activity as an input that allows the device to infer or anticipate user needs. The five questions for sensing systems of Bellotti et al. [2002] proposed as a framework for this special issue, primarily address the foreground, but neglect critical issues with background sensing. To support our perspective, we discuss a variety of foreground and background sensing techniques that we have implemented for sensor-enhanced mobile devices, such as powering on the device when the user picks it up, sensing when the user is holding the device to his ear, automatically switching between portrait and landscape display orientations depending on how the user is holding the device, and scrolling the display using tilt. We also contribute system architecture issues, such as using the foreground/background model to handle cross-talk between multiple sensor-based interaction techniques, and theoretical perspectives, such as a classification of recognition errors based on explicitly considering transitions between the foreground and background. Based on our experiences, we propose design issues and lessons learned for foreground/background sensing systems.
Ken Hinckley, Jeffrey S. Pierce, Eric Horvitz, Mike Sinclair
ACM Trans. Comput. Hum. Interact.3
2004 The Backdoor Key: A Path to Understanding Problem Hardness
Yongshao Ruan, Henry A. Kautz, Eric Horvitz
AAAI3
2004 ZoneZoom: map navigation for smartphones with recursive view segmentation
abstract
ZoneZoom is an input technique that lets users traverse large information spaces on smartphones. Our technique ZoneZoom, segments a given view of an information space into nine sub-segments, each of which is mapped to a key on the number keypad of the smartphone. This segmentation can be hand-crafted by the information space author or dynamically created at run-time. ZoneZoom supports "spring-loaded" view shifting which allows users to easily "glance" at nearby areas and then quickly return to their current view. Our ZoneZoom technique lets users gain an overview and compare information from different parts of a dataset. SmartPhlow is an optimized application for browsing a map of local-area road traffic conditions.
Daniel C. Robbins, Edward Cutrell, Raman Sarin, Eric Horvitz
AVI4
2004 Scalable Fabric: flexible task management
abstract
Our studies have shown that as displays become larger, users leave more windows open for easy multitasking. A larger number of windows, however, may increase the time that users spend arranging and switching between tasks. We present Scalable Fabric, a task management system designed to address problems with the proliferation of open windows on the PC desktop. Scalable Fabric couples window management with a flexible visual representation to provide a focus-plus-context solution to desktop complexity. Users interact with windows in a central focus region of the display in a normal manner, but when a user moves a window into the periphery, it shrinks in size, getting smaller as it nears the edge of the display. The window "minimize" action is redefined to return the window to its preferred location in the periphery, allowing windows to remain visible when not in use. Windows in the periphery may be grouped together into named tasks, and task switching is accomplished with a single mouse click. The spatial arrangement of tasks leverages human spatial memory to make task switching easier. We review the evolution of Scalable Fabric over three design iterations, including discussion of results from two user studies that were performed to compare the experience with Scalable Fabric to that of the Microsoft Windows XP TaskBar.
George G. Robertson, Eric Horvitz, Mary Czerwinski, Patrick Baudisch, Dugald Ralph Hutchings, Brian Meyers, Daniel C. Robbins, Greg Smith
AVI2
2004 A diary study of task switching and interruptions
abstract
We report on a diary study of the activities of information workers aimed at characterizing how people interleave multiple tasks amidst interruptions. The week-long study revealed the type and complexity of activities performed, the nature of the interruptions experienced, and the difficulty of shifting among numerous tasks. We present key findings from the diary study and discuss implications of the findings. Finally, we describe promising directions in the design of software tools for task management, motivated by the findings. Author Keywords Multitasking, diary study, task switching, interruptions, information worker, office and workplace. ACM Classification Keywords H5.m. Information interfaces and presentation (e.g., HCI).
Mary Czerwinski, Eric Horvitz, Susan Wilhite
CHI2
2004 BusyBody: creating and fielding personalized models of the cost of interruption
abstract
Interest has been growing in opportunities to build and deploy statistical models that can infer a computer user's current interruptability from computer activity and relevant contextual information. We describe a system that intermittently asks users to assess their perceived interruptability during a training phase and that builds decision-theoretic models with the ability to predict the cost of interrupting the user. The models are used at run-time to compute the expected cost of interruptions, providing a mediator for incoming notifications, based on a consideration of a user's current and recent history of computer activity, meeting status, location, time of day, and whether a conversation is detected.
Eric Horvitz, Paul Koch, Johnson Apacible
CSCW1
2004 LOCADIO: Inferring Motion and Location from Wi-Fi Signal Strengths
abstract
Context is a critical ingredient of ubiquitous computing. While it is possible to use specialized sensors and beacons to measure certain aspects of a user's context, we are interested in what we can infer from using the existing 802.11 wireless network infrastructure that already exists in many places. The context parameters we infer are the location of a client (with a median error of 1.5 meters) and an indicator of whether or not the client is in motion (with a classification accuracy of 87%). Our system, called LOCADIO, uses Wi-Fi signal strengths from existing access points measured on the client to infer both pieces of context. For motion, we measure the variance of the signal strength of the strongest access point as input to a simple two-state hidden Markov model (HMM) for smoothing transitions between the inferred states of "still" and "moving". For location, we exploit the fact that Wi-Fi signal strengths vary with location, and we use another HMM on a graph of location nodes whose transition probabilities are a function of the building's floor plan, expected pedestrian speeds, and our still/moving inference. Our probabilistic approach to inferring context gives a convenient way of balancing noisy measured data such as signal strengths against our a priori assumptions about a user's behavior.
John Krumm, Eric Horvitz
MobiQuitous2
2004 Optimizing Automated Call Routing by Integrating Spoken Dialog Models with Queuing Models
Tim Paek, Eric Horvitz
HLT-NAACL2
2004 Implicit queries (IQ) for contextualized search
abstract
The Implicit Query (IQ) prototype is a system which automatically generates context-sensitive searches based on a user's current computing activities. In the demo, we show IQ running when users are reading or composing email. Queries are automatically generated by analyzing the email message, and results are presented in a small pane adjacent to the current window to provide peripheral awareness of related information.
Susan T. Dumais, Edward Cutrell, Raman Sarin, Eric Horvitz
SIGIR4
2004 Newsjunkie: providing personalized newsfeeds via analysis of information novelty
abstract
We present a principled methodology for filtering news stories by formal measures of information novelty, and show how the techniques can be usedto custom-tailor news feeds based on information that a user has already reviewed. We review methods for analyzing novelty and then describe Newsjunkie, a system that personalizes news for users by identifying the novelty of stories in the context of stories they have already reviewed. Newsjunkie employs novelty-analysis algorithms that represent articles as words and named entities. The algorithms analyze inter-andintra-document dynamics by considering how information evolves over timefrom article to article, as well as within individual articles. We review the results of a user study undertaken to gauge the value of the approachover legacy time-based review of newsfeeds, and also to compare the performance of alternate distance metrics that are used to estimate the dissimilarity between candidate new articles and sets of previously reviewed articles.
Evgeniy Gabrilovich, Susan T. Dumais, Eric Horvitz
WWW3
2004 Layered representations for learning and inferring office activity from multiple sensory channels
Nuria Oliver, Ashutosh Garg 0001, Eric Horvitz
Comput. Vis. Image Underst.3
2004 Actions, answers, and uncertainty: a decision-making perspective on Web-based question answering
David Azari, Eric Horvitz, Susan T. Dumais, Eric Brill
Inf. Process. Manag.2
2003 RightSPOT: A Novel Sense of Location for a Smart Personal Object
John Krumm, Gerry Cermak, Eric Horvitz
UbiComp3
2003 Learning and reasoning about interruption
abstract
We present methods for inferring the cost of interrupting users based on multiple streams of events including information generated by interactions with computing devices, visual and acoustical analyses, and data drawn from online calendars. Following a review of prior work on techniques for deliberating about the cost of interruption associated with notifications, we introduce methods for learning models from data that can be used to compute the expected cost of interruption for a user. We describe the Interruption Workbench, a set of event-capture and modeling tools. Finally, we review experiments that characterize the accuracy of the models for predicting interruption cost and discuss research directions.
Eric Horvitz, Johnson Apacible
ICMI1
2003 Selective perception policies for guiding sensing and computation in multimodal systems: a comparative analysis
abstract
Intensive computations required for sensing and processing perceptual information can impose significant burdens on personal computer systems. We explore several policies for selective perception in SEER, a multimodal system for recognizing office activity that relies on a layered Hidden Markov Model representation. We review our efforts to employ expected-value-of-information (EVI) computations to limit sensing and analysis in a context-sensitive manner. We discuss an implementation of a one-step myopic EVI analysis and compare the results of using the myopic EVI with a heuristic sensing policy that makes observations at different frequencies. Both policies are then compared to a random perception policy, where sensors are selected at random. Finally, we discuss the sensitivity of ideal perceptual actions to preferences encoded in utility models about information value and the cost of sensing.
Nuria Oliver, Eric Horvitz
ICMI2
2003 Milestones in Time: The Value of Landmarks in Retrieving Information from Personal Stores
Meredith Ringel Morris, Edward Cutrell, Susan T. Dumais, Eric Horvitz
INTERACT4
2003 Web-Based Question Answering: A Decision-Making Perspective
David Azari, Eric Horvitz, Susan T. Dumais, Eric Brill
UAI2
2002 Scope: providing awareness of multiple notifications at a glance
abstract
We describe the design and functionality of the Scope, a glanceable notification summarizer. The Scope is an information visualization designed to unify notifications and minimize distractions. It allows users to remain aware of notifications from multiple sources of information, including e-mail, instant messaging, information alerts, and appointments. The design employs a circular radar-like screen divided into sectors that group different kinds of notifications. The more urgent a notification is, the more centrally it is placed. Visual emphasis and annotation is used to reveal important properties of notifications. Several natural gestures allow users to zoom in on particular regions and to selectively drill down on items. We present key aspects of the Scope design, review the results of an initial user study, and describe the motivation and outcome of an iteration on the visual design.
Maarten van Dantzich, Daniel C. Robbins, Eric Horvitz, Mary Czerwinski
AVI3
2002 Restart Policies with Dependence among Runs: A Dynamic Programming Approach
Yongshao Ruan, Eric Horvitz, Henry A. Kautz
CP2
2002 Layered Representations for Human Activity Recognition
abstract
We present the use of layered probabilistic representations using hidden Markov models for performing sensing, learning, and inference at multiple levels of temporal granularity We describe the use of representation in a system that diagnoses states of a user's activity based on real-time streams of evidence from video, acoustic, and computer interactions. We review the representation, present an implementation, and report on experiments with the layered representation in an office-awareness application.
Nuria Oliver, Eric Horvitz, Ashutosh Garg 0001
ICMI2
2002 Uncertainty, intelligence, and interaction
abstract
Uncertainty about a user's knowledge, intentions, and attention is inescapable in human-computer interaction. I will survey challenges and opportunities of harnessing explicit representations of uncertainty and preferences in intelligent user interfaces. After reviewing representative projects at Microsoft, I will describe longer-term research directions aimed at embedding representation, inference, and learning under uncertainty more deeply into the fabric of computer systems and interfaces.
Eric Horvitz
IUI1
2002 Probabilistic combination of text classifiers using reliability indicators: models and results
abstract
The intuition that different text classifiers behave in qualitatively different ways has long motivated attempts to build a better metaclassifier via some combination of classifiers. We introduce a probabilistic method for combining classifiers that considers the context-sensitive reliabilities of contributing classifiers. The method harnesses reliability indicators---variables that provide a valuable signal about the performance of classifiers in different situations. We provide background, present procedures for building metaclassifiers that take into consideration both reliability indicators and classifier outputs, and review a set of comparative studies undertaken to evaluate the methodology.
Paul N. Bennett, Susan T. Dumais, Eric Horvitz
SIGIR3
2002 Coordinates: Probabilistic Forecasting of Presence and Availability
Eric Horvitz, Paul Koch, Carl Myers Kadie, Andy Jacobs
UAI1
2002 Web montage: a dynamic personalized start page
abstract
Despite the connotation of the words "browsing" and "surfing," web usage often follows routine patterns of access. However, few mechanisms exist to assist users with these routine tasks; bookmarks or portal sites must be maintained manually and are insensitive to the user's browsing context. To fill this void, we designed and implemented the montage system. A web montage is an ensemble of links and content fused into a single view. Such a coalesced view can be presented to the user whenever he or she opens the browser or returns to the start page. We pose a number of hypotheses about how users would interact with such a system, and test these hypotheses with a fielded user study. Our findings support some design decisions, such as using browsing context to tailor the montage, raise questions about others, and point the way toward future work.
Corin R. Anderson, Eric Horvitz
WWW2
2001 Using Machine Learning Techniques to Interpret WH-questions
abstract
We describe a set of supervised machine learning experiments centering on the construction of statistical models of WH-questions. These models, which are built from shallow linguistic features of questions, are employed to predict target variables which represent a user's informational goals. We report on different aspects of the predictive performance of our models, including the influence of various training and testing factors on predictive performance, and examine the relationships among the target variables.
Ingrid Zuckerman, Eric Horvitz
ACL2
2001 Notification, Disruption, and Memory: Effects of Messaging Interruptions on Memory and Performance
Edward Cutrell, Mary Czerwinski, Eric Horvitz
INTERACT3
2001 A Bayesian Approach to Tackling Hard Computational Problems
Eric Horvitz, Yongshao Ruan, Carla P. Gomes, Henry A. Kautz, Bart Selman, David Maxwell Chickering
UAI1
2001 Toward more sensitive mobile phones
abstract
Although cell phones are extremely useful, they can be annoying and distracting to owners and others nearby. We describe sensing techniques intended to help make mobile phones more polite and less distracting. For example, our phone's ringing quiets as soon as the user responds to an incoming call, and the ring mutes if the user glances at the caller ID and decides not to answer. We also eliminate the need to press a TALK button to answer an incoming call by recognizing if the user picks up the phone and listens to it.
Ken Hinckley, Eric Horvitz
UIST2
2001 Principles and applications of continual computation
Eric Horvitz
Artif. Intell.1
2001 Computational tradeoffs under bounded resources
Eric Horvitz, Shlomo Zilberstein
Artif. Intell.1
2000 A Normative Examination of Ensemble Learning Algorithms
David M. Pennock, Pedrito Maynard-Reid II, C. Lee Giles, Eric Horvitz
ICML4
2000 Deeplistener: harnessing expected utility to guide clarification dialog in spoken language systems
abstract
We describe research on endowing spoken language systems with the ability to consider the cost of misrecognition, and using that knowledge to guide clarification dialog about a user’s intentions. Our approach relies on coupling utility-directed policies for dialog with the ongoing Bayesian fusion of evidence obtained from multiple utterances recognized during an interaction. After describing the methodology, we review the operation of a prototype system called DeepListener. DeepListener considers evidence gathered about utterances over time to make decisions about the optimal dialog strategy or realworld action to take given uncertainties about a user’s intentions and the costs and benefits of different outcomes.
Eric Horvitz, Tim Paek
INTERSPEECH1
2000 Continuous listening for unconstrained spoken dialog
abstract
A major hindrance to rendering spoken dialog systems capable of ongoing, continuous listening without requiring a push-to-talk device is the problem of distinguishing speech which is intended for the system from that which is overheard. We present a decision-theoretic approach to this problem that exploits Bayesian models of spoken dialog at four levels of analysis within a domain-independent, multi-modal computational architecture called Quartet. We applied Quartet to the task of navigating PowerPoint slide shows during a spoken presentation in a prototype system called Presenter. We describe the runtime behavior of Presenter as well as the results of an experimental study comparing the performance of Presenter to human subjects in discriminating arbitrarily formed spoken requests for slide navigation during a recorded lecture.
Tim Paek, Eric Horvitz, Eric K. Ringger
INTERSPEECH2
2000 Uncertainty, Utility, and Understanding
Eric Horvitz
Intelligent Tutoring Systems1
2000 Conversation as Action Under Uncertainty
Tim Paek, Eric Horvitz
UAI2
2000 Collaborative Filtering by Personality Diagnosis: A Hybrid Memory and Model-Based Approach
David M. Pennock, Eric Horvitz, Steve Lawrence, C. Lee Giles
UAI2
2000 Sensing techniques for mobile interaction
abstract
We describe sensing techniques motivated by unique aspects of human-computer interaction with handheld devices in mobile settings.Special features of mobile interaction include changing orientation and position, changing venues, the use of computing as auxiliary to ongoing, real-world activities like talking to a colleague, and the general intimacy of use for such devices.We introduce and integrate a set of sensors into a handheld device, and demonstrate several new functionalities engendered by the sensors, such as recording memos when the device is held like a cell phone, switching between portrait and landscape display modes by holding the device in the desired orientation, automatically powering up the device when the user picks it up the device to start using it, and scrolling the display using tilt.We present an informal experiment, initial usability testing results, and user reactions to these techniques.
Ken Hinckley, Jeffrey S. Pierce, Mike Sinclair, Eric Horvitz
UIST4
1999 Principles of Mixed-Initiative User Interfaces
abstract
Recent debate has centered on the relative promise of focusing user-interface research on developing new metaphors and tools that enhance users' abilities to directly manipulate objects versus directing effort toward developing interface agents that provide automation. In this paper, we review principles that show promise for allowing engineers to enhance human---computer interaction through an elegant coupling of automated services with direct manipulation. Key ideas will be highlighted in terms of the LookOut system for scheduling and meeting management. Keywords Intelligent agents, direct manipulation, user modeling, probability, decision theory, UI design INTRODUCTION There has been debate among researchers about where great opportunities lay for innovating in the realm of human--- computer interaction [10]. One group of researchers has expressed enthusiasm for the development and application of new kinds of automated services, often referred to as interface "agents." The effo...
Eric Horvitz
CHI1
1999 Continual Computation Policies for Allocating Offline and Real-Time Resources
Eric Horvitz
IJCAI1
1999 Bridging Science and Applications (Panel)
abstract
No abstract available.
Jude W. Shavlik, Lawrence Birnbaum, William R. Swartout, Eric Horvitz, Barbara Hayes-Roth
IUI4
1999 Attention-Sensitive Alerting
Eric Horvitz, Andy Jacobs, David Hovel
UAI1
1998 Continual Computation Policies for Utility-Directed Prefetching
abstract
Article Free Access Share on Continual computation policies for utility-directed prefetching Author: Eric Horvitz Microsoft Research, One Microsoft way, Redmond, WA Microsoft Research, One Microsoft way, Redmond, WAView Profile Authors Info & Claims CIKM '98: Proceedings of the seventh international conference on Information and knowledge managementNovember 1998Pages 175–184https://doi.org/10.1145/288627.288655Published:01 November 1998Publication History 16citation275DownloadsMetricsTotal Citations16Total Downloads275Last 12 Months7Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Eric Horvitz
CIKM1
1998 Inferring Informational Goals from Free-Text Queries: A Bayesian Approach
David Heckerman, Eric Horvitz
UAI2
1998 The Lumière Project: Bayesian User Modeling for Inferring the Goals and Needs of Software Users
Eric Horvitz, John S. Breese, David Heckerman, David Hovel, Koos Rommelse
UAI1
1997 Compelling Intelligent User Interfaces - How Much AI?
abstract
Article Compelling intelligent user interfaces—how much AI? Share on Authors: Larry Birnbaum ILS/Northwestern, 1890 Maple Avenue, Evanston IL ILS/Northwestern, 1890 Maple Avenue, Evanston ILView Profile , Eric Horvitz Microsoft Research, One Microsoft Way, Redmond WA Microsoft Research, One Microsoft Way, Redmond WAView Profile , David Kurlander Microsoft Research, One Microsoft Way, Redmond WA Microsoft Research, One Microsoft Way, Redmond WAView Profile , Henry Lieberman MIT Media Laboratory, 20 Ames Street, Cambridge MA MIT Media Laboratory, 20 Ames Street, Cambridge MAView Profile , Joe Marks MERL, 201 Broadway, Cambridge MA MERL, 201 Broadway, Cambridge MAView Profile , Steve Roth Robotics Institute, Carnegie Mellon University, Pittsburgh PA Robotics Institute, Carnegie Mellon University, Pittsburgh PAView Profile Authors Info & Claims IUI '97: Proceedings of the 2nd international conference on Intelligent user interfacesJanuary 1997 Pages 173–175https://doi.org/10.1145/238218.238319Online:06 January 1997Publication History 16citation637DownloadsMetricsTotal Citations16Total Downloads637Last 12 Months41Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Lawrence Birnbaum, Eric Horvitz, David Kurlander, Henry Lieberman, Joe Marks, Steven F. Roth
IUI2
1997 Perception, Attention, and Resources: A Decision-Theoretic Approach to Graphics Rendering
Eric Horvitz, Jed Lengyel
UAI1
1997 Time-Critical Action: Representations and Application
Eric Horvitz, Adam Seiver
UAI1
1996 Decision-Theoretic Reasoning and the Human-Computer Interface: Advances in Embedded Intelligent Agents (abstract)
Eric Horvitz
SOFSEM1
1996 A Graph-Theoretic Analysis of Information Value
Kim-Leng Poh, Eric Horvitz
UAI2
1995 Display of Information for Time-Critical Decision Making
Eric Horvitz, Matthew Barry
UAI1
1995 Reasoning, Metareasoning, and Mathematical Truth: Studies of Theorem Proving under Limited Resources
Eric Horvitz, Adrian C. Klein
UAI1
1995 Exploiting System Hierarchy to Compute Repair Plans in Probabilistic Model-Based Diagnosis
Sampath Srinivas, Eric Horvitz
UAI2
1994 Dynamic Construction and Refinement of Utility-Based Categorization Models
abstract
The actions taken by an automated decision-making agent can be enhanced by including mechanisms that enable the agent to categorize concepts effectively. We pose a utility-based approach to categorization based on the idea that categorization should be carried out in the service of action. The choice of concepts is critical in the effective selection of actions under resource constraints. We propose a decision-theoretic framework for categorization which involves reasoning about alternative categorization models consisting of sets of interrelated concepts at varying levels of abstraction. Categorization models that are too abstract may overlook details that are critical for selecting the most appropriate actions. Categorization models that are too detailed, however, may be too expensive to process and may contain irrelevant information. Categorization models are therefore evaluated on the basis of the expected value of their recommended action, taking into account the resource cost of their evaluation. A knowledge representation scheme, known as probabilistic conceptual networks, has been developed to support the dynamic construction of models at varying levels of abstraction. This scheme combines the formalisms of influence diagrams from decision analysis and inheritance/abstraction hierarchies from AI. We also propose an incremental approach to categorical reasoning. By applying decision-theoretic control of model refinement, a resource-constrained actor iteratively decides between continuing to improve the current level of abstraction in the model, or to act immediately.>
Kim-Leng Poh, Michael R. Fehling, Eric Horvitz
IEEE Trans. Syst. Man Cybern. Syst.3
1993 Utility-Based Abstraction and Categorization
Eric Horvitz, Adrian C. Klein
UAI1
1993 Reasoning about the Value of Decision-Model Refinement: Methods and Application
Kim-Leng Poh, Eric Horvitz
UAI2
1993 A Bayesian analysis of simulation algorithms for inference in belief networks
abstract
Abstract A belief network is a graphical representation of the underlying probabilistic relationships in a complex system. Belief networks have been employed as a representation of uncertain relationships in computer‐based diagnostic systems. These diagnostic systems provide assistance by assigning likelihoods to alternative explanatory hypotheses in response to a set of findings or observations. Approximation algorithms have been used to compute likelihoods of hypotheses in large networks. We analyze the performance of leading Monte Carlo approximation algorithms for computing posterior probabilities in belief networks. The analysis differs from earlier attempts to characterize the behavior of simulation algorithms in our explicit use of Bayesian statistics: We update a probability distribution over target probabilities of interest with information from randomized trials. For real ϵ, δ < 1 and for a probabilistic inference Pr[x|e], the output of an inference approximation algorithm in an (ϵ, δ)‐estimate of Pr[x|e] if with probability at least 1 – δ the output is within relative error ϵ of Pr[x|e]. We construct a stopping rule for the number of simulations required by logic sampling, randomized approximation schemes, and likelihood weighting to provide (ϵ, δ)‐estimates of Pr[x|e]. With Probability 1 – δ, the stopping rule is optimal in the sense that the algorithm performs the minimum number of required simulations. We prove that our stopping rules are insensitive to the prior probability distribution on Pr[x|e]. © 1993 by John Wiley & Sons, Inc.
Paul Dagum, Eric Horvitz
Networks2
1993 An Approximate Nonmyopic Computation for Value of Information
abstract
It is argued that decision analysts and expert-system designers have avoided the intractability of exact computation of the value of information by relying on a myopic assumption that only one additional test will be performed, even when there is an opportunity to make large number of observations. An alternative to the myopic analysis is presented. In particular, an approximate method for computing the value of information of a set of tests, which exploits the statistical properties of large samples, is given. The approximation is linear in the number of tests, in contrast with the exact computation, which is exponential in the number of tests. The approach is not as general as in a complete nonmyopic analysis, in which all possible sequences of observations are considered. In addition, the approximation is limited to specific classes of dependencies among evidence and to binary hypothesis and decision variables. Nonetheless, as demonstrated with a simple application, the approach can offer an improvement over the myopic analysis.>
David Heckerman, Eric Horvitz, Blackford Middleton
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 Dynamic Network Models for Forecasting
Paul Dagum, Adam Galper, Eric Horvitz
UAI3
1992 Reformulating Inference Problems Through Selective Conditioning
Paul Dagum, Eric Horvitz
UAI2
1991 An Approximate Nonmyopic Computation for Value of Information
David Heckerman, Eric Horvitz, Blackford Middleton
UAI2
1991 Time-Dependent Utility and Action Under Uncertainty
Eric Horvitz, Geoffrey Rutledge
UAI1
1990 Ideal reformulation of belief networks
John S. Breese, Eric Horvitz
UAI2
1990 Problem formulation as the reduction of a decision model
David Heckerman, Eric Horvitz
UAI2
1989 Reflection and Action Under Scarce Resources: Theoretical Principles and Empirical Study
Eric Horvitz, Gregory F. Cooper, David Heckerman
IJCAI1
1988 Reasoning under Varying and Uncertain Resource Constraints
Eric Horvitz
AAAI1
1988 Reasoning about beliefs and actions under computational resource constraints
Eric Horvitz
Int. J. Approx. Reason.1
1988 Decision theory in expert systems and artificial intelligenc
abstract
Despite their different perspectives, artificial intelligence (AI) and the disciplines of decision science have common roots and strive for similar goals. This paper surveys the potential for addressing problems in representation, inference, knowledge engineering, and explanation within the decision-theoretic framework. Recent analyses of the restrictions of several traditional AI reasoning techniques, coupled with the development of more tractable and expressive decision-theoretic representation and inference strategies, have stimulated renewed interest in decision theory and decision analysis. We describe early experience with simple probabilistic schemes for automated reasoning, review the dominant expert-system paradigm, and survey some recent research at the crossroads of AI and decision science. In particular, we present the belief network and influence diagram representations. Finally, we discuss issues that have not been studied in detail within the expert-systems setting, yet are crucial for developing theoretical methods and computational architectures for automated reasoners.
Eric Horvitz, John S. Breese, Max Henrion
Int. J. Approx. Reason.1
1987 On the Expressiveness of Rule-based Systems for Reasoning with Uncertainty
David Heckerman, Eric Horvitz
AAAI2
1987 Reasoning about Beliefs and Actions Under Computational Resource Constraints
Eric Horvitz
UAI1
1986 A Framework for Comparing Alternative Formalisms for Plausible Reasoning
Eric Horvitz, David Heckerman, Curt Langlotz
AAAI1
1986 The myth of modularity in rule-based systems for reasoning with uncertainty
David Heckerman, Eric Horvitz
UAI2
1985 The Inconsistent Use of Measures of Certainty in Artificial Intelligence Research
Eric Horvitz, David Heckerman
UAI1