EDBT 2026 Demo / reviewers in the wild / expert
Jennifer Wortman Vaughan
dblp:w/JenniferWortman · also Jennifer Wortman
· DBLP profile ↗
70ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0002-7807-2018ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 20 · 15 since 2021Theory of computation · 15 · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Applied, interdisciplinary, general and emerging computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | "Helping Me Versus Doing It for Me": Designing for Agency in LLM-Infused Writing Tools for Science JournalismabstractJournalists rely on their agency—the ability to exercise independent judgment in alignment with their values—to fulfill their democratic social role. In this study, we investigate how LLM-infused writing tools reshape journalists’ agency in editorial decision making. In interviews with 20 science journalists, we presented four hypothetical LLM-infused writing tools representing a range of possible design space configurations. We find that journalists are selectively willing to cede control: they view AI that gathers information or offers feedback as supporting their efficiency by automating execution while leaving decision making intact. In contrast, they see AI that generates core ideas or drafts as a threat to their autonomy, skill development, self-fulfillment, and professional relationships. This sensitivity extends to seemingly automatable tasks such as manipulating writing voice with AI, which are seen as reducing opportunities for reflection and critical thinking. We discuss the implications of these findings for design that preserves journalistic agency in the moment, and over the long term. Sachita Nishal, Mina Lee 0002, Nicholas Diakopoulos, Jennifer Wortman Vaughan |
CHI | 4 |
| 2025 | Canvil: Designerly Adaptation for LLM-Powered User Experiences
K. J. Kevin Feng, Qingzi Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan, Amy X. Zhang, David W. McDonald |
CHI | 4 |
| 2025 | Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesabstractLarge language models (LLMs) can produce erroneous responses that sound fluent and convincing, raising the risk that users will rely on these responses as if they were correct. Mitigating such overreliance is a key challenge. Through a think-aloud study in which participants use an LLM-infused application to answer objective questions, we identify several features of LLM responses that shape users' reliance: explanations (supporting details for answers), inconsistencies in explanations, and sources. Through a large-scale, pre-registered, controlled experiment (N=308), we isolate and study the effects of these features on users' reliance, accuracy, and other measures. We find that the presence of explanations increases reliance on both correct and incorrect responses. However, we observe less reliance on incorrect responses when sources are provided or when explanations exhibit inconsistencies. We discuss the implications of these findings for fostering appropriate reliance on LLMs. Sunnie S. Y. Kim, Jennifer Wortman Vaughan, Qingzi Vera Liao, Tania Lombrozo, Olga Russakovsky |
CHI | 2 |
| 2025 | Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Researchabstract"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific information from a generative-AI model's parameters, e.g., a particular individual's personal data or the inclusion of copyrighted content in the model's training data. Unlearning is also proposed as a way to prevent a model from generating targeted types of information in its outputs, e.g., generations that closely resemble a particular individual's data or reflect the concept of "Spiderman." Both of these goals--the targeted removal of information from a model and the targeted suppression of information from a model's outputs--present various technical and substantive challenges. We provide a framework for ML researchers and policymakers to think rigorously about these challenges, identifying several mismatches between the goals of unlearning and feasible implementations. These mismatches explain why unlearning is not a general-purpose solution for circumscribing generative-AI model behavior in service of broader positive impact. A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen, Kevin Klyman, Matthew Jagielski, Katja Filippova, Ziyu Liu 0002, Alexandra Chouldechova, Jamie Hayes, Yangsibo Huang, Eleni Triantafillou, Peter Kairouz, Nicole Mitchell, Niloofar Mireshghallah, Abigail Z. Jacobs, James Grimmelmann, Vitaly Shmatikov, Christopher De Sa, Ilia Shumailov, Andreas Terzis, Solon Barocas, Jennifer Wortman Vaughan, danah boyd, Yejin Choi 0001, Oluwasanmi Koyejo, Fernando A. Delgado, Percy Liang, Daniel E. Ho, Pamela Samuelson, Miles Brundage, David Bau, Seth Neel, Hanna M. Wallach, Amy Cyphert, Mark A. Lemley, Nicolas Papernot, Katherine Lee |
NeurIPS | 22 |
| 2025 | Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their WorkabstractRecent years have witnessed increasing calls for computing researchers to grapple with the societal impacts of their work. Tools such as impact assessments have gained prominence as a method to uncover potential impacts, and a number of publication venues now encourage authors to include an impact statement in their submissions. Despite this recent push, little is known about the way researchers go about assessing, articulating, and addressing the potential negative societal impact of their work --- especially in industry settings, where research outcomes are often quickly integrated into products and services. In addition, while there are nascent efforts to support researchers in this task, there remains a dearth of empirically-informed tools and processes. Through interviews with 25 industry computing researchers across different companies and research areas, we identify four key factors that influence how they grapple with (or choose not to grapple with) the societal impact of their research: the relationship between industry researchers and product teams; organizational dynamics and cultures that prioritize innovation and speed; misconceptions about societal impact; and a lack of sufficient infrastructure to support researchers. To develop an effective impact assessment template tailored to industry computing researchers' needs, we conduct an iterative co-design process with these 25 industry researchers, along with an additional 16 researchers and practitioners with prior experience and expertise in reviewing and developing impact assessments or responsible computing practices more broadly. Through the co-design process, we develop 10 design considerations to facilitate the effective design, implementation, and adaptation of an impact assessment template for use in industry research settings and beyond, as well as our own "Societal Impact Assessment" template with concrete scaffolds. We explore the effectiveness of this template through a user study with 15 industry research interns, revealing both its strengths and limitations. Finally, we discuss the implications for future researchers, organizations, and policymakers seeking to foster more responsible research practices. Wesley Deng, Solon Barocas, Jennifer Wortman Vaughan |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2025 | Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code CompletionsabstractLarge-scale generative models have enabled the development of AI-powered code completion tools to assist programmers in writing code. Like all AI-powered tools, these code completion tools are not always accurate and can introduce bugs or even security vulnerabilities into code if not properly detected and corrected by a human programmer. One technique that has been proposed and implemented to help programmers locate potential errors is to highlight uncertain tokens. However, little is known about the effectiveness of this technique. Through a mixed-methods study with 30 programmers, we compare three conditions: providing the AI system's code completion alone, highlighting tokens with the lowest likelihood of being generated by the underlying generative model, and highlighting tokens with the highest predicted likelihood of being edited by a programmer. We find that highlighting tokens with the highest predicted likelihood of being edited leads to faster task completion and more targeted edits, and is subjectively preferred by study participants. In contrast, highlighting tokens according to their probability of being generated does not provide any benefit over the baseline with no highlighting. We further explore the design space of how to convey uncertainty in AI-powered code completion tools and find that programmers prefer highlights that are granular, informative, interpretable, and not overwhelming. This work contributes to building an understanding of what uncertainty means for generative models and how to convey it effectively. Helena Vasconcelos, Gagan Bansal, Adam Fourney, Qingzi Vera Liao, Jennifer Wortman Vaughan |
ACM Trans. Comput. Hum. Interact. | 5 |
| 2024 | (De)Noise: Moderating the Inconsistency Between Human Decision-MakersabstractPrior research in psychology has found that people's decisions are often inconsistent. An individual's decisions vary across time, and decisions vary even more across people. Inconsistencies have been identified not only in subjective matters, like matters of taste, but also in settings one might expect to be more objective, such as sentencing, job performance evaluations, or real estate appraisals. In our study, we explore whether algorithmic decision aids can be used to moderate the degree of inconsistency in human decision-making in the context of real estate appraisal. In a large-scale human-subject experiment, we study how different forms of algorithmic assistance influence the way that people review and update their estimates of real estate prices. We find that both (i) asking respondents to review their estimates in a series of algorithmically chosen pairwise comparisons and (ii) providing respondents with traditional machine advice are effective strategies for influencing human responses. Compared to simply reviewing initial estimates one by one, the aforementioned strategies lead to (i) a higher propensity to update initial estimates, (ii) a higher accuracy of post-review estimates, and (iii) a higher degree of consistency between the post-review estimates of different respondents. While these effects are more pronounced with traditional machine advice, the approach of reviewing algorithmically chosen pairs can be implemented in a wider range of settings, since it does not require access to ground truth data. Nina Grgic-Hlaca, Junaid Ali 0001, Krishna P. Gummadi, Jennifer Wortman Vaughan |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2024 | Tinker, Tailor, Configure, Customize: The Articulation Work of Contextualizing an AI Fairness ChecklistabstractMany responsible AI resources, such as toolkits, playbooks, and checklists, have been developed to support AI practitioners in identifying, measuring, and mitigating potential fairness-related harms. These resources are often designed to be general purpose in order to be applicable to a variety of use cases, domains, and deployment contexts. However, this can lead to decontextualization, where such resources lack the level of relevance or specificity needed to use them. To understand how AI practitioners might contextualize one such resource, an AI fairness checklist, for their particular use cases, domains, and deployment contexts, we conducted a retrospective contextual inquiry with 13 AI practitioners from seven organizations. We identify how contextualizing this checklist introduces new forms of work for AI practitioners and other stakeholders, as well as opening up new sites for negotiation and contestation of values in AI. We also identify how the contextualization process may help AI practitioners develop a shared language around AI fairness, and we identify tensions related to ownership over this process that suggest larger issues of accountability in responsible AI work. Michael A. Madaio, Hanna M. Wallach, Jennifer Wortman Vaughan |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2023 | Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User ExperienceabstractDespite the widespread use of artificial intelligence (AI), designing user experiences (UX) for AI-powered systems remains challenging. UX designers face hurdles understanding AI technologies, such as pre-trained language models, as design materials. This limits their ability to ideate and make decisions about whether, where, and how to use AI. To address this problem, we bridge the literature on AI design and AI transparency to explore whether and how frameworks for transparent model reporting can support design ideation with pre-trained models. By interviewing 23 UX practitioners, we find that practitioners frequently work with pre-trained models, but lack support for UX-led ideation. Through a scenario-based design task, we identify common goals that designers seek model understanding for and pinpoint their model transparency information needs. Our study highlights the pivotal role that UX designers can play in Responsible AI and calls for supporting their understanding of AI limitations through model transparency and interrogation. Qingzi Vera Liao, Hariharan Subramonyam, Jennifer Wortman Vaughan |
CHI | 4 |
| 2023 | GAM Coach: Towards Interactive and User-centered Algorithmic RecourseabstractMachine learning (ML) recourse techniques are increasingly used in high-stakes domains, providing end users with actions to alter ML predictions, but they assume ML developers understand what input variables can be changed. However, a recourse plan’s actionability is subjective and unlikely to match developers’ expectations completely. We present GAM Coach, a novel open-source system that adapts integer linear programming to generate customizable counterfactual explanations for Generalized Additive Models (GAMs), and leverages interactive visualizations to enable end users to iteratively generate recourse plans meeting their needs. A quantitative user study with 41 participants shows our tool is usable and useful, and users prefer personalized recourse plans over generic plans. Through a log analysis, we explore how users discover satisfactory recourse plans, and provide empirical evidence that transparency can lead to more opportunities for everyday users to discover counterintuitive patterns in ML models. GAM Coach is available at: https://poloclub.github.io/gam-coach/. Zijie J. Wang, Jennifer Wortman Vaughan, Rich Caruana, Polo Chau |
CHI | 2 |
| 2023 | Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with ExplanationsabstractAI explanations are often mentioned as a way to improve human-AI decision-making, but empirical studies have not found consistent evidence of explanations' effectiveness and, on the contrary, suggest that they can increase overreliance when the AI system is wrong. While many factors may affect reliance on AI support, one important factor is how decision-makers reconcile their own intuition---beliefs or heuristics, based on prior knowledge, experience, or pattern recognition, used to make judgments---with the information provided by the AI system to determine when to override AI predictions. We conduct a think-aloud, mixed-methods study with two explanation types (feature- and example-based) for two prediction tasks to explore how decision-makers' intuition affects their use of AI predictions and explanations, and ultimately their choice of when to rely on AI. Our results identify three types of intuition involved in reasoning about AI predictions and explanations: intuition about the task outcome, features, and AI limitations. Building on these, we summarize three observed pathways for decision-makers to apply their own intuition and override AI predictions. We use these pathways to explain why (1) the feature-based explanations we used did not improve participants' decision outcomes and increased their overreliance on AI, and (2) the example-based explanations we used improved decision-makers' performance over feature-based explanations and helped achieve complementary human-AI performance. Overall, our work identifies directions for further development of AI decision-support systems and explanation methods that help decision-makers effectively apply their intuition to achieve appropriate reliance on AI. Valerie Chen, Qingzi Vera Liao, Jennifer Wortman Vaughan, Gagan Bansal |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2023 | Greedy Algorithm Almost Dominates in Smoothed Contextual BanditsabstractAbstract. Online learning algorithms, widely used to power search and content optimization on the web, must balance exploration and exploitation, potentially sacrificing the experience of current users in order to gain information that will lead to better decisions in the future. While necessary in the worst case, explicit exploration has a number of disadvantages compared to the greedy algorithm that always “exploits” by choosing an action that currently looks optimal. We determine under what conditions inherent diversity in the data makes explicit exploration unnecessary. We build on a recent line of work on the smoothed analysis of the greedy algorithm in the linear contextual bandits model. We improve on prior results to show that the greedy algorithm almost matches the best possible Bayesian regret rate of any other algorithm on the same problem instance whenever the diversity conditions hold. The key technical finding is that data collected by the greedy algorithm suffices to simulate a run of any other algorithm. Further, we prove that under a particular smoothness assumption, the Bayesian regret of the greedy algorithm is at most [Formula: see text] in the worst case, where [Formula: see text] is the time horizon. Manish Raghavan, Aleksandrs Slivkins, Jennifer Wortman Vaughan, Steven Z. Wu |
SIAM J. Comput. | 3 |
| 2022 | Interpretability, Then What? Editing Machine Learning Models to Reflect Human Knowledge and ValuesabstractMachine learning (ML) interpretability techniques can reveal undesirable patterns in data that models exploit to make predictions-potentially causing harms once deployed. However, how to take action to address these patterns is not always clear. In a collaboration between ML and human-computer interaction researchers, physicians, and data scientists, we develop GAM Changer, the first interactive system to help domain experts and data scientists easily and responsibly edit Generalized Additive Models (GAMs) and fix problematic patterns. With novel interaction techniques, our tool puts interpretability into action-empowering users to analyze, validate, and align model behaviors with their knowledge and values. Physicians have started to use our tool to investigate and fix pneumonia and sepsis risk prediction models, and an evaluation with 7 data scientists working in diverse domains highlights that our tool is easy to use, meets their model editing needs, and fits into their current workflows. Built with modern web technologies, our tool runs locally in users' web browsers or computational notebooks, lowering the barrier to use. GAM Changer is available at the following public demo link: https://interpret.ml/gam-changer. Zijie J. Wang, Alex Kale, Harsha Nori, Peter Stella, Mark E. Nunnally, Polo Chau, Mihaela Vorvoreanu, Jennifer Wortman Vaughan, Rich Caruana |
KDD | 8 |
| 2022 | Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataabstractData is central to the development and evaluation of machine learning (ML) models. However, the use of problematic or inappropriate datasets can result in harms when the resulting models are deployed. To encourage responsible AI practice through more deliberate reflection on datasets and transparency around the processes by which they are created, researchers and practitioners have begun to advocate for increased data documentation and have proposed several data documentation frameworks. However, there is little research on whether these data documentation frameworks meet the needs of ML practitioners, who both create and consume datasets. To address this gap, we set out to understand ML practitioners' data documentation perceptions, needs, challenges, and desiderata, with the ultimate goal of deriving design requirements that can inform future data documentation frameworks. We conducted a series of semi-structured interviews with 14 ML practitioners at a single large, international technology company. We had them answer a list of questions taken from datasheets for datasets~\citegebru2018datasheets. Our findings show that current approaches to data documentation are largely ad hoc and myopic in nature. Participants expressed needs for data documentation frameworks to be adaptable to their contexts, integrated into their existing tools and workflows, and automated wherever possible. Despite the fact that data documentation frameworks are often motivated from the perspective of responsible AI, participants did not make the connection between the questions that they were asked to answer and their responsible AI implications. In addition, participants often had difficulties prioritizing the needs of dataset consumers and providing information that someone unfamiliar with their datasets might need to know. Based on these findings, we derive seven design requirements for future data documentation frameworks such as more actionable guidance on how the characteristics of datasets might result in harms and how these harms might be mitigated, more explicit prompts for reflection, automated adaptation to different contexts, and integration into ML practitioners' existing tools and workflows. Amy Heger, Elizabeth B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach, Jennifer Wortman Vaughan |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2022 | Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for SupportabstractVarious tools and practices have been developed to support practitioners in identifying, assessing, and mitigating fairness-related harms caused by AI systems. However, prior research has highlighted gaps between the intended design of these tools and practices and their use within particular contexts, including gaps caused by the role that organizational factors play in shaping fairness work. In this paper, we investigate these gaps for one such practice: disaggregated evaluations of AI systems, intended to uncover performance disparities between demographic groups. By conducting semi-structured interviews and structured workshops with thirty-three AI practitioners from ten teams at three technology companies, we identify practitioners' processes, challenges, and needs for support when designing disaggregated evaluations. We find that practitioners face challenges when choosing performance metrics, identifying the most relevant direct stakeholders and demographic groups on which to focus, and collecting datasets with which to conduct disaggregated evaluations. More generally, we identify impacts on fairness work stemming from a lack of engagement with direct stakeholders or domain experts, business imperatives that prioritize customers over marginalized groups, and the drive to deploy AI systems at scale. Michael A. Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan, Hanna M. Wallach |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2021 | Designing Disaggregated Evaluations of AI Systems: Choices, Considerations, and TradeoffsabstractDisaggregated evaluations of AI systems, in which system performance is assessed and reported separately for different groups of people, are conceptually simple. However, their design involves a variety of choices. Some of these choices influence the results that will be obtained, and thus the conclusions that can be drawn; others influence the impacts---both beneficial and harmful---that a disaggregated evaluation will have on people, including the people whose data is used to conduct the evaluation. We argue that a deeper understanding of these choices will enable researchers and practitioners to design careful and conclusive disaggregated evaluations. We also argue that better documentation of these choices, along with the underlying considerations and tradeoffs that have been made, will help others when interpreting an evaluation's results and conclusions. Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, W. Duncan Wadsworth, Hanna M. Wallach |
AIES | 6 |
| 2021 | Manipulating and Measuring Model InterpretabilityabstractWith machine learning models being increasingly used to aid decision making even in high-stakes domains, there has been a growing interest in developing interpretable models. Although many supposedly interpretable models have been proposed, there have been relatively few experimental studies investigating whether these models achieve their intended effects, such as making people more closely follow a model’s predictions when it is beneficial for them to do so or enabling them to detect when a model has made a mistake. We present a sequence of pre-registered experiments (N = 3, 800) in which we showed participants functionally identical models that varied only in two factors commonly thought to make machine learning models more or less interpretable: the number of features and the transparency of the model (i.e., whether the model internals are clear or black box). Predictably, participants who saw a clear model with few features could better simulate the model’s predictions. However, we did not find that participants more closely followed its predictions. Furthermore, showing participants a clear model meant that they were less able to detect and correct for the model’s sizable mistakes, seemingly due to information overload. These counterintuitive findings emphasize the importance of testing over intuition when developing interpretable models. Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, Hanna M. Wallach |
CHI | 4 |
| 2021 | From Human Explanation to Model Interpretability: A Framework Based on Weight of EvidenceabstractWe take inspiration from the study of human explanation to inform the design and evaluation of interpretability methods in machine learning. First, we survey the literature on human explanation in philosophy, cognitive science, and the social sciences, and propose a list of design principles for machine-generated explanations that are meaningful to humans. Using the concept of weight of evidence from information theory, we develop a method for generating explanations that adhere to these principles. We show that this method can be adapted to handle high-dimensional, multi-class settings, yielding a flexible framework for generating explanations. We demonstrate that these explanations can be estimated accurately from finite samples and are robust to small perturbations of the inputs. We also evaluate our method through a qualitative user study with machine learning practitioners, where we observe that the resulting explanations are usable despite some participants struggling with background concepts like prior class probabilities. Finally, we conclude by surfacing design implications for interpretability tools in general. David Alvarez-Melis, Harmanpreet Kaur, Hal Daumé III, Hanna M. Wallach, Jennifer Wortman Vaughan |
HCOMP | 5 |
| 2020 | Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningabstractMachine learning (ML) models are now routinely deployed in domains ranging from criminal justice to healthcare. With this newfound ubiquity, ML has moved beyond academia and grown into an engineering discipline. To that end, interpretability tools have been designed to help data scientists and machine learning practitioners better understand how ML models work. However, there has been little evaluation of the extent to which these tools achieve this goal. We study data scientists' use of two existing interpretability tools, the InterpretML implementation of GAMs and the SHAP Python package. We conduct a contextual inquiry (N=11) and a survey (N=197) of data scientists to observe how they use interpretability tools to uncover common issues that arise when building and evaluating ML models. Our results indicate that data scientists over-trust and misuse interpretability tools. Furthermore, few of our participants were able to accurately describe the visualizations output by these tools. We highlight qualitative themes for data scientists' mental models of interpretability tools. We conclude with implications for researchers and tool designers, and contextualize our findings in the social science literature. Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna M. Wallach, Jennifer Wortman Vaughan |
CHI | 6 |
| 2020 | Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIabstractMany organizations have published principles intended to guide the ethical development and deployment of AI systems; however, their abstract nature makes them difficult to operationalize. Some organizations have therefore produced AI ethics checklists, as well as checklists for more specific concepts, such as fairness, as applied to AI systems. But unless checklists are grounded in practitioners' needs, they may be misused. To understand the role of checklists in AI ethics, we conducted an iterative co-design process with 48 practitioners, focusing on fairness. We co-designed an AI fairness checklist and identified desiderata and concerns for AI fairness checklists in general. We found that AI fairness checklists could provide organizational infrastructure for formalizing ad-hoc processes and empowering individual advocates. We highlight aspects of organizational culture that may impact the efficacy of AI fairness checklists, and suggest future design directions. Michael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. Wallach |
CHI | 3 |
| 2020 | No-Regret and Incentive-Compatible Online LearningabstractWe study online learning settings in which experts act strategically to maximize their influence on the learning algorithm’s predictions by potentially misreporting their beliefs about a sequence of binary events. Our goal is twofold. First, we want the learning algorithm to be no-regret with respect to the best-fixed expert in hindsight. Second, we want incentive compatibility, a guarantee that each expert’s best strategy is to report his true beliefs about the realization of each event. To achieve this goal, we build on the literature on wagering mechanisms, a type of multi-agent scoring rule. We provide algorithms that achieve no regret and incentive compatibility for myopic experts for both the full and partial information settings. In experiments on datasets from FiveThirtyEight, our algorithms have regret comparable to classic no-regret algorithms, which are not incentive-compatible. Finally, we identify an incentive-compatible algorithm for forward-looking strategic agents that exhibits diminishing regret in practice. Rupert Freeman, David M. Pennock, Chara Podimata, Jennifer Wortman Vaughan |
ICML | 4 |
| 2020 | Oracle-efficient Online Learning and Auction Design
Miroslav Dudík, Nika Haghtalab, Robert E. Schapire, Vasilis Syrgkanis, Jennifer Wortman Vaughan |
J. ACM | 6 |
| 2019 | Group Fairness for the Allocation of Indivisible GoodsabstractWe consider the problem of fairly dividing a collection of indivisible goods among a set of players. Much of the existing literature on fair division focuses on notions of individual fairness. For instance, envy-freeness requires that no player prefer the set of goods allocated to another player to her own allocation. We observe that an algorithm satisfying such individual fairness notions can still treat groups of players unfairly, with one group desiring the goods allocated to another. Our main contribution is a notion of group fairness, which implies most existing notions of individual fairness. Group fairness (like individual fairness) cannot be satisfied exactly with indivisible goods. Thus, we introduce two “up to one good” style relaxations. We show that, somewhat surprisingly, certain local optima of the Nash welfare function satisfy both relaxations and can be computed in pseudo-polynomial time by local search. Our experiments reveal faster computation and stronger fairness guarantees in practice. Vincent Conitzer, Rupert Freeman, Nisarg Shah 0001, Jennifer Wortman Vaughan |
AAAI | 4 |
| 2019 | An Equivalence between Wagering and Fair-Division Mechanisms
Rupert Freeman, David M. Pennock, Jennifer Wortman Vaughan |
AAAI | 3 |
| 2019 | Improving Fairness in Machine Learning Systems: What Do Industry Practitioners Need?abstractThe potential for machine learning (ML) systems to amplify social inequities and unfairness is receiving increasing popular and academic attention. A surge of recent work has focused on the development of algorithmic tools to assess and mitigate such unfairness. If these tools are to have a positive impact on industry practice, however, it is crucial that their design be informed by an understanding of real-world needs. Through 35 semi-structured interviews and an anonymous survey of 267 ML practitioners, we conduct the first systematic investigation of commercial product teams' challenges and needs for support in developing fairer ML systems. We identify areas of alignment and disconnect between the challenges faced by teams in practice and the solutions proposed in the fair ML research literature. Based on these findings, we highlight directions for future ML and HCI research that will better address practitioners' needs. Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miroslav Dudík, Hanna M. Wallach |
CHI | 2 |
| 2019 | Understanding the Effect of Accuracy on Trust in Machine Learning ModelsabstractWe address a relatively under-explored aspect of human-computer interaction: people's abilities to understand the relationship between a machine learning model's stated performance on held-out data and its expected performance post deployment. We conduct large-scale, randomized human-subject experiments to examine whether laypeople's trust in a model, measured in terms of both the frequency with which they revise their predictions to match those of the model and their self-reported levels of trust in the model, varies depending on the model's stated accuracy on held-out data and on its observed accuracy in practice. We find that people's trust in a model is affected by both its stated accuracy and its observed accuracy, and that the effect of stated accuracy can change depending on the observed accuracy. Our work relates to recent research on interpretable machine learning, but moves beyond the typical focus on model internals, exploring a different component of the machine learning pipeline. Ming Yin 0001, Jennifer Wortman Vaughan, Hanna M. Wallach |
CHI | 2 |
| 2019 | Using Search Queries to Understand Health Information Needs in Africa
Rediet Abebe, Shawndra Hill, Jennifer Wortman Vaughan, Peter M. Small, H. Andrew Schwartz |
ICWSM | 3 |
| 2018 | Incentive-Compatible Forecasting CompetitionsabstractWe consider the design of forecasting competitions in which multiple forecasters make predictions about one or more independent events and compete for a single prize. We have two objectives: (1) to award the prize to the most accurate forecaster, and (2) to incentivize forecasters to report truthfully, so that forecasts are informative and forecasters need not spend any cognitive effort strategizing about reports. Proper scoring rules incentivize truthful reporting if all forecasters are paid according to their scores. However, incentives become distorted if only the best-scoring forecaster wins a prize, since forecasters can often increase their probability of having the highest score by reporting extreme beliefs. Even if forecasters do report truthfully, awarding the prize to the forecaster with highest score does not guarantee that high-accuracy forecasters are likely to win; in extreme cases, it can result in a perfect forecaster having zero probability of winning. In this paper, we introduce a truthful forecaster selection mechanism. We lower-bound the probability that our mechanism selects the most accurate forecaster, and give rates for how quickly this bound approaches 1 as the number of events grows. Our techniques can be generalized to the related problems of outputting a ranking over forecasters and hiring a forecaster with high accuracy on future events. Jens Witkowski, Rupert Freeman, Jennifer Wortman Vaughan, David M. Pennock, Andreas Krause 0001 |
AAAI | 3 |
| 2018 | The Externalities of Exploration and How Data Diversity Helps ExploitationabstractOnline learning algorithms, widely used to power search and content optimization on the web, must balance exploration and exploitation, potentially sacrificing the experience of current users in order to gain information that will lead to better decisions in the future. Recently, concerns have been raised about whether the process of exploration could be viewed as unfair, placing too much burden on certain individuals or groups. Motivated by these concerns, we initiate the study of the externalities of exploration—the undesirable side effects that the presence of one party may impose on another—under the linear contextual bandits model. We introduce the notion of a group externality, measuring the extent to which the presence of one population of users (the majority) impacts the rewards of another (the minority). We show that this impact can, in some cases, be negative, and that, in a certain sense, no algorithm can avoid it. We then move on to study externalities at the individual level, interpreting the act of exploration as an externality imposed on the current user of a system by future users. This drives us to ask under what conditions inherent diversity in the data makes explicit exploration unnecessary. We build on a recent line of work on the smoothed analysis of the greedy algorithm that always chooses the action that currently looks optimal. We improve on prior results to show that a greedy approach almost matches the best possible Bayesian regret rate of any other algorithm on the same problem instance whenever the diversity conditions hold, and that this regret is at most $\tilde{O}(T^{1/3})$. Returning to group-level effects, we show that under the same conditions, negative group externalities essentially vanish if one runs the greedy algorithm. Together, our results uncover a sharp contrast between the high externalities that exist in the worst case, and the ability to remove all externalities if the data is sufficiently diverse. Manish Raghavan, Aleksandrs Slivkins, Jennifer Wortman Vaughan, Steven Z. Wu |
COLT | 3 |
| 2017 | Oracle-Efficient Online Learning and Auction DesignabstractWe consider the design of computationally efficient online learning algorithms in an adversarial setting in which the learner has access to an offline optimization oracle. We present an algorithm called Generalized Follow-the-Perturbed-Leader and provide conditions under which it is oracle-efficient while achieving vanishing regret. Our results make significant progress on an open problem raised by Hazan and Koren [31], who showed that oracle-efficient algorithms do not exist in general [30] and asked whether one can identify properties under which oracle-efficient online learning may be possible. Our auction-design framework considers an auctioneer learning an optimal auction for a sequence of adversarially selected valuations with the goal of achieving revenue that is almost as good as the optimal auction in hindsight, among a class of auctions. We give oracle-efficient learning results for: (1) VCG auctions with bidder-specific reserves in single-parameter settings, (2) envy-free item pricing in multi-item auctions, and (3) s-level auctions of Morgenstern and Roughgarden [43] for single-item settings. The last result leads to an approximation of the overall optimal Myerson auction when bidders’ valuations are drawn according to a fast-mixing Markov process, extending prior work that only gave such guarantees for the i.i.d. setting. Finally, we derive various extensions, including: (1) oracle-efficient algorithms for the contextual learning setting in which the learner has access to side information (such as bidder demographics), (2) learning with approximate oracles such as those based on Maximal-in-Range algorithms, and (3) no-regret bidding in simultaneous auctions, resolving an open problem of Daskalakis and Syrgkanis [14]. Miroslav Dudík, Nika Haghtalab, Robert E. Schapire, Vasilis Syrgkanis, Jennifer Wortman Vaughan |
FOCS | 6 |
| 2017 | A Decomposition of Forecast Error in Prediction MarketsabstractWe analyze sources of error in prediction market forecasts in order to bound the difference between a security's price and the ground truth it estimates. We consider cost-function-based prediction markets in which an automated market maker adjusts security prices according to the history of trade. We decompose the forecasting error into three components: sampling error, arising because traders only possess noisy estimates of ground truth; market-maker bias, resulting from the use of a particular market maker (i.e., cost function) to facilitate trade; and convergence error, arising because, at any point in time, market prices may still be in flux. Our goal is to make explicit the tradeoffs between these error components, influenced by design decisions such as the functional form of the cost function and the amount of liquidity in the market. We consider a specific model in which traders have exponential utility and exponential-family beliefs representing noisy estimates of ground truth. In this setting, sampling error vanishes as the number of traders grows, but there is a tradeoff between the other two components. We provide both upper and lower bounds on market-maker bias and convergence error, and demonstrate via numerical simulations that these bounds are tight. Our results yield new insights into the question of how to set the market's liquidity parameter and into the forecasting benefits of enforcing coherent prices across securities. Miroslav Dudík, Sébastien Lahaie, Ryan Rogers 0002, Jennifer Wortman Vaughan |
NIPS | 4 |
| 2017 | The Double Clinching Auction for WageringabstractWe develop the first incentive compatible and near-Pareto-optimal wagering mechanism. Wagering mechanisms can be used to elicit predictions from agents who reveal their beliefs by placing bets. Lambert et al. [20, 21] introduced weighted score wagering mechanisms, a class of budget-balanced wagering mechanisms under which agents with immutable beliefs truthfully report their predictions. However, we demonstrate that these and other existing incentive compatible wagering mechanisms are not Pareto optimal: agents have significant budget left over even when additional trade would be mutually beneficial. Motivated by this observation, we design a new wagering mechanism, the double clinching auction, a two-sided version of the adaptive clinching auction [9]. We show that no wagering mechanism can simultaneously satisfy weak budget balance, individual rationality, weak incentive compatibility, and Pareto optimality. However, we prove that the double clinching auction attains the first three and show in a series of simulations using real contest data that it comes much closer to Pareto optimality than previously known incentive compatible wagering mechanisms, in some cases almost matching the efficiency of the Pareto optimal (but not incentive compatible) parimutuel consensus mechanism. When the goal of wagering is to crowdsource probabilities, Pareto optimality drives participation and incentive compatibility drives accuracy, making the double clinching auction an attractive and practical choice. Our mechanism may be of independent interest as the first two-sided version of the adaptive clinching auction. Rupert Freeman, David M. Pennock, Jennifer Wortman Vaughan |
EC | 3 |
| 2017 | Making Better Use of the Crowd: How Crowdsourcing Can Advance Machine Learning Research
Jennifer Wortman Vaughan |
J. Mach. Learn. Res. | 1 |
| 2016 | The Possibilities and Limitations of Private Prediction MarketsabstractWe consider the design of private prediction markets, financial markets designed to elicit predictions about uncertain events without revealing too much information about market participants' actions or beliefs. Our goal is to design market mechanisms in which participants' trades or wagers influence the market's behavior in a way that leads to accurate predictions, yet no single participant has too much influence over what others are able to observe. We study the possibilities and limitations of such mechanisms using tools from differential privacy. We begin by designing a private one-shot wagering mechanism in which bettors specify a belief about the likelihood of a future event and a corresponding monetary wager. Wagers are redistributed among bettors in a way that more highly rewards those with accurate predictions. We provide a class of wagering mechanisms that are guaranteed to satisfy truthfulness, budget balance on expectation, and other desirable properties while additionally guaranteeing epsilon-joint differential privacy in the bettors' reported beliefs, and analyze the trade-off between the achievable level of privacy and the sensitivity of a bettor's payment to her own report. We then ask whether it is possible to obtain privacy in dynamic prediction markets, focusing our attention on the popular cost-function framework in which securities with payments linked to future events are bought and sold by an automated market maker. We show that under general conditions, it is impossible for such a market maker to simultaneously achieve bounded worst-case loss and epsilon-differential privacy without allowing the privacy guarantee to degrade extremely quickly as the number of trades grows, making such markets impractical in settings in which privacy is valued. We conclude by suggesting several avenues for potentially circumventing this lower bound. Rachel Cummings, David M. Pennock, Jennifer Wortman Vaughan |
EC | 3 |
| 2016 | Bounded Rationality in Wagering Mechanisms
David M. Pennock, Vasilis Syrgkanis, Jennifer Wortman Vaughan |
UAI | 3 |
| 2016 | The Communication Network Within the CrowdabstractSince its inception, crowdsourcing has been considered a black-box approach to solicit labor from a crowd of workers. Furthermore, the "crowd" has been viewed as a group of independent workers dispersed all over the world. Recent studies based on in-person interviews have opened up the black box and shown that the crowd is not a collection of independent workers, but instead that workers communicate and collaborate with each other. Put another way, prior work has shown the existence of edges between workers. We build on and extend this discovery by mapping the entire communication network of workers on Amazon Mechanical Turk, a leading crowdsourcing platform. We execute a task in which over 10,000 workers from across the globe self-report their communication links to other workers, thereby mapping the communication network among workers. Our results suggest that while a large percentage of workers indeed appear to be independent, there is a rich network topology over the rest of the population. That is, there is a substantial communication network within the crowd. We further examine how online forum usage relates to network topology, how workers communicate with each other via this network, how workers' experience levels relate to their network positions, and how U.S. workers differ from international workers in their network characteristics. We conclude by discussing the implications of our findings for requesters, workers, and platform providers like Amazon. Ming Yin 0001, Mary L. Gray, Siddharth Suri, Jennifer Wortman Vaughan |
WWW | 4 |
| 2016 | Adaptive Contract Design for Crowdsourcing Markets: Bandit Algorithms for Repeated Principal-Agent ProblemsabstractCrowdsourcing markets have emerged as a popular platform for matching available workers with tasks to complete. The payment for a particular task is typically set by the task's requester, and may be adjusted based on the quality of the completed work, for example, through the use of "bonus" payments. In this paper, we study the requester's problem of dynamically adjusting quality-contingent payments for tasks. We consider a multi-round version of the well-known principal-agent model, whereby in each round a worker makes a strategic choice of the effort level which is not directly observable by the requester. In particular, our formulation significantly generalizes the budget-free online task pricing problems studied in prior work. We treat this problem as a multi-armed bandit problem, with each "arm" representing a potential contract. To cope with the large (and in fact, infinite) number of arms, we propose a new algorithm, AgnosticZooming, which discretizes the contract space into a finite number of regions, effectively treating each region as a single arm. This discretization is adaptively refined, so that more promising regions of the contract space are eventually discretized more finely. We analyze this algorithm, showing that it achieves regret sublinear in the time horizon and substantially improves over non-adaptive discretization (which is the only competing approach in the literature). Our results advance the state of art on several different topics: the theory of crowdsourcing markets, principal-agent problems, multi-armed bandits, and dynamic pricing. Chien-Ju Ho, Aleksandrs Slivkins, Jennifer Wortman Vaughan |
J. Artif. Intell. Res. | 3 |
| 2015 | Integrating Market Makers, Limit Orders, and Continuous Trade in Prediction Marketsabstractresearch-article Share on Integrating Market Makers, Limit Orders, and Continuous Trade in Prediction Markets Authors: Hoda Heidari University of Pennsylvania, Philadelphia, PA, USA University of Pennsylvania, Philadelphia, PA, USAView Profile , Sebastien Lahaie Microsoft research, New York, NY, USA Microsoft research, New York, NY, USAView Profile , David M. Pennock Microsoft research, New York, NY, USA Microsoft research, New York, NY, USAView Profile , Jennifer Wortman Vaughan Microsoft Research, New York, NY, USA Microsoft Research, New York, NY, USAView Profile Authors Info & Claims EC '15: Proceedings of the Sixteenth ACM Conference on Economics and ComputationJune 2015 Pages 583–600https://doi.org/10.1145/2764468.2764532Published:15 June 2015Publication History 1citation131DownloadsMetricsTotal Citations1Total Downloads131Last 12 Months3Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Hoda Heidari, Sébastien Lahaie, David M. Pennock, Jennifer Wortman Vaughan |
EC | 4 |
| 2015 | Incentivizing High Quality CrowdworkabstractWe study the causal effects of financial incentives on the quality of crowdwork. We focus on performance-based payments (PBPs), bonus payments awarded to workers for producing high quality work. We design and run randomized behavioral experiments on the popular crowdsourcing platform Amazon Mechanical Turk with the goal of understanding when, where, and why PBPs help, identifying properties of the payment, payment structure, and the task itself that make them most effective. We provide examples of tasks for which PBPs do improve quality. For such tasks, the effectiveness of PBPs is not too sensitive to the threshold for quality required to receive the bonus, while the magnitude of the bonus must be large enough to make the reward salient. We also present examples of tasks for which PBPs do not improve quality. Our results suggest that for PBPs to improve quality, the task must be effort-responsive: the task must allow workers to produce higher quality work by exerting more effort. We also give a simple method to determine if a task is effort-responsive a priori. Furthermore, our experiments suggest that all payments on Mechanical Turk are, to some degree, implicitly performance-based in that workers believe their work may be rejected if their performance is sufficiently poor. Finally, we propose a new model of worker behavior that extends the standard principal-agent model from economics to include a worker's subjective beliefs about his likelihood of being paid, and show that the predictions of this model are in line with our experimental findings. This model may be useful as a foundation for theoretical studies of incentives in crowdsourcing markets. Chien-Ju Ho, Aleksandrs Slivkins, Siddharth Suri, Jennifer Wortman Vaughan |
WWW | 4 |
| 2014 | A general volume-parameterized market making frameworkabstractWe introduce a framework for automated market making for prediction markets, the volume parameterized market (VPM), in which securities are priced based on the market maker's current liabilities as well as the total volume of trade in the market. We provide a set of mathematical tools that can be used to analyze markets in this framework, and show that many existing market makers (including cost-function based markets [Chen and Pennock 2007; Abernethy et al. 2011, 2013], profit-charging markets [Othman and Sandholm 2012], and buy-only markets [Li and Vaughan 2013]) all fall into this framework as special cases. Using the framework, we design a new market maker, the perspective market, that satisfies four desirable properties (worst-case loss, no arbitrage, increasing liquidity, and shrinking spread) in the complex market setting, but fails to satisfy information incorporation. However, we show that the sacrifice of information incorporation is unavoidable: we prove an impossibility result showing that any market maker that prices securities based only on the trade history cannot satisfy all five properties simultaneously. Instead, we show that perspective markets may satisfy a weaker notion that we call center-price information incorporation. Jacob D. Abernethy, Rafael M. Frongillo, Jennifer Wortman Vaughan |
EC | 4 |
| 2014 | Removing arbitrage from wagering mechanismsabstractWe observe that Lambert et al.'s [2008] family of weighted score wagering mechanisms admit arbitrage: participants can extract a guaranteed positive payoff by betting on any prediction within a certain range. In essence, participants leave free money on the table when they ``agree to disagree,'' and as a result, rewards don't necessarily go to the most informed and accurate participants. This observation suggests that when participants have immutable beliefs, it may be possible to design alternative mechanisms in which the center can make a profit by removing this arbitrage opportunity without sacrificing incentive properties such as individual rationality, incentive compatibility, and sybilproofness. We introduce a new family of wagering mechanisms called no-arbitrage wagering mechanisms that retain many of the positive properties of weighted score wagering mechanisms, but with the arbitrage opportunity removed. We show several structural results about the class of mechanisms that satisfy no-arbitrage in conjunction with other properties, and provide examples of no-arbitrage wagering mechanisms with interesting properties. Yiling Chen 0001, Nikhil R. Devanur, David M. Pennock, Jennifer Wortman Vaughan |
EC | 4 |
| 2014 | Adaptive contract design for crowdsourcing markets: bandit algorithms for repeated principal-agent problemsabstractCrowdsourcing markets have emerged as a popular platform for matching available workers with tasks to complete. The payment for a particular task is typically set by the task's requester, and may be adjusted based on the quality of the completed work, for example, through the use of 'bonus' payments. In this paper, we study the requester's problem of dynamically adjusting quality-contingent payments for tasks. We consider a multi-round version of the well-known principal-agent model, whereby in each round a worker makes a strategic choice of the effort level which is not directly observable by the requester. In particular, our formulation significantly generalizes the budget-free online task pricing problems studied in prior work. Chien-Ju Ho, Aleksandrs Slivkins, Jennifer Wortman Vaughan |
EC | 3 |
| 2014 | Market Making with Decreasing Utility for Information
Miroslav Dudík, Rafael M. Frongillo, Jennifer Wortman Vaughan |
UAI | 3 |
| 2014 | Computational social science and social computing
Winter A. Mason, Jennifer Wortman Vaughan, Hanna M. Wallach |
Mach. Learn. | 2 |
| 2013 | Adaptive Task Assignment for Crowdsourced ClassificationabstractCrowdsourcing markets have gained popularity as a tool for inexpensively collecting data from diverse populations of workers. Classification tasks, in which workers provide labels (such as “offensive” or “not offensive”) for instances (such as websites), are among the most common tasks posted, but due to a mix of human error and the overwhelming prevalence of spam, the labels collected are often noisy. This problem is typically addressed by collecting labels for each instance from multiple workers and combining them in a clever way. However, the question of how to choose which tasks to assign to each worker is often overlooked. We investigate the problem of task assignment and label inference for heterogeneous classification tasks. By applying online primal-dual techniques, we derive a provably near-optimal adaptive assignment algorithm. We show that adaptively assigning workers to tasks can lead to more accurate predictions at a lower cost when the available workers are diverse. Chien-Ju Ho, Shahin Jabbari, Jennifer Wortman Vaughan |
ICML (1) | 3 |
| 2013 | Cost function market makers for measurable spacesabstractWe characterize cost function market makers designed to elicit traders' beliefs about the expectations of an infinite set of random variables or the full distribution of a continuous random variable. This characterization is derived from a duality perspective that associates the market maker's liabilities with market beliefs, generalizing the framework of [11,13], but relies on a new subdifferential analysis. It differs from prior approaches in that it allows arbitrary market beliefs, not just those that admit density functions. This allows us to overcome the impossibility results of [10] and design the first automated market maker for betting on the realization of a continuous random variable taking values in {0,1} that has bounded loss without resorting to discretization. Additionally, we show that scoring rules are derived from the same duality and share a close connection with cost functions for eliciting beliefs. Yiling Chen 0001, Michael Ruberry, Jennifer Wortman Vaughan |
EC | 3 |
| 2013 | An axiomatic characterization of adaptive-liquidity market makersabstractPrediction markets offer contingent securities with payoffs linked to future events. The market price of a security reveals information about traders' beliefs, but prediction markets often suffer from low liquidity, which can prevent trades from occurring. One way to inject liquidity into a market is through the use of an automated market maker, an algorithmic agent willing to accept some risk in order to facilitate trades. Abernethy et al. [2013] proposed a general framework for the design of automated market makers, defining a set of axioms that a market should satisfy, and characterizing the class of market makers that satisfy these axioms. However, the liquidity of any market in their class, quantified in terms of the rate at which prices adapt to trades, is fixed a priori and does not change as the volume of trade increases. Othman and Sandholm [2011] proposed a class of liquidity-adaptive markets, but gave little guidance for how to choose a market from within this class. Combining ideas from Abernethy et al. [2013] and Othman and Sandholm [2011], we provide an axiomatic characterization of a parameterized class of automated market makers with adaptive liquidity. A primary advantage of our framework is the ability to analyze important market properties, such as its ability to aggregate information or make a profit, in terms of market parameters, which we do using techniques from convex analysis and geometry. For example, we show that the curvature of the price space can be used to manage a trade-off between information loss and profit. This analysis offers guidance for market designers who wish to choose a particular market maker to implement. Jennifer Wortman Vaughan |
EC | 2 |
| 2012 | Online Task Assignment in Crowdsourcing MarketsabstractWe explore the problem of assigning heterogeneous tasks to workers with different, unknown skill sets in crowdsourcing markets such as Amazon Mechanical Turk. We first formalize the online task assignment problem, in which a requester has a fixed set of tasks and a budget that specifies how many times he would like each task completed. Workers arrive one at a time (with the same worker potentially arriving multiple times), and must be assigned to a task upon arrival. The goal is to allocate workers to tasks in a way that maximizes the total benefit that the requester obtains from the completed work. Inspired by recent research on the online adwords problem, we present a two-phase exploration-exploitation assignment algorithm and prove that it is competitive with respect to the optimal offline algorithm which has access to the unknown skill levels of each worker. We empirically evaluate this algorithm using data collected on Mechanical Turk and show that it performs better than random assignment or greedy algorithms. To our knowledge, this is the first work to extend the online primal-dual technique used in the online adwords problem to a scenario with unknown parameters, and the first to offer an empirical validation of an online primal-dual algorithm. Chien-Ju Ho, Jennifer Wortman Vaughan |
AAAI | 2 |
| 2012 | Designing Informative Securities
Yiling Chen 0001, Michael Ruberry, Jennifer Wortman Vaughan |
UAI | 3 |
| 2011 | An optimization-based framework for automated market-makingabstractWe propose a general framework for the design of securities markets over combinatorial or infinite state or outcome spaces. The framework enables the design of computationally efficient markets tailored to an arbitrary, yet relatively small, space of securities with bounded payoff. We prove that any market satisfying a set of intuitive conditions must price securities via a convex cost function, which is constructed via conjugate duality. Rather than deal with an exponentially large or infinite outcome space directly, our framework only requires optimization over a convex hull. By reducing the problem of automated market making to convex optimization, where many efficient algorithms exist, we arrive at a range of new polynomial-time pricing mechanisms for various problems. We demonstrate the advantages of this framework with the design of some particular markets. We also show that by relaxing the convex hull we can gain computational tractability without compromising the market institution’s bounded budget. Jacob D. Abernethy, Yiling Chen 0001, Jennifer Wortman Vaughan |
EC | 3 |
| 2010 | Regret Minimization With Concept Drift
Koby Crammer, Yishay Mansour, Eyal Even-Dar, Jennifer Wortman Vaughan |
COLT | 4 |
| 2010 | Evolution with Drifting Targets
Varun Kanade, Leslie G. Valiant, Jennifer Wortman Vaughan |
COLT | 3 |
| 2010 | A new understanding of prediction markets via no-regret learningabstractWe explore the striking mathematical connections that exist between market scoring rules, cost function based prediction markets, and no-regret learning. We first show that any cost function based prediction market can be interpreted as an algorithm for the commonly studied problem of learning from expert advice by equating the set of outcomes on which bets are placed in the market with the set of experts in the learning setting, and equating trades made in the market with losses observed by the learning algorithm. If the loss of the market organizer is bounded, this bound can be used to derive an O(√T) regret bound for the corresponding learning algorithm. We then show that the class of markets with convex cost functions exactly corresponds to the class of Follow the Regularized Leader learning algorithms, with the choice of a cost function in the market corresponding to the choice of a regularizer in the learning problem. Finally, we show an equivalence between market scoring rules and prediction markets with convex cost functions. This implies both that any market scoring rule can be implemented as a cost function based market maker, and that market scoring rules can be interpreted naturally as Follow the Regularized Leader algorithms. These connections provide new insight into how it is that commonly studied markets, such as the Logarithmic Market Scoring Rule, can aggregate opinions into accurate estimates of the likelihood of future events. Yiling Chen 0001, Jennifer Wortman Vaughan |
EC | 2 |
| 2010 | Maintaining Equilibria During Exploration in Sponsored Search Auctions
John Langford 0001, Lihong Li 0001, Yevgeniy Vorobeychik, Jennifer Wortman Vaughan |
Algorithmica | 4 |
| 2010 | The true sample complexity of active learning
Maria-Florina Balcan, Steve Hanneke, Jennifer Wortman Vaughan |
Mach. Learn. | 3 |
| 2010 | A theory of learning from different domainsabstractDiscriminative learning methods for classification perform well when training and test data are drawn from the same distribution. Often, however, we have plentiful labeled training data from a source domain but wish to learn a classifier which performs well on a target domain with a different distribution and little or no labeled training data. In this work we investigate two questions. First, under what conditions can a classifier trained from source data be expected to perform well on target data? Second, given a small amount of labeled target data, how should we combine it during training with the large amount of labeled source data to achieve the lowest target error at test time? We address the first question by bounding a classifier’s target error in terms of its source error and the divergence between the two domains. We give a classifier-induced divergence measure that can be estimated from finite, unlabeled samples from the domains. Under the assumption that there exists some hypothesis that performs well in both domains, we show that this quantity together with the empirical source error characterize the target error of a source-trained classifier. We answer the second question by bounding the target error of a model which minimizes a convex combination of the empirical source and target errors. Previous theoretical work has considered minimizing just the source error, just the target error, or weighting instances from the two domains equally. We show how to choose the optimal combination of source and target error as a function of the divergence, the sample sizes of both domains, and the complexity of the hypothesis class. The resulting bound generalizes the previously studied cases and is always at least as tight as a bound which considers minimizing only the target error or an equal weighting of source and target errors. Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira 0003, Jennifer Wortman Vaughan |
Mach. Learn. | 6 |
| 2009 | Censored Exploration and the Dark Pool Problem
Kuzman Ganchev, Michael Kearns, Yuriy Nevmyvaka, Jennifer Wortman Vaughan |
UAI | 4 |
| 2008 | The True Sample Complexity of Active Learning
Maria-Florina Balcan, Steve Hanneke, Jennifer Wortman Vaughan |
COLT | 3 |
| 2008 | Learning from Collective Behavior
Michael Kearns, Jennifer Wortman Vaughan |
COLT | 2 |
| 2008 | Exploration scavengingabstractWe examine the problem of evaluating a policy in the contextual bandit setting using only observations collected during the execution of another policy. We show that policy evaluation can be impossible if the exploration policy chooses actions based on the side information provided at each time step. We then propose and prove the correctness of a principled method for policy evaluation which works when this is not the case, even when the exploration policy is deterministic, as long as each action is explored sufficiently often. We apply this general technique to the problem of offline evaluation of internet advertising policies. Although our theoretical results hold only when the exploration policy chooses ads independent of side information, an assumption that is typically violated by commercial systems, we show how clever uses of the theory provide non-trivial and realistic applications. We also provide an empirical demonstration of the effectiveness of our techniques on real ad placement data. John Langford 0001, Alexander L. Strehl, Jennifer Wortman Vaughan |
ICML | 3 |
| 2008 | Complexity of combinatorial market makersabstractWe analyze the computational complexity of market maker pricing algorithms for combinatorial prediction markets. We focus on Hanson's popular logarithmic market scoring rule market maker (LMSR). Our goal is to implicitly maintain correct LMSR prices across an exponentially large outcome space. We examine both permutation combinatorics, where outcomes are permutations of objects, and Boolean combinatorics, where outcomes are combinations of binary events. We look at three restrictive languages that limit what traders can bet on. Even with severely limited languages, we find that LMSR pricing is #P-hard, even when the same language admits polynomial-time matching without the market maker. We then propose an approximation technique for pricing permutation markets based on an algorithm for online permutation learning. The connections we draw between LMSR pricing and the literature on online learning with expert advice may be of independent interest. Yiling Chen 0001, Lance Fortnow, Nicolas S. Lambert, David M. Pennock, Jennifer Wortman Vaughan |
EC | 5 |
| 2008 | Self-financed wagering mechanisms for forecastingabstractWe examine a class of wagering mechanisms designed to elicit truthful predictions from a group of people without requiring any outside subsidy. We propose a number of desirable properties for wagering mechanisms, identifying one mechanism - weighted-score wagering - that satisfies all of the properties. Moreover, we show that a single-parameter generalization of weighted-score wagering is the only mechanism that satisfies these properties. We explore some variants of the core mechanism based on practical considerations. Nicolas S. Lambert, John Langford 0001, Jennifer Wortman Vaughan, Yiling Chen 0001, Daniel M. Reeves, Yoav Shoham, David M. Pennock |
EC | 3 |
| 2008 | Learning from Multiple Sources
Koby Crammer, Michael Kearns, Jennifer Wortman Vaughan |
J. Mach. Learn. Res. | 3 |
| 2008 | Regret to the best vs. regret to the average
Eyal Even-Dar, Michael Kearns, Yishay Mansour, Jennifer Wortman Vaughan |
Mach. Learn. | 4 |
| 2007 | Regret to the Best vs. Regret to the Average
Eyal Even-Dar, Michael Kearns, Yishay Mansour, Jennifer Wortman Vaughan |
COLT | 4 |
| 2007 | Learning Bounds for Domain AdaptationabstractEmpirical risk minimization offers well-known learning guarantees when training and test data come from the same domain. In the real world, though, we often wish to adapt a classifier from a source domain with a large amount of training data to different target domain with very little training data. In this work we give uniform convergence bounds for algorithms that minimize a convex combination of source and target empirical risk. The bounds explicitly model the inherent trade-off between training on a large but inaccurate source data set and a small but accurate target training set. Our theory also gives results when we have multiple source domains, each of which may have a different number of instances, and we exhibit cases in which minimizing a non-uniform combination of source risks can achieve much lower target error than standard empirical risk minimization. John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira 0003, Jennifer Wortman Vaughan |
NIPS | 5 |
| 2007 | Privacy-Preserving Belief Propagation and SamplingabstractWe provide provably privacy-preserving versions of belief propagation, Gibbs sampling, and other local algorithms — distributed multiparty protocols in which each party or vertex learns only its final local value, and absolutely nothing else. Michael Kearns, Jinsong Tan, Jennifer Wortman Vaughan |
NIPS | 3 |
| 2006 | Risk-Sensitive Online Learning
Eyal Even-Dar, Michael Kearns, Jennifer Wortman Vaughan |
ALT | 3 |
| 2006 | Learning from Multiple SourcesabstractWe consider the problem of learning accurate models from multiple sources of "nearby" data. Given distinct samples from multiple data sources and estimates of the dissimilarities between these sources, we provide a general theory of which samples should be used to learn models for each source. This theory is applicable in a broad decision-theoretic learning framework, and yields results for classification and regression generally, and for density estimation within the exponential family. A key component of our approach is the development of approximate triangle inequalities for expected loss, which may be of independent interest. Koby Crammer, Michael Kearns, Jennifer Wortman Vaughan |
NIPS | 3 |
| 2005 | Learning from Data of Variable QualityabstractWe initiate the study of learning from multiple sources of limited data, each of which may be corrupted at a different rate. We develop a com- plete theory of which data sources should be used for two fundamental problems: estimating the bias of a coin, and learning a classifier in the presence of label noise. In both cases, efficient algorithms are provided for computing the optimal subset of data. Koby Crammer, Michael Kearns, Jennifer Wortman Vaughan |
NIPS | 3 |