Daniel S. Weld

dblp:w/DanielSWeld · also Daniel Sabby Weld · DBLP profile ↗
← Back
172ranked-venue papers
17as first author
26since 2021 · last 2026
0000-0002-3255-0109ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 106 · 14 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 49 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 29 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6Theory of computation · 4
YearPublicationVenuePosition
2026 Generating Literature-Driven Scientific Theories at Scale
abstract
Contemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theory building remain underexplored.In this work, we formulate the problem of synthesizing theories consisting of qualitative and quantitative laws from large corpora of scientific literature.We study theory generation at scale, using 13.7k source papers to synthesize 2.9k theories, examining how generation using literaturegrounding versus parametric knowledge, and accuracy-focused versus novelty-focused generation objectives change theory properties.Our experiments show that, compared to using parametric LLM memory for generation, our literature-supported method creates theories that are significantly better at both matching existing evidence and at predicting future results from 4.6k subsequently-written papers. 1
Peter A. Jansen, Peter Clark, Doug Downey, Daniel S. Weld
ACL (1)4
2026 Cocoa: Co-Planning and Co-Execution with AI Agents
abstract
As AI agents take on increasingly long-running tasks involving sophisticated planning and execution, there is a corresponding need for novel interaction designs that enable deeper human-agent collaboration. However, most prior works leverage human interaction to fix “autonomous” workflows that have yet to become fully autonomous or rigidly treat planning and execution as separate stages. Based on a formative study with 9 researchers using AI to support their work, we propose a design that affords greater flexibility in collaboration, so that users can 1) delegate agency to the user or agent via a collaborative plan where individual steps can be assigned; and 2) interleave planning and execution so that plans can adjust after partial execution. We introduce Cocoa, a system that takes design inspiration from computational notebooks to support complex research tasks. A lab study (n = 16) found that Cocoa enabled steerability without sacrificing ease-of-use, and a week-long field deployment (n = 7) showed how researchers collaborated with Cocoa to accomplish real-world tasks.
K. J. Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S. Weld, Amy X. Zhang, Joseph Chee Chang
CHI7
2026 RDoFlow: Automatically assessing under-specified statistical analyses in HCI
abstract
When designing and analyzing a study, researchers must navigate a large space of methodological decisions, or “researcher degrees of freedom.” If these choices are not preregistered or transparently reported, they can increase the risk of inflated false-positive rates and exaggerated effect sizes, undermining scientific credibility. Drawing on psychology research that characterizes these degrees of freedom, we create a protocol for scoring how hypotheses are reported in the HCI literature (ReportDoF). We manually apply ReportDoF to 100 hypotheses from HCI texts authored between 2015-2025, including both preregistrations and papers. Based on this experience, we contribute an LLM workflow and proof-of-concept interactive interface (RDoFlow) that applies ReportDoF to new texts, enabling large-scale analysis of the composition and quality of reported analysis specifications. For example, RDoFlow reveals that HCI research more frequently tests multiple dependent variables for a single hypothesis than psychology research does—a practice that increases the risk of false positives.
Madeleine Grunde-McLaughlin, Weixuan Liu, Ria Patil, Nino Migineishvili, Emily Reif, Ranjay Krishna, Daniel S. Weld, Jeffrey Heer
IUI7
2025 Toward Living Narrative Reviews: An Empirical Study of the Processes and Challenges in Updating Survey Articles in Computing Research
Raymond Fok, Alexa F. Siu, Daniel S. Weld
CHI3
2025 Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows
abstract
LLM chains enable complex tasks by decomposing work into a sequence of subtasks. Similarly, the more established techniques of crowdsourcing workflows decompose complex tasks into smaller tasks for human crowdworkers. Chains address LLM errors analogously to the way crowdsourcing workflows address human error. To characterize opportunities for LLM chaining, we survey 107 papers across the crowdsourcing and chaining literature to construct a design space for chain development. The design space covers a designer’s objectives and the tactics used to build workflows. We then surface strategies that mediate how workflows use tactics to achieve objectives. To explore how techniques from crowdsourcing may apply to chaining, we adapt crowdsourcing workflows to implement LLM chains across three case studies: creating a taxonomy, shortening text, and writing a short story. From the design space and our case studies, we identify takeaways for effective chain design and raise implications for future research and development.
Madeleine Grunde-McLaughlin, Michelle S. Lam, Ranjay Krishna, Daniel S. Weld, Jeffrey Heer
ACM Trans. Comput. Hum. Interact.4
2024 ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models
abstract
Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim, Daniel S Weld, Joseph Chee Chang, Kyle Lo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim 0001, Daniel S. Weld, Joseph Chee Chang, Kyle Lo
EMNLP7
2024 Qlarify: Recursively Expandable Abstracts for Dynamic Information Retrieval over Scientific Papers
abstract
Navigating the vast scientific literature often starts with browsing a paper’s abstract. However, when a reader seeks additional information, not present in the abstract, they face a costly cognitive chasm during their dive into the full text. To bridge this gap, we introduce recursively expandable abstracts, a novel interaction paradigm that dynamically expands abstracts by progressively incorporating additional information from the papers’ full text. This lightweight interaction allows scholars to specify their information needs by quickly brushing over the abstract or selecting AI-suggested expandable entities. Relevant information is synthesized using a retrieval-augmented generation approach, presented as a fluid, threaded expansion of the abstract, and made efficiently verifiable via attribution to relevant source-passages in the paper. Through a series of user studies, we demonstrate the utility of recursively expandable abstracts and identify future opportunities to support low-effort and just-in-time exploration of long-form information contexts through LLM-powered interactions.
Raymond Fok, Joseph Chee Chang, Tal August, Amy X. Zhang, Daniel S. Weld
UIST5
2024 Accelerating Scientific Paper Skimming with Augmented Intelligence Through Customizable Faceted Highlights
abstract
Scholars need to keep up with an exponentially increasing flood of scientific papers. To aid this challenge, we introduce Scim , a novel intelligent interface that helps scholars skim papers to rapidly review and gain a cursory understanding of its contents. Scim supports the skimming process by highlighting salient content within a paper, directing a scholar’s attention. These automatically-extracted highlights are faceted by content type, evenly-distributed across a paper, and have a density configurable by scholars. We evaluate Scim with an in-lab usability study and a longitudinal diary study, revealing how its highlights facilitate the more efficient construction of a conceptualization of a paper. Finally, we describe the process of scaling highlights from their conception within Scim , a research prototype, to production on over 521,000 papers within the Semantic Reader, a publicly-available augmented reading interface for scientific papers. We conclude by discussing design considerations and tensions for the design of future skimming tools with augmented intelligence.
Raymond Fok, Luca Soldaini, Cassidy Trier, Erin Bransom, Kelsey MacMillan, Evie Yu-Yen Cheng, Hita Kambhamettu, Jonathan Bragg, Kyle Lo, Marti A. Hearst, Andrew Head, Daniel S. Weld
ACM Trans. Interact. Intell. Syst.12
2023 CiteSee: Augmenting Citations in Scientific Papers with Persistent and Personalized Historical Context
abstract
When reading a scholarly article, inline citations help researchers contextualize the current article and discover relevant prior work. However, it can be challenging to prioritize and make sense of the hundreds of citations encountered during literature reviews. This paper introduces CiteSee, a paper reading tool that leverages a user’s publishing, reading, and saving activities to provide personalized visual augmentations and context around citations. First, CiteSee connects the current paper to familiar contexts by surfacing known citations a user had cited or opened. Second, CiteSee helps users prioritize their exploration by highlighting relevant but unknown citations based on saving and reading history. We conducted a lab study that suggests CiteSee is significantly more effective for paper discovery than three baselines. A field deployment study shows CiteSee helps participants keep track of their explorations and leads to better situational awareness and increased paper discovery via inline citation when conducting real-world literature reviews.
Joseph Chee Chang, Amy X. Zhang, Jonathan Bragg, Andrew Head, Kyle Lo, Doug Downey, Daniel S. Weld
CHI7
2023 Scim: Intelligent Skimming Support for Scientific Papers
abstract
Scholars need to keep up with an exponentially increasing flood of scientific papers. To aid this challenge, we introduce Scim, a novel intelligent interface that helps experienced researchers skim – or rapidly review – a paper to attain a cursory understanding of its contents. Scim supports the skimming process by highlighting salient paper contents in order to direct a reader’s attention. The system’s highlights are faceted by content type, evenly distributed across a paper, and have a density configurable by readers at both the global and local level. We evaluate Scim with both an in-lab usability study and a longitudinal diary study, revealing how its highlights facilitate the more efficient construction of a conceptualization of a paper. We conclude by discussing design considerations and tensions for the design of future intelligent skimming tools.
Raymond Fok, Hita Kambhamettu, Luca Soldaini, Jonathan Bragg, Kyle Lo, Marti A. Hearst, Andrew Head, Daniel S. Weld
IUI8
2023 ScatterShot: Interactive In-context Example Curation for Text Transformation
abstract
The in-context learning capabilities of LLMs like GPT-3 allow annotators to customize an LLM to their specific tasks with a small number of examples. However, users tend to include only the most obvious patterns when crafting examples, resulting in underspecified in-context functions that fall short on unseen cases. Further, it is hard to know when “enough” examples have been included even for known patterns. In this work, we present ScatterShot, an interactive system for building high-quality demonstration sets for in-context learning. ScatterShot iteratively slices unlabeled data into task-specific patterns, samples informative inputs from underexplored or not-yet-saturated slices in an active learning manner, and helps users label more efficiently with the help of an LLM and the current example set. In simulation studies on two text perturbation scenarios, ScatterShot sampling improves the resulting few-shot functions by 4-5 percentage points over random sampling, with less variance as more examples are added. In a user study, ScatterShot greatly helps users in covering different patterns in the input space and labeling in-context examples more efficiently, resulting in better in-context learning and less user effort.
Sherry Tongshuang Wu, Hua Shen 0005, Daniel S. Weld, Jeffrey Heer, Marco Túlio Ribeiro
IUI3
2023 LIMEADE: From AI Explanations to Advice Taking
abstract
Research in human-centered AI has shown the benefits of systems that can explain their predictions. Methods that allow AI to take advice from humans in response to explanations are similarly useful. While both capabilities are well developed for transparent learning models (e.g., linear models and GA 2 Ms) and recent techniques (e.g., LIME and SHAP) can generate explanations for opaque models, little attention has been given to advice methods for opaque models. This article introduces LIMEADE, the first general framework that translates both positive and negative advice (expressed using high-level vocabulary such as that employed by post hoc explanations) into an update to an arbitrary, underlying opaque model. We demonstrate the generality of our approach with case studies on 70 real-world models across two broad domains: image classification and text recommendation. We show that our method improves accuracy compared to a rigorous baseline on the image classification domains. For the text modality, we apply our framework to a neural recommender system for scientific papers on a public website; our user study shows that our framework leads to significantly higher perceived user control, trust, and satisfaction.
Benjamin Charles Germain Lee, Doug Downey, Kyle Lo, Daniel S. Weld
ACM Trans. Interact. Intell. Syst.4
2022 A Search Engine for Discovery of Scientific Challenges and Directions
abstract
Keeping track of scientific challenges, advances and emerging directions is a fundamental part of research. However, researchers face a flood of papers that hinders discovery of important knowledge. In biomedicine, this directly impacts human lives. To address this problem, we present a novel task of extraction and search of scientific challenges and directions, to facilitate rapid knowledge discovery. We construct and release an expert-annotated corpus of texts sampled from full-length papers, labeled with novel semantic categories that generalize across many types of challenges and directions. We focus on a large corpus of interdisciplinary work relating to the COVID-19 pandemic, ranging from biomedicine to areas such as AI and economics. We apply a model trained on our data to identify challenges and directions across the corpus and build a dedicated search engine. In experiments with 19 researchers and clinicians using our system, we outperform a popular scientific search engine in assisting knowledge discovery. Finally, we show that models trained on our resource generalize to the wider biomedical domain and to AI papers, highlighting its broad utility. We make our data, model and search engine publicly available.
Dan Lahav, Jon Saad-Falcon, Bailey Kuehl, Sophie Johnson, Sravanthi Parasa, Noam Shomron, Polo Chau, Diyi Yang, Eric Horvitz, Daniel S. Weld, Tom Hope
AAAI10
2022 From Who You Know to What You Read: Augmenting Scientific Recommendations with Implicit Social Networks
abstract
The ever-increasing pace of scientific publication necessitates methods for quickly identifying relevant papers. While neural recommenders trained on user interests can help, they still result in long, monotonous lists of suggested papers. To improve the discovery experience we introduce multiple new methods for augmenting recommendations with textual relevance messages that highlight knowledge-graph connections between recommended papers and a user’s publication and interaction history. We explore associations mediated by author entities and those using citations alone. In a large-scale, real-world study, we show how our approach significantly increases engagement—and future engagement when mediated by authors—without introducing bias towards highly-cited authors. To expand message coverage for users with less publication or interaction history, we develop a novel method that highlights connections with proxy authors of interest to users and evaluate it in a controlled lab study. Finally, we synthesize design implications for future graph-based messages.
Hyeonsu B. Kang, Rafal Kocielnik, Andrew Head, Jiangjiang Yang, Matt Latzke, Aniket Kittur, Daniel S. Weld, Doug Downey, Jonathan Bragg
CHI7
2022 Bursting Scientific Filter Bubbles: Boosting Innovation via Novel Author Discovery
abstract
Isolated silos of scientific research and the growing challenge of information overload limit awareness across the literature and hinder innovation. Algorithmic curation and recommendation, which often prioritize relevance, can further reinforce these informational “filter bubbles.” In response, we describe Bridger, a system for facilitating discovery of scholars and their work. We construct a faceted representation of authors with information gleaned from their papers and inferred author personas, and use it to develop an approach that locates commonalities and contrasts between scientists to balance relevance and novelty. In studies with computer science researchers, this approach helps users discover authors considered useful for generating novel research directions. We also demonstrate an approach for displaying information about authors, boosting the ability to understand the work of new, unfamiliar scholars. Our analysis reveals that Bridger connects authors who have different citation profiles and publish in different venues, raising the prospect of bridging diverse scientific communities.
Jason Portenoy, Marissa Radensky, Jevin D. West, Eric Horvitz, Daniel S. Weld, Tom Hope
CHI5
2022 GENIE: Toward Reproducible and Standardized Human Evaluation for Text Generation
abstract
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, Daniel Weld. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi 0001, Noah A. Smith, Daniel S. Weld
EMNLP8
2022 CiteRead: Integrating Localized Citation Contexts into Scientific Paper Reading
abstract
When reading a scholarly paper, scientists oftentimes wish to understand how follow-on work has built on or engages with what they are reading. While a paper itself can only discuss prior work, some scientific search engines can provide a list of all subsequent citing papers; unfortunately, they are undifferentiated and disconnected from the contents of the original reference paper. In this work, we introduce a novel paper reading experience that integrates relevant information about follow-on work directly into a paper, allowing readers to learn about newer papers and see how a paper is discussed by its citing papers in the context of the reference paper. We built a tool, called CiteRead, that implements the following three contributions: 1) automated techniques for selecting important citing papers, building on results from a formative study we conducted, 2) an automated process for localizing commentary provided by citing papers to a place in the reference paper, and 3) an interactive experience that allows readers to seamlessly alternate between the reference paper and information from citing papers (e.g., citation sentences), placed in the margins. Based on a user study with 12 scientists, we found that in comparison to having just a list of citing papers and their citation sentences, the use of CiteRead while reading allows for better comprehension and retention of information about follow-on work.
Napol Rachatasumrit, Jonathan Bragg, Amy X. Zhang, Daniel S. Weld
IUI4
2022 FeedLens: Polymorphic Lenses for Personalizing Exploratory Search over Knowledge Graphs
abstract
The vast scale and open-ended nature of knowledge graphs (KGs) make exploratory search over them cognitively demanding for users. We introduce a new technique, polymorphic lenses, that improves exploratory search over a KG by obtaining new leverage from the existing preference models that KG-based systems maintain for recommending content. The approach is based on a simple but powerful observation: in a KG, preference models can be re-targeted to recommend not only entities of a single base entity type (e.g., papers in the scientific literature KG, products in an e-commerce KG), but also all other types (e.g., authors, conferences, institutions; sellers, buyers). We implement our technique in a novel system, FeedLens, which is built over Semantic Scholar, a production system for navigating the scientific literature KG. FeedLens reuses the existing preference models on Semantic Scholar—people’s curated research feeds—as lenses for exploratory search. Semantic Scholar users can curate multiple feeds/lenses for different topics of interest, e.g., one for human-centered AI and another for document embeddings. Although these lenses are defined in terms of papers, FeedLens re-purposes them to also guide search over authors, institutions, venues, etc. Our system design is based on feedback from intended users via two pilot surveys (n = 17 and n = 13, respectively). We compare FeedLens and Semantic Scholar via a third (within-subjects) user study (n = 15) and find that FeedLens increases user engagement while reducing the cognitive effort required to complete a short literature review task. Our qualitative results also highlight people’s preference for this more effective exploratory search experience enabled by FeedLens.
Harmanpreet Kaur, Doug Downey, Amanpreet Singh, Evie Yu-Yen Cheng, Daniel S. Weld, Jonathan Bragg
UIST5
2022 VILA: Improving Structured Content Extraction from Scientific PDFs Using Visual Layout Groups
abstract
Abstract Accurately extracting structured content from PDFs is a critical first step for NLP over scientific papers. Recent work has improved extraction accuracy by incorporating elementary layout information, for example, each token’s 2D position on the page, into language model pretraining. We introduce new methods that explicitly model VIsual LAyout (VILA) groups, that is, text lines or text blocks, to further improve performance. In our I-VILA approach, we show that simply inserting special tokens denoting layout group boundaries into model inputs can lead to a 1.9% Macro F1 improvement in token classification. In the H-VILA approach, we show that hierarchical encoding of layout-groups can result in up to 47% inference time reduction with less than 0.8% Macro F1 loss. Unlike prior layout-aware approaches, our methods do not require expensive additional pretraining, only fine-tuning, which we show can reduce training cost by up to 95%. Experiments are conducted on a newly curated evaluation suite, S2-VLUE, that unifies existing automatically labeled datasets and includes a new dataset of manual annotations covering diverse papers from 19 scientific disciplines. Pre-trained weights, benchmark datasets, and source code are available at https://github.com/allenai/VILA.
Shannon Shen 0001, Kyle Lo, Lucy Lu Wang, Bailey Kuehl, Daniel S. Weld, Doug Downey
Trans. Assoc. Comput. Linguistics5
2021 Is the Most Accurate AI the Best Teammate? Optimizing AI for Teamwork
abstract
AI practitioners typically strive to develop the most accurate systems, making an implicit assumption that the AI system will function autonomously. However, in practice, AI systems often are used to provide advice to people in domains ranging from criminal justice and finance to healthcare. In such AI-advised decision making, humans and machines form a team, where the human is responsible for making final decisions. But is the most accurate AI the best teammate? We argue "not necessarily" --- predictable performance may be worth a slight sacrifice in AI accuracy. Instead, we argue that AI systems should be trained in a human-centered manner, directly optimized for team performance. We study this proposal for a specific type of human-AI teaming, where the human overseer chooses to either accept the AI recommendation or solve the task themselves. To optimize the team performance for this setting we maximize the team's expected utility, expressed in terms of the quality of the final decision, cost of verifying, and individual accuracies of people and machines. Our experiments with linear and non-linear models on real-world, high-stakes datasets show that the most accuracy AI may not lead to highest team performance and show the benefit of modeling teamwork during training through improvements in expected team utility across datasets, considering parameters such as human skill and the cost of mistakes. We discuss the shortcoming of current optimization approaches beyond well-studied loss functions such as log-loss, and encourage future work on AI optimization problems motivated by human-AI collaboration.
Gagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz, Daniel S. Weld
AAAI5
2021 Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models
abstract
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, Daniel Weld. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Sherry Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld
ACL/IJCNLP (1)4
2021 SciA11y: Converting Scientific Papers to Accessible HTML
abstract
We present SciA11y, a system that renders inaccessible scientific paper PDFs into HTML. SciA11y uses machine learning models to extract and understand the content of scientific PDFs, and reorganizes the resulting paper components into a form that better supports skimming and scanning for blind and low vision (BLV) readers. SciA11y adds navigation features such as tagged headings, a table of contents, and bidirectional links between inline citations and references, which allow readers to resolve citations without losing their context. A set of 1.5 million open access papers are processed and available at https://scia11y.org/. This system is a first step in addressing scientific PDF accessibility, and may significantly improve the experience of paper reading for BLV users.
Lucy Lu Wang, Isabel Cachola, Jonathan Bragg, Evie Yu-Yen Cheng, Chelsea Haupt, Matt Latzke, Bailey Kuehl, Madeleine van Zuylen, Linda Wagner, Daniel S. Weld
ASSETS10
2021 Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
abstract
Many researchers motivate explainable AI with studies showing that human-AI team performance on decision-making tasks improves when the AI explains its recommendations. However, prior studies observed improvements from explanations only when the AI, alone, outperformed both the human and the best team. Can explanations help lead to complementary performance, where team accuracy is higher than either the human or the AI working solo? We conduct mixed-method user studies on three datasets, where an AI with accuracy comparable to humans helps participants solve a task (explaining itself in some conditions). While we observed complementary improvements from AI augmentation, they were not increased by explanations. Rather, explanations increased the chance that humans will accept the AI’s recommendation, regardless of its correctness. Our result poses new challenges for human-centered AI: Can we develop explanatory approaches that encourage appropriate trust in AI, and therefore help generate (or improve) complementary performance?
Gagan Bansal, Sherry Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Túlio Ribeiro, Daniel S. Weld
CHI8
2021 Augmenting Scientific Papers with Just-in-Time, Position-Sensitive Definitions of Terms and Symbols
abstract
Despite the central importance of research papers to scientific progress, they can be difficult to read. Comprehension is often stymied when the information needed to understand a passage resides somewhere else—in another section, or in another paper. In this work, we envision how interfaces can bring definitions of technical terms and symbols to readers when and where they need them most. We introduce ScholarPhi, an augmented reading interface with four novel features: (1) tooltips that surface position-sensitive definitions from elsewhere in a paper, (2) a filter over the paper that “declutters” it to reveal how the term or symbol is used across the paper, (3) automatic equation diagrams that expose multiple definitions in parallel, and (4) an automatically generated glossary of important terms and symbols. A usability study showed that the tool helps researchers of all experience levels read papers. Furthermore, researchers were eager to have ScholarPhi’s definitions available to support their everyday reading.
Andrew Head, Kyle Lo, Dongyeop Kang, Raymond Fok, Sam Skjonsberg, Daniel S. Weld, Marti A. Hearst
CHI6
2021 Extracting a Knowledge Base of Mechanisms from COVID-19 Papers
abstract
Tom Hope, Aida Amini, David Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel Weld, Roy Schwartz, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tom Hope, Aida Amini, Dave Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel S. Weld, Roy Schwartz 0001, Hannaneh Hajishirzi
NAACL-HLT7
2021 Goldilocks: Consistent Crowdsourced Scalar Annotations with Relative Uncertainty
abstract
Human ratings have become a crucial resource for training and evaluating machine learning systems. However, traditional elicitation methods for absolute and comparative rating suffer from issues with consistency and often do not distinguish between uncertainty due to disagreement between annotators and ambiguity inherent to the item being rated. In this work, we present Goldilocks, a novel crowd rating elicitation technique for collecting calibrated scalar annotations that also distinguishes inherent ambiguity from inter-annotator disagreement. We introduce two main ideas: grounding absolute rating scales with examples and using a two-step bounding process to establish a range for an item's placement. We test our designs in three domains: judging toxicity of online comments, estimating satiety of food depicted in images, and estimating age based on portraits. We show that (1) Goldilocks can improve consistency in domains where interpretation of the scale is not universal, and that (2) representing items with ranges lets us simultaneously capture different sources of uncertainty leading to better estimates of pairwise relationship distributions.
Quanze Chen, Daniel S. Weld, Amy X. Zhang
Proc. ACM Hum. Comput. Interact.2
2020 SPECTER: Document-level Representation Learning using Citation-informed Transformers
abstract
Representation learning is a critical ingredient for natural language processing systems.Recent Transformer language models like BERT learn powerful textual representations, but these models are targeted towards token-and sentence-level training objectives and do not leverage information on inter-document relatedness, which limits their document-level representation power.For applications on scientific documents, such as classification and recommendation, the embeddings power strong performance on end tasks.We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph.Unlike existing pretrained language models, SPECTER can be easily applied to downstream applications without task-specific fine-tuning.Additionally, to encourage further research on document-level models, we introduce SCIDOCS, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation.We show that SPECTER outperforms a variety of competitive baselines on the benchmark.1
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, Daniel S. Weld
ACL5
2020 S2ORC: The Semantic Scholar Open Research Corpus
abstract
We introduce S2ORC, 1 a large corpus of 81.1M English-language academic papers spanning many academic disciplines.The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for 8.1M open access papers.Full text is annotated with automaticallydetected inline mentions of citations, figures, and tables, each linked to their corresponding paper objects.In S2ORC, we aggregate papers from hundreds of academic publishers and digital archives into a unified source, and create the largest publicly-available collection of machine-readable academic text to date.We hope this resource will facilitate research and development of tools and tasks for text mining over academic text.
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, Daniel S. Weld
ACL5
2020 No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive ML
abstract
Automatically generated explanations of how machine learning (ML) models reason can help users understand and accept them. However, explanations can have unintended consequences: promoting over-reliance or undermining trust. This paper investigates how explanations shape users' perceptions of ML models with or without the ability to provide feedback to them: (1) does revealing model flaws increase users' desire to "fix" them; (2) does providing explanations cause users to believe - wrongly - that models are introspective, and will thus improve over time. Through two controlled experiments - varying model quality - we show how the combination of explanations and user feedback impacted perceptions, such as frustration and expectations of model improvement. Explanations without opportunity for feedback were frustrating with a lower quality model, while interactions between explanation and feedback for the higher quality model suggest that detailed feedback should not be requested without explanation. Users expected model correction, regardless of whether they provided feedback or received explanations.
Alison Smith-Renner, Ron Fan, Melissa Birchfield, Sherry Tongshuang Wu, Jordan L. Boyd-Graber, Daniel S. Weld, Leah Findlater
CHI6
2020 The Newspaper Navigator Dataset: Extracting Headlines and Visual Content from 16 Million Historic Newspaper Pages in Chronicling America
abstract
Chronicling America is a product of the National Digital Newspaper Program, a partnership between the Library of Congress and the National Endowment for the Humanities to digitize historic American newspapers. Over 16 million pages have been digitized to date, complete with high-resolution images and machine-readable METS/ALTO OCR. Of considerable interest to Chronicling America users is a semantified corpus, complete with extracted visual content and headlines. To accomplish this, we introduce a visual content recognition model trained on bounding box annotations collected as part of the Library of Congress's Beyond Words crowdsourcing initiative and augmented with additional annotations including those of headlines and advertisements. We describe our pipeline that utilizes this deep learning model to extract 7 classes of visual content: headlines, photographs, illustrations, maps, comics, editorial cartoons, and advertisements, complete with textual content such as captions derived from the METS/ALTO OCR, as well as image embeddings. We report the results of running the pipeline on 16.3 million pages from the Chronicling America corpus and describe the resulting Newspaper Navigator dataset, the largest dataset of extracted visual content from historic newspapers ever produced. The Newspaper Navigator dataset, finetuned visual content recognition model, and all source code are placed in the public domain for unrestricted re-use.
Benjamin Charles Germain Lee, Jaime Mears, Eileen Jakeway, Meghan Ferriter, Chris Adams, Nathan Yarasavage, Deborah Thomas, Kate Zwaard, Daniel S. Weld
CIKM9
2020 High-Precision Extraction of Emerging Concepts from Scientific Literature
abstract
Identification of new concepts in scientific literature can help power faceted search, scientific trend analysis, knowledge-base construction, and more, but current methods are lacking. Manual identification can't keep up with the torrent of new publications, while the precision of existing automatic techniques is too low for many applications. We present an unsupervised concept extraction method for scientific literature that achieves much higher precision than previous work. Our approach relies on a simple but novel intuition: each scientific concept is likely to be introduced or popularized by a single paper that is disproportionately cited by subsequent papers mentioning the concept. From a corpus of computer science papers on arXiv, we find that our method achieves a [email protected] of 99%, compared to 86% for prior work, and a substantially better precision-yield trade-off across the top 15,000 extractions. To stimulate research in this area, we release our code and data.
Daniel King, Doug Downey, Daniel S. Weld
SIGIR3
2020 SpanBERT: Improving Pre-training by Representing and Predicting Spans
abstract
We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the span boundary representations to predict the entire content of the masked span, without relying on the individual token representations within it. SpanBERT consistently outperforms BERT and our better-tuned baselines, with substantial gains on span selection tasks such as question answering and coreference resolution. In particular, with the same training data and model size as BERT large , our single model obtains 94.6% and 88.7% F1 on SQuAD 1.1 and 2.0 respectively. We also achieve a new state of the art on the OntoNotes coreference resolution task (79.6% F1), strong performance on the TACRED relation extraction benchmark, and even gains on GLUE. 1
Mandar Joshi, Danqi Chen 0001, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, Omer Levy
Trans. Assoc. Comput. Linguistics4
2019 Updates in Human-AI Teams: Understanding and Addressing the Performance/Compatibility Tradeoff
abstract
AI systems are being deployed to support human decision making in high-stakes domains such as healthcare and criminal justice. In many cases, the human and AI form a team, in which the human makes decisions after reviewing the AI’s inferences. A successful partnership requires that the human develops insights into the performance of the AI system, including its failures. We study the influence of updates to an AI system in this setting. While updates can increase the AI’s predictive performance, they may also lead to behavioral changes that are at odds with the user’s prior experiences and confidence in the AI’s inferences. We show that updates that increase AI performance may actually hurt team performance. We introduce the notion of the compatibility of an AI update with prior user experience and present methods for studying the role of compatibility in human-AI teams. Empirical results on three high-stakes classification tasks show that current machine learning algorithms do not produce compatible updates. We propose a re-training objective to improve the compatibility of an update by penalizing new errors. The objective offers full leverage of the performance/compatibility tradeoff across different datasets, enabling more compatible yet accurate updates.
Gagan Bansal, Besmira Nushi, Ece Kamar, Daniel S. Weld, Walter S. Lasecki, Eric Horvitz
AAAI4
2019 Errudite: Scalable, Reproducible, and Testable Error Analysis
abstract
Though error analysis is crucial to understanding and improving NLP models, the common practice of manual, subjective categorization of a small sample of errors can yield biased and incomplete conclusions.This paper codifies model and task agnostic principles for informative error analysis, and presents Errudite, an interactive tool for better supporting this process.First, error groups should be precisely defined for reproducibility; Errudite supports this with an expressive domainspecific language.Second, to avoid spurious conclusions, a large set of instances should be analyzed, including both positive and negative examples; Errudite enables systematic grouping of relevant instances with filtering queries.Third, hypotheses about the cause of errors should be explicitly tested; Errudite supports this via automated counterfactual rewriting.We validate our approach with a user study, finding that Errudite (1) enables users to perform high quality and reproducible error analyses with less effort, (2) reveals substantial ambiguities in prior published error analyses practices, and (3) enhances the error analysis experience by allowing users to test and revise prior beliefs.
Sherry Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld
ACL (1)4
2019 Guidelines for Human-AI Interaction
abstract
Advances in artificial intelligence (AI) frame opportunities and challenges for user interface design. Principles for human-AI interaction have been discussed in the human-computer interaction community for over two decades, but more study and innovation are needed in light of advances in AI and the growing uses of AI technologies in human-facing applications. We propose 18 generally applicable design guidelines for human-AI interaction. These guidelines are validated through multiple rounds of evaluation including a user study with 49 design practitioners who tested the guidelines against 20 popular AI-infused products. The results verify the relevance of the guidelines over a spectrum of interaction scenarios and reveal gaps in our knowledge, highlighting opportunities for further research. Based on the evaluations, we believe the set of design guidelines can serve as a resource to practitioners working on the design of applications and features that harness AI technologies, and to researchers interested in the further development of human-AI interaction design principles.
Saleema Amershi, Daniel S. Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi T. Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, Eric Horvitz
CHI2
2019 Cicero: Multi-Turn, Contextual Argumentation for Accurate Crowdsourcing
abstract
Traditional approaches for ensuring high quality crowdwork have failed to achieve high-accuracy on difficult problems. Aggregating redundant answers often fails on the hardest problems when the majority is confused. Argumentation has been shown to be effective in mitigating these drawbacks. However, existing argumentation systems only support limited interactions and show workers general justifications, not context-specific arguments targeted to their reasoning. This paper presents Cicero, a new workflow that improves crowd accuracy on difficult tasks by engaging workers in multi-turn, contextual discussions through real-time, synchronous argumentation. Our experiments show that compared to previous argumentation systems which only improve the average individual worker accuracy by 6.8 percentage points on the Relation Extraction domain, our workflow achieves 16.7 percentage point improvement. Furthermore, previous argumentation approaches don't apply to tasks with many possible answers; in contrast, Cicero works well in these cases, raising accuracy from 66.7% to 98.8% on the Codenames domain.
Quanze Chen, Jonathan Bragg, Lydia B. Chilton, Daniel S. Weld
CHI4
2019 Pretrained Language Models for Sequential Sentence Classification
abstract
Arman Cohan, Iz Beltagy, Daniel King, Bhavana Dalvi, Dan Weld. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arman Cohan, Iz Beltagy, Daniel King, Bhavana Dalvi, Daniel S. Weld
EMNLP/IJCNLP (1)5
2019 BERT for Coreference Resolution: Baselines and Analysis
abstract
Mandar Joshi, Omer Levy, Luke Zettlemoyer, Daniel Weld. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mandar Joshi, Omer Levy, Luke Zettlemoyer, Daniel S. Weld
EMNLP/IJCNLP (1)4
2019 Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance
abstract
Decisions made by human-AI teams (e.g., AI-advised humans) are increasingly common in high-stakes domains such as healthcare, criminal justice, and finance. Achieving high team performance depends on more than just the accuracy of the AI system: Since the human and the AI may have different expertise, the highest team performance is often reached when they both know how and when to complement one another. We focus on a factor that is crucial to supporting such complementary: the human’s mental model of the AI capabilities, specifically the AI system’s error boundary (i.e. knowing “When does the AI err?”). Awareness of this lets the human decide when to accept or override the AI’s recommendation. We highlight two key properties of an AI’s error boundary, parsimony and stochasticity, and a property of the task, dimensionality. We show experimentally how these properties affect humans’ mental models of AI capabilities and the resulting team performance. We connect our evaluations to related work and propose goals, beyond accuracy, that merit consideration during model selection and optimization to improve overall human-AI team performance.
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S. Lasecki, Daniel S. Weld, Eric Horvitz
HCOMP5
2019 Local Decision Pitfalls in Interactive Machine Learning: An Investigation into Feature Selection in Sentiment Analysis
abstract
Tools for Interactive Machine Learning (IML) enable end users to update models in a “rapid, focused, and incremental”—yet local—manner. In this work, we study the question of local decision making in an IML context around feature selection for a sentiment classification task. Specifically, we characterize the utility of interactive feature selection through a combination of human-subjects experiments and computational simulations. We find that, in expectation, interactive modification fails to improve model performance and may hamper generalization due to overfitting. We examine how these trends are affected by the dataset, learning algorithm, and the training set size. Across these factors we observe consistent generalization issues. Our results suggest that rapid iterations with IML systems can be dangerous if they encourage local actions divorced from global context, degrading overall model performance. We conclude by discussing the implications of our feature selection results to the broader area of IML systems and research.
Sherry Tongshuang Wu, Daniel S. Weld, Jeffrey Heer
ACM Trans. Comput. Hum. Interact.2
2018 A Coverage-Based Utility Model for Identifying Unknown Unknowns
abstract
A classifier’s low confidence in prediction is often indicative of whether its prediction will be wrong; in this case, inputs are called known unknowns. In contrast, unknown unknowns (UUs) are inputs on which a classifier makes a high confidence mistake. Identifying UUs is especially important in safety-critical domains like medicine (diagnosis) and law (recidivism prediction). Previous work by Lakkaraju et al. (2017) on identifying unknown unknowns assumes that the utility of each revealed UU is independent of the others, rather than considering the set holistically. While this assumption yields an efficient discovery algorithm, we argue that it produces an incomplete understanding of the classifier’s limitations. In response, this paper proposes a new class of utility models that rewards how well the discovered UUs cover (or "explain") a sample distribution of expected queries. Although choosing an optimal cover is intractable, even if the UUs were known, our utility model is monotone submodular, affording a greedy discovery strategy. Experimental results on four datasets show that our method outperforms bandit-based approaches and achieves within 60.9% utility of an omniscient, tractable upper bound.
Gagan Bansal, Daniel S. Weld
AAAI2
2018 Active Learning with Unbalanced Classes and Example-Generation Queries
abstract
Machine learning in real-world high-skew domains is difficult, because traditional strategies for crowdsourcing labeled training examples are ineffective at locating the scarce minority-class examples. For example, both random sampling and traditional active learning (which reduces to random sampling when just starting) will most likely recover very few minority-class examples. To bootstrap the machine learning process, researchers have proposed tasking the crowd with finding or generating minority-class examples, but such strategies have their weaknesses as well. They are unnecessarily expensive in well-balanced domains, and they often yield samples from a biased distribution that is unrepresentative of the one being learned.This paper extends the traditional active learning framework by investigating the problem of intelligently switching between various crowdsourcing strategies for obtaining labeled training examples in order to optimally train a classifier. We start by analyzing several such strategies (e.g., annotate an example, generate a minority-class example, etc.), and then develop a novel, skew-robust algorithm, called MB-CB, for the control problem. Experiments show that our method outperforms state-of-the-art GL-Hybrid by up to 14.3 points in F1 AUC, across various domains and class-frequency settings.
Christopher H. Lin, Mausam, Daniel S. Weld
HCOMP3
2018 Sprout: Crowd-Powered Task Design for Crowdsourcing
abstract
While crowdsourcing enables data collection at scale, ensuring high-quality data remains a challenge. In particular, effective task design underlies nearly every reported crowdsourcing success, yet remains difficult to accomplish. Task design is hard because it involves a costly iterative process: identifying the kind of work output one wants, conveying this information to workers, observing worker performance, understanding what remains ambiguous, revising the instructions, and repeating the process until the resulting output is satisfactory. To facilitate this process, we propose a novel meta-workflow that helps requesters optimize crowdsourcing task designs and Sprout, our open-source tool, which implements this workflow. Sprout improves task designs by (1) eliciting points of confusion from crowd workers, (2) enabling requesters to quickly understand these misconceptions and the overall space of questions, and (3) guiding requesters to improve the task design in response. We report the results of a user study with two labeling tasks demonstrating that requesters strongly prefer Sprout and produce higher-rated instructions compared to current best practices for creating gated instructions (instructions plus a workflow for training and testing workers). We also offer a set of design recommendations for future tools that support crowdsourcing task design.
Jonathan Bragg, Mausam, Daniel S. Weld
UIST3
2018 StaQC: A Systematically Mined Question-Code Dataset from Stack Overflow
abstract
Stack Overflow (SO) has been a great source of natural language questions and their code solutions (i.e., question-code pairs), which are critical for many tasks including code retrieval and annotation. In most existing research, question-code pairs were collected heuristically and tend to have low quality. In this paper, we investigate a new problem of systematically mining question-code pairs from Stack Overflow (in contrast to heuristically collecting them). It is formulated as predicting whether or not a code snippet is a standalone solution to a question. We propose a novel Bi-View Hierarchical Neural Network which can capture both the programming content and the textual context of a code snippet (i.e., two views) to make a prediction. On two manually annotated datasets in Python and SQL domain, our framework substantially outperforms heuristic methods with at least 15% higher F1 and accuracy. Furthermore, we present StaQC (Stack Overflow Question-Code pairs), the largest dataset to date of ~148K Python and ~120K SQL question-code pairs, automatically mined from SO using our framework. Under various case studies, we demonstrate that StaQC can greatly help develop data-hungry models for associating natural language with programming language
Ziyu Yao 0002, Daniel S. Weld, Wei-Peng Chen, Huan Sun 0001
WWW2
2017 TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
abstract
We present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples.TriviaQA includes 95K questionanswer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions.We show that, in comparison to other recently introduced large-scale datasets, TriviaQA (1) has relatively complex, compositional questions, (2) has considerable syntactic and lexical variability between questions and corresponding answer-evidence sentences, and (3) requires more cross sentence reasoning to find answers.We also present two baseline algorithms: a featurebased classifier and a state-of-the-art neural network, that performs well on SQuAD reading comprehension.Neither approach comes close to human performance (23% and 40% vs. 80%), suggesting that Trivi-aQA is a challenging testbed that is worth significant future study. 1
Mandar Joshi, Eunsol Choi, Daniel S. Weld, Luke Zettlemoyer
ACL (1)3
2016 Re-Active Learning: Active Learning with Relabeling
abstract
Active learning seeks to train the best classifier at the lowest annotation cost by intelligently picking the best examples to label. Traditional algorithms assume there is a single annotator and disregard the possibility of requesting additional independent annotations for a previously labeled example. However, relabeling examples is important, because all annotators make mistakes — especially crowdsourced workers, who have become a common source of training data. This paper seeks to understand the difference in marginal value between decreasing the noise of the training set via relabeling and increasing the size and diversity of the (noisier) training set by labeling new examples. We use the term re-active learning to denote this generalization of active learning. We show how traditional active learning methods perform poorly at re-active learning, present new algorithms designed for this important problem, formally characterize their behavior, and empirically show that our methods effectively make this tradeoff.
Christopher H. Lin, Mausam, Daniel S. Weld
AAAI3
2016 Toward Automatic Bootstrapping of Online Communities Using Decision-theoretic Optimization
abstract
Successful online communities (e.g., Wikipedia, Yelp, and StackOverflow) can produce valuable content. However, many communities fail in their initial stages. Starting an online community is challenging because there is not enough content to attract a critical mass of active members. This paper examines methods for addressing this cold-start problem in datamining-bootstrappable communities by attracting non-members to contribute to the community. We make four contributions: 1) we characterize a set of communities that are “datamining-bootstrappable” and define the bootstrapping problem in terms of decision-theoretic optimization, 2) we estimate the model parameters in a case study involving the Open AI Resources website, 3) we demonstrate that non-members' predicted interest levels and request design are important features that can significantly affect the contribution rate, and 4) we ran a simulation experiment using data generated with the learned parameters and show that our decision-theoretic optimization algorithm can generate as much community utility when bootstrapping the community as our strongest baseline while issuing only 55% as many contribution requests.
Shih-Wen Huang, Jonathan Bragg, Isaac Cowhey, Oren Etzioni, Daniel S. Weld
CSCW5
2016 MicroTalk: Using Argumentation to Improve Crowdsourcing Accuracy
abstract
Crowd workers are human and thus sometimes make mistakes. In order to ensure the highest quality output, requesters often issue redundant jobs with gold test questions and sophisticated aggregation mechanisms based on expectation maximization (EM). While these methods yield accurate results in many cases, they fail on extremely difficult problems with local minima, such as situations where the majority of workers get the answer wrong. Indeed, this has caused some researchers to conclude that on some tasks crowdsourcing can never achieve high accuracies, no matter how many workers are involved. This paper presents a new quality-control workflow, called MicroTalk, that requires some workers to Justify their reasoning and asks others to Reconsider their decisions after reading counter-arguments from workers with opposing views. Experiments on a challenging NLP annotation task with workers from Amazon Mechanical Turk show that (1) argumentation improves the accuracy of individual workers by 20%, (2) restricting consideration to workers with complex explanations improves accuracy even more, and (3) our complete MicroTalk aggregation workflow produces much higher accuracy than simpler voting approaches for a range of budgets.
Ryan Drapeau, Lydia B. Chilton, Jonathan Bragg, Daniel S. Weld
HCOMP4
2016 Effective Crowd Annotation for Relation Extraction
abstract
Angli Liu, Stephen Soderland, Jonathan Bragg, Christopher H. Lin, Xiao Ling, Daniel S. Weld. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Angli Liu, Stephen Soderland, Jonathan Bragg, Christopher H. Lin, Daniel S. Weld
HLT-NAACL6
2015 Intelligent Control of Crowdsourcing
abstract
Crowd-sourcing labor markets (e.g., Amazon Mechanical Turk) are booming, because they enable rapid construction of complex workflows that seamlessly mix human computation with computer automation. Example applications range from photo tagging to audio-visual transcription and interlingual translation. Similarly, workflows on citizen science sites (e.g. GalaxyZoo) have allowed ordinary people to pool their effort and make interesting discoveries. Unfortunately, constructing a good workflow is difficult, be- cause the quality of the work performed by humans is highly variable. Typically, a task designer will experiment with several alternative workflows to accomplish a task, varying the amount of redundant labor, until she devises a control strategy that delivers acceptable performance. Fortunately, this control challenge can often be formulated as an automated planning problem ripe for algorithms from the probabilistic planning and reinforcement learning literature. I describe our recent work on the decision-theoretic control of crowd sourcing and suggest open problems for future research.
Daniel S. Weld
IUI1
2015 Design Challenges for Entity Linking
abstract
Recent research on entity linking (EL) has introduced a plethora of promising techniques, ranging from deep neural networks to joint inference. But despite numerous papers there is surprisingly little understanding of the state of the art in EL. We attack this confusion by analyzing differences between several versions of the EL problem and presenting a simple yet effective, modular, unsupervised system, called Vinculum, for entity linking. We conduct an extensive evaluation on nine data sets, comparing Vinculum with two state-of-the-art systems, and elucidate key aspects of the system that include mention extraction, candidate generation, entity type prediction, entity coreference, and coherence.
Sameer Singh 0001, Daniel S. Weld
Trans. Assoc. Comput. Linguistics3
2015 Exploiting Parallel News Streams for Unsupervised Event Extraction
abstract
Most approaches to relation extraction, the task of extracting ground facts from natural language text, are based on machine learning and thus starved by scarce training data. Manual annotation is too expensive to scale to a comprehensive set of relations. Distant supervision, which automatically creates training data, only works with relations that already populate a knowledge base (KB). Unfortunately, KBs such as FreeBase rarely cover event relations ( e.g. “person travels to location”). Thus, the problem of extracting a wide range of events — e.g., from news streams — is an important, open challenge. This paper introduces NewsSpike-RE, a novel, unsupervised algorithm that discovers event relations and then learns to extract them. NewsSpike-RE uses a novel probabilistic graphical model to cluster sentences describing similar events from parallel news streams. These clusters then comprise training data for the extractor. Our evaluation shows that NewsSpike-RE generates high quality training sentences and learns extractors that perform much better than rival approaches, more than doubling the area under a precision-recall curve compared to Universal Schemas.
Congle Zhang, Stephen Soderland, Daniel S. Weld
Trans. Assoc. Comput. Linguistics3
2014 Frenzy: collaborative data organization for creating conference sessions
abstract
Organizing conference sessions around themes improves the experience for attendees. However, the session creation process can be difficult and time-consuming due to the amount of expertise and effort required to consider alternative paper groupings. We present a collaborative web application called Frenzy to draw on the efforts and knowledge of an entire program committee. Frenzy comprises (a) interfaces to support large numbers of experts working collectively to create sessions, and (b) a two-stage process that decomposes the session-creation problem into meta-data elicitation and global constraint satisfaction. Meta-data elicitation involves a large group of experts working simultaneously, while global constraint satisfaction involves a smaller group that uses the meta-data to form sessions.
Lydia B. Chilton, Juho Kim 0001, Paul André, Felicia Cordeiro, James A. Landay, Daniel S. Weld, Steven Dow, Rob Miller 0001
CHI6
2014 Type-Aware Distantly Supervised Relation Extraction with Linked Arguments
abstract
Distant supervision has become the lead-ing method for training large-scale rela-tion extractors, with nearly universal adop-tion in recent TAC knowledge-base pop-ulation competitions. However, there are still many questions about the best way to learn such extractors. In this paper we investigate four orthogonal improvements: integrating named entity linking (NEL) and coreference resolution into argument identification for training and extraction, enforcing type constraints of linked argu-ments, and partitioning the model by rela-tion type signature. We evaluate sentential extraction perfor-mance on two datasets: the popular set of NY Times articles partially annotated by Hoffmann et al. (2011) and a new dataset, called GORECO, that is comprehensively annotated for 48 common relations. We find that using NEL for argument identi-fication boosts performance over the tra-ditional approach (named entity recogni-tion with string match), and there is further improvement from using argument types. Our best system boosts precision by 44% and recall by 70%. 1
Mitchell Koch, John Gilmer, Stephen Soderland, Daniel S. Weld
EMNLP4
2014 Parallel Task Routing for Crowdsourcing
abstract
An ideal crowdsourcing or citizen-science system would route tasks to the most appropriate workers, but the best assignment is unclear because workers have varying skill, tasks have varying difficulty, and assigning several workers to a single task may significantly improve output quality. This paper defines a space of task routing problems, proves that even the simplest is NP-hard, and develops several approximation algorithms for parallel routing problems. We show that an intuitive class of requesters' utility functions is submodular, which lets us provide iterative methods for dynamically allocating batches of tasks that make near-optimal use of available workers in each round. Experiments with live oDesk workers show that our task routing algorithm uses only 48% of the human labor compared to the commonly used round-robin strategy. Further, we provide versions of our task routing algorithm which enable it to scale to large numbers of workers and questions and to handle workers with variable response times while still providing significant benefit over common baselines.
Jonathan Bragg, Andrey Kolobov, Mausam, Daniel S. Weld
HCOMP4
2014 To Re(label), or Not To Re(label)
abstract
One of the most popular uses of crowdsourcing is to provide training data for supervised machine learning algorithms. Since human annotators often make errors, requesters commonly ask multiple workers to label each example. But is this strategy always the most cost effective use of crowdsourced workers? We argue "No" --- often classifiers can achieve higher accuracies when trained with noisy "unilabeled" data. However, in some cases relabeling is extremely important. We discuss three factors that may make relabeling an effective strategy: classifier expressiveness, worker accuracy, and budget.
Christopher H. Lin, Mausam, Daniel S. Weld
HCOMP3
2013 Cascade: crowdsourcing taxonomy creation
abstract
Taxonomies are a useful and ubiquitous way of organizing information. However, creating organizational hierarchies is difficult because the process requires a global understanding of the objects to be categorized. Usually one is created by an individual or a small group of people working together for hours or even days. Unfortunately, this centralized approach does not work well for the large, quickly changing datasets found on the web. Cascade is an automated workflow that allows crowd workers to spend as little at 20 seconds each while collectively making a taxonomy. We evaluate Cascade and show that on three datasets its quality is 80-90% of that of experts. Cascade has a competitive cost to expert information architects, despite taking six times more human labor. Fortunately, this labor can be parallelized such that Cascade will run in as fast as four minutes instead of hours or days.
Lydia B. Chilton, Greg Little, Darren Edge, Daniel S. Weld, James A. Landay
CHI4
2013 Joint Coreference Resolution and Named-Entity Linking with Multi-Pass Sieves
abstract
Many errors in coreference resolution come from semantic mismatches due to inadequate world knowledge.Errors in named-entity linking (NEL), on the other hand, are often caused by superficial modeling of entity context.This paper demonstrates that these two tasks are complementary.We introduce NECO, a new model for named entity linking and coreference resolution, which solves both problems jointly, reducing the errors made on each.NECO extends the Stanford deterministic coreference system by automatically linking mentions to Wikipedia and introducing new NEL-informed mention-merging sieves.Linking improves mention-detection and enables new semantic attributes to be incorporated from Freebase, while coreference provides better context modeling by propagating named-entity links within mention clusters.Experiments show consistent improvements across a number of datasets and experimental conditions, including over 11% reduction in MUC coreference error and nearly 21% reduction in F1 NEL error on ACE 2004 newswire data.
Hannaneh Hajishirzi, Leila Zilles, Daniel S. Weld, Luke Zettlemoyer
EMNLP3
2013 Harvesting Parallel News Streams to Generate Paraphrases of Event Relations
abstract
The distributional hypothesis, which states that words that occur in similar contexts tend to have similar meanings, has inspired several Web mining algorithms for paraphrasing semantically equivalent phrases.Unfortunately, these methods have several drawbacks, such as confusing synonyms with antonyms and causes with effects.This paper introduces three Temporal Correspondence Heuristics, that characterize regularities in parallel news streams, and shows how they may be used to generate high precision paraphrases for event relations.We encode the heuristics in a probabilistic graphical model to create the NEWSSPIKE algorithm for mining news streams.We present experiments demonstrating that NEWSSPIKE significantly outperforms several competitive baselines.In order to spur further research, we provide a large annotated corpus of timestamped news articles as well as the paraphrases produced by NEWSSPIKE.
Congle Zhang, Daniel S. Weld
EMNLP2
2013 Crowdsourcing Multi-Label Classification for Taxonomy Creation
abstract
Recent work has introduced CASCADE, an algorithm for creating a globally-consistent taxonomy by crowdsourcing microwork from many individuals, each of whom may see only a tiny fraction of the data (Chilton et al. 2013). While CASCADE needs only unskilled labor and produces taxonomies whose quality approaches that of human experts, it uses significantly more labor than experts. This paper presents DELUGE, an improved workflow that produces taxonomies with comparable quality using significantly less crowd labor. Specifically, our method for crowdsourcing multi-label classification optimizes CASCADE’s most costly step (categorization) using less than 10% of the labor required by the original approach. DELUGE’s savings come from the use of decision theory and machine learning, which allow it to pose microtasks that aim to maximize information gain.
Jonathan Bragg, Mausam, Daniel S. Weld
HCOMP3
2013 POMDP-based control of workflows for crowdsourcing
Peng Dai 0001, Christopher H. Lin, Mausam, Daniel S. Weld
Artif. Intell.4
2012 LRTDP Versus UCT for Online Probabilistic Planning
abstract
UCT, the premier method for solving games such as Go, is also becoming the dominant algorithm for probabilistic planning. Out of the five solvers at the International Probabilistic Planning Competition (IPPC) 2011, four were based on the UCT algorithm. However, while a UCT-based planner, PROST, won the contest, an LRTDP-based system, Glutton, came in a close second, outperforming other systems derived from UCT. These results raise a question: what are the strengths and weaknesses of LRTDP and UCT in practice? This paper starts answering this question by contrasting the two approaches in the context of finite-horizon MDPs. We demonstrate that in such scenarios, UCT's lack of a sound termination condition is a serious practical disadvantage. In order to handle an MDP with a large finite horizon under a time constraint, UCT forces an expert to guess a non-myopic lookahead value for which it should be able to converge on the encountered states. Mistakes in setting this parameter can greatly hurt UCT's performance. In contrast, LRTDP's convergence criterion allows for an iterative deepening strategy. Using this strategy, LRTDP automatically finds the largest lookahead value feasible under the given time constraint. As a result, LRTDP has better performance and stronger theoretical properties. We present an online version of Glutton, named Gourmand, that illustrates this analysis and outperforms PROST on the set of IPPC-2011 problems.
Andrey Kolobov, Mausam, Daniel S. Weld
AAAI3
2012 Dynamically Switching between Synergistic Workflows for Crowdsourcing
abstract
To ensure quality results from unreliable crowdsourced workers, task designers often construct complex workflows and aggregate worker responses from redundant runs. Frequently, they experiment with several alternative workflows to accomplish the task, and eventually deploy the one that achieves the best performance during early trials. Surprisingly, this seemingly natural design paradigm does not achieve the full potential of crowdsourcing. In particular, using a single workflow (even the best) to accomplish a task is suboptimal. We show that alternative workflows can compose synergistically to yield much higher quality output. We formalize the insight with a novel probabilistic graphical model. Based on this model, we design and implement AGENTHUNT, a POMDP-based controller that dynamically switches between these workflows to achieve higher returns on investment. Additionally, we design offline and online methods for learning model parameters. Live experiments on Amazon Mechanical Turk demonstrate the superiority of AGENTHUNT for the task of generating NLP training data, yielding up to 50% error reduction and greater net utility compared to previous methods.
Christopher H. Lin, Mausam, Daniel S. Weld
AAAI3
2012 Fine-Grained Entity Recognition
abstract
Entity Recognition (ER) is a key component of relation extraction systems and many other natural-language processing applications. Unfortunately, most ER systems are restricted to produce labels from to a small set of entity classes, e.g., person, organization, location or miscellaneous. In order to intelligently understand text and extract a wide range of information, it is useful to more precisely determine the semantic classes of entities mentioned in unstructured text. This paper defines a fine-grained set of 112 tags, formulates the tagging problem as multi-class, multi-label classification, describes an unsupervised method for collecting training data, and presents the FIGER implementation. Experiments show that the system accurately predicts the tags for entities. Moreover, it provides useful information for a relation extraction system, increasing the F1 score by 93%. We make FIGER and its data available as a resource for future work.
Daniel S. Weld
AAAI2
2012 Ontological Smoothing for Relation Extraction with Minimal Supervision
abstract
Relation extraction, the process of converting natural language text into structured knowledge, is increasingly important. Most successful techniques use supervised machine learning to generate extractors from sentences that have been manually labeled with the relations' arguments. Unfortunately, these methods require numerous training examples, which are expensive and time-consuming to produce. This paper presents ontological smoothing, a semi-supervisedtechnique that learns extractors for a set of minimally-labeledrelations. Ontological smoothing has three phases. First, itgenerates a mapping between the target relations and a backgroundknowledge-base. Second, it uses distant supervision toheuristically generate new training examples for the targetrelations. Finally, it learns an extractor from a combination of theoriginal and newly-generated examples. Experiments on 65 relationsacross three target domains show that ontological smoothing candramatically improve precision and recall, even rivaling fully supervisedperformance in many cases.
Congle Zhang, Raphael Hoffmann, Daniel S. Weld
AAAI3
2012 Regroup: interactive machine learning for on-demand group creation in social networks
abstract
We present ReGroup, a novel end-user interactive machine learning system for helping people create custom, on demand groups in online social networks. As a person adds members to a group, ReGroup iteratively learns a probabilistic model of group membership specific to that group. ReGroup then uses its currently learned model to suggest additional members and group characteristics for filtering. Our evaluation shows that ReGroup is effective for helping people create large and varied groups, whereas traditional methods (searching by name or selecting from an alphabetical list) are better suited for small groups whose members can be easily recalled by name. By facilitating on demand group creation, ReGroup can enable in-context sharing and potentially encourage better online privacy practices. In addition, applying interactive machine learning to social network group creation introduces several challenges for designing effective end-user interaction with machine learning. We identify these challenges and discuss how we address them in ReGroup.
Saleema Amershi, James Fogarty, Daniel S. Weld
CHI3
2012 A Theory of Goal-Oriented MDPs with Dead Ends
Andrey Kolobov, Mausam, Daniel S. Weld
UAI3
2012 Crowdsourcing Control: Moving Beyond Multiple Choice
Christopher H. Lin, Mausam, Daniel S. Weld
UAI3
2012 Discovering hidden structure in factored MDPs
Andrey Kolobov, Mausam, Daniel S. Weld
Artif. Intell.3
2011 Artificial Intelligence for Artificial Artificial Intelligence
abstract
Crowdsourcing platforms such as Amazon Mechanical Turk have become popular for a wide variety of human intelligence tasks; however, quality control continues to be a significant challenge. Recently, we propose TurKontrol, a theoretical model based on POMDPs to optimize iterative, crowd-sourced workflows. However, they neither describe how to learn the model parameters, nor show its effectiveness in a real crowd-sourced setting. Learning is challenging due to the scale of the model and noisy data: there are hundreds of thousands of workers with high-variance abilities. This paper presents an end-to-end system that first learns TurKontrol's POMDP parameters from real Mechanical Turk data, and then applies the model to dynamically optimize live tasks. We validate the model and use it to control a successive-improvement process on Mechanical Turk. By modeling worker accuracy and voting patterns, our system produces significantly superior artifacts compared to those generated through nonadaptive workflows using the same amount of money.
Peng Dai 0001, Mausam, Daniel S. Weld
AAAI3
2011 Knowledge-Based Weak Supervision for Information Extraction of Overlapping Relations
Raphael Hoffmann, Congle Zhang, Luke Zettlemoyer, Daniel S. Weld
ACL5
2011 Towards Scalable MDP Algorithms
Andrey Kolobov, Mausam, Daniel S. Weld
IJCAI3
2011 Topological Value Iteration Algorithms
Peng Dai 0001, Mausam, Daniel S. Weld, Judy Goldsmith
J. Artif. Intell. Res.3
2010 Decision-Theoretic Control of Crowd-Sourced Workflows
abstract
Crowd-sourcing is a recent framework in which human intelligence tasks are outsourced to a crowd of unknown people ("workers") as an open call (e.g., on Amazon's Mechanical Turk). Crowd-sourcing has become immensely popular with hoards of employers ("requesters"), who use it to solve a wide variety of jobs, such as dictation transcription, content screening, etc. In order to achieve quality results, requesters often subdivide a large task into a chain of bite-sized subtasks that are combined into a complex, iterative workflow in which workers check and improve each other's results. This paper raises an exciting question for AI — could an autonomous agent control these workflows without human intervention, yielding better results than today's state of the art, a fixed control program? We describe a planner, TurKontrol, that formulates workflow control as a decision-theoretic optimization problem, trading off the implicit quality of a solution artifact against the cost for workers to achieve it. We lay the mathematical framework to govern the various decisions at each point in a popular class of workflows. Based on our analysis we implement the workflow control algorithm and present experiments demonstrating that TurKontrol obtains much higher utilities than popular fixed policies.
Peng Dai 0001, Mausam, Daniel S. Weld
AAAI3
2010 SixthSense: Fast and Reliable Recognition of Dead Ends in MDPs
abstract
The results of the latest International Probabilistic Planning Competition (IPPC-2008) indicate that the presence of dead ends, states with no trajectory to the goal, makes MDPs hard for modern probabilistic planners. Implicit dead ends, states with executable actions but no path to the goal, are particularly challenging; existing MDP solvers spend much time and memory identifying these states. As a first attempt to address this issue, we propose a machine learning algorithm called SIXTHSENSE. SIXTHSENSE helps existing MDP solvers by finding nogoods, conjunctions of literals whose truth in a state implies that the state is a dead end. Importantly, our learned nogoods are sound, and hence the states they identify are true dead ends. SIXTHSENSE is very fast, needs little training data, and takes only a small fraction of total planning time. While IPPC problems may have millions of dead ends, they may typically be represented with only a dozen or two no-goods. Thus, nogood learning efficiently produces a quick and reliable means for dead-end recognition. Our experiments show that the nogoods found by SIXTHSENSE routinely reduce planning space and time on IPPC domains, enabling some planners to solve problems they could not previously handle.
Andrey Kolobov, Mausam, Daniel S. Weld
AAAI3
2010 Temporal Information Extraction
abstract
Research on information extraction (IE) seeks to distill relational tuples from natural language text, such as the contents of the WWW. Most IE work has focussed on identifying static facts, encoding them as binary relations. This is unfortunate, because the vast majority of facts are fluents, only holding true during an interval of time. It is less helpful to extract PresidentOf(Bill-Clinton, USA) without the temporal scope 1/20/93 — 1/20/01. This paper presents TIE, a novel, information-extraction system, which distills facts from text while inducing as much temporal information as possible. In addition to recognizing temporal relations between times and events, TIE performs global inference, enforcing transitivity to bound the start and ending times for each event. We introduce the notion of temporal entropy as a way to evaluate the performance of temporal IE systems and present experiments showing that TIE outperforms three alternative approaches.
Daniel S. Weld
AAAI2
2010 Learning 5000 Relational Extractors
Raphael Hoffmann, Congle Zhang, Daniel S. Weld
ACL3
2010 Open Information Extraction Using Wikipedia
Fei Wu 0003, Daniel S. Weld
ACL2
2010 Learning First-Order Horn Clauses from Web Text
Stefan Schoenmackers, Jesse Davis, Oren Etzioni, Daniel S. Weld
EMNLP4
2010 Automatically generating personalized user interfaces with Supple
Krzysztof Z. Gajos, Daniel S. Weld, Jacob O. Wobbrock
Artif. Intell.2
2010 Panlingual lexical translation via probabilistic inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Kobi Reiter, Michael Skinner, Marcus Sammer, Jeff A. Bilmes
Artif. Intell.4
2009 Compiling a Massive, Multilingual Dictionary via Probabilistic Inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Michael Skinner, Jeff A. Bilmes
ACL/IJCNLP4
2009 Amplifying community content creation with mixed initiative information extraction
abstract
Although existing work has explored both information extraction and community content creation, most research has focused on them in isolation. In contrast, we see the greatest leverage in the synergistic pairing of these methods as two interlocking feedback cycles. This paper explores the potential synergy promised if these cycles can be made to accelerate each other by exploiting the same edits to advance both community content creation and learning-based information extraction. We examine our proposed synergy in the context of Wikipedia infoboxes and the Kylin information extraction system. After developing and refining a set of interfaces to present the verification of Kylin extractions as a non primary task in the context of Wikipedia articles, we develop an innovative use of Web search advertising services to study people engaged in some other primary task. We demonstrate our proposed synergy by analyzing our deployment from two complementary perspectives: (1) we show we accelerate community content creation by using Kylin's information extraction to significantly increase the likelihood that a person visiting a Wikipedia article as a part of some other primary task will spontaneously choose to help improve the article's infobox, and (2) we show we accelerate information extraction by using contributions collected from people interacting with our designs to significantly improve Kylin's extraction performance.
Raphael Hoffmann, Saleema Amershi, Kayur Patel, Fei Wu 0003, James Fogarty, Daniel S. Weld
CHI6
2009 Domain-Independent, Automatic Partitioning for Probabilistic Planning
Peng Dai 0001, Mausam, Daniel S. Weld
IJCAI3
2009 ReTrASE: Integrating Paradigms for Approximate Probabilistic Planning
Andrey Kolobov, Mausam, Daniel S. Weld
IJCAI3
2009 Information arbitrage across multi-lingual Wikipedia
abstract
The rapid globalization of Wikipedia is generating a parallel, multi-lingual corpus of unprecedented scale. Pages for the same topic in many different languages emerge both as a result of manual translation and independent development. Unfortunately, these pages may appear at different times, vary in size, scope, and quality. Furthermore, differential growth rates cause the conceptual mapping between articles in different languages to be both complex and dynamic. These disparities provide the opportunity for a powerful form of information arbitrage--leveraging articles in one or more languages to improve the content in another. Analyzing four large language domains (English, Spanish, French, and German), we present Ziggurat, an automated system for aligning Wikipedia infoboxes, creating new infoboxes as necessary, filling in missing information, and detecting discrepancies between parallel pages. Our method uses self-supervised learning and our experiments demonstrate the method's feasibility, even in the absence of dictionaries.
Eytan Adar, Michael Skinner, Daniel S. Weld
WSDM3
2008 Partitioned External-Memory Value Iteration
Peng Dai 0001, Mausam, Daniel S. Weld
AAAI3
2008 Decision-Theoretic User Interface Generation
Krzysztof Z. Gajos, Daniel S. Weld, Jacob O. Wobbrock
AAAI2
2008 Intelligence in Wikipedia
Daniel S. Weld, Fei Wu 0003, Eytan Adar, Saleema Amershi, James Fogarty, Raphael Hoffmann, Kayur Patel, Michael Skinner
AAAI1
2008 Predictability and accuracy in adaptive user interfaces
abstract
While proponents of adaptive user interfaces tout potential performance gains, critics argue that adaptation's unpredictability may disorient users, causing more harm than good. We present a study that examines the relative effects of predictability and accuracy on the usability of adaptive UIs. Our results show that increasing predictability and accuracy led to strongly improved satisfaction. Increasing accuracy also resulted in improved performance and higher utilization of the adaptive interface. Contrary to our expectations, improvement in accuracy had a stronger effect on performance, utilization and some satisfaction ratings than the improvement in predictability.
Krzysztof Z. Gajos, Katherine Everitt, Desney S. Tan, Mary Czerwinski, Daniel S. Weld
CHI5
2008 Improving the performance of motor-impaired users with automatically-generated, ability-based interfaces
abstract
We evaluate two systems for automatically generating personalized interfaces adapted to the individual motor capabilities of users with motor impairments. The first system, SUPPLE, adapts to users' capabilities indirectly by first using the ARNAULD preference elicitation engine to model a user's preferences regarding how he or she likes the interfaces to be created. The second system, SUPPLE++, models a user's motor abilities directly from a set of one-time motor performance tests. In a study comparing these approaches to baseline interfaces, participants with motor impairments were 26.4% faster using ability-based user interfaces generated by SUPPLE++. They also made 73% fewer errors, strongly preferred those interfaces to the manufacturers' defaults, and found them more efficient, easier to use, and much less physically tiring. These findings indicate that rather than requiring some users with motor impairments to adapt themselves to software using separate assistive technologies, software can now adapt itself to the capabilities of its users.
Krzysztof Z. Gajos, Jacob O. Wobbrock, Daniel S. Weld
CHI3
2008 Evaluating visual cues for window switching on large screens
abstract
An increasing number of users are adopting large, multi-monitor displays. The resulting setups cover such a broad viewing angle that users can no longer simultaneously perceive all parts of the screen. Changes outside the user's visual field often go unnoticed. As a result, users sometimes have trouble locating the active window, for example after switching focus. This paper surveys graphical cues designed to direct visual attention and adapts them to window switching. Visual cues include five types of frames and mask around the target window and four trails leading to the window. We report the results of two user studies. The first evaluates each cue in isolation. The second evaluates hybrid techniques created by combining the most successful candidates from the first study. The best cues were visually sparse --- combinations of curved frames which use color to pop-out and tapered trails with predictable origin.
Raphael Hoffmann, Patrick Baudisch, Daniel S. Weld
CHI3
2008 Scaling Textual Inference to the Web
Stefan Schoenmackers, Oren Etzioni, Daniel S. Weld
EMNLP3
2008 Recovering from errors during programming by demonstration
abstract
Many end-users wish to customize their applications, automating common tasks and routines. Unfortunately, this automation is difficult today — users must choose between brittle macros and complex scripting languages. Programming by demonstration (PBD) offers a middle ground, allowing users to demonstrate a procedure multiple times and generalizing the requisite behavior with machine learning. Unfortunately, many PBD systems are almost as brittle as macro recorders, offering few ways for a user to control the learning process or correct the demonstrations used as training examples. This paper presents CHINLE, a system which automatically constructs PBD systems for applications based on their interface specification. The resulting PBD systems have novel interaction and visualization methods, which allow the user to easily monitor and guide the learning process, facilitating error recovery during training. CHINLE-constructed PBD systems learn procedures with conditionals and perform partial learning if the procedure is too complex to learn completely. ACM Classification D.2.2 [Design Tools and Techniques]: User Interfaces, H1.2. [Models and principles]: User/Machine
Jiun-Hung Chen, Daniel S. Weld
IUI2
2008 Information extraction from Wikipedia: moving down the long tail
abstract
Not only is Wikipedia a comprehensive source of quality information, it has several kinds of internal structure (e.g., relational summaries known as infoboxes), which enable self-supervised information extraction. While previous efforts at extraction from Wikipedia achieve high precision and recall on well-populated classes of articles, they fail in a larger number of cases, largely because incomplete articles and infrequent use of infoboxes lead to insufficient training data. This paper presents three novel techniques for increasing recall from Wikipedia's long tail of sparse classes: (1) shrinkage over an automatically-learned subsumption taxonomy, (2) a retraining technique for improving the training data, and (3) supplementing results by extracting from the broader Web. Our experiments compare design variations and show that, used in concert, these techniques increase recall by a factor of 1.76 to 8.71 while maintaining or increasing precision.
Fei Wu 0003, Raphael Hoffmann, Daniel S. Weld
KDD3
2008 Zoetrope: interacting with the ephemeral web
abstract
The Web is ephemeral. Pages change frequently, and it is nearly impossible to find data or follow a link after the underlying page evolves. We present Zoetrope, a system that enables interaction with the historicalWeb (pages, links, and embedded data) that would otherwise be lost to time. Using a number of novel interactions, the temporal Web can be manipulated, queried, and analyzed from the context of familar pages. Zoetrope is based on a set of operators for manipulating content streams. We describe these primitives and the associated indexing strategies for handling temporal Web data. They form the basis of Zoetrope and enable our construction of new temporal interactions and visualizations.
Eytan Adar, Mira Dontcheva, James Fogarty, Daniel S. Weld
UIST4
2008 Automatically refining the wikipedia infobox ontology
abstract
The combined efforts of human volunteers have recently extracted numerous facts from Wikipedia, storing them as machine-harvestable object-attribute-value triples in Wikipedia infoboxes. Machine learning systems, such as Kylin, use these infoboxes as training data, accurately extracting even more semantic knowledge from natural language text. But in order to realize the full power of this information, it must be situated in a cleanly-structured ontology. This paper introduces KOG, an autonomous system for refining Wikipedia's infobox-class ontology towards this end. We cast the problem of ontology refinement as a machine learning problem and solve it using both SVMs and a more powerful joint-inference approach expressed in Markov Logic Networks. We present experiments demonstrating the superiority of the joint-inference approach and evaluating other aspects of our system. Using these techniques, we build a rich ontology, integrating Wikipedia's infobox-class schemata with WordNet. We demonstrate how the resulting ontology may be used to enhance Wikipedia with improved query processing and other features.
Fei Wu 0003, Daniel S. Weld
WWW2
2008 Planning with Durative Actions in Stochastic Domains
abstract
Probabilistic planning problems are typically modeled as a Markov Decision Process (MDP). MDPs, while an otherwise expressive model, allow only for sequential, non-durative actions. This poses severe restrictions in modeling and solving a real world planning problem. We extend the MDP model to incorporate 1) simultaneous action execution, 2) durative actions, and 3) stochastic durations. We develop several algorithms to combat the computational explosion introduced by these features. The key theoretical ideas used in building these algorithms are -- modeling a complex problem as an MDP in extended state/action space, pruning of irrelevant actions, sampling of relevant actions, using informed heuristics to guide the search, hybridizing different planners to achieve benefits of both, approximating the problem and replanning. Our empirical evaluation illuminates the different merits in using various algorithms, viz., optimality, empirical closeness to optimality, theoretical error bounds, and speed.
Mausam, Daniel S. Weld
J. Artif. Intell. Res.2
2007 Autonomously semantifying wikipedia
abstract
Berners-Lee's compelling vision of a Semantic Web is hindered by a chicken-and-egg problem, which can be best solved by a bootstrapping method - creating enough structured data to motivate the development of applications. This paper argues that autonomously "Semantifying Wikipedia" is the best way to solve the problem. We choose Wikipedia as an initial data source, because it is comprehensive, not too large, high-quality, and contains enough manually-derived structure to bootstrap an autonomous, self-supervised process. We identify several types of structures which can be automatically enhanced in Wikipedia (e.g., link structure, taxonomic data, infoboxes, etc.), and we describea prototype implementation of a self-supervised, machine learning system which realizes our vision. Preliminary experiments demonstrate the high precision of our system's extracted data - in one case equaling that of humans.
Fei Wu 0003, Daniel S. Weld
CIKM2
2007 When is Temporal Planning Really Temporal?
William Cushing, Subbarao Kambhampati, Mausam, Daniel S. Weld
IJCAI4
2007 A Hybridized Planner for Stochastic Domains
Mausam, Piergiorgio Bertoli, Daniel S. Weld
IJCAI3
2007 Automatically generating user interfaces adapted to users' motor and vision capabilities
abstract
Most of today's GUIs are designed for the typical, able-bodied user; atypical users are, for the most part, left to adapt as best they can, perhaps using specialized assistive technologies as an aid. In this paper, we present an alternative approach: SUPPLE++ automatically generates interfaces which are tailored to an individual's motor capabilities and can be easily adjusted to accommodate varying vision capabilities. SUPPLE++ models users. motor capabilities based on a onetime motor performance test and uses this model in an optimization process, generating a personalized interface. A preliminary study indicates that while there is still room for improvement, SUPPLE++ allowed one user to complete tasks that she could not perform using a standard interface, while for the remaining users it resulted in an average time savings of 20%, ranging from an slowdown of 3% to a speedup of 43%.
Krzysztof Z. Gajos, Jacob O. Wobbrock, Daniel S. Weld
UIST3
2007 Assieme: finding and leveraging implicit references in a web search interface for programmers
abstract
Programmers regularly use search as part of the development process, attempting to identify an appropriate API for a problem, seeking more information about an API, and seeking samples that show how to use an API. However, neither general-purpose search engines nor existing code search engines currently fit their needs, in large part because the information programmers need is distributed across many pages. We present Assieme, a Web search interface that effectively supports common programming search tasks by combining information from Web-accessible Java Archive (JAR) files, API documentation, and pages that include explanatory text and sample code. Assieme uses a novel approach to finding and resolving implicit references to Java packages, types, and members within sample code on the Web. In a study of programmers performing searches related to common programming tasks, we show that programmers obtain better solutions, using fewer queries, in the same amount of time spent using a general Web search interface.
Raphael Hoffmann, James Fogarty, Daniel S. Weld
UIST3
2007 Why we search: visualizing and predicting user behavior
abstract
The aggregation and comparison of behavioral patterns on the WWW represent a tremendous opportunity for understanding past behaviors and predicting future behaviors. In this paper, we take a first step at achieving this goal. We present a large scale study correlating the behaviors of Internet users on multiple systems ranging in size from 27 million queries to 14 million blog posts to 20,000 news articles. We formalize a model for events in these time-varying datasets and study their correlation. We have created an interface for analyzing the datasets, which includes a novel visual artifact, the DTWRadar, for summarizing differences between time series. Using our tool we identify a number of behavioral properties that allow us to understand the predictive power of patterns of use.
Eytan Adar, Daniel S. Weld, Brian N. Bershad, Steve D. Gribble
WWW2
2006 Probabilistic Temporal Planning with Uncertain Durations
Mausam, Daniel S. Weld
AAAI2
2006 Automatically generating custom user interfaces for users with physical disabilities
abstract
No abstract available.
Krzysztof Z. Gajos, Jing Jing Long, Daniel S. Weld
ASSETS3
2006 Exploring the design space for adaptive graphical user interfaces
abstract
For decades, researchers have presented different adaptive user interfaces and discussed the pros and cons of adaptation on task performance and satisfaction. Little research, however, has been directed at isolating and understanding those aspects of adaptive interfaces which make some of them successful and others not. We have designed and implemented three adaptive graphical interfaces and evaluated them in two experiments along with a non-adaptive baseline. In this paper we synthesize our results with previous work and discuss how different design choices and interactions affect the success of adaptive graphical user interfaces.
Krzysztof Z. Gajos, Mary Czerwinski, Desney S. Tan, Daniel S. Weld
AVI4
2005 Extending Continuous Time Bayesian Networks
Karthik Gopalratnam, Henry A. Kautz, Daniel S. Weld
AAAI3
2005 Fast and Robust Interface Generation for Ubiquitous Applications
Krzysztof Z. Gajos, David B. Christianson, Raphael Hoffmann, Tal Shaked, Kiera Henning, Jing Jing Long, Daniel S. Weld
UbiComp7
2005 Preference elicitation for interface optimization
abstract
Decision-theoretic optimization is becoming a popular tool in the user interface community, but creating accurate cost (or utility) functions has become a bottleneck --- in most cases the numerous parameters of these functions are chosen manually, which is a tedious and error-prone process. This paper describes ARNAULD, a general interactive tool for eliciting user preferences concerning concrete outcomes and using this feedback to automatically learn a factored cost function. We empirically evaluate our machine learning algorithm and two automatic query generation approaches and report on an informal user study.
Krzysztof Z. Gajos, Daniel S. Weld
UIST2
2005 Unsupervised named-entity extraction from the Web: An experimental study
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates
Artif. Intell.7
2004 Methods for Domain-Independent Information Extraction from the Web: An Experimental Comparison
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates
AAAI7
2004 Solving Concurrent Markov Decision Processes
Mausam, Daniel S. Weld
AAAI2
2004 SUPPLE: automatically generating user interfaces
abstract
In order to give people ubiquitous access to software applications, device controllers, and Internet services, it will be necessary to automatically adapt user interfaces to the computational devices at hand (eg, cell phones, PDAs, touch panels, etc.). While previous researchers have proposed solutions to this problem, each has limitations. This paper proposes a novel solution based on treating interface adaptation as an optimization problem. When asked to render an interface on a specific device, our supple system searches for the rendition that meets the device's constraints and minimizes the estimated effort for the user's expected interface actions. We make several contributions: 1) precisely defining the interface rendition problem, 2) demonstrating how user traces can be used to customize interface rendering to particular user's usage pattern, 3) presenting an efficient interface rendering algorithm, 4) performing experiments that demonstrate the utility of our approach.
Krzysztof Z. Gajos, Daniel S. Weld
IUI2
2004 Adapting to Source Properties in Processing Data Integration Queries
abstract
An effective query optimizer finds a query plan that exploits the characteristics of the source data. In data integration, little is known in advance about sources' properties, which necessitates the use of adaptive query processing techniques to adjust query processing on-the-fly. Prior work in adaptive query processing has focused on compensating for delays and adjusting for mis-estimated cardinality or selectivity values. In this paper, we present a generalized architecture for adaptive query processing and introduce a new technique, called adaptive data partitioning (ADP), which is based on the idea of dividing the source data into regions, each executed by different, complementary plans. We show how this model can be applied in novel ways to not only correct for underestimated selectivity and cardinality values, but also to discover and exploit order in the source data, and to detect and exploit source data that can be effectively pre-aggregated. We experimentally compare a number of alternative strategies and show that our approach is effective.
Zachary G. Ives, Alon Y. Halevy, Daniel S. Weld
SIGMOD Conference3
2004 Web-scale information extraction in knowitall: (preliminary results)
abstract
Manually querying search engines in order to accumulate a large bodyof factual information is a tedious, error-prone process of piecemealsearch. Search engines retrieve and rank potentially relevantdocuments for human perusal, but do not extract facts, assessconfidence, or fuse information from multiple documents. This paperintroduces KnowItAll, a system that aims to automate the tedious process ofextracting large collections of facts from the web in an autonomous,domain-independent, and scalable manner.The paper describes preliminary experiments in which an instance of KnowItAll, running for four days on a single machine, was able to automatically extract 54,753 facts. KnowItAll associates a probability with each fact enabling it to trade off precision and recall. The paper analyzes KnowItAll's architecture and reports on lessons learned for the design of large-scale information extraction systems.
Oren Etzioni, Michael J. Cafarella, Doug Downey, Stanley Kok, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates
WWW8
2003 Automatically Personalizing User Interfaces
Daniel S. Weld, Corin R. Anderson, Pedro M. Domingos, Oren Etzioni, Krzysztof Z. Gajos, Tessa A. Lau, Steven A. Wolfman
IJCAI1
2003 What users want
abstract
Todays computer interfaces are one size fits all. Users with little programming experience have only limited opportunities to customize their interface to their task and work habits (e.g., adding buttons to a toolbar). Furthermore, the overhead induced by generic interfaces will be proportionately greater on small form-factor PDAs, embedded applications and wearable devices. Searching for a solution, researchers argue that productivity can be greatly enhanced if interfaces anticipated their users, adapted to their preferences, and reacted to high-level customization requests. But realizing these benefits is tricky, because there is an inherent tension between the dynamism implied by automatic interface adaptation and the stability required in order for the user to maintain an accurate mental model, predict the computers behavior, and feel in control. In this talk, I discuss several principles governing effective adaptation, describe algorithms for data mining user action traces, and suggest mechanisms for dynamically transforming interfaces
Daniel S. Weld
IUI1
2003 A reliable natural language interface to household appliances
abstract
As household appliances grow in complexity and sophistication, they become harder and harder to use, particularly because of their tiny display screens and limited keyboards. This paper describes a strategy for building natural language interfaces to appliances that circumvents these problems. Our approach leverages decades of research on planning and natural language interfaces to databases by reducing the appliance problem to the database problem; the reduction provably maintains desirable properties of the database interface. The paper goes on to describe the implementation and evaluation of the EXACT interface to appliances, which is based on this reduction. EXACT maps each English user request to an SQL query, which is transformed to create a PDDL goal, and uses the Blackbox planner [13] to map the planning problem to a sequence of appliance commands that satisfy the original request. Both theoretical arguments and experimental evaluation show that EXACT is highly reliable
Alexander Yates, Oren Etzioni, Daniel S. Weld
IUI3
2003 Learning programs from traces using version space algebra
abstract
While existing learning techniques can be viewed as inducing programs from examples, most research has focused on rather narrow classes of programs, e.g., decision trees or logic rules. In contrast, most of today's programs are written in languages such as C++ or Java. Thus, many tasks we wish to automate (e.g. programming by demonstration and software reverse engineering) might be best formulated as induction of code in a procedural language. In this paper we apply version space algebra [10] to learn such procedural programs given execution traces. We consider two variants of the problem (whether or not program-step information is included in the traces) and evaluate our implementation on a corpus of programs drawn from introductory computer science textbooks. We show that our system can learn correct programs from few traces.
Tessa A. Lau, Pedro M. Domingos, Daniel S. Weld
K-CAP3
2003 Programming by Demonstration Using Version Space Algebra
Tessa A. Lau, Steven A. Wolfman, Pedro M. Domingos, Daniel S. Weld
Mach. Learn.4
2002 Relational Markov models and their application to adaptive web navigation
abstract
Relational Markov models (RMMs) are a generalization of Markov models where states can be of different types, with each type described by a different set of variables. The domain of each variable can be hierarchically structured, and shrinkage is carried out over the cross product of these hierarchies. RMMs make effective learning possible in domains with very large and heterogeneous state spaces, given only sparse data. We apply them to modeling the behavior of web site users, improving prediction in our PROTEUS architecture for personalizing web sites. We present experiments on an e-commerce and an academic web site showing that RMMs are substantially more accurate than alternative methods, and make good predictions even when applied to previously-unvisited parts of the site.
Corin R. Anderson, Pedro M. Domingos, Daniel S. Weld
KDD3
2002 An XML query engine for network-bound data
Zachary G. Ives, Alon Y. Halevy, Daniel S. Weld
VLDB J.3
2001 The Nimble XML Data Integration System
abstract
For better or for worse, XML has emerged as a de facto standard for data interchange. This consensus is likely to lead to increased demand for technology that allows users to integrate data from a variety of applications, repositories, and partners, which are located across the corporate intranet or on the Internet. Nimble Technology has spent two years developing a product to service this market. Originally conceived after decades of person-years of research on data integration, the product is now being deployed at several Fortune-500 beta-customer sites. The article reports on the key challenges faced in the design of our product and highlights some issues which require more attention from the research community. In particular we address architectural issues arising from designing a product to support XML as its core representation, choices in the design of the underlying algebra, on-the-fly data cleaning and caching and materialization policies.
Denise Draper, Alon Y. Halevy, Daniel S. Weld
ICDE3
2001 Adaptive Web Navigation for Wireless Devices
Corin R. Anderson, Pedro M. Domingos, Daniel S. Weld
IJCAI3
2001 Mixed initiative interfaces for learning tasks: SMARTedit talks back
abstract
Applications of machine learning can be viewed as teacherstudent interactions in which the teacher provides training examples and the student learns a generalization of the training examples. One such application of great interest to the IUI community is adaptive user interfaces. In the traditional learning interface, the scope of teacher-student interactions consists solely of the teacher/user providing some number of training examples to the student/learner and testing the learned model on new examples. Active learning approaches go one step beyond the traditional interaction model and allow the student to propose new training examples that are then solved by the teacher. In this paper, we propose that interfaces for machine learning should even more closely resemble human teacher-student relationships. A teacher's time and attention are precious resources. An intelligent studentmust proactively contribute to the learning process, by reasoning about the quality of its knowledge, collaborating with the teacher, and suggesting new examples for her to solve. The paper describes a varietyof richinteraction modes that enhance the learning process and presents a decision-theoretic framework, called DIAManD, for choosing the best interaction. We apply the framework to the SMARTedit programming by demonstration system and describe experimental validation and preliminary user feedback.
Steven A. Wolfman, Tessa A. Lau, Pedro M. Domingos, Daniel S. Weld
IUI4
2001 The Nimble Integration Engine
abstract
The consensus that XML has become the de facto standard for data interchange will spur demand for technology that allows users to integrate data from a variety of applications, repositories, and legacy systems which are located across the corporate intranet or at partner companies on the Internet. In the past two years, Nimble Technology has developed a product for this market. Spawned from over a persondecade of data integration research, the product has been deployed at several Fortune-500 beta-customer sites. This abstract reports on the key challenges we faced in the design of our product and highlights some issues we think require more attention from the research community. 1.
Denise Draper, Alon Y. Halevy, Daniel S. Weld
SIGMOD Conference3
2001 Updating XML
abstract
As XML has developed over the past few years, its role has expanded beyond its original domain as a semantics-preserving markup language for online documents, and it is now also the de facto format for interchanging data between heterogeneous systems. Data sources expert XML “views” over their data, and other system can directly import or query these views. As a result, there has been great interest in languages and systems for expressing queries over XML data, whether the XML is stored in a repository or generated as a view over some other data storage format.
Igor Tatarinov, Zachary G. Ives, Alon Y. Halevy, Daniel S. Weld
SIGMOD Conference4
2001 Personalizing Web Sites for Mobile Users
abstract
The fastest growing communityofweb users is that of mobile visitors who browse with wireless PDAs, cell phones, and pagers. Unfortunately,mostweb sites today are optimized exclusively for desktop, broadband clients, and deliver content poorly suited for mobile devices | devices that can display only a few lines of text, are on slow wireless network connections, and cannot run client-side programs or scripts. To best serve the needs of this growing community,wepropose building web site personalizers that observethebehavior of web visitors and automatically customize and adapt web sites for each individual mobile visitor. In this paper, welay the theoretical foundations for web site personalization, discuss our implementation of the web site personalizer #######, and present experiments evaluating its behavior on a number of academic and commercial web sites. Our initial results indicate that automatically adapting web content for mobile visitors saves a considerable amount of time and eort when seeking information \\on the go." Keywords Adaptiveweb sites, personalization, wireless web 1.
Corin R. Anderson, Pedro M. Domingos, Daniel S. Weld
WWW3
2001 Scaling question answering to the Web
abstract
The wealth of information on the web makes it an attractive resource for seeking quick answers to simple, factual questions such as "who was the first American in space?" or "what is the second tallest mountain in the world?" Yet today's most advanced web search services (e.g., Google and AskJeeves) make it surprisingly tedious to locate answers to such questions. In this paper, we extend question-answering techniques, first studied in the information retrieval literature, to the web and experimentally evaluate their performance. First we introduce MULDER, which we believe to be the first general-purpose, fully-automated question-answering system available on the web. Second, we describe MULDER's architecture, which relies on multiple search-engine queries, natural-language parsing, and a novel voting procedure to yield reliable answers coupled with high recall. Finally, we compare MULDER's performance to that of Google and AskJeeves on questions drawn from the TREC-8 question track. We find that MULDER's recall is more than a factor of three higher than that of AskJeeves. In addition, we find that Google requires 6.6 times as much user effort to achieve the same level of recall as MULDER. 1.
Cody C. T. Kwok, Oren Etzioni, Daniel S. Weld
WWW3
2001 Scaling question answering to the web
abstract
The wealth of information on the web makes it an attractive resource for seeking quick answers to simple, factual questions such as “who was the first American in space?” or “what is the second tallest mountain in the world?” Yet today's most advanced web search services (e.g., Google and AskJeeves) make it surprisingly tedious to locate answers to such questions. In this paper, we extend question-answering techniques, first studied in the information retrieval literature, to the web and experimentally evaluate their performance.First we introduce Mulder, which we believe to be the first general-purpose, fully-automated question-answering system available on the web. Second, we describe Mulder's architecture, which relies on multiple search-engine queries, natural-language parsing, and a novel voting procedure to yield reliable answers coupled with high recall. Finally, we compare Mulder's performance to that of Google and AskJeeves on questions drawn from the TREC-8 question answering track. We find that Mulder's recall is more than a factor of three higher than that of AskJeeves. In addition, we find that Google requires 6.6 times as much user effort to achieve the same level of recall as Mulder.
Cody C. T. Kwok, Oren Etzioni, Daniel S. Weld
ACM Trans. Inf. Syst.3
2000 Version Space Algebra and its Application to Programming by Demonstration
Tessa A. Lau, Pedro M. Domingos, Daniel S. Weld
ICML3
2000 Intelligent Internet systems
Alon Y. Halevy, Daniel S. Weld
Artif. Intell.2
1999 Temporal Planning with Mutual Exclusion Reasoning
David E. Smith 0001, Daniel S. Weld
IJCAI2
1999 The LPSAT Engine & Its Application to Resource Planning
Steven A. Wolfman, Daniel S. Weld
IJCAI2
1999 Programming by Demonstration: An Inductive Learning Formulation
abstract
Although Programmingby Demonstration (PBD) has the potential to improve the productivity of unsophisticated users, previous PBD systems have used brittle, heuristic, domain-specific approaches to execution-trace generalization.In this paper we define two applicationindependent methods for performing generalization that are based on well-understood machine learning technology.TGENV~ uses version-space generalization, and TGENFOIL is based on the FOIL inductive logic programming algorithm.We analyze each method both theoretically and empirically, arguing that TGENVS has lower sample complexity, but TGENFOIL can learn a much more interesting class of programs.
Tessa A. Lau, Daniel S. Weld
IUI2
1999 An Adaptive Query Execution System for Data Integration
abstract
Query processing in data integration occurs over network-bound, autonomous data sources. This requires extensions to traditional optimization and execution techniques for three reasons: there is an absence of quality statistics about the data, data transfer rates are unpredictable and bursty, and slow or unavailable data sources can often be replaced by overlapping or mirrored sources. This paper presents the Tukwila data integration system, designed to support adaptivity at its core using a two-pronged approach. Interleaved planning and execution with partial optimization allows Tukwila to quickly recover from decisions based on inaccurate estimates. During execution, Tukwila uses adaptive query operators such as the double pipelined hash join, which produces answers quickly, and the dynamic collector, which robustly and efficiently computes unions across overlapping data sources. We demonstrate that the Tukwila architecture extends previous innovations in adaptive execution (such as query scrambling, mid-execution re-optimization, and choose nodes), and we present experimental evidence that our techniques result in behavior desirable for a data integration system.
Zachary G. Ives, Daniela Florescu, Marc T. Friedman, Alon Y. Halevy, Daniel S. Weld
SIGMOD Conference5
1997 Automatic SAT-Compilation of Planning Problems
Michael D. Ernst, Todd D. Millstein, Daniel S. Weld
IJCAI3
1997 Efficiently Executing Information-Gathering Plans
Marc T. Friedman, Daniel S. Weld
IJCAI (1)2
1997 Wrapper Induction for Information Extraction
Nicholas Kushmerick, Daniel S. Weld, Robert B. Doorenbos
IJCAI (1)2
1997 Sound and Efficient Closed-World Reasoning for Planning
Oren Etzioni, Keith Golden, Daniel S. Weld
Artif. Intell.3
1997 Learning to Understand Information on the Internet: An Example-Based Approach
Mike Perkowitz, Robert B. Doorenbos, Oren Etzioni, Daniel S. Weld
J. Intell. Inf. Syst.4
1996 Representing Sensing Actions: The Middle Ground Revisited
Keith Golden, Daniel S. Weld
KR2
1995 An Algorithm for Probabilistic Planning
Nicholas Kushmerick, Steve Hanks, Daniel S. Weld
Artif. Intell.3
1995 A Domain-Independent Algorithm for Plan Adaptation
abstract
The paradigms of transformational planning, case-based planning, and plan debugging all involve a process known as plan adaptation - modifying or repairing an old plan so it solves a new problem. In this paper we provide a domain-independent algorithm for plan adaptation, demonstrate that it is sound, complete, and systematic, and compare it to other adaptation algorithms in the literature. Our approach is based on a view of planning as searching a graph of partial plans. Generative planning starts at the graph's root and moves from node to node using plan-refinement operators. In planning by adaptation, a library plan - an arbitrary node in the plan graph - is the starting point for the search, and the plan-adaptation algorithm can apply both the same refinement operators available to a generative planner and can also retract constraints and steps from the plan. Our algorithm's completeness ensures that the adaptation algorithm will eventually search the entire graph and its systematicity ensures that it will do so without redundantly searching any parts of the graph.
Steve Hanks, Daniel S. Weld
J. Artif. Intell. Res.2
1994 Task-Decomposition via Plan Parsing
Anthony Barrett, Daniel S. Weld
AAAI2
1994 Omnipotence Without Omniscience: Efficient Sensor Management for Planning
Keith Golden, Oren Etzioni, Daniel S. Weld
AAAI3
1994 An Algorithm for Probabilistic Least-Commitment Planning
Nicholas Kushmerick, Steve Hanks, Daniel S. Weld
AAAI3
1994 Temporal Planning with Continuous Change
J. Scott Penberthy, Daniel S. Weld
AAAI2
1994 The First Law of Robotics (A Call to Arms)
Daniel S. Weld, Oren Etzioni
AAAI1
1994 Tractable Closed World Reasoning with Updates
Oren Etzioni, Keith Golden, Daniel S. Weld
KR3
1994 A Probablistic Model of Action for Least-Commitment Planning with Information Gathering
Denise Draper, Steve Hanks, Daniel S. Weld
UAI3
1994 Partial-Order Planning: Evaluating Possible Efficiency Gains
Anthony Barrett, Daniel S. Weld
Artif. Intell.2
1993 Real-Time Self-Explanatory Simulation
Franz G. Amador, Adam Finkelstein, Daniel S. Weld
AAAI3
1993 Innovative Design as Systematic Search
Dorothy Neville, Daniel S. Weld
AAAI2
1993 A Framework for Model-Based Repair
Daniel S. Weld
AAAI2
1993 Characterizing Subgoal Interactions for Planning
Anthony Barrett, Daniel S. Weld
IJCAI2
1993 Ernest Davis, Representations of Commonsense Knowledge
Daniel S. Weld
Artif. Intell.1
1993 Electronic "How Things Work" Articles: Two Early Prototypes
abstract
The electronic encyclopedia exploratorium (E/sup 3/) is a vision of a future computer system-an electronic book describing how thing work. Typical articles in E/sup 3/ will describe such mechanisms as compression refrigerators, engines, telescopes, and mechanical linkages. Each article will provide simulations, three-dimensional animated graphics that the user can manipulate, laboratory areas that allow a user to modify the device or experiment with related artifacts, and a facility for asking questions and receiving customized, computer-generated English-language explanations. Some of the foundational technology is discussed, focusing on topics in artificial intelligence, graphics, and user interfaces. The initial prototype system and the technical lessons learned from it, as well as the second prototype currently under construction, are described.>
Franz G. Amador, Deborah Berman, Alan Borning, Tony DeRose, Adam Finkelstein, Dorothy Neville, David Notkin, David Salesin, Michael Salisbury, Joe Sherman, Daniel S. Weld, Georges Winkenbach
IEEE Trans. Knowl. Data Eng.12
1992 An Approach to Planning with Incomplete Information
Oren Etzioni, Steve Hanks, Daniel S. Weld, Denise Draper, Neal Lesh, Mike Williamson
KR3
1992 UCPOP: A Sound, Complete, Partial Order Planner for ADL
J. Scott Penberthy, Daniel S. Weld
KR2
1992 Reasoning about Model Accuracy
Daniel S. Weld
Artif. Intell.1
1992 Qualitative Physics: Albatross or Eagle?
Daniel S. Weld
Comput. Intell.1
1990 Approximation Reformulations
Daniel S. Weld
AAAI1
1990 Exaggeration
Daniel S. Weld
Artif. Intell.1
1989 Donald A. Norman, The Psychology of Everyday Things
Daniel S. Weld
Artif. Intell.1
1988 Choices for comparative analysis: DQ analysis or exaggeration?
Daniel S. Weld
Artif. Intell. Eng.1
1988 Comparative Analysis
Daniel S. Weld
Artif. Intell.1
1988 George Lakoff, Women, Fire, and Dangerous Things
Daniel S. Weld
Artif. Intell.1
1987 Comparative Analysis
Daniel S. Weld
IJCAI1
1986 The Use of Aggregation in Causal Simulation
Daniel S. Weld
Artif. Intell.1
1985 Combining Discrete and Continuous Process Models
Daniel S. Weld
IJCAI1