VLDB 2026 Research / reviewers in the wild / expert
Doug Downey
dblp:57/5363 · also Douglas Downey
· DBLP profile ↗
64ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0002-4737-8444ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-authorHuman-computer interaction and ubiquitous computing · 6 · 4 since 2021Computer networks · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating Literature-Driven Scientific Theories at ScaleabstractContemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theory building remain underexplored.In this work, we formulate the problem of synthesizing theories consisting of qualitative and quantitative laws from large corpora of scientific literature.We study theory generation at scale, using 13.7k source papers to synthesize 2.9k theories, examining how generation using literaturegrounding versus parametric knowledge, and accuracy-focused versus novelty-focused generation objectives change theory properties.Our experiments show that, compared to using parametric LLM memory for generation, our literature-supported method creates theories that are significantly better at both matching existing evidence and at predicting future results from 4.6k subsequently-written papers. 1 Peter A. Jansen, Peter Clark, Doug Downey, Daniel S. Weld |
ACL (1) | 3 |
| 2025 | SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific LiteratureabstractDavid Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, Shannon Zejiang Shen, Doug Downey, Hannaneh Hajishirzi, Arman Cohan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Dave Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh 0001, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, Shannon Shen 0001, Doug Downey, Hannaneh Hajishirzi, Arman Cohan |
EMNLP | 12 |
| 2025 | SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded TasksabstractWe present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons.By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses.The platform currently supports 44 open-source and proprietary foundation models and has collected over 19,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality.We discuss the results and insights based on the model ranking leaderboard.To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on our collected preference data. The benchmark measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark’s challenges and emphasize the need for more reliable automated evaluation methods. Yilun Zhao 0001, Tiansheng Hu, Sihong Wu, Ronan Le Bras 0001, Yixin Liu 0003, Robert Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao 0013, Hannaneh Hajishirzi, Doug Downey, Arman Cohan |
NeurIPS | 13 |
| 2024 | SciMON: Scientific Inspiration Machines Optimized for NoveltyabstractWe explore and enhance the ability of neural language models to generate novel scientific directions grounded in literature.Work on literature-based hypothesis generation has traditionally focused on binary link predictionseverely limiting the expressivity of hypotheses.This line of work also does not focus on optimizing novelty.We take a dramatic departure with a novel setting in which models use as input background contexts (e.g., problems, experimental settings, goals), and output natural language ideas grounded in literature.We present SCIMON, a modeling framework that uses retrieval of "inspirations" from past scientific papers, and explicitly optimizes for novelty by iteratively comparing to prior papers and updating idea suggestions until sufficient novelty is achieved.Comprehensive evaluations reveal that GPT-4 tends to generate ideas with overall low technical depth and novelty, while our methods partially mitigate this issue.Our work represents a first step toward evaluating and developing language models that generate new ideas derived from the scientific literature 1 . Qingyun Wang 0005, Doug Downey, Heng Ji 0001, Tom Hope |
ACL (1) | 2 |
| 2024 | ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer ReviewsabstractMike D’Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, Doug Downey. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mike D'Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, Doug Downey |
ACL (1) | 7 |
| 2023 | I2D2: Inductive Knowledge Distillation with NeuroLogic and Self-ImitationabstractChandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras, Ximing Lu, Lianhui Qin, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, Yejin Choi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras 0001, Ximing Lu, Lianhui Qin, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, Yejin Choi 0001 |
ACL (1) | 3 |
| 2023 | CiteSee: Augmenting Citations in Scientific Papers with Persistent and Personalized Historical ContextabstractWhen reading a scholarly article, inline citations help researchers contextualize the current article and discover relevant prior work. However, it can be challenging to prioritize and make sense of the hundreds of citations encountered during literature reviews. This paper introduces CiteSee, a paper reading tool that leverages a user’s publishing, reading, and saving activities to provide personalized visual augmentations and context around citations. First, CiteSee connects the current paper to familiar contexts by surfacing known citations a user had cited or opened. Second, CiteSee helps users prioritize their exploration by highlighting relevant but unknown citations based on saving and reading history. We conducted a lab study that suggests CiteSee is significantly more effective for paper discovery than three baselines. A field deployment study shows CiteSee helps participants keep track of their explorations and leads to better situational awareness and increased paper discovery via inline citation when conducting real-world literature reviews. Joseph Chee Chang, Amy X. Zhang, Jonathan Bragg, Andrew Head, Kyle Lo, Doug Downey, Daniel S. Weld |
CHI | 6 |
| 2023 | Relatedly: Scaffolding Literature Reviews with Existing Related Work SectionsabstractScholars who want to research a scientific topic must take time to read, extract meaning, and identify connections across many papers. As scientific literature grows, this becomes increasingly challenging. Meanwhile, authors summarize prior research in papers’ related work sections, though this is scoped to support a single paper. A formative study found that while reading multiple related work paragraphs helps overview a topic, it is hard to navigate overlapping and diverging references and research foci. In this work, we design a system, Relatedly, that scaffolds exploring and reading multiple related work paragraphs on a topic, with features including dynamic re-ranking and highlighting to spotlight unexplored dissimilar information, auto-generated descriptive paragraph headings, and low-lighting of redundant information. From a within-subjects user study (n=15), we found that scholars generate more coherent, insightful, and comprehensive topic outlines using Relatedly compared to a baseline paper list. Srishti Palani, Aakanksha Naik, Doug Downey, Amy X. Zhang, Jonathan Bragg, Joseph Chee Chang |
CHI | 3 |
| 2023 | Penguins Don't Fly: Reasoning about Generics through Instantiations and ExceptionsabstractEmily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathleen McKeown, Doug Downey, Yejin Choi. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Emily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathy McKeown, Doug Downey, Yejin Choi 0001 |
EACL | 5 |
| 2023 | S2abEL: A Dataset for Entity Linking from Scientific TablesabstractEntity linking (EL) is the task of linking a textual mention to its corresponding entry in a knowledge base, and is critical for many knowledge-intensive NLP applications.When applied to tables in scientific papers, EL is a step toward large-scale scientific knowledge bases that could enable advanced scientific question answering and analytics.We present the first dataset for EL in scientific tables.EL for scientific tables is especially challenging because scientific knowledge bases can be very incomplete, and disambiguating table mentions typically requires understanding the paper's text in addition to the table.Our dataset, Scientific Table Entity Linking (S2abEL), focuses on EL in machine learning results tables and includes hand-labeled cell types, attributed sources, and entity links from the PaperswithCode taxonomy for 8,429 cells from 732 tables.We introduce a neural baseline method designed for EL on scientific tables containing many out-of-knowledge-base mentions, and show that it significantly outperforms a state-of-the-art generic table EL method.The best baselines fall below human performance, and our analysis highlights avenues for improvement. Yuze Lou, Bailey Kuehl, Erin Bransom, Sergey Feldman, Aakanksha Naik, Doug Downey |
EMNLP | 6 |
| 2023 | SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsabstractLearned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning.However, existing benchmarks for evaluating these representations fail to capture the diversity of relevant tasks.In response, we introduce SciRepEval, the first comprehensive benchmark for training and evaluating scientific document representations.It includes 24 challenging and realistic tasks, 8 of which are new, across four formats: classification, regression, ranking and search.We then use this benchmark to study and improve the generalization ability of scientific document representation models.We show how state-of-the-art models like SPECTER and SciNCL struggle to generalize across the task formats, and that simple multi-task training fails to improve them.However, a new approach that learns multiple embeddings per document, each tailored to a different format, can improve performance.We experiment with task-format-specific control codes and adapters and find they outperform the existing single-embedding state-of-the-art by over 2 points absolute.We release the resulting family of multi-format models, called SPECTER2, for the community to use and build on. Amanpreet Singh, Mike D'Arcy, Arman Cohan, Doug Downey, Sergey Feldman |
EMNLP | 4 |
| 2023 | LIMEADE: From AI Explanations to Advice TakingabstractResearch in human-centered AI has shown the benefits of systems that can explain their predictions. Methods that allow AI to take advice from humans in response to explanations are similarly useful. While both capabilities are well developed for transparent learning models (e.g., linear models and GA 2 Ms) and recent techniques (e.g., LIME and SHAP) can generate explanations for opaque models, little attention has been given to advice methods for opaque models. This article introduces LIMEADE, the first general framework that translates both positive and negative advice (expressed using high-level vocabulary such as that employed by post hoc explanations) into an update to an arbitrary, underlying opaque model. We demonstrate the generality of our approach with case studies on 70 real-world models across two broad domains: image classification and text recommendation. We show that our method improves accuracy compared to a rigorous baseline on the image classification domains. For the text modality, we apply our framework to a neural recommender system for scientific papers on a public website; our user study shows that our framework leads to significantly higher perceived user control, trust, and satisfaction. Benjamin Charles Germain Lee, Doug Downey, Kyle Lo, Daniel S. Weld |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2022 | From Who You Know to What You Read: Augmenting Scientific Recommendations with Implicit Social NetworksabstractThe ever-increasing pace of scientific publication necessitates methods for quickly identifying relevant papers. While neural recommenders trained on user interests can help, they still result in long, monotonous lists of suggested papers. To improve the discovery experience we introduce multiple new methods for augmenting recommendations with textual relevance messages that highlight knowledge-graph connections between recommended papers and a user’s publication and interaction history. We explore associations mediated by author entities and those using citations alone. In a large-scale, real-world study, we show how our approach significantly increases engagement—and future engagement when mediated by authors—without introducing bias towards highly-cited authors. To expand message coverage for users with less publication or interaction history, we develop a novel method that highlights connections with proxy authors of interest to users and evaluate it in a controlled lab study. Finally, we synthesize design implications for future graph-based messages. Hyeonsu B. Kang, Rafal Kocielnik, Andrew Head, Jiangjiang Yang, Matt Latzke, Aniket Kittur, Daniel S. Weld, Doug Downey, Jonathan Bragg |
CHI | 8 |
| 2022 | Building a Shared Conceptual Model of Complex, Heterogeneous Data Systems: A Demonstration
Michael R. Anderson, Yuze Lou, Jiayun Zou, Michael J. Cafarella, Sarah E. Chasins, Doug Downey, Dinghao Shen, Jenny M. Vo-Phamhi, Anna Zeng |
CIDR | 6 |
| 2022 | Multi-LexSum: Real-world Summaries of Civil Rights Lawsuits at Multiple GranularitiesabstractWith the advent of large language models, methods for abstractive summarization have made great strides, creating potential for use in applications to aid knowledge workers processing unwieldy document collections. One such setting is the Civil Rights Litigation Clearinghouse (CRLC, https://clearinghouse.net), which posts information about large-scale civil rights lawsuits, serving lawyers, scholars, and the general public. Today, summarization in the CRLC requires extensive training of lawyers and law students who spend hours per case understanding multiple relevant documents in order to produce high-quality summaries of key events and outcomes. Motivated by this ongoing real-world summarization effort, we introduce Multi-LexSum, a collection of 9,280 expert-authored summaries drawn from ongoing CRLC writing. Multi-LexSum presents a challenging multi-document summarization task given the length of the source documents, often exceeding two hundred pages per case. Furthermore, Multi-LexSum is distinct from other datasets in its multiple target summaries, each at a different granularity (ranging from one-sentence "extreme" summaries to multi-paragraph narrations of over five hundred words). We present extensive analysis demonstrating that despite the high-quality summaries in the training data (adhering to strict content and style guidelines), state-of-the-art summarization models perform poorly on this task. We release Multi-LexSum for further summarization research and to facilitate the development of applications to assist in the CRLC's mission at https://multilexsum.github.io. Shannon Shen 0001, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, Doug Downey |
NeurIPS | 6 |
| 2022 | FeedLens: Polymorphic Lenses for Personalizing Exploratory Search over Knowledge GraphsabstractThe vast scale and open-ended nature of knowledge graphs (KGs) make exploratory search over them cognitively demanding for users. We introduce a new technique, polymorphic lenses, that improves exploratory search over a KG by obtaining new leverage from the existing preference models that KG-based systems maintain for recommending content. The approach is based on a simple but powerful observation: in a KG, preference models can be re-targeted to recommend not only entities of a single base entity type (e.g., papers in the scientific literature KG, products in an e-commerce KG), but also all other types (e.g., authors, conferences, institutions; sellers, buyers). We implement our technique in a novel system, FeedLens, which is built over Semantic Scholar, a production system for navigating the scientific literature KG. FeedLens reuses the existing preference models on Semantic Scholar—people’s curated research feeds—as lenses for exploratory search. Semantic Scholar users can curate multiple feeds/lenses for different topics of interest, e.g., one for human-centered AI and another for document embeddings. Although these lenses are defined in terms of papers, FeedLens re-purposes them to also guide search over authors, institutions, venues, etc. Our system design is based on feedback from intended users via two pilot surveys (n = 17 and n = 13, respectively). We compare FeedLens and Semantic Scholar via a third (within-subjects) user study (n = 15) and find that FeedLens increases user engagement while reducing the cognitive effort required to complete a short literature review task. Our qualitative results also highlight people’s preference for this more effective exploratory search experience enabled by FeedLens. Harmanpreet Kaur, Doug Downey, Amanpreet Singh, Evie Yu-Yen Cheng, Daniel S. Weld, Jonathan Bragg |
UIST | 2 |
| 2022 | ABNIRML: Analyzing the Behavior of Neural IR ModelsabstractAbstract Pretrained contextualized language models such as BERT and T5 have established a new state-of-the-art for ad-hoc search. However, it is not yet well understood why these methods are so effective, what makes some variants more effective than others, and what pitfalls they may have. We present a new comprehensive framework for Analyzing the Behavior of Neural IR ModeLs (ABNIRML), which includes new types of diagnostic probes that allow us to test several characteristics—such as writing styles, factuality, sensitivity to paraphrasing and word order—that are not addressed by previous techniques. To demonstrate the value of the framework, we conduct an extensive empirical study that yields insights into the factors that contribute to the neural model’s gains, and identify potential unintended biases the models exhibit. Some of our results confirm conventional wisdom, for example, that recent neural ranking models rely less on exact term overlap with the query, and instead leverage richer linguistic information, evidenced by their higher sensitivity to word and sentence order. Other results are more surprising, such as that some models (e.g., T5 and ColBERT) are biased towards factually correct (rather than simply relevant) texts. Further, some characteristics vary even for the same base language model, and other characteristics can appear due to random variations during model training.1 Sean MacAvaney, Sergey Feldman, Nazli Goharian, Doug Downey, Arman Cohan |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | VILA: Improving Structured Content Extraction from Scientific PDFs Using Visual Layout GroupsabstractAbstract Accurately extracting structured content from PDFs is a critical first step for NLP over scientific papers. Recent work has improved extraction accuracy by incorporating elementary layout information, for example, each token’s 2D position on the page, into language model pretraining. We introduce new methods that explicitly model VIsual LAyout (VILA) groups, that is, text lines or text blocks, to further improve performance. In our I-VILA approach, we show that simply inserting special tokens denoting layout group boundaries into model inputs can lead to a 1.9% Macro F1 improvement in token classification. In the H-VILA approach, we show that hierarchical encoding of layout-groups can result in up to 47% inference time reduction with less than 0.8% Macro F1 loss. Unlike prior layout-aware approaches, our methods do not require expensive additional pretraining, only fine-tuning, which we show can reduce training cost by up to 95%. Experiments are conducted on a newly curated evaluation suite, S2-VLUE, that unifies existing automatically labeled datasets and includes a new dataset of manual annotations covering diverse papers from 19 scientific disciplines. Pre-trained weights, benchmark datasets, and source code are available at https://github.com/allenai/VILA. Shannon Shen 0001, Kyle Lo, Lucy Lu Wang, Bailey Kuehl, Daniel S. Weld, Doug Downey |
Trans. Assoc. Comput. Linguistics | 6 |
| 2021 | Who's on First?: Probing the Learning and Representation Capabilities of Language Models on Deterministic Closed DomainsabstractThe capabilities of today’s natural language processing systems are typically evaluated using large datasets of curated questions and answers. While these are critical benchmarks of progress, they also suffer from weakness due to artificial distributions and incomplete knowledge. Artifacts arising from artificial distributions can overstate language model performance, while incomplete knowledge limits fine-grained analysis. In this work, we introduce a complementary benchmarking approach based on SimPlified Language Activity Traces (SPLAT). SPLATs are corpora of language encodings of activity in some closed domain (we study traces from chess and baseball games in this work). SPLAT datasets use naturally-arising distributions, allow the generation of question-answer pairs at scale, and afford complete knowledge in their closed domains. We show that language models of three different architectures can answer questions about world states using only verb-like encodings of activity. Our approach is extensible to new language models and additional question-answering tasks. David Demeter, Doug Downey |
CoNLL | 2 |
| 2021 | "It doesn't look good for a date": Transforming Critiques into Preferences for Conversational Recommendation SystemsabstractConversations aimed at determining good recommendations are iterative in nature.People often express their preferences in terms of a critique of the current recommendation (e.g., "It doesn't look good for a date"), requiring some degree of common sense for a preference to be inferred.In this work, we present a method for transforming a user critique into a positive preference (e.g., "I prefer more romantic") in order to retrieve reviews pertaining to potentially better recommendations (e.g., "Perfect for a romantic dinner").We leverage a large neural language model (LM) in a fewshot setting to perform critique-to-preference transformation, and we test two methods for retrieving recommendations: one that matches embeddings, and another that fine-tunes an LM for the task.We instantiate this approach in the restaurant domain and evaluate it using a new dataset of restaurant critiques.In an ablation study, we show that utilizing critiqueto-preference transformation improves recommendations, and that there are at least three general cases that explain this improved performance. Victor S. Bursztyn, Jennifer A. Healey, Nedim Lipka, Eunyee Koh, Doug Downey, Lawrence Birnbaum |
EMNLP (1) | 5 |
| 2021 | Simplified Data Wrangling with ir_datasetsabstractManaging the data for Information Retrieval (IR) experiments can be challenging. Dataset documentation is scattered across the Internet and once one obtains a copy of the data, there are numerous different data formats to work with. Even basic formats can have subtle dataset-specific nuances that need to be considered for proper use. To help mitigate these challenges, we introduce a new robust and lightweight tool (ir_datasets) for acquiring, managing, and performing typical operations over datasets used in IR. We primarily focus on textual datasets used for ad-hoc search. This tool provides both a Python and command line interface to numerous IR datasets and benchmarks. To our knowledge, this is the most extensive tool of its kind. Integrations with popular IR indexing and experimentation toolkits demonstrate the tool's utility. We also provide documentation of these datasets through the \sys catalog: https://ir-datasets.com/. The catalog acts as a hub for information on datasets used in IR, providing core information about what data each benchmark provides as well as links to more detailed information. We welcome community contributions and intend to continue to maintain and grow this tool. Sean MacAvaney, Andrew Yates, Sergey Feldman, Doug Downey, Arman Cohan, Nazli Goharian |
SIGIR | 4 |
| 2020 | Just Add Functions: A Neural-Symbolic Language Model
David Demeter, Doug Downey |
AAAI | 2 |
| 2020 | SPECTER: Document-level Representation Learning using Citation-informed TransformersabstractRepresentation learning is a critical ingredient for natural language processing systems.Recent Transformer language models like BERT learn powerful textual representations, but these models are targeted towards token-and sentence-level training objectives and do not leverage information on inter-document relatedness, which limits their document-level representation power.For applications on scientific documents, such as classification and recommendation, the embeddings power strong performance on end tasks.We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph.Unlike existing pretrained language models, SPECTER can be easily applied to downstream applications without task-specific fine-tuning.Additionally, to encourage further research on document-level models, we introduce SCIDOCS, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation.We show that SPECTER outperforms a variety of competitive baselines on the benchmark.1 Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, Daniel S. Weld |
ACL | 4 |
| 2020 | Stolen Probability: A Structural Weakness of Neural Language ModelsabstractNeural Network Language Models (NNLMs) generate probability distributions by applying a softmax function to a distance metric formed by taking the dot product of a prediction vector with all word vectors in a high-dimensional embedding space.The dot-product distance metric forms part of the inductive bias of NNLMs.Although NNLMs optimize well with this inductive bias, we show that this results in a sub-optimal ordering of the embedding space that structurally impoverishes some words at the expense of others when assigning probability.We present numerical, theoretical and empirical analyses showing that words on the interior of the convex hull in the embedding space have their probability bounded by the probabilities of the words on the hull. David Demeter, Gregory Kimmel, Doug Downey |
ACL | 3 |
| 2020 | Don't Stop Pretraining: Adapt Language Models to Domains and TasksabstractLanguage models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of these broad-coverage models, we investigate whether it is still helpful to tailor a pretrained model to the domain of a target task. We present a study across four domains (biomedical and computer science publications, news, and reviews) and eight classification tasks, showing that a second phase of pretraining in-domain (domain-adaptive pretraining) leads to performance gains, under both high- and low-resource settings. Moreover, adapting to the task's unlabeled data (task-adaptive pretraining) improves performance even after domain-adaptive pretraining. Finally, we show that adapting to a task corpus augmented using simple data selection strategies is an effective alternative, especially when resources for domain-adaptive pretraining might be unavailable. Overall, we consistently find that multi-phase adaptive pretraining offers large gains in task performance. Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith |
ACL | 6 |
| 2020 | Abductive Commonsense Reasoning
Chandra Bhagavatula, Ronan Le Bras 0001, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Yih, Yejin Choi 0001 |
ICLR | 7 |
| 2020 | High-Precision Extraction of Emerging Concepts from Scientific LiteratureabstractIdentification of new concepts in scientific literature can help power faceted search, scientific trend analysis, knowledge-base construction, and more, but current methods are lacking. Manual identification can't keep up with the torrent of new publications, while the precision of existing automatic techniques is too low for many applications. We present an unsupervised concept extraction method for scientific literature that achieves much higher precision than previous work. Our approach relies on a simple but novel intuition: each scientific concept is likely to be introduced or popularized by a single paper that is disproportionately cited by subsequent papers mentioning the concept. From a corpus of computer science papers on arXiv, we find that our method achieves a [email protected] of 99%, compared to 86% for prior work, and a substantially better precision-yield trade-off across the top 15,000 extractions. To stimulate research in this area, we release our code and data. Daniel King, Doug Downey, Daniel S. Weld |
SIGIR | 2 |
| 2020 | Practical Methods for Semi-automated Peer Grading in a Classroom SettingabstractPeer grading, in which students grade each other's work, can provide an educational opportunity for students and reduce grading effort for instructors. A variety of methods have been proposed for synthesizing peer-assigned grades into accurate submission grades. However, when the assumptions behind these methods are not met, they may underperform a simple baseline of averaging the peer grades. We introduce SABTXT, which improves over previous work through two mechanisms. First, SABTXT uses a limited amount of historical instructor ground truth to model and correct for each peer's grading bias. Secondly, SABTXT models the thoroughness of a peer review based on its textual content, and puts more weight on the more thorough peer reviews when computing submission grades. In our experiments with over ten thousand peer reviews collected over four courses, we show that SABTXT outperforms existing approaches on our collected data, and achieves a mean squared error that is 6% lower than the strongest baseline on average. Zheng Yuan 0004, Doug Downey |
UMAP | 2 |
| 2019 | A new evaluation framework for topic modeling algorithms based on synthetic corporaabstractTopic models are in widespread use in natural language processing and beyond. Here, we propose a new framework for the evaluation of topic modeling algorithms based on synthetic corpora containing an unambiguously defined ground truth topic structure. The major innovation of our approach is the ability to quantify the agreement between the planted and inferred topic structures by comparing the assigned topic labels at the level of the tokens. In experiments, our approach yields novel insights about the relative strengths of topic models as corpus characteristics vary, and the first evidence of an “undetectable phase” for topic models when the planted structure is weak. We also establish the practical relevance of the insights gained for synthetic corpora by predicting the performance of topic modeling algorithms in classification tasks in real-world corpora. Hanyu Shi 0003, Martin Gerlach, Isabel Diersen, Doug Downey, Luis A. Nunes Amaral |
AISTATS | 4 |
| 2018 | Controlling Global Statistics in Recurrent Neural Network Text GenerationabstractRecurrent neural network language models (RNNLMs) are an essential component for many language generation tasks such as machine translation, summarization, and automated conversation. Often, we would like to subject the text generated by the RNNLM to constraints, in order to overcome systemic errors (e.g. word repetition) or achieve application-specific goals (e.g. more positive sentiment). In this paper, we present a method for training RNNLMs to simultaneously optimize likelihood and follow a given set of statistical constraints on text generation. The problem is challenging because the statistical constraints are defined over aggregate model behavior, rather than model parameters, meaning that a straightforward parameter regularization approach is insufficient. We solve this problem using a dynamic regularizer that updates as training proceeds, based on the generative behavior of the RNNLMs. Our experiments show that the dynamic regularizer outperforms both generic training and a static regularization baseline. The approach is successful at improving word-level repetition statistics by a factor of four in RNNLMs on a definition modeling task. It also improves model perplexity when the statistical constraints are $n$-gram statistics taken from a large corpus. Thanapon Noraset, David Demeter, Doug Downey |
AAAI | 3 |
| 2018 | OTyper: A Neural Architecture for Open Named Entity TypingabstractNamed Entity Typing (NET) is valuable for many natural language processing tasks, such as relation extraction, question answering, knowledge base population, and co-reference resolution. Classical NET targeted a few coarse-grained types, but the task has expanded to sets of hundreds of types in recent years. Existing work in NET assumes that the target types are specified in advance, and that hand-labeled examples of each type are available. In this work, we introduce the task of Open Named Entity Typing (ONET), which is NET when the set of target types is not known in advance. We propose a neural network architecture for ONET, called OTyper, and evaluate its ability to tag entities with types not seen in training. On the benchmark FIGER(GOLD) dataset, OTyper achieves a weighted AUC-ROC score of 0.870 on unseen types, substantially outperforming pattern- and embedding-based baselines. Zheng Yuan 0004, Doug Downey |
AAAI | 2 |
| 2018 | Estimating Marginal Probabilities of n-grams for Recurrent Neural Language ModelsabstractRecurrent neural network language models (RNNLMs) are the current standard-bearer for statistical language modeling.However, RNNLMs only estimate probabilities for complete sequences of text, whereas some applications require context-independent phrase probabilities instead.In this paper, we study how to compute an RNNLM's marginal probability: the probability that the model assigns to a short sequence of text when the preceding context is not known.We introduce a simple method of altering the RNNLM training to make the model more accurate at marginal estimation.Our experiments demonstrate that the technique is effective compared to baselines including the traditional RNNLM probability and an importance sampling approach.Finally, we show how we can use the marginal estimation to improve an RNNLM by training the marginals to match n-gram probabilities from a larger corpus. Thanapon Noraset, Doug Downey, Lidong Bing |
EMNLP | 2 |
| 2017 | Definition Modeling: Learning to Define Word Embeddings in Natural LanguageabstractDistributed representations of words have been shown to capture lexical semantics, based on their effectiveness in word similarity and analogical relation tasks. But, these tasks only evaluate lexical semantics indirectly. In this paper, we study whether it is possible to utilize distributed representations to generate dictionary definitions of words, as a more direct and transparent representation of the embeddings' semantics. We introduce definition modeling, the task of generating a definition for a given word and its embedding. We present different definition model architectures based on recurrent neural networks, and experiment with the models over multiple data sets. Our results show that a model that controls dependencies between the word being defined and the definition words performs significantly better, and that a character-level convolution layer that leverages morphology can complement word-level embeddings. Our analysis reveals which components of our models contribute to accuracy. Finally, the errors made by a definition model may provide insight into the shortcomings of word embeddings. Thanapon Noraset, Lawrence Birnbaum, Doug Downey |
AAAI | 4 |
| 2017 | PAG2ADMG: A Novel Methodology to Enumerate Causal Graph StructuresabstractCausal graphs, such as directed acyclic graphs (DAGs) and partial ancestral graphs (PAGs), represent causal relationships among variables in a model. Methods exist for learning DAGs and PAGs from data and for converting DAGs to PAGs. However, these methods only output a single causal graph consistent with the independencies/dependencies (the Markov equivalence class M) estimated from the data. However, many distinct graphs may be consistent with M, and a data modeler may wish to select among these using domain knowledge. In this paper, we present a method that makes this possible. We introduce PAG2ADMG, the first method for enumerating all causal graphs consistent with M, under certain assumptions. PAG2ADMG converts a given PAG into a set of acyclic directed mixed graphs (ADMGs). We prove the correctness of the approach and demonstrate its efficiency relative to brute-force enumeration. Nishant Subramani, Doug Downey |
AAAI | 2 |
| 2017 | VecShare: A Framework for Sharing Word Representation VectorsabstractMany Natural Language Processing (NLP) models rely on distributed vector representations of words.Because the process of training word vectors can require large amounts of data and computation, NLP researchers and practitioners often utilize pre-trained embeddings downloaded from the Web.However, finding the best embeddings for a given task is difficult, and can be computationally prohibitive.We present a framework, called VecShare, that makes it easy to share and retrieve word embeddings on the Web.The framework leverages a public data-sharing infrastructure to host embedding sets, and provides automated mechanisms for retrieving the embeddings most similar to a given corpus.We perform an experimental evaluation of VecShare's similarity strategies, and show that they are effective at efficiently retrieving embeddings that boost accuracy in a document classification task.Finally, we provide an open-source Python library for using the VecShare framework.1 Jared Fernandez, Zhaocheng Yu, Doug Downey |
EMNLP | 3 |
| 2016 | Learning Hierarchically Decomposable Concepts with Active Over-LabelingabstractMany classification tasks target high-level concepts that can be decomposed into a hierarchy of finer-grained sub-concepts. For example, some string entities that are Locations are also Attractions, some Attractions are Museums, etc. Such hierarchies are common in named entity recognition (NER), document classification, and biological sequence analysis. We present a new approach for learning hierarchically decomposable concepts. The approach learns a high-level classifier (e.g., location vs. non-location) by seperately learning multiple finer-grained classifiers (e.g., museum vs. non-museum), and then combining the results. Soliciting labels at a finer level of granularity than that of the target concept is a new approach to active learning, which we term active over-labeling. In experiments in NER and document classification tasks, we show that active over-labeling substantially improves area under the precision-recall curve when compared with standard passive or active learning. Finally, because finer-grained labels may be more expensive to obtain, we also present a cost-sensitive active learner that uses a multi-armed bandit approach to dynamically choose the label granularity to target, and show that the bandit-based learner is robust to differences in label cost and labeling budget. Yuji Mo, Stephen D. Scott 0001, Doug Downey |
ICDM | 3 |
| 2016 | Beating the Artificial Chaos: Fighting OSN Spam Using Its Own TemplatesabstractOnline social networks (OSNs) are extremely popular among Internet users. However, spam originating from friends and acquaintances not only reduces the joy of Internet surfing but also causes damage to less security-savvy users. Prior countermeasures combat OSN spam from different angles. Due to the diversity of spam, there is hardly any existing method that can independently detect the majority or most of OSN spam. In this paper, we empirically analyze the textual pattern of a large collection of OSN spam. An inspiring finding is that the majority (e.g., 76.4% in 2015) of the collected spam is generated with underlying templates. Based on the analysis, we propose tangram, an OSN spam filtering system that performs online inspection on the stream of user-generated messages. Tangram extracts the templates of spam detected by existing methods and then matching messages against the templates toward the accurate and the fast spam detection. It automatically divides the OSN spam into segments and uses the segments to construct templates to filter future spam. Experimental results on Twitter and Facebook data sets show that tangram is highly accurate and can rapidly generate templates to throttle newly emerged campaigns. Furthermore, we analyze the behavior of detected OSN spammers. We find a series of spammer properties-such as spamming accounts are created in bursts and a single active organization orchestrates more spam than all other spammers combined-that promise more comprehensive spam countermeasures. Tiantian Zhu 0001, Yi Yang 0042, Kai Bu, Yan Chen 0004, Doug Downey, Kathy Lee, Alok N. Choudhary |
IEEE/ACM Trans. Netw. | 6 |
| 2015 | Efficient Methods for Inferring Large Sparse Topic HierarchiesabstractDoug Downey, Chandra Bhagavatula, Yi Yang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Doug Downey, Chandra Bhagavatula, Yi Yang 0042 |
ACL (1) | 1 |
| 2015 | Efficient Methods for Incorporating Knowledge into Topic ModelsabstractLatent Dirichlet allocation (LDA) is a popular topic modeling technique for exploring hidden topics in text corpora.Increasingly, topic modeling needs to scale to larger topic spaces and use richer forms of prior knowledge, such as word correlations or document labels.However, inference is cumbersome for LDA models with prior knowledge.As a result, LDA models that use prior knowledge only work in small-scale scenarios.In this work, we propose a factor graph framework, Sparse Constrained LDA (SC-LDA), for efficiently incorporating prior knowledge into LDA.We evaluate SC-LDA's ability to incorporate word correlation knowledge and document label knowledge on three benchmark datasets.Compared to several baseline methods, SC-LDA achieves comparable performance but is significantly faster. Yi Yang 0042, Doug Downey, Jordan L. Boyd-Graber |
EMNLP | 2 |
| 2015 | TabEL: Entity Linking in Web Tables
Chandra Bhagavatula, Thanapon Noraset, Doug Downey |
ISWC (1) | 3 |
| 2014 | Spam ain't as diverse as it seems: throttling OSN spam with templates underneathabstractIn online social networks (OSNs), spam originating from friends and acquaintances not only reduces the joy of Internet surfing but also causes damage to less security-savvy users. Prior countermeasures combat OSN spam from different angles. Due to the diversity of spam, there is hardly any existing method that can independently detect the majority or most of OSN spam. In this paper, we empirically analyze the textual pattern of a large collection of OSN spam. An inspiring finding is that the majority (63.0%) of the collected spam is generated with underlying templates. We therefore propose extracting templates of spam detected by existing methods and then matching messages against the templates toward accurate and fast spam detection. We implement this insight through Tangram, an OSN spam filtering system that performs online inspection on the stream of user-generated messages. Tangram automatically divides OSN spam into segments and uses the segments to construct templates to filter future spam. Experimental results show that Tangram is highly accurate and can rapidly generate templates to throttle newly emerged campaigns. Specifically, Tangram detects the most prevalent template-based spam with 95.7% true positive rate, whereas the existing template generation approach detects only 32.3%. The integration of Tangram and its auxiliary spam filter achieves an overall accuracy of 85.4% true positive rate and 0.33% false positive rate. Yi Yang 0042, Kai Bu, Yan Chen 0004, Doug Downey, Kathy Lee, Alok N. Choudhary |
ACSAC | 5 |
| 2014 | Adding High-Precision Links to WikipediaabstractWikipedia's link structure is a valuable resource for natural language processing tasks, but only a fraction of the concepts mentioned in each article are annotated with hyperlinks.In this paper, we study how to augment Wikipedia with additional high-precision links.We present 3W, a system that identifies concept mentions in Wikipedia text, and links each mention to its referent page.3W leverages rich semantic information present in Wikipedia to achieve high precision.Our experiments demonstrate that 3W can add an average of seven new links to each Wikipedia article, at a precision of 0.98. Thanapon Noraset, Chandra Bhagavatula, Doug Downey |
EMNLP | 3 |
| 2014 | Analyzing the content emphasis of web search enginesabstractMillions of people search the Web each day. As a consequence, the ranking algorithms employed by Web search engines have a profound influence on which pages users visit. Characterizing this influence, and informing users when different engines favor certain sites or points of view, enables more transparent access to the Web's information. We present PAWS, a platform for analyzing differences among Web search engines. PAWS measures content emphasis: the degree to which differences across search engines' rankings correlate with features of the ranked content, including point of view (e.g., positive or negative orientation toward their company's products) and advertisements. We propose an approach for identifying the orientations in search results at scale, through a novel technique that minimizes the expected number of human judgments required. We apply PAWS to news search on Google and Bing, and find no evidence that the engines emphasize results that express positive orientation toward the engine company's products. We do find that the engines emphasize particular news sites, and that they also favor pages containing their company's advertisements, as opposed to competitor advertisements. Mohammed A. Alam, Doug Downey |
SIGIR | 2 |
| 2014 | Learning Representations for Weakly Supervised Natural Language Processing TasksabstractFinding the right representations for words is critical for building accurate NLP systems when domain-specific labeled data for the task is scarce. This article investigates novel techniques for extracting features from n-gram models, Hidden Markov Models, and other statistical language models, including a novel Partial Lattice Markov Random Field model. Experiments on part-of-speech tagging and information extraction, among other tasks, indicate that features taken from statistical language models, in combination with more traditional features, outperform traditional representations alone, and that graphical model representations outperform n-gram models, especially on sparse and polysemous words. Fei Huang 0008, Arun Ahuja, Doug Downey, Yi Yang 0042, Yuhong Guo, Alexander Yates |
Comput. Linguistics | 3 |
| 2013 | Scaling Semi-supervised Naive Bayes with Feature Marginals
Michael Lucas, Doug Downey |
ACL (1) | 2 |
| 2013 | A probabilistic graphical model for brand reputation assessment in social networksabstractSocial media has become a popular platform that connects people who share information, in particular personal opinions. Through such a fast information exchange mechanism, reputation of individuals, consumer products, or business companies can be quickly built up within a social network. Recently, applications mining social network data start emerging to find the communities sharing the same interests for marketing purposes. Knowing the reputation of social network entities, such as celebrities or business companies, can help develop better strategies for election campaigns or new product advertisements. In this paper, we propose a probabilistic graphical model to collectively measure reputations of entities in social networks. By collecting and analyzing large amount of user activities on Facebook, our model can effectively and efficiently rank entities, such as presidential candidates, professional sport teams, musician bands, and companies, based on their social reputation. The proposed model produces results largely consistent with the two publicly available systems - movie ranking in Internet Movie Database and business school ranking by the US news & World Report - with the correlation coefficients of 0.75 and -0.71, respectively. Kunpeng Zhang 0001, Doug Downey, Zhengzhang Chen, Yusheng Xie, Yu Cheng 0001, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
ASONAM | 2 |
| 2013 | Overcoming the Memory Bottleneck in Distributed Training of Latent Variable Models of Text
Yi Yang 0042, Alexander Yates, Doug Downey |
HLT-NAACL | 3 |
| 2012 | Explanatory semantic relatedness and explicit spatialization for exploratory searchabstractExploratory search, in which a user investigates complex concepts, is cumbersome with today's search engines. We present a new exploratory search approach that generates interactive visualizations of query concepts using thematic cartography (e.g. choropleth maps, heat maps). We show how the approach can be applied broadly across both geographic and non-geographic contexts through explicit spatialization, a novel method that leverages any figure or diagram -- from a periodic table, to a parliamentary seating chart, to a world map -- as a spatial search environment. We enable this capability by introducing explanatory semantic relatedness measures. These measures extend frequently-used semantic relatedness measures to not only estimate the degree of relatedness between two concepts, but also generate human-readable explanations for their estimates by mining Wikipedia's text, hyperlinks, and category structure. We implement our approach in a system called Atlasify, evaluate its key components, and present several use cases. Brent J. Hecht, Samuel Carton, Mahmood Quaderi, Johannes Schöning, Martin Raubal, Darren Gergle, Doug Downey |
SIGIR | 7 |
| 2012 | Sentiment identification by incorporating syntax, semantics and context informationabstractThis paper proposes a method based on conditional random fields to incorporate sentence structure (syntax and semantics) and context information to identify sentiments of sentences within a document. It also proposes and evaluates two different active learning strategies for labeling sentiment data. The experiments with the proposed approach demonstrate a 5-15% improvement in accuracy on Amazon customer reviews compared to existing supervised learning and rule-based methods. Kunpeng Zhang 0001, Yusheng Xie, Yu Cheng 0001, Daniel Honbo, Doug Downey, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
SIGIR | 5 |
| 2011 | Local and Global Algorithms for Disambiguation to Wikipedia
Lev-Arie Ratinov, Dan Roth 0001, Doug Downey |
ACL | 3 |
| 2011 | Language Models as Representations for Weakly Supervised NLP Tasks
Fei Huang 0008, Alexander Yates, Arun Ahuja, Doug Downey |
CoNLL | 4 |
| 2010 | Improved Extraction Assessment through Better Language Models
Arun Ahuja, Doug Downey |
HLT-NAACL | 2 |
| 2010 | Analysis of a probabilistic model of redundancy in unsupervised information extraction
Doug Downey, Oren Etzioni, Stephen Soderland |
Artif. Intell. | 1 |
| 2008 | Understanding the relationship between searchers' queries and information goalsabstractWe describe results from Web search log studies aimed at elucidating user behaviors associated with queries and destination URLs that appear with different frequencies. We note the diversity of information goals that searchers have and the differing ways that goals are specified. We examine rare and common information goals that are specified using rare or common queries. We identify several significant differences in user behavior depending on the rarity of the query and the destination URL. We find that searchers are more likely to be successful when the frequencies of the query and destination URL are similar. We also establish that the behavioral differences observed for queries and goals of varying rarity persist even after accounting for potential confounding variables, including query length, search engine ranking, session duration, and task difficulty. Finally, using an information-theoretic measure of search difficulty, we show that the benefits obtained by search and navigation actions depend on the frequency of the information goal. Doug Downey, Susan T. Dumais, Daniel J. Liebling, Eric Horvitz |
CIKM | 1 |
| 2008 | It's a Contradiction - no, it's not: A Case Study using Functional Relations
Alan Ritter, Stephen Soderland, Doug Downey, Oren Etzioni |
EMNLP | 3 |
| 2008 | Look Ma, No Hands: Analyzing the Monotonic Feature Abstraction for Text ClassificationabstractIs accurate classification possible in the absence of hand-labeled data? This paper introduces the Monotonic Feature (MF) abstraction--where the probability of class membership increases monotonically with the MF's value. The paper proves that when an MF is given, PAC learning is possible with no hand-labeled data under certain assumptions. We argue that MFs arise naturally in a broad range of textual classification applications. On the classic "20 Newsgroups" data set, a learner given an MF and unlabeled data achieves classification accuracy equal to that of a state-of-the-art semi-supervised learner relying on 160 hand-labeled examples. Even when MFs are not given as input, their presence or absence can be determined from a small amount of hand-labeled data, which yields a new semi-supervised learning method that reduces error by 15% on the 20 Newsgroups data. Doug Downey, Oren Etzioni |
NIPS | 1 |
| 2007 | Sparse Information Extraction: Unsupervised Language Models to the Rescue
Doug Downey, Stefan Schoenmackers, Oren Etzioni |
ACL | 1 |
| 2007 | Locating Complex Named Entities in Web Text
Doug Downey, Matthew Broadhead, Oren Etzioni |
IJCAI | 1 |
| 2007 | Models of Searching and Browsing: Languages, Studies, and Application
Doug Downey, Susan T. Dumais, Eric Horvitz |
IJCAI | 1 |
| 2007 | Heads and tails: studies of web search with common and rare queriesabstractA large fraction of queries submitted to Web search enginesoccur very infrequently. We describe search log studiesaimed at elucidating behaviors associated with rare andcommon queries. We present several analyses and discussresearch directions. Doug Downey, Susan T. Dumais, Eric Horvitz |
SIGIR | 1 |
| 2005 | A Probabilistic Model of Redundancy in Information Extraction
Doug Downey, Oren Etzioni, Stephen Soderland |
IJCAI | 1 |
| 2005 | Unsupervised named-entity extraction from the Web: An experimental study
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates |
Artif. Intell. | 3 |
| 2004 | Methods for Domain-Independent Information Extraction from the Web: An Experimental Comparison
Oren Etzioni, Michael J. Cafarella, Doug Downey, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates |
AAAI | 3 |
| 2004 | Web-scale information extraction in knowitall: (preliminary results)abstractManually querying search engines in order to accumulate a large bodyof factual information is a tedious, error-prone process of piecemealsearch. Search engines retrieve and rank potentially relevantdocuments for human perusal, but do not extract facts, assessconfidence, or fuse information from multiple documents. This paperintroduces KnowItAll, a system that aims to automate the tedious process ofextracting large collections of facts from the web in an autonomous,domain-independent, and scalable manner.The paper describes preliminary experiments in which an instance of KnowItAll, running for four days on a single machine, was able to automatically extract 54,753 facts. KnowItAll associates a probability with each fact enabling it to trade off precision and recall. The paper analyzes KnowItAll's architecture and reports on lessons learned for the design of large-scale information extraction systems. Oren Etzioni, Michael J. Cafarella, Doug Downey, Stanley Kok, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates |
WWW | 3 |