EDBT 2026 Demo / reviewers in the wild / expert
Chris Callison-Burch
dblp:78/1408 · also Christopher Callison-Burch
· DBLP profile ↗
117ranked-venue papers
12as first author
46since 2021 · last 2026
0000-0001-8196-1943ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 106 · 11 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 11 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LaTeX2Layout: High-Fidelity, Scalable Document Layout Annotation Pipeline for Layout DetectionabstractGeneral-purpose Vision-Language Models (VLMs) are increasingly integral to modern AI systems for document understanding, yet their ability to perform fine-grained layout analysis remains severely underdeveloped. Overcoming this limitation requires large-scale, high-fidelity training datasets. However, current annotation methods that rely on parsing rendered PDFs are costly, error-prone, and difficult to scale. We propose a different paradigm: extracting ground-truth layout directly from the LaTeX compilation process rather than the final PDF. We present LaTeX2Layout, a generalizable procedural pipeline that recovers pixel-accurate bounding boxes and reading order from compiler traces. This enables the generation of a 140K-page dataset, including 120K programmatically generated synthetic variants that more than double the layout diversity of real-world data. Using this dataset, we fine-tune an efficient 3B-parameter VLM with an easy-to-hard curriculum that accelerates convergence. Our model achieves Kendall's tau=0.95 for reading order and mAP@50=0.91 for element grounding, delivering nearly 200% relative improvement over strong zero-shot baselines such as GPT-4o and Claude-3.7. Feijiang Han, Skyler Cheung, Delip Rao, Chris Callison-Burch, Lyle H. Ungar |
AAAI | 7 |
| 2026 | NSF-SciFy: Mining the NSF Awards Database for Scientific ClaimsabstractWe introduce NSF-SciFy, a comprehensive dataset of scientific claims and investigation proposals extracted from National Science Foundation award abstracts. While previous scientific claim verification datasets have been limited in size and scope, NSF-SciFy represents a significant advance with 2.8 million claims from 400,000 abstracts spanning all science and mathematics disciplines. We present two focused subsets: NSF-SciFy-MatSci with 114,000 claims from materials science awards, and NSF-SciFy-20K with 135,000 claims across five NSF directorates. Using zero-shot prompting, we develop a scalable approach for joint extraction of scientific claims and investigation proposals. We demonstrate the dataset’s utility through three downstream tasks: non-technical abstract generation, claim extraction, and investigation proposal extraction. Fine-tuning language models on our dataset yields substantial improvements, with relative gains often exceeding 100%, particularly for claim and proposal extraction tasks. Our error analysis reveals that extracted claims exhibit high precision but lower recall, suggesting opportunities for further methodological refinement. NSF-SciFy enables new research directions in large-scale claim verification, scientific discovery tracking, and meta-scientific analysis. Delip Rao, Weiqiu You, Eric Wong 0001, Chris Callison-Burch |
ACL (1) | 4 |
| 2026 | What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
Delip Rao, Chris Callison-Burch |
NLDB | 2 |
| 2026 | ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
Delip Rao, Feijiang Han, Chris Callison-Burch |
NLDB | 3 |
| 2025 | Calibrating Large Language Models with Sample ConsistencyabstractAccurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF. Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch |
AAAI | 9 |
| 2025 | Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data GenerationabstractYue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi, Christopher Clark. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yue Yang 0006, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi |
ACL (1) | 8 |
| 2025 | Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News CoverageabstractMainstream media, through their decisions on what to cover and how to frame the stories they cover, can mislead readers without using outright falsehoods.Therefore, it is crucial to have tools that expose these editorial choices underlying media bias.In this paper, we introduce the Media Bias Detector, a tool for researchers, journalists, and news consumers.By integrating large language models, we provide near real-time granular insights into the topics, tone, political lean, and facts of news articles aggregated to the publisher level.We assessed the tool's impact by interviewing 13 experts from journalism, communications, and political science, revealing key insights into usability and functionality, practical applications, and AI's role in powering media bias tools.We explored this in more depth with a follow-up survey of 150 news consumers.This work highlights opportunities for AI-driven tools that empower users to critically engage with media content, particularly in politically charged environments. Jenny S. Wang, Samar Haider, Amir Tohidi, Anushkaa Gupta, Chris Callison-Burch, David M. Rothschild, Duncan J. Watts |
CHI | 6 |
| 2025 | Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsabstractToday’s most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well- tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog. Matt Deitke, Sangho Lee 0008, Rohun Tripathi, Yue Yang 0006, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, Yen-Sung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch, Andrew Head, Rose Hendrix, Favyen Bastani, Eli VanderBilt, Nathan Lambert 0001, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz 0002, Aaron Sarnat, Byron Bischoff, Pete Walsh 0001, Chris Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hannaneh Hajishirzi, Ross B. Girshick, Ali Farhadi, Aniruddha Kembhavi |
CVPR | 19 |
| 2025 | Concept Lancet: Image Editing with Compositional Representation TransplantabstractDiffusion models are widely used for image editing tasks. Existing editing methods often design a representation manipulation procedure by curating an edit direction in the text embedding or score space. However, such a procedure faces a key challenge: overestimating the edit strength harms visual consistency while underestimating it fails the editing task. Notably, each source image may require a different editing strength, and it is costly to search for an appropriate strength via trial-and-error. To address this challenge, we propose ${\boldsymbol{Co}}{\text{ncept}}\,{\boldsymbol{Lan}}{\text{cet}}$ (CoLan), a zero-shot plug-and-play framework for principled representation manipulation in diffusion-based image editing. At inference time, we decompose the source input in the latent (text embedding or diffusion score) space as a sparse linear combination of the representations of the collected visual concepts. This allows us to accurately estimate the presence of concepts in each image, which informs the edit. Based on the editing task (replace/add/remove), we perform a customized concept transplant process to impose the corresponding editing direction. To sufficiently model the concept space, we curate a conceptual representation dataset, CoLan-150K, which contains diverse descriptions and scenarios of visual terms and phrases for the latent dictionary. Experiments on multiple diffusion-based image editing baselines show that methods equipped with CoLan achieve state-of-the-art performance in editing effectiveness and consistency preservation. Jinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Hancheng Min, Chris Callison-Burch, René Vidal |
CVPR | 5 |
| 2025 | ViUniT: Visual Unit Tests for More Robust Visual ProgrammingabstractProgramming based approaches to reasoning tasks have substantially expanded the types of questions models can answer about visual scenes. Yet on benchmark visual reasoning data, when models answer correctly, they produce incorrect programs 33% of the time. These models are often right for the wrong reasons and risk unexpected failures on new data. Unit tests play a foundational role in ensuring code correctness and could be used to repair such failures. We propose Visual Unit Testing (ViUniT), a framework to improve the reliability of visual programs by automatically generating unit tests. In our framework, a unit test is represented as a novel image and answer pair meant to verify the logical correctness of a program produced for a given query. Our method leverages a language model to create unit tests in the form of image descriptions and expected answers, followed by image synthesis to produce corresponding images. We conduct a comprehensive analysis of what constitutes an effective visual unit test suite, exploring unit test generation, sampling strategies, image generation methods, and varying the number of programs and unit tests. Additionally, we introduce four applications of visual unit tests: best program selection, answer refusal, re-prompting, and unsupervised reward formulations for reinforcement learning. Experiments with two models across three datasets in visual question answering and image-text matching demonstrate that ViUniT improves model performance by 11.4 points in accuracy. Notably, it enables 7B open-source language models to outperform gpt-4o-mini in visual program generation by an average of 7.7 points and reduces the occurrence of programs that are correct for the wrong reasons by 40%. Artemis Panagopoulou, Honglu Zhou, Silvio Savarese, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles |
CVPR | 5 |
| 2025 | Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3DabstractArtemis Panagopoulou, Le Xue, Honglu Zhou, Silvio Savarese, Ran Xu, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Artemis Panagopoulou, Le Xue, Honglu Zhou, Silvio Savarese, Ran Xu 0001, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles |
EMNLP | 7 |
| 2025 | Probabilistic Soundness Guarantees in LLM Reasoning ChainsabstractIn reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion.Current LLM-based error detection methods often fail to detect propagated errors because earlier errors can corrupt judgments of downstream reasoning.To better detect such errors, we introduce Autoregressive Reasoning Entailment Stability (ARES), a probabilistic framework that evaluates each reasoning step based solely on previously-verified premises.This inductive method yields a nuanced score for each step and provides certified statistical guarantees of its soundness, rather than a brittle binary label.ARES achieves state-of-the-art performance across four benchmarks (72.1% Macro-F1, +8.2 points) and demonstrates superior robustness on very long synthetic reasoning chains, where it excels at detecting propagated errors (90.3% F1, +27.6 points). 1 Weiqiu You, Anton Xue, Shreya Havaldar, Delip Rao, Helen Jin, Chris Callison-Burch, Eric Wong 0001 |
EMNLP | 6 |
| 2025 | StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel ExamplesabstractAjay Patel, Jiacheng Zhu, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathleen McKeown, Chris Callison-Burch. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ajay Patel, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathy McKeown, Chris Callison-Burch |
NAACL (Long Papers) | 7 |
| 2025 | Predicting explainable dementia types with LLM-aided feature engineeringabstractMOTIVATION: The integration of Machine Learning and Artificial Intelligence (AI) into healthcare has immense potential due to the rapidly growing volume of clinical data. However, existing AI models, particularly Large Language Models (LLMs) like GPT-4, face significant challenges in terms of explainability and reliability, particularly in high-stakes domains like healthcare. RESULTS: This paper proposes a novel LLM-aided feature engineering approach that enhances interpretability by extracting clinically relevant features from the Oxford Textbook of Medicine. By converting clinical notes into concept vector representations and employing a linear classifier, our method achieved an accuracy of 0.72, outperforming a traditional n-gram Logistic Regression baseline (0.64) and the GPT-4 baseline (0.48), while focusing on high-level clinical features. We also explore using Text Embeddings to reduce the overall time and cost of our approach by 97%. AVAILABILITY AND IMPLEMENTATION: All code relevant to this paper is available at: https://github.com/AdityaKashyap423/Dementia_LLM_Feature_Engineering/tree/main. Aditya Kashyap, Delip Rao, Mary Regina Boland, Li Shen 0001, Chris Callison-Burch |
Bioinform. | 5 |
| 2024 | ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style TransferabstractTextual style transfer is the task of transforming stylistic properties of text while preserving meaning. Target "styles" can be defined in numerous ways, ranging from single attributes (e.g. formality) to authorship (e.g. Shakespeare). Previous unsupervised style-transfer approaches generally rely on significant amounts of labeled data for only a fixed set of styles or require large language models. In contrast, we introduce a novel diffusion-based framework for general-purpose style transfer that can be flexibly adapted to arbitrary target styles at inference time. Our parameter-efficient approach, ParaGuide, leverages paraphrase-conditioned diffusion models alongside gradient-based guidance from both off-the-shelf classifiers and strong existing style embedders to transform the style of text while preserving semantic information. We validate the method on the Enron Email Corpus, with both human and automatic evaluations, and find that it outperforms strong baselines on formality, sentiment, and even authorship style transfer. Zachary Horvitz, Ajay Patel, Chris Callison-Burch, Kathy McKeown |
AAAI | 3 |
| 2024 | RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsabstractLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch |
ACL (1) | 8 |
| 2024 | DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM WorkflowsabstractLarge language models (LLMs) have become a dominant and important tool for NLP researchers in a wide range of tasks.Today, many researchers use LLMs in synthetic data generation, task evaluation, fine-tuning, distillation, and other model-in-the-loop research workflows.However, challenges arise when using these models that stem from their scale, their closed source nature, and the lack of standardized tooling for these new and emerging workflows.The rapid rise to prominence of these models and these unique challenges has had immediate adverse impacts on open science and on the reproducibility of work that uses them.In this ACL 2024 theme track paper, we introduce DataDreamer, an open source Python library that allows researchers to write simple code to implement powerful LLM workflows.DataDreamer also helps researchers adhere to best practices that we propose to encourage open science and reproducibility.The library and documentation are available at: https://github.com/datadreamer-dev /DataDreamer. Ajay Patel, Colin Raffel, Chris Callison-Burch |
ACL (1) | 3 |
| 2024 | Choice-75: A Dataset on Decision Branching in Script LearningabstractScript learning studies how daily events unfold. It enables machines to reason about narratives with implicit information. Previous works mainly consider a script as a linear sequence of events while ignoring the potential branches that arise due to people’s circumstantial choices. We hence propose Choice-75, the first benchmark that challenges intelligent systems to make decisions given descriptive scenarios, containing 75 scripts and more than 600 scenarios. We also present preliminary results with current large language models (LLM). Although they demonstrate overall decent performances, there is still notable headroom in hard scenarios. Zhaoyi Hou, Li Zhang 0039, Chris Callison-Burch |
LREC/COLING | 3 |
| 2024 | Holodeck: Language Guided Generation of 3D Embodied AI Environmentsabstract3D simulated environments play a critical role in Embodied AI, but their creation requires expertise and extensive manual effort, restricting their diversity and scope. To miti-gate this limitation, we present Holodeck, a system that generates 3D environments to match a user-supplied prompt fullyautomatedly. Holodeck can generate diverse scenes, e.g., arcades, spas, and museums, adjust the designs for styles, and can capture the semantics of complex queries such as “apartment for a researcher with a cat” and “office of a professor who is a fan of Star Wars”. Holodeck leverages a large language model (i.e., GPT-4) for common sense knowledge about what the scene might look like and uses a large collection of 3D assets from Objaverse to populate the scene with diverse objects. To address the challenge of positioning objects correctly, we prompt GPT-4 to generate spatial relational constraints between objects and then optimize the layout to satisfy those constraints. Our large-scale human evaluation shows that annotators prefer Holodeck over manually designed procedural baselines in residential scenes and that Holodeck can produce high-quality outputs for diverse scene types. We also demonstrate an exciting application of Holodeck in Embodied AI, training agents to navigate in novel scenes like music rooms and daycares without human-constructed data, which is a significant step forward in developing general-purpose embodied agents. Yue Yang 0006, Fan-Yun Sun, Luca Weihs, Eli VanderBilt, Alvaro Herrasti, Winson Han, Jiajun Wu 0001, Nick Haber, Ranjay Krishna, Lingjie Liu, Chris Callison-Burch, Mark Yatskar, Aniruddha Kembhavi |
CVPR | 11 |
| 2024 | OpenPI2.0: An Improved Dataset for Entity Tracking in TextsabstractLi Zhang, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Li Zhang 0039, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon |
EACL (1) | 4 |
| 2024 | CoMo: Controllable Motion Generation Through Language Guided Pose Code Editing
Yiming Huang 0011, Weilin Wan 0001, Yue Yang 0006, Chris Callison-Burch, Mark Yatskar, Lingjie Liu |
ECCV (29) | 4 |
| 2024 | This Land is Your, My Land: Evaluating Geopolitical Bias in Language Models through Territorial DisputesabstractBryan Li, Samar Haider, Chris Callison-Burch. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Bryan Li, Samar Haider, Chris Callison-Burch |
NAACL-HLT | 3 |
| 2024 | A Textbook Remedy for Domain Shifts: Knowledge Priors for Medical Image AnalysisabstractWhile deep networks have achieved broad success in analyzing natural images, when applied to medical scans, they often fail in unexcepted situations. We investigate this challenge and focus on model sensitivity to domain shifts, such as data sampled from different hospitals or data confounded by demographic variables such as sex, race, etc, in the context of chest X-rays and skin lesion images. A key finding we show empirically is that existing visual backbones lack an appropriate prior from the architecture for reliable generalization in these settings. Taking inspiration from medical training, we propose giving deep networks a prior grounded in explicit medical knowledge communicated in natural language. To this end, we introduce Knowledge-enhanced Bottlenecks (KnoBo), a class of concept bottleneck models that incorporates knowledge priors that constrain it to reason with clinically relevant factors found in medical textbooks or PubMed. KnoBo uses retrieval-augmented language models to design an appropriate concept space paired with an automatic training procedure for recognizing the concept. We evaluate different resources of knowledge and recognition architectures on a broad range of domain shifts across 20 datasets. In our comprehensive evaluation with two imaging modalities, KnoBo outperforms fine-tuned models on confounded datasets by 32.4% on average. Finally, evaluations reveal that PubMed is a promising resource for making medical models less sensitive to domain shift, outperforming other resources on both diversity of information and final prediction performance. Yue Yang 0006, Mona Gandhi, Michael S. Yao, Chris Callison-Burch, James C. Gee, Mark Yatskar |
NeurIPS | 6 |
| 2024 | PaCE: Parsimonious Concept Engineering for Large Language ModelsabstractLarge Language Models (LLMs) are being used for a wide variety of tasks. While they are capable of generating human-like responses, they can also produce undesirable output including potentially harmful information, racist or sexist language, and hallucinations. Alignment methods are designed to reduce such undesirable output, via techniques such as fine-tuning, prompt engineering, and representation engineering. However, existing methods face several challenges: some require costly fine-tuning for every alignment task; some do not adequately remove undesirable concepts, failing alignment; some remove benign concepts, lowering the linguistic capabilities of LLMs. To address these issues, we propose Parsimonious Concept Engineering (PaCE), a novel activation engineering framework for alignment. First, to sufficiently model the concepts, we construct a large-scale concept dictionary in the activation space, in which each atom corresponds to a semantic concept. Given any alignment task, we instruct a concept partitioner to efficiently annotate the concepts as benign or undesirable. Then, at inference time, we decompose the LLM activations along the concept dictionary via sparse coding, to accurately represent the activations as linear combinations of benign and undesirable components. By removing the latter ones from the activations, we reorient the behavior of the LLM towards the alignment goal. We conduct experiments on tasks such as response detoxification, faithfulness enhancement, and sentiment revising, and show that PaCE achieves state-of-the-art alignment performance while maintaining linguistic capabilities. Jinqi Luo, Tianjiao Ding, Kwan Ho Ryan Chan, Darshan Thaker, Aditya Chattopadhyay, Chris Callison-Burch, René Vidal |
NeurIPS | 6 |
| 2024 | Towards Faithful Model Explanation in NLP: A SurveyabstractAbstract End-to-end neural Natural Language Processing (NLP) models are notoriously difficult to understand. This has given rise to numerous efforts towards model explainability in recent years. One desideratum of model explanation is faithfulness, that is, an explanation should accurately represent the reasoning process behind the model’s prediction. In this survey, we review over 110 model explanation methods in NLP through the lens of faithfulness. We first discuss the definition and evaluation of faithfulness, as well as its significance for explainability. We then introduce recent advances in faithful explanation, grouping existing approaches into five categories: similarity-based methods, analysis of model-internal structures, backpropagation-based methods, counterfactual intervention, and self-explanatory models. For each category, we synthesize its representative studies, strengths, and weaknesses. Finally, we summarize their common virtues and remaining challenges, and reflect on future work directions towards faithful explainability in NLP. Qing Lyu 0001, Marianna Apidianaki, Chris Callison-Burch |
Comput. Linguistics | 3 |
| 2023 | Rewriting the Script: Adapting Text Instructions for Voice InteractionabstractVoice assistants have sharply risen in popularity in recent years, but their use has been limited mostly to simple applications like music, hands-free search, or control of internet-of-things devices. What would it take for voice assistants to guide people through more complex tasks? In our work, we study the limitations of the dominant approach voice assistants take to complex task guidance: reading aloud written instructions. Using recipes as an example, we observe twelve participants cook at home with a state-of-the-art voice assistant. We learn that the current approach leads to nine challenges, including obscuring the bigger picture, overwhelming users with too much information, and failing to communicate affordances. Instructions delivered by a voice assistant are especially difficult because they cannot be skimmed as easily as written instructions. Alexa in particular did not surface crucial details to the user or answer questions well. We draw on our observations to propose eight ways in which voice assistants can “rewrite the script”—summarizing, signposting, splitting, elaborating, volunteering, reordering, redistributing, and visualizing—to transform written sources into forms that are readily communicated through spoken conversation. We conclude with a vision of how modern advancements in natural language processing can be leveraged for intelligent agents to guide users effectively through complex tasks. Alyssa Hwang, Natasha Oza, Chris Callison-Burch, Andrew Head |
Conference on Designing Interactive Systems | 3 |
| 2023 | Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextabstractAs text generated by large language models proliferates, it becomes vital to understand how humans engage with such text, and whether or not they are able to detect when the text they are reading did not originate with a human writer. Prior work on human detection of generated text focuses on the case where an entire passage is either human-written or machine-generated. In this paper, we study a more realistic setting where text begins as human-written and transitions to being generated by state-of-the-art neural language models. We show that, while annotators often struggle at this task, there is substantial variance in annotator skill and that given proper incentives, annotators can improve at this task over time. Furthermore, we conduct a detailed comparison study and analyze how a variety of variables (model size, decoding strategy, fine-tuning, prompt genre, etc.) affect human detection performance. Finally, we collect error annotations from our participants and use them to show that certain textual genres influence models to make different types of errors and that certain sentence-level features correlate highly with annotator selection. We release the RoFT dataset: a collection of over 21,000 human annotations paired with error classifications to encourage future work in human detection and evaluation of generated text. Liam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi, Chris Callison-Burch |
AAAI | 5 |
| 2023 | Open-Domain Hierarchical Event Schema Induction by Incremental Prompting and VerificationabstractEvent schemas are a form of world knowledge about the typical progression of events.Recent methods for event schema induction use information extraction systems to construct a large number of event graph instances from documents, and then learn to generalize the schema from such instances.In contrast, we propose to treat event schemas as a form of commonsense knowledge that can be derived from large language models (LLMs).This new paradigm greatly simplifies the schema induction process and allows us to handle both hierarchical relations and temporal relations between events in a straightforward way.Since event schemas have complex graph structures, we design an incremental prompting and verification method INCSCHEMA to break down the construction of a complex event graph into three stages: event skeleton construction, event expansion, and event-event relation verification.Compared to directly using LLMs to generate a linearized graph, INCSCHEMA can generate large and complex schemas with 7.2% F1 improvement in temporal relations and 31.0%F1 improvement in hierarchical relations.In addition, compared to the previous state-of-the-art closed-domain schema induction model, human assessors were able to cover ∼10% more events when translating the schemas into coherent stories and rated our schemas 1.3 points higher (on a 5-point scale) in terms of readability. 1 Ruining Zhao, Manling Li, Heng Ji 0001, Chris Callison-Burch, Jiawei Han 0001 |
ACL (1) | 5 |
| 2023 | Explanation-based Finetuning Makes Models More Robust to Spurious CuesabstractJosh Magnus Ludan, Yixuan Meng, Tai Nguyen, Saurabh Shah, Qing Lyu, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Josh Magnus Ludan, Yixuan Meng, Tai Nguyen 0005, Saurabh Shah, Qing Lyu 0001, Marianna Apidianaki, Chris Callison-Burch |
ACL (1) | 7 |
| 2023 | I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and DragonsabstractPei Zhou, Andrew Zhu, Jennifer Hu, Jay Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, Prithviraj Ammanabrolu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Andrew Zhu, Jennifer Hu 0001, Jay Pujara, Xiang Ren 0001, Chris Callison-Burch, Yejin Choi 0001, Prithviraj Ammanabrolu |
ACL (1) | 6 |
| 2023 | FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State InformationabstractDungeons & Dragons (D&D) is a tabletop roleplaying game with complex natural language interactions between players and hidden state information.Recent work has shown that large language models (LLMs) that have access to state information can generate higher quality game turns than LLMs that use dialog history alone.However, previous work used game state information that was heuristically created and was not a true gold standard game state.We present FIREBALL, a large dataset containing nearly 25,000 unique sessions from real D&D gameplay on Discord with true game state info.We recorded game play sessions of players who used the Avrae bot, which was developed to aid people in playing D&D online, capturing language, game commands and underlying game state information.We demonstrate that FIRE-BALL can improve natural language generation (NLG) by using Avrae state information, improving both automated metrics and human judgments of quality.Additionally, we show that LLMs can generate executable Avrae commands, particularly after finetuning. Andrew Zhu, Karmanya Aggarwal, Alexander H. Feng, Lara J. Martin, Chris Callison-Burch |
ACL (1) | 5 |
| 2023 | Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image ClassificationabstractConcept Bottleneck Models (CBM) are inherently interpretable models that factor model decisions into humanreadable concepts. They allow people to easily understand why a model is failing, a critical feature for high-stakes applications. CBMs require manually specified concepts and often under-perform their black box counterparts, preventing their broad adoption. We address these shortcomings and are first to show how to construct high-performance CBMs without manual specification of similar accuracy to black box models. Our approach, Language Guided Bottlenecks (LaBo), leverages a language model, GPT-3, to define a large space of possible bottlenecks. Given a problem domain, LaBo uses GPT-3 to produce factual sentences about categories to form candidate concepts. LaBo efficiently searches possible bottlenecks through a novel submodular utility that promotes the selection of discriminative and diverse information. Ultimately, GPT-3's sentential concepts can be aligned to images using CLIP, to form a bottleneck layer. Experiments demonstrate that LaBo is a highly effective prior for concepts important to visual recognition. In the evaluation with 11 diverse datasets, LaBo bottlenecks excel at few-shot classification: they are 11.7% more accurate than black box linear probes at 1 shot and comparable with more data. Overall, LaBo demonstrates that inherently interpretable models can be widely applied at similar, or better, performance than black box approaches.11Code and data are available at https://github.com/YueYANG1996/LaBo Yue Yang 0006, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, Mark Yatskar |
CVPR | 5 |
| 2023 | Bidirectional Language Models Are Also Few-shot Learners
Ajay Patel, Bryan Li, Mohammad Sadegh Rasooli, Noah Constant, Colin Raffel, Chris Callison-Burch |
ICLR | 6 |
| 2023 | Faithful Chain-of-Thought ReasoningabstractQing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qing Lyu 0001, Shreya Havaldar, Adam Stein, Li Zhang 0039, Delip Rao, Eric Wong 0001, Marianna Apidianaki, Chris Callison-Burch |
IJCNLP (1) | 8 |
| 2023 | Learning When to Speak: Latency and Quality Trade-offs for Simultaneous Speech-to-Speech Translation with Offline Models
Liam Dugan, Anshul Wadhawan, Kyle Spence, Chris Callison-Burch, Morgan McGuire, Victor B. Zordan |
INTERSPEECH | 4 |
| 2022 | Deduplicating Training Data Makes Language Models BetterabstractKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini |
ACL (1) | 6 |
| 2022 | Show Me More Details: Discovering Hierarchies of Procedures from Semi-structured Web DataabstractShuyan Zhou, Li Zhang, Yue Yang, Qing Lyu, Pengcheng Yin, Chris Callison-Burch, Graham Neubig. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shuyan Zhou, Li Zhang 0039, Yue Yang 0006, Qing Lyu 0001, Chris Callison-Burch, Graham Neubig |
ACL (1) | 6 |
| 2022 | Dungeons and Dragons as a Dialog Challenge for Artificial IntelligenceabstractAI researchers have posited Dungeons and Dragons (D&D) as a challenge problem to test systems on various language-related capabilities.In this paper, we frame D&D specifically as a dialogue system challenge, where the tasks are to both generate the next conversational turn in the game and predict the state of the game given the dialogue history.We create a gameplay dataset consisting of nearly 900 games, with a total of 7,000 players, 800,000 dialogue turns, 500,000 dice rolls, and 58 million words.We automatically annotate the data with partial state information about the game play.We train a large language model (LM) to generate the next game turn, conditioning it on different information.The LM can respond as a particular character or as the player who runs the game-i.e., the Dungeon Master (DM).It is trained to produce dialogue that is either in-character (roleplaying in the fictional world) or out-of-character (discussing rules or strategy).We perform a human evaluation to determine what factors make the generated output plausible and interesting.We further perform an automatic evaluation to determine how well the model can predict the game state given the history and examine how well tracking the game state improves its ability to produce plausible conversational output. Chris Callison-Burch, Gaurav Tomar, Lara J. Martin, Daphne Ippolito, Suma Bailis, David Reitter |
EMNLP | 1 |
| 2022 | Unsupervised Entity Linking with Guided Summarization and Multiple-Choice SelectionabstractEntity linking, the task of linking potentially ambiguous mentions in texts to corresponding knowledge-base entities, is an important component for language understanding.We address two challenge in entity linking: how to leverage wider contexts surrounding a mention, and how to deal with limited training data.We propose a fully unsupervised model called SumMC that first generates a guided summary of the contexts conditioning on the mention, and then casts the task to a multiple-choice problem where the model chooses an entity from a list of candidates.In addition to evaluating our model on existing datasets that focus on named entities, we create a new dataset that links noun phrases from WikiHow to Wikidata.We show that our SumMC model achieves stateof-the-art unsupervised performance on our new dataset and on existing datasets. Li Zhang 0039, Chris Callison-Burch |
EMNLP | 3 |
| 2022 | Did that happen? Predicting Social Media Posts that are Indicative of what happened in a scene: A case study of a TV showabstractWhile popular Television (TV) shows are airing, some users interested in these shows publish social media posts about the show. Analyzing social media posts related to a TV show can be beneficial for gaining insights about what happened during scenes of the show. This is a challenging task partly because a significant number of social media posts associated with a TV show or event may not clearly describe what happened during the event. In this work, we propose a method to predict social media posts (associated with scenes of a TV show) that are indicative of what transpired during the scenes of the show. We evaluate our method on social media (Twitter) posts associated with an episode of a popular TV show, Game of Thrones. We show that for each of the identified scenes, with high AUC’s, our method can predict posts that are indicative of what happened in a scene from those that are not-indicative. Based on Twitters policy, we will make the Tweeter ID’s of the Twitter posts used for this work publicly available. Anietie Andy, Reno Kriz, Sharath Chandra Guntuku, Derry Wijaya, Chris Callison-Burch |
LREC | 5 |
| 2022 | Is "My Favorite New Movie" My Favorite Movie? Probing the Understanding of Recursive Noun PhrasesabstractQing Lyu, Zheng Hua, Daoxin Li, Li Zhang, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Qing Lyu 0001, Daoxin Li, Li Zhang 0039, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 6 |
| 2021 | BiSECT: Learning to Split and Rephrase Sentences with BitextsabstractAn important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary.We introduce a novel dataset and a new model for this 'split and rephrase' task.Our BISECT training data consists of 1 million long English sentences paired with shorter, meaning-equivalent English sentences.We obtain these by extracting 1-2 sentence alignments in bilingual parallel corpora and then using machine translation to convert both sides of the corpus into the same language.BISECT contains higher quality training examples than previous Split and Rephrase corpora, with sentence splits that require more significant modifications.We categorize examples in our corpus, and use these categories in a novel model that allows us to target specific regions of the input sentence to be split and edited.Moreover, we show that models trained on BISECT can perform a wider variety of split operations and improve upon previous state-of-the-art approaches in automatic and human evaluations.1 Joongwon Kim, Mounica Maddela, Reno Kriz, Wei Xu 0004, Chris Callison-Burch |
EMNLP (1) | 5 |
| 2021 | "Wikily" Supervised Neural Translation Tailored to Cross-Lingual TasksabstractWe present a simple but effective approach for leveraging Wikipedia for neural machine translation as well as cross-lingual tasks of image captioning and dependency parsing without using any direct supervision from external parallel data or supervised models in the target language.We show that first sentences and titles of linked Wikipedia pages, as well as crosslingual image captions, are strong signals for a seed parallel data to extract bilingual dictionaries and cross-lingual word embeddings for mining parallel text from Wikipedia.Our final model achieves high BLEU scores that are close to or sometimes higher than strong supervised baselines in low-resource languages; e.g.supervised BLEU of 4.0 versus 12.1 from our model in English-to-Kazakh.Moreover, we tailor our "wikily" supervised translation models to unsupervised image captioning, and cross-lingual dependency parser transfer.In image captioning, we train a multitasking machine translation and image captioning pipeline for Arabic and English from which the Arabic training data is a translated version of the English captioning data, using our wikily-supervised translation models.Our captioning results on Arabic are slightly better than that of its supervised model.In dependency parsing, we translate a large amount of monolingual text, and use it as artificial training data in an annotation projection framework.We show that our model outperforms recent work on cross-lingual transfer of dependency parsers. Mohammad Sadegh Rasooli, Chris Callison-Burch, Derry Wijaya |
EMNLP (1) | 2 |
| 2021 | Visual Goal-Step Inference using wikiHowabstractUnderstanding what sequence of steps are needed to complete a goal can help artificial intelligence systems reason about human activities.Past work in NLP has examined the task of goal-step inference for text.We introduce the visual analogue.We propose the Visual Goal-Step Inference (VGSI) task, where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal.With a new dataset harvested from wikiHow consisting of 772,277 images representing human actions, we show that our task is challenging for state-of-theart multimodal models.Moreover, the multimodal representation learned from our data can be effectively transferred to other datasets like HowTo100m, increasing the VGSI accuracy by 15 -20%.Our task will facilitate multimodal reasoning about procedural events. Yue Yang 0006, Artemis Panagopoulou, Qing Lyu 0001, Li Zhang 0039, Mark Yatskar, Chris Callison-Burch |
EMNLP (1) | 6 |
| 2021 | Goal-Oriented Script ConstructionabstractThe knowledge of scripts, common chains of events in stereotypical scenarios, is a valuable asset for task-oriented natural language understanding systems.We propose the Goal-Oriented Script Construction task, where a model produces a sequence of steps to accomplish a given goal.We pilot our task on the first multilingual script learning dataset supporting 18 languages collected from wikiHow, a website containing half a million how-to articles.For baselines, we consider both a generationbased approach using a language model and a retrieval-based approach by first retrieving the relevant steps from a large candidate pool and then ordering them.We show that our task is practical, feasible but challenging for state-of-the-art Transformer models, and that our methods can be readily deployed for various other datasets and domains with decent zero-shot performance 1 . * Equal contribution. Qing Lyu 0001, Li Zhang 0039, Chris Callison-Burch |
INLG | 3 |
| 2021 | Cultural and Geographical Influences on Image Translatability of Words across LanguagesabstractNikzad Khani, Isidora Tourni, Mohammad Sadegh Rasooli, Chris Callison-Burch, Derry Tanti Wijaya. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nikzad Khani, Isidora Chara Tourni, Mohammad Sadegh Rasooli, Chris Callison-Burch, Derry Wijaya |
NAACL-HLT | 4 |
| 2020 | Automatic Detection of Generated Text is Easiest when Humans are FooledabstractRecent advancements in neural language modelling make it possible to rapidly generate vast amounts of human-sounding text.The capabilities of humans and automatic discriminators to detect machine-generated text have been a large source of research interest, but humans and machines rely on different cues to make their decisions.Here, we perform careful benchmarking and analysis of three popular sampling-based decoding strategies-topk, nucleus sampling, and untruncated random sampling-and show that improvements in decoding methods have primarily optimized for fooling humans.This comes at the expense of introducing statistical abnormalities that make detection easy for automatic systems.We also show that though both human and automatic detector performance improve with longer excerpt length, even multi-sentence excerpts can fool expert human raters over 30% of the time.Our findings reveal the importance of using both human and automatic detectors to assess the humanness of text generation systems. Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas Eck |
ACL | 3 |
| 2020 | Toward Better Storylines with Sentence-Level Language ModelsabstractWe propose a sentence-level language model which selects the next sentence in a story from a finite set of fluent alternatives.Since it does not need to model fluency, the sentence-level language model can focus on longer range dependencies, which are crucial for multisentence coherence.Rather than dealing with individual words, our method treats the story so far as a list of pre-trained sentence embeddings and predicts an embedding for the next sentence, which is more efficient than predicting word embeddings.Notably this allows us to consider a large number of candidates for the next sentence during training.We demonstrate the effectiveness of our approach with state-of-the-art accuracy on the unsupervised Story Cloze task and with promising results on larger-scale next sentence prediction tasks. Daphne Ippolito, David Grangier, Douglas Eck, Chris Callison-Burch |
ACL | 4 |
| 2020 | Reasoning about Goals, Steps, and Temporal Ordering with WikiHowabstractWe propose a suite of reasoning tasks on two types of relations between procedural events: goal-step relations ("learn poses" is a step in the larger goal of "doing yoga") and step-step temporal relations ("buy a yoga mat" typically precedes "learn poses"). We introduce a dataset targeting these two relations based on wikiHow, a website of instructional how-to articles. Our human-validated test set serves as a reliable benchmark for commonsense inference, with a gap of about 10% to 20% between the performance of state-of-the-art transformer models and human performance. Our automatically-generated training set allows models to effectively transfer to out-of-domain tasks requiring knowledge of procedural events, with greatly improved performances on SWAG, Snips, and the Story Cloze Test in zero- and few-shot settings. Li Zhang 0039, Qing Lyu 0001, Chris Callison-Burch |
EMNLP (1) | 3 |
| 2019 | Comparison of Diverse Decoding Methods from Conditional Language ModelsabstractWhile conditional language models have greatly improved in their ability to output high-quality natural language, many NLP applications benefit from being able to generate a diverse set of candidate sequences.Diverse decoding strategies aim to, within a givensized candidate list, cover as much of the space of high-quality outputs as possible, leading to improvements for tasks that re-rank and combine candidate outputs.Standard decoding methods, such as beam search, optimize for generating high likelihood sequences rather than diverse ones, though recent work has focused on increasing diversity in these methods.In this work, we perform an extensive survey of decoding-time strategies for generating diverse outputs from conditional language models.We also show how diversity can be improved without sacrificing quality by oversampling additional candidates, then filtering to the desired number. Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, Chris Callison-Burch |
ACL (1) | 5 |
| 2019 | Paraphrase-Sense-Tagged SentencesabstractMany natural language processing tasks require discriminating the particular meaning of a word in context, but building corpora for developing sense-aware models can be a challenge. We present a large resource of example usages for words having a particular meaning, called Paraphrase-Sense-Tagged Sentences (PSTS). Built on the premise that a word’s paraphrases instantiate its fine-grained meanings (i.e., bug has different meanings corresponding to its paraphrases fly and microbe) the resource contains up to 10,000 sentences for each of 3 million target-paraphrase pairs where the target word takes on the meaning of the paraphrase. We describe an automatic method based on bilingual pivoting used to enumerate sentences for PSTS, and present two models for ranking PSTS sentences based on their quality. Finally, we demonstrate the utility of PSTS by using it to build a dataset for the task of hypernym prediction in context. Training a model on this automatically generated dataset produces accuracy that is competitive with a model trained on smaller datasets crafted with some manual effort. Anne Cocos, Chris Callison-Burch |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | Learning Translations via Images with a Massively Multilingual Image DatasetabstractJohn Hewitt, Daphne Ippolito, Brendan Callahan, Reno Kriz, Derry Tanti Wijaya, Chris Callison-Burch. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. John Hewitt, Daphne Ippolito, Brendan Callahan, Reno Kriz, Derry Wijaya, Chris Callison-Burch |
ACL (1) | 6 |
| 2018 | A Data-Driven Analysis of Workers' Earnings on Amazon Mechanical TurkabstractA growing number of people are working as part of on-line crowd work. Crowd work is often thought to be low wage work. However, we know little about the wage distribution in practice and what causes low/high earnings in this setting. We recorded 2,676 workers performing 3.8 million tasks on Amazon Mechanical Turk. Our task-level analysis revealed that workers earned a median hourly wage of only ~$2/h, and only 4% earned more than $7.25/h. While the average requester pays more than $11/h, lower-paying requesters post much more work. Our wage calculations are influenced by how unpaid work is accounted for, e.g., time spent searching for tasks, working on tasks that are rejected, and working on tasks that are ultimately not submitted. We further explore the characteristics of tasks and working patterns that yield higher hourly wages. Our analysis informs platform design and worker tools to create a more positive future for crowd work. Kotaro Hara, Abi Adams, Kristy Milland, Saiph Savage, Chris Callison-Burch, Jeffrey P. Bigham |
CHI | 5 |
| 2018 | Learning Scalar Adjective Intensity from ParaphrasesabstractAdjectives like warm, hot, and scalding all describe temperature but differ in intensity.Understanding these differences between adjectives is a necessary part of reasoning about natural language.We propose a new paraphrasebased method to automatically learn the relative intensity relation that holds between a pair of scalar adjectives.Our approach analyzes over 36k adjectival pairs from the Paraphrase Database under the assumption that, for example, paraphrase pair really hot ↔ scalding suggests that hot < scalding.We show that combining this paraphrase evidence with existing, complementary pattern-and lexicon-based approaches improves the quality of systems for automatically ordering sets of scalar adjectives and inferring the polarity of indirect answers to yes/no questions. Anne Cocos, Veronica Wharton, Ellie Pavlick, Marianna Apidianaki, Chris Callison-Burch |
EMNLP | 5 |
| 2018 | Introducing NIEUW: Novel Incentives and Workflows for Eliciting Linguistic Data
Christopher Cieri, James Fiumara, Mark Y. Liberman, Chris Callison-Burch, Jonathan Wright |
LREC | 4 |
| 2018 | Comparing Constraints for Taxonomic OrganizationabstractAnne Cocos, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Anne Cocos, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 3 |
| 2018 | Simplification Using Paraphrases and Context-Based Lexical SubstitutionabstractReno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Reno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 4 |
| 2017 | Learning Translations via Matrix CompletionabstractDerry Tanti Wijaya, Brendan Callahan, John Hewitt, Jie Gao, Xiao Ling, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. Derry Wijaya, Brendan Callahan, John Hewitt, Marianna Apidianaki, Chris Callison-Burch |
EMNLP | 7 |
| 2017 | A Comprehensive Analysis of Bilingual Lexicon InductionabstractBilingual lexicon induction is the task of inducing word translations from monolingual corpora in two languages. In this article we present the most comprehensive analysis of bilingual lexicon induction to date. We present experiments on a wide range of languages and data sizes. We examine translation into English from 25 foreign languages: Albanian, Azeri, Bengali, Bosnian, Bulgarian, Cebuano, Gujarati, Hindi, Hungarian, Indonesian, Latvian, Nepali, Romanian, Serbian, Slovak, Somali, Spanish, Swedish, Tamil, Telugu, Turkish, Ukrainian, Uzbek, Vietnamese, and Welsh. We analyze the behavior of bilingual lexicon induction on low-frequency words, rather than testing solely on high-frequency words, as previous research has done. Low-frequency words are more relevant to statistical machine translation, where systems typically lack translations of rare words that fall outside of their training data. We systematically explore a wide range of features and phenomena that affect the quality of the translations discovered by bilingual lexicon induction. We provide illustrative examples of the highest ranking translations for orthogonal signals of translation equivalence like contextual similarity and temporal similarity. We analyze the effects of frequency and burstiness, and the sizes of the seed bilingual dictionaries and the monolingual training corpora. Additionally, we introduce a novel discriminative approach to bilingual lexicon induction. Our discriminative model is capable of combining a wide variety of features that individually provide only weak indications of translation equivalence. When feature weights are discriminatively set, these signals produce dramatically higher translation quality than previous approaches that combined signals in an unsupervised fashion (e.g., using minimum reciprocal rank). We also directly compare our model's performance against a sophisticated generative approach, the matching canonical correlation analysis (MCCA) algorithm used by Haghighi et al. ( 2008 ). Our algorithm achieves an accuracy of 42% versus MCCA's 15%. Ann Irvine, Chris Callison-Burch |
Comput. Linguistics | 2 |
| 2017 | Crowd control: Effectively utilizing unscreened crowd workers for biomedical data annotation
Anne Cocos, Ting Qian, Chris Callison-Burch, Aaron J. Masino |
J. Biomed. Informatics | 3 |
| 2016 | Most "babies" are "little" and most "problems" are "huge": Compositional Entailment in Adjective-NounsabstractWe examine adjective-noun (AN) composition in the task of recognizing textual entailment (RTE).We analyze behavior of ANs in large corpora and show that, despite conventional wisdom, adjectives do not always restrict the denotation of the nouns they modify.We use natural logic to characterize the variety of entailment relations that can result from AN composition.Predicting these relations depends on context and on commonsense knowledge, making AN composition especially challenging for current RTE systems.We demonstrate the inability of current stateof-the-art systems to handle AN composition in a simplified RTE task which involves the insertion of only a single word. Ellie Pavlick, Chris Callison-Burch |
ACL (1) | 2 |
| 2016 | Tense Manages to Predict Implicative Behavior in VerbsabstractImplicative verbs (e.g.manage) entail their complement clauses, while non-implicative verbs (e.g.want) do not.For example, while managing to solve the problem entails solving the problem, no such inference follows from wanting to solve the problem.Differentiating between implicative and non-implicative verbs is therefore an essential component of natural language understanding, relevant to applications such as textual entailment and summarization.We present a simple method for predicting implicativeness which exploits known constraints on the tense of implicative verbs and their complements.We show that this yields an effective, data-driven way of capturing this nuanced property in verbs.(0.14) UFJ wants to merge with Mitsubishi, a combination that'd surpass Citigroup as the world's biggest bank.⇒ The merger of Japanese Banks creates the world's biggest bank.(0.55)After graduating, Gallager chose to accept a full scholarship to play football for Temple University.⇒ Gallager attended Temple University.(0.68) Wilkins was allowed to leave in 1987 to join French outfit Paris Saint-Germain.⇒ Wilkins departed Milan in 1987. Ellie Pavlick, Chris Callison-Burch |
EMNLP | 2 |
| 2016 | The Gun Violence Database: A new task and data set for NLPabstractWe argue that NLP researchers are especially well-positioned to contribute to the national discussion about gun violence.Reasoning about the causes and outcomes of gun violence is typically dominated by politics and emotion, and data-driven research on the topic is stymied by a shortage of data and a lack of federal funding.However, data abounds in the form of unstructured text from news articles across the country.This is an ideal application of NLP technologies, such as relation extraction, coreference resolution, and event detection.We introduce a new and growing dataset, the Gun Violence Database, in order to facilitate the adaptation of current NLP technologies to the domain of gun violence, thus enabling better social science research on this important and under-resourced problem. Ellie Pavlick, Heng Ji 0001, Xiaoman Pan, Chris Callison-Burch |
EMNLP | 4 |
| 2016 | Clustering Paraphrases by Word Sense
Anne Cocos, Chris Callison-Burch |
HLT-NAACL | 2 |
| 2016 | End-to-end statistical machine translation with zero or small parallel textsabstractAbstract We use bilingual lexicon induction techniques, which learn translations from monolingual texts in two languages, to build an end-to-end statistical machine translation (SMT) system without the use of any bilingual sentence-aligned parallel corpora. We present detailed analysis of the accuracy of bilingual lexicon induction, and show how a discriminative model can be used to combine various signals of translation equivalence (like contextual similarity, temporal similarity, orthographic similarity and topic similarity). Our discriminative model produces higher accuracy translations than previous bilingual lexicon induction techniques. We reuse these signals of translation equivalence as features on a phrase-based SMT system. These monolingually estimated features enhance low resource SMT systems in addition to allowing end-to-end machine translation without parallel corpora. Ann Irvine, Chris Callison-Burch |
Nat. Lang. Eng. | 2 |
| 2016 | Optimizing Statistical Machine Translation for Text SimplificationabstractMost recent sentence simplification systems use basic machine translation models to learn lexical and syntactic paraphrases from a manually simplified parallel corpus. These methods are limited by the quality and quantity of manually simplified corpora, which are expensive to build. In this paper, we conduct an in-depth adaptation of statistical machine translation to perform text simplification, taking advantage of large-scale paraphrases learned from bilingual texts and a small amount of manual simplifications with multiple references. Our work is the first to design automatic metrics that are effective for tuning and evaluating simplification systems, which will facilitate iterative development for this task. Wei Xu 0004, Courtney Napoles, Ellie Pavlick, Quanze Chen, Chris Callison-Burch |
Trans. Assoc. Comput. Linguistics | 5 |
| 2015 | Adding Semantics to Data-Driven ParaphrasingabstractEllie Pavlick, Johan Bos, Malvina Nissim, Charley Beller, Benjamin Van Durme, Chris Callison-Burch. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Ellie Pavlick, Johan Bos, Malvina Nissim, Charley Beller, Benjamin Van Durme, Chris Callison-Burch |
ACL (1) | 6 |
| 2015 | Extracting Structured Information via Automatic + Human ComputationabstractWe present a system for extracting structured information from unstructured text using a combination of information retrieval, natural language processing, machine learning, and crowdsourcing. We test our pipeline by building a structured database of gun violence incidents in the United States. The results of our pilot study demonstrate that the proposed methodology is a viable way of collecting large-scale, up-to-date data for public health, public policy, and social science research. Ellie Pavlick, Chris Callison-Burch |
HCOMP | 2 |
| 2015 | Crowdsourcing for NLPabstractCrowdsourced applications to scientific problems is a hot research area, with over 10,000 publications in the past five years. Platforms such as Amazons Mechanical Turk and CrowdFlower provide researchers with easy access to large numbers of workers. The crowds vast supply of inexpensive, intelligent labor allows people to attack problems that were previously impractical and gives potential for detailed scientific inquiry of social, psychological, economic, and linguistic phenomena via massive sample sizes of human annotated data. We introduce crowdsourcing and describe how it is being used in both industry and academia. Crowdsourcing is valuable to computational linguists both (a) as a source of labeled training data for use in machine learning and (b) as a means of collecting computational social science data that link language use to underlying beliefs and behavior. We present case studies for both categories: (a) collecting labeled data for use in natural language processing tasks such as word sense disambiguation and machine translation and (b) collecting experimental data in the context of psychology; e.g. finding how word use varies with age, sex, personality, health, and happiness. We will also cover tools and techniques for crowdsourcing. Effectively collecting crowdsourced data requires careful attention to the collection process, through selection of appropriately qualified workers, giving clear instructions that are understandable to non-?experts, and performing quality control on the results to eliminate spammers who complete tasks randomly or carelessly in order to collect the small financial reward. We will introduce different crowdsourcing platforms, review privacy and institutional review board issues, and provide rules of thumb for cost and time estimates. Crowdsourced data also has a particular structure that raises issues in statistical analysis; we describe some of the key methods to address these issues. No prior exposure to the area is required. Chris Callison-Burch, Lyle H. Ungar, Ellie Pavlick |
HLT-NAACL | 1 |
| 2015 | Cost Optimization in Crowdsourcing Translation: Low cost translations made even cheaper
Mingkun Gao, Wei Xu 0004, Chris Callison-Burch |
HLT-NAACL | 3 |
| 2015 | Problems in Current Text Simplification Research: New Data Can HelpabstractSimple Wikipedia has dominated simplification research in the past 5 years. In this opinion paper, we argue that focusing on Wikipedia limits simplification research. We back up our arguments with corpus analysis and by highlighting statements that other researchers have made in the simplification literature. We introduce a new simplification dataset that is a significant improvement over Simple Wikipedia, and present a novel quantitative-comparative approach to study the quality of simplification data resources. Wei Xu 0004, Chris Callison-Burch, Courtney Napoles |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Are Two Heads Better than One? Crowdsourced Translation via a Two-Step Collaboration of Non-Professional Translators and EditorsabstractCrowdsourcing is a viable mechanism for creating training data for machine translation.It provides a low cost, fast turnaround way of processing large volumes of data.However, when compared to professional translation, naive collection of translations from non-professionals yields low-quality results.Careful quality control is necessary for crowdsourcing to work well.In this paper, we examine the challenges of a two-step collaboration process with translation and post-editing by non-professionals.We develop graphbased ranking models that automatically select the best output from multiple redundant versions of translations and edits, and improves translation quality closer to professionals. Mingkun Gao, Ellie Pavlick, Chris Callison-Burch |
ACL (1) | 4 |
| 2014 | Hallucinating Phrase Translations for Low Resource MTabstractWe demonstrate that "hallucinating" phrasal translations can significantly improve the quality of machine translation in low resource conditions.Our hallucinated phrase tables consist of entries composed from multiple unigram translations drawn from the baseline phrase table and from translations that are induced from monolingual corpora.The hallucinated phrase table is very noisy.Its translations are low precision but high recall.We counter this by introducing 30 new feature functions (including a variety of monolinguallyestimated features) and by aggressively pruning the phrase table.Our analysis evaluates the intrinsic quality of our hallucinated phrase pairs as well as their impact in end-to-end Spanish-English and Hindi-English MT. Ann Irvine, Chris Callison-Burch |
CoNLL | 2 |
| 2014 | PARADIGM: Paraphrase Diagnostics through Grammar MatchingabstractParaphrase evaluation is typically done either manually or through indirect, taskbased evaluation.We introduce an intrinsic evaluation PARADIGM which measures the goodness of paraphrase collections that are represented using synchronous grammars.We formulate two measures that evaluate these paraphrase grammars using gold standard sentential paraphrases drawn from a monolingual parallel corpus.The first measure calculates how often a paraphrase grammar is able to synchronously parse the sentence pairs in the corpus.The second measure enumerates paraphrase rules from the monolingual parallel corpus and calculates the overlap between this reference paraphrase collection and the paraphrase resource being evaluated.We demonstrate the use of these evaluation metrics on paraphrase collections derived from three different data types: multiple translations of classic French novels, comparable sentence pairs drawn from different newspapers, and bilingual parallel corpora.We show that PARADIGM correlates with human judgments more strongly than BLEU on a task-based evaluation of paraphrase quality. Jonathan Weese, Juri Ganitkevitch, Chris Callison-Burch |
EACL | 3 |
| 2014 | Crowd-Workers: Aggregating Information Across Turkers to Help Them Find Higher Paying WorkabstractThe Mechanical Turk crowdsourcing platform currently fails to provide the most basic piece of information to enable workers to make informed decisions about which tasks to undertake: what is the expected hourly pay? Mechanical Turk advertises a reward amount per assignment, but does not give any indication of how long each assignment will take. We have developed a browser plugin that tracks the length of time it takes to complete a task, and a web service that aggregates the information across many workers. Our web service, crowd-workers.com, allows workers to discovery higher paying work by sorting tasks by estimated hourly rate. Chris Callison-Burch |
HCOMP | 1 |
| 2014 | Poetry of the Crowd: A Human Computation Algorithm to Convert Prose into Rhyming VerseabstractPoetry composition is a very complex task that requires a poet to satisfy multiple constraints concurrently. We believe that the task can be augmented by combining the creative abilities of humans with computational algorithms that efficiently constrain and permute available choices. We present a hybrid method for generating poetry from prose that combines crowdsourcing with natural language processing (NLP) machinery. We test the ability of crowd workers to accomplish the technically challenging and creative task of composing poems. Quanze Chen, Chenyang Lei, Wei Xu 0004, Ellie Pavlick, Chris Callison-Burch |
HCOMP | 5 |
| 2014 | A Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic
Ryan Cotterell, Chris Callison-Burch |
LREC | 2 |
| 2014 | The Multilingual Paraphrase Database
Juri Ganitkevitch, Chris Callison-Burch |
LREC | 2 |
| 2014 | The American Local News Corpus
Ann Irvine, Joshua Langfus, Chris Callison-Burch |
LREC | 3 |
| 2014 | Arabic Dialect IdentificationabstractThe written form of the Arabic language, Modern Standard Arabic (MSA), differs in a non-trivial manner from the various spoken regional dialects of Arabic—the true “native languages” of Arabic speakers. Those dialects, in turn, differ quite a bit from each other. However, due to MSA's prevalence in written form, almost all Arabic data sets have predominantly MSA content. In this article, we describe the creation of a novel Arabic resource with dialect annotations. We have created a large monolingual data set rich in dialectal Arabic content called the Arabic On-line Commentary Data set (Zaidan and Callison-Burch 2011). We describe our annotation effort to identify the dialect level (and dialect itself) in each of more than 100,000 sentences from the data set by crowdsourcing the annotation task, and delve into interesting annotator behaviors (like over-identification of one's own dialect). Using this new annotated data set, we consider the task of Arabic dialect identification: Given the word sequence forming an Arabic sentence, determine the variety of Arabic in which it is written. We use the data to train and evaluate automatic classifiers for dialect identification, and establish that classifiers using dialectal data significantly and dramatically outperform baselines that use MSA-only data, achieving near-human classification accuracy. Finally, we apply our classifiers to discover dialectical data from a large Web crawl consisting of 3.5 million pages mined from on-line Arabic newspapers. Omar Zaidan, Chris Callison-Burch |
Comput. Linguistics | 2 |
| 2014 | The Language Demographics of Amazon Mechanical TurkabstractWe present a large scale study of the languages spoken by bilingual workers on Mechanical Turk (MTurk). We establish a methodology for determining the language skills of anonymous crowd workers that is more robust than simple surveying. We validate workers’ self-reported language skill claims by measuring their ability to correctly translate words, and by geolocating workers to see if they reside in countries where the languages are likely to be spoken. Rather than posting a one-off survey, we posted paid tasks consisting of 1,000 assignments to translate a total of 10,000 words in each of 100 languages. Our study ran for several months, and was highly visible on the MTurk crowdsourcing platform, increasing the chances that bilingual workers would complete it. Our study was useful both to create bilingual dictionaries and to act as census of the bilingual speakers on MTurk. We use this data to recommend languages with the largest speaker populations as good candidates for other researchers who want to develop crowdsourced, multilingual technologies. To further demonstrate the value of creating data via crowdsourcing, we hire workers to create bilingual parallel corpora in six Indian languages, and use them to train statistical machine translation systems. Ellie Pavlick, Matt Post, Ann Irvine, Dmitry Kachaev, Chris Callison-Burch |
Trans. Assoc. Comput. Linguistics | 5 |
| 2014 | Extracting Lexically Divergent Paraphrases from TwitterabstractWe present MultiP (Multi-instance Learning Paraphrase Model), a new model suited to identify paraphrases within the short messages on Twitter. We jointly model paraphrase relations between word and sentence pairs and assume only sentence-level annotations during learning. Using this principled latent variable model alone, we achieve the performance competitive with a state-of-the-art method which combines a latent space model with a feature-based supervised classifier. Our model also captures lexically divergent paraphrases that differ from yet complement previous methods; combining our model with previous work significantly outperforms the state-of-the-art. In addition, we present a novel annotation methodology that has allowed us to crowdsource a paraphrase corpus from Twitter. We make this new dataset available to the research community. Wei Xu 0004, Alan Ritter, Chris Callison-Burch, William B. Dolan, Yangfeng Ji |
Trans. Assoc. Comput. Linguistics | 3 |
| 2013 | Dirt Cheap Web-Scale Parallel Text from the Common Crawl
Jason Smith 0006, Herve Saint-Amand, Magdalena Plamada, Philipp Koehn, Chris Callison-Burch, Adam Lopez |
ACL (1) | 5 |
| 2013 | Semi-Markov Phrase-Based Monolingual AlignmentabstractWe introduce a novel discriminative model for phrase-based monolingual alignment using a semi-Markov CRF.Our model achieves stateof-the-art alignment accuracy on two phrasebased alignment datasets (RTE and paraphrase), while doing significantly better than other strong baselines in both non-identical alignment and phrase-only alignment.Additional experiments highlight the potential benefit of our alignment model to RTE, paraphrase identification and question answering, where even a naive application of our model's alignment score approaches the state of the art. Xuchen Yao, Benjamin Van Durme, Chris Callison-Burch, Peter Clark |
EMNLP | 3 |
| 2013 | PPDB: The Paraphrase Database
Juri Ganitkevitch, Benjamin Van Durme, Chris Callison-Burch |
HLT-NAACL | 3 |
| 2013 | Supervised Bilingual Lexicon Induction with Multiple Monolingual Signals
Ann Irvine, Chris Callison-Burch |
HLT-NAACL | 2 |
| 2013 | Answer Extraction as Sequence Tagging with Tree Edit Distance
Xuchen Yao, Benjamin Van Durme, Chris Callison-Burch, Peter Clark |
HLT-NAACL | 3 |
| 2013 | Learning to translate with products of novices: a suite of open-ended challenge problems for teaching MTabstractMachine translation (MT) draws from several different disciplines, making it a complex subject to teach. There are excellent pedagogical texts, but problems in MT and current algorithms for solving them are best learned by doing. As a centerpiece of our MT course, we devised a series of open-ended challenges for students in which the goal was to improve performance on carefully constrained instances of four key MT tasks: alignment, decoding, evaluation, and reranking. Students brought a diverse set of techniques to the problems, including some novel solutions which performed remarkably well. A surprising and exciting outcome was that student solutions or their combinations fared competitively on some tasks, demonstrating that even newcomers to the field can help improve the state-of-the-art on hard NLP problems while simultaneously learning a great deal. The problems, baseline code, and results are freely available. Adam Lopez, Matt Post, Chris Callison-Burch, Jonathan Weese, Juri Ganitkevitch, Narges Ahmidi, Olivia Buzek, Leah Hanson, Beaniesh Jamil, Matthias A. Lee, Ya-Ting Lin, Henry Pao, Fatima Rivera, Leili Shahriyari, Debu Sinha, Adam R. Teichert, Stephen Wampler, Michael Weinberger, Daguang Xu, Lin Yang 0002, Shang Zhao 0002 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2012 | Semi-supervised discriminative language modeling for Turkish ASRabstractWe present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results. Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 15 |
| 2012 | Hallucinated n-best lists for discriminative language modelingabstractThis paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods. Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 12 |
| 2012 | Continuous space discriminative language modelingabstractDiscriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well. Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 12 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 13 |
| 2012 | Expectations of Word Sense in Parallel Corpora
Xuchen Yao, Benjamin Van Durme, Chris Callison-Burch |
HLT-NAACL | 3 |
| 2012 | Machine Translation of Arabic Dialects
Rabih Zbib, Erika Malchiodi, Jacob Devlin, David Stallard, Spyridon Matsoukas, Richard M. Schwartz, John Makhoul, Omar Zaidan, Chris Callison-Burch |
HLT-NAACL | 9 |
| 2012 | Modality and Negation in SIMT Use of Modality and Negation in Semantically-Informed Syntactic MTabstractThis article describes the resource- and system-building efforts of an 8-week Johns Hopkins University Human Language Technology Center of Excellence Summer Camp for Applied Language Exploration (SCALE-2009) on Semantically Informed Machine Translation (SIMT). We describe a new modality/negation (MN) annotation scheme, the creation of a (publicly available) MN lexicon, and two automated MN taggers that we built using the annotation scheme and lexicon. Our annotation scheme isolates three components of modality and negation: a trigger (a word that conveys modality or negation), a target (an action associated with modality or negation), and a holder (an experiencer of modality). We describe how our MN lexicon was semi-automatically produced and we demonstrate that a structure-based MN tagger results in precision around 86% (depending on genre) for tagging of a standard LDC data set. We apply our MN annotation scheme to statistical machine translation using a syntactic framework that supports the inclusion of semantic annotations. Syntactic tags enriched with semantic annotations are assigned to parse trees in the target-language training texts through a process of tree grafting. Although the focus of our work is modality and negation, the tree grafting procedure is general and supports other types of semantic information. We exploit this capability by including named entities, produced by a pre-existing tagger, in addition to the MN elements produced by the taggers described here. The resulting system significantly outperformed a linguistically naive baseline model (Hiero), and reached the highest scores yet reported on the NIST 2009 Urdu–English test set. This finding supports the hypothesis that both syntactic and semantic information can improve translation quality. Kathrin Baker, Michael Bloodgood, Bonnie J. Dorr, Chris Callison-Burch, Nathaniel Wesley Filardo, Christine D. Piatko, Lori S. Levin |
Comput. Linguistics | 4 |
| 2011 | Incremental Syntactic Language Models for Phrase-based Translation
Lane Schwartz, Chris Callison-Burch, William Schuler, Stephen T. Wu |
ACL | 2 |
| 2011 | Crowdsourcing Translation: Professional Quality from Non-Professionals
Omar Zaidan, Chris Callison-Burch |
ACL | 2 |
| 2011 | Learning Sentential Paraphrases from Bilingual Parallel Corpora for Text-to-Text Generation
Juri Ganitkevitch, Chris Callison-Burch, Courtney Napoles, Benjamin Van Durme |
EMNLP | 2 |
| 2010 | Bucking the Trend: Large-Scale Cost-Focused Active Learning for Statistical Machine Translation
Michael Bloodgood, Chris Callison-Burch |
ACL | 2 |
| 2010 | Stream-based Translation Models for Statistical Machine Translation
Abby D. Levenberg, Chris Callison-Burch, Miles Osborne |
HLT-NAACL | 2 |
| 2010 | Cheap, Fast and Good Enough: Automatic Speech Recognition with Non-Expert Transcription
Scott Novotney, Chris Callison-Burch |
HLT-NAACL | 2 |
| 2010 | Predicting Human-Targeted Translation Edit Rate via Untrained Human Annotators
Omar Zaidan, Chris Callison-Burch |
HLT-NAACL | 2 |
| 2009 | Improving Translation Lexicon Induction from Monolingual Corpora via Dependency Contexts and Part-of-Speech Equivalences
Nikesh Garera, Chris Callison-Burch, David Yarowsky |
CoNLL | 2 |
| 2009 | Fast, Cheap, and Creative: Evaluating Translation Quality Using Amazon's Mechanical Turk
Chris Callison-Burch |
EMNLP | 1 |
| 2009 | Improved Statistical Machine Translation Using Monolingually-Derived Paraphrases
Yuval Marton, Chris Callison-Burch, Philip Resnik |
EMNLP | 2 |
| 2009 | Feasibility of Human-in-the-loop Minimum Error Rate Training
Omar Zaidan, Chris Callison-Burch |
EMNLP | 2 |
| 2008 | ParaMetric: An Automatic Evaluation Metric for Paraphrasing
Chris Callison-Burch, Trevor Cohn, Mirella Lapata |
COLING | 1 |
| 2008 | Syntactic Constraints on Paraphrases Extracted from Parallel Corpora
Chris Callison-Burch |
EMNLP | 1 |
| 2008 | Constructing Corpora for the Development and Evaluation of Paraphrase SystemsabstractAutomatic paraphrasing is an important component in many natural language processing tasks. In this article we present a new parallel corpus with paraphrase annotations. We adopt a definition of paraphrase based on word alignments and show that it yields high inter-annotator agreement. As Kappa is suited to nominal data, we employ an alternative agreement statistic which is appropriate for structured alignment tasks. We discuss how the corpus can be usefully employed in evaluating paraphrase systems automatically (e.g., by measuring precision, recall, and F1) and also in developing linguistically rich paraphrase models based on syntactic structure. Trevor Cohn, Chris Callison-Burch, Mirella Lapata |
Comput. Linguistics | 2 |
| 2007 | Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst |
ACL | 4 |
| 2006 | Re-evaluation the Role of Bleu in Machine Translation Research
Chris Callison-Burch, Miles Osborne, Philipp Koehn |
EACL | 1 |
| 2006 | Improved Statistical Machine Translation Using Paraphrases
Chris Callison-Burch, Philipp Koehn, Miles Osborne |
HLT-NAACL | 1 |
| 2005 | Paraphrasing with Bilingual Parallel CorporaabstractPrevious work has used monolingual parallel corpora to extract and generate paraphrases. We show that this task can be done using bilingual parallel corpora, a much more commonly available resource. Using alignment techniques from phrase-based statistical machine translation, we show how paraphrases in one language can be identified using a phrase in another language as a pivot. We define a paraphrase probability that allows paraphrases extracted from a bilingual parallel corpus to be ranked using translation probabilities, and show how it can be refined to take contextual information into account. We evaluate our paraphrase extraction and ranking methods using a set of manual word alignments, and contrast the quality with paraphrases extracted from automatic alignments. Colin J. Bannard, Chris Callison-Burch |
ACL | 2 |
| 2005 | Scaling Phrase-Based Statistical Machine Translation to Larger Corpora and Longer PhrasesabstractIn this paper we describe a novel data structure for phrase-based statistical machine translation which allows for the retrieval of arbitrarily long phrases while simultaneously using less memory than is required by current decoder implementations. We detail the computational complexity and average retrieval times for looking up phrase translations in our suffix array-based data structure. We show how sampling can be used to reduce the retrieval time by orders of magnitude with no loss in translation quality. Chris Callison-Burch, Colin J. Bannard, Josh Schroeder |
ACL | 1 |
| 2005 | A compact data structure for searchable translation memories
Chris Callison-Burch, Colin J. Bannard, Josh Schroeder |
EAMT | 1 |
| 2004 | Statistical Machine Translation with Word- and Sentence-Aligned Parallel CorporaabstractThe parameters of statistical translation models are typically estimated from sentence-aligned parallel corpora. We show that significant improvements in the alignment and translation quality of such models can be achieved by additionally including word-aligned data during training. Incorporating word-level alignments into the parameter estimation of the IBM models reduces alignment error rate and increases the Bleu score when compared to training the same models only on sentence-aligned data. On the Verbmobil data set, we attain a 38% reduction in the alignment error rate and a higher Bleu score with half as many training examples. We discuss how varying the ratio of word-aligned to sentence-aligned data affects the expected performance gain. Chris Callison-Burch, David Talbot, Miles Osborne |
ACL | 1 |
| 2001 | A program for automatically selecting the best output from multiple machine translation enginesabstractThis paper describes a program that automatically selects the best translation from a set of translations produced by multiple commercial machine translation engines. The program is simplified by assuming that the most fluent item in the set is the best translation. Fluency is determined using a trigram language model. Results are provided illustrating how well the program performs for human ranked data as compared to each of its constituent engines. Chris Callison-Burch, Raymond S. Flournoy |
MTSummit | 1 |