EDBT 2026 Demo / reviewers in the wild / expert
Dan Roth 0001
dblp:r/DanRoth
· DBLP profile ↗
379ranked-venue papers
23as first author
93since 2021 · last 2026
0009-0002-1447-5173ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 342 · 18 first-author · 90 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 36 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 3 since 2021Theory of computation · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REaR : Retrieve, Expand and Refine for Effective Multitable RetrievalabstractRishita Agarwal, Himanshu Singhal, Peter Baile Chen, Manan Roy Choudhury, Dan Roth, Vivek Gupta. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rishita Agarwal, Himanshu Singhal, Peter Baile Chen, Manan Roy Choudhury, Dan Roth 0001, Vivek Gupta 0001 |
ACL (1) | 5 |
| 2026 | ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware FilteringabstractMarianne Menglin Liu, Daniel Garcia, Fjona Parllaku, Vikas Upadhyay, Fahad Shah, Dan Roth. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Marianne Menglin Liu, Fjona Parllaku, Vikas Upadhyay, Fahad Shah, Dan Roth 0001 |
ACL (1) | 6 |
| 2026 | Toward Robust Evaluation for Multilingual Grammatical Error Correction: Can Large Language Models Replace Human References?abstractA standard method for evaluating grammatical error correction systems severely underestimates performance, as it compares outputs against a small, fixed set of human references, despite the large space of possible valid corrections.Prior research has shown that using a closest-gold reference -i.e., a human reference generated with respect to the system output rather than the original text -yields more accurate performance estimates.Yet, producing such references for each system individually is costly.We introduce an automated method for generating closest-gold references by prompting a large language model (LLM) with system outputs.We find that performance scores computed using automatic closest-gold references correlate well with human closest-golds, whereas standard reference-based evaluations show weak or no correlation.Building on this insight, we use both fixed human references and closest-gold references generated by Claude and Llama to compare the performance of supervised models and GPT-4 across 14 benchmarks spanning 12 languages.Consequently, while prior work has shown that GPT-4 appears to lag behind traditional models, we demonstrate that this is due to the failures of the standard evaluation method that systematically underestimates GPT-4 performance more severely than that of supervised models.We show that a more appropriate evaluation approach, based on the closest gold method, reveals that GPT-4 outperforms traditional stateof-the-art models on almost all languages.1 Alla Rozovskaya, Dan Roth 0001 |
ACL (1) | 2 |
| 2026 | SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL BenchmarksabstractMohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooy, Graham Horwood, Dan Roth. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooyjadid, Graham Horwood, Dan Roth 0001 |
ACL (1) | 5 |
| 2026 | LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document UnderstandingabstractZhivar Sourati, Zheng Wang, Marianne Menglin Liu, Yazhe Hu, Mengqing Guo, Sujeeth Bharadwaj, Kyu J. Han, Tao Sheng, Sujith Ravi, Morteza Dehghani, Dan Roth. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhivar Sourati, Marianne Menglin Liu, Yazhe Hu, Mengqing Guo, Sujeeth Bharadwaj, Kyu J. Han, Sujith Ravi, Morteza Dehghani, Dan Roth 0001 |
ACL (1) | 11 |
| 2026 | When Vision-Language Models Judge Without Seeing: Exposing Informativeness BiasabstractThe reliability of VLM-as-a-Judge is critical for the automatic evaluation of vision-language models (VLMs).Despite recent progress, our analysis reveals that VLM-as-a-Judge often pays limited attention to the image when making decisions.Instead, they often blindly favor the more informative answer, even when they can recognize it conflicts with the image content.We call this problem informativeness bias, which significantly undermines judge reliability.To address it, we propose BIRCH (Balanced Informativeness and CoRrectness with a Truthful AnCHor), a judging paradigm that first corrects inconsistencies with the image content in candidate answers, and then compares the answers against this corrected version.This shifts the judge's focus from informativeness to image-grounded correctness.Experiments on multiple models and benchmarks show that BIRCH reduces informativeness bias by up to 17%, resulting in performance gains of up to 9.8%.Our work reveals an overlooked but fundamental flaw in current VLM-as-a-Judge systems and highlights the need for more principled designs.Preference: Answer A Answer A: The image does not provide explicit information about the name of the park.Answer B: The benches reside in Concord Pacific Place Park.The park's name is inscribed on the bench. Xiaohan Zou, Roshan Sridhar, Mohammadtaher Safarzadeh, Dan Roth 0001 |
ACL (1) | 4 |
| 2026 | MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of DocumentsabstractAbstract Automated agents, powered by large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve— far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks—with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts, and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco. Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth 0001, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval MethodabstractReal-world open-domain questions can be complex, especially when answering them requires integrating information from multiple sources.Effectively identifying the necessary information involves aligning it with the available data and its organization.However, existing RAG solutions address the alignment problem in a limited manner.Using off-the-shelf LLMs for question decomposition lacks awareness of the available data and its structure, often resulting in suboptimal retrieval performance.Alternatively, iteratively generating follow-up queries and interacting with the data collection, as explored in agentic RAG approaches, shows potential but is often inefficient since each successive query depends on previous results rather than being guided by the overall organization of the available data.To address the alignment problem, we introduce an LLM-based retrieval method -ARM, designed to better align questions with the organization of the data collection.Instead of solely matching query utterance, ARM explores relationships among data objects, enabling a retrieve-all-atonce solution for complex queries.Experimental results demonstrate that ARM significantly outperforms existing RAG methods on various complex open-domain QA tasks across multiple modalities, achieving superior retrieval performance and downstream accuracy while significantly lowering monetary costs. 1 Peter Baile Chen, Yi Zhang 0001, Michael J. Cafarella, Dan Roth 0001 |
ACL (1) | 4 |
| 2025 | DeAL: Decoding-time Alignment for Large Language ModelsabstractJames Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas, Saab Mansour, Katrin Kirchhoff, Dan Roth. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas 0004, Saab Mansour, Katrin Kirchhoff, Dan Roth 0001 |
ACL (1) | 9 |
| 2025 | Weaver: Interweaving SQL and LLM for Table ReasoningabstractQuerying tables with unstructured data is challenging due to the presence of text (or image), either embedded in the table or in external paragraphs, which traditional SQL struggles to process, especially for tasks requiring semantic reasoning.While Large Language Models (LLMs) excel at understanding context, they face limitations with long input sequences.Existing approaches that combine SQL and LLM typically rely on rigid, predefined workflows, limiting their adaptability to complex queries.To address these issues, we introduce Weaver , a modular pipeline that dynamically integrates SQL and LLM for table-based question answering (Table QA).Weaver generates a flexible, step-by-step plan that combines SQL for structured data retrieval with LLMs for semantic processing.By decomposing complex queries into manageable subtasks, Weaver improves accuracy and generalization.Our experiments show that Weaver consistently outperforms state-ofthe-art methods across four Table QA datasets, reducing both API calls and error rates. Rohit Khoja, Devanshu Gupta, Yanjie Fu, Dan Roth 0001, Vivek Gupta 0001 |
EMNLP | 4 |
| 2025 | AutoCT: Automating Interpretable Clinical Trial Prediction with LLM AgentsabstractClinical trials are critical for advancing medical treatments but remain prohibitively expensive and time-consuming.Accurate prediction of clinical trial outcomes can significantly reduce research and development costs and accelerate drug discovery.While recent deep learning models have shown promise by leveraging unstructured data, their black-box nature, lack of interpretability, and vulnerability to label leakage limit their practical use in high-stakes biomedical contexts.In this work, we propose AUTOCT 1 , a novel framework that combines the reasoning capabilities of large language models with the explainability of classical machine learning.AUTOCT autonomously generates, evaluates, and refines tabular features based on public information without human input.Our method uses Monte Carlo Tree Search to iteratively optimize predictive performance.Experimental results show that AUTOCT performs on par with or better than state-of-the-art methods on clinical trial prediction tasks within only a limited number of self-refinement iterations, establishing a new paradigm for scalable, interpretable, and cost-efficient clinical trial prediction. Fengze Liu, Haoyu Wang 0005, Joonhyuk Cho, Dan Roth 0001, Andrew Lo |
EMNLP | 4 |
| 2025 | LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense RetrievalabstractWhile significant progress has been made with dual-and bi-encoder dense retrievers, they often struggle on queries with logical connectives, a use case often overlooked yet important in downstream applications.In this paper, we introduce LOGICOL, a logically informed contrastive learning objective for dense retrievers.LOGICOL builds upon in-batch supervised contrastive learning and learns dense retrievers to respect the subset and mutually exclusive set relation between query results.We evaluated the effectiveness of LOGICOL in the entity retrieval task, where the model is expected to retrieve a set of Wikipedia entities that satisfy the implicit logical constraints of the query.We show that models trained with LOGICOL show improvements both in terms of retrieval performance and logical consistency in the results.We provide detailed analysis and insights to uncover why queries with logical connectives are challenging for dense retrievers and why LOGI-COL is effective.Our codes and data are available at https://github.com/yanzhen4/LogiCoL. A not B: Species of orchids inMalaysia but not Thailand. Yanzhen Shen, Xueqiang Xu, Yunyi Zhang 0001, Chaitanya Malaviya, Dan Roth 0001 |
EMNLP | 6 |
| 2025 | MrGuard: A Multilingual Reasoning Guardrail for Universal LLM SafetyabstractLarge Language Models (LLMs) are susceptible to adversarial attacks such as jailbreaking, which can elicit harmful or unsafe behaviors.This vulnerability is exacerbated in multilingual settings, where multilingual safetyaligned data is often limited.Thus, developing a guardrail capable of detecting and filtering unsafe content across diverse languages is critical for deploying LLMs in real-world applications.In this work, we introduce a multilingual guardrail with reasoning for prompt classification.Our method consists of: (1) synthetic multilingual data generation incorporating culturally and linguistically nuanced variants, (2) supervised fine-tuning, and (3) a curriculum-based Group Relative Policy Optimization (GRPO) framework that further improves performance.Experimental results demonstrate that our multilingual guardrail, Mr-Guard, consistently outperforms recent baselines across both in-domain and out-of-domain languages by more than 15%.We also evaluate MrGuard's robustness to multilingual variations, such as code-switching and low-resource language distractors in the prompt, and demonstrate that it preserves safety judgments under these challenging conditions.The multilingual reasoning capability of our guardrail enables it to generate explanations, which are particularly useful for understanding languagespecific risks and ambiguities in multilingual content moderation. Yahan Yang, Soham Dan, Dan Roth 0001, Insup Lee 0001 |
EMNLP | 4 |
| 2025 | BIRD: A Trustworthy Bayesian Inference Framework for Large Language ModelsabstractPredictive models often need to work with incomplete information in real-world tasks. Consequently, they must provide reliable probability or confidence estimation, especially in large-scale decision-making and planning tasks. Current large language models (LLMs) are insufficient for accurate estimations, but they can generate relevant factors that may affect the probabilities, produce coarse-grained probabilities when the information is more complete, and help determine which factors are relevant to specific downstream contexts. In this paper, we make use of these capabilities of LLMs to provide a significantly more accurate probabilistic estimation. We propose BIRD, a novel probabilistic inference framework that aligns a Bayesian network with LLM abductions and then estimates more accurate probabilities in a deduction step. We show BIRD provides reliable probability estimations that are 30% better than those provided directly by LLM baselines. These estimates further contribute to better and more trustworthy decision making. Yu Feng 0013, Ben Zhou, Dan Roth 0001 |
ICLR | 4 |
| 2025 | Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judgeabstractThe effectiveness of automatic evaluation of generative models is typically measured by comparing the labels generated via automation with human labels using correlation metrics.
However, metrics like Krippendorff's $\alpha$ and Randolph's $\kappa$ were originally designed to measure the reliability of human labeling, thus make assumptions about typical human labeling behavior, and these assumptions may not be applicable to machine generated labels.
In this paper, we show how *relying on a single aggregate correlation score* can obscure fundamental differences between human labels and those from automatic evaluation, including LLM-as-a-Judge.
Specifically, we demonstrate that when the proportion of samples with variation or uncertainty in human assigned labels is relatively high, machine labels (generated by automatic evaluation methods) may superficially appear to have similar or better correlation with the human majority label compared to the human-to-human (HH) correlation.
This can create the illusion that labels from automatic evaluation approximates the human majority label.
However, as the proportion of samples with consistent human labels increases, the correlation between machine and human labels fall well below HH correlation.
Based on these findings, we first propose *stratifying data by human label uncertainty* to provide a more robust analysis of automatic evaluation performance. Second, recognizing that uncertainty and variation are inherent in perception-based human evaluations, such as those involving attitudes or preferences, we introduce a new metric -*binned Jensen-Shannon Divergence for perception* for such scenarios to better measure the effectiveness of automatic evaluations. Third, we present visualization techniques -- *perception charts*, to contextualize correlation measures appropriately and to show the strengths and limitations of automatic evaluation. We have open-sourced our analysis and visualization tools at https://github.com/amazon-science/BeyondCorrelation. Aparna Elangovan, Lei Xu 0040, Jongwoo Ko, Mahsa Elyasi, Sravan Babu Bodapati, Dan Roth 0001 |
ICLR | 7 |
| 2025 | MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingabstractWe introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal relations). Comprising 11,264 images and 2,600 multiple-choice questions, MuirBench is created in a pairwise manner, where each standard instance is paired with an unanswerable variant that has minimal semantic differences, in order for a reliable assessment. Evaluated upon 20 recent multi-modal LLMs, our results reveal that even the best-performing models like GPT-4o and Gemini Pro find it challenging to solve MuirBench, achieving 68.0% and 49.3% in accuracy. Open-source multimodal LLMs trained on single images can hardly generalize to multi-image questions, hovering below 33.3% in accuracy. These results highlight the importance of MuirBench in encouraging the community to develop multimodal LLMs that can look beyond a single image, suggesting potential pathways for future improvements. Fei Wang 0060, James Y. Huang, Zekun Li 0007, Qin Liu 0010, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu 0014, Wenxuan Zhou 0002, Kai Zhang 0008, Tianyi Lorena Yan, Wenjie Mo 0001, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang 0001, Dan Roth 0001, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
ICLR | 18 |
| 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image UnderstandingabstractStructured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. However, current multimodal large language models (LLMs) lack this multihop selective attention capability. In this work, we introduce ReFocus, a simple yet effective framework that equips multimodal LLMs with the ability to generate ``visual thoughts'' by performing visual editing on the input image through code, shifting and refining their visual focuses. Specifically, ReFocus enables multimodal LLMs to generate Python codes to call tools and modify the input image, sequentially drawing boxes, highlighting sections, and masking out areas, thereby enhancing the visual reasoning process. We experiment upon a wide range of structured image understanding tasks involving tables and charts. ReFocus largely improves performance on all tasks over GPT-4o without visual editing, yielding an average gain of 11.0% on table tasks and 6.8% on chart tasks. We present an in-depth analysis of the effects of different visual edits, and reasons why ReFocus can improve the performance without introducing additional information. Further, we collect a 14k training set using ReFocus, and prove that such visual chain-of-thought with intermediate information offers a better supervision than standard VQA data, reaching a 8.0% average gain over the same model trained with QA pairs and 2.6% over CoT. Minqian Liu, Zhengyuan Yang, John Corring, Yijuan Lu, Dan Roth 0001, Dinei A. F. Florêncio, Cha Zhang |
ICML | 7 |
| 2025 | GIVE: Structured Reasoning of Large Language Models with Knowledge Graph Inspired Veracity ExtrapolationabstractExisting approaches based on context prompting or reinforcement learning (RL) to improve the reasoning capacities of large language models (LLMs) depend on the LLMs’ internal knowledge to produce reliable Chain-Of-Thought (CoT). However, no matter the size of LLMs, certain problems cannot be resolved in a single forward pass. Meanwhile, agent-based reasoning systems require access to a comprehensive nonparametric knowledge base, which is often costly or not feasible for use in scientific and niche domains. We present Graph Inspired Veracity Extrapolation (GIVE), a novel reasoning method that merges parametric and non-parametric memories to improve accurate reasoning with minimal external input. GIVE guides the LLM agent to select the most pertinent expert data ($\textbf{observe}$), engage in query-specific associative thinking ($\textbf{reflect}$), and then synthesize this information to produce the final output ($\textbf{speak}$). Extensive experiments demonstrated the following benefits of our framework: (1) GIVE increases the performance of LLMs across various sizes. (2) In some scenarios, GIVE allows smaller LLMs to surpass larger, more sophisticated ones in scientific tasks ($\textbf{GPT3.5T + GIVE > GPT4}$). (3) GIVE is effective on scientific and open-domain assessments. (4) GIVE is a training-free method that enables LLMs to tackle new problems that extend beyond their training data (up to $\textbf{43.5}$% $\rightarrow$ $\textbf{88.2}$% accuracy improvement). (5) GIVE allows LLM agents to reason using both restricted (very small) and noisy (very large) knowledge sources, accommodating knowledge graphs (KG) ranging from $\textbf{135}$ to more than $\textbf{840k}$ nodes. (6) The reasoning process involved in GIVE is fully interpretable. Our code is available at https://github.com/Jason-Tree/GIVE Jiashu He, Mingyu Derek Ma, Jinxuan Fan, Dan Roth 0001, Wei Wang 0010, Alejandro Ribeiro |
ICML | 4 |
| 2025 | On Reasoning LLMs: Myths, Merits, and How to Move ForwardabstractThe rapid progress made over the last few years in generating linguistically coherent natural language has blurred, in the mind of many, the difference between natural language generation, understanding, knowledge retrieval and use, and the ability to reason with respect to the world. Nevertheless, reliably and consistently supporting high-level decisions that depend on natural language understanding and heterogenous information retrieval is still difficult, mostly, but not only, since most of these tasks are computationally more complex than language models can support. I will discuss some of the challenges underlying reasoning and information access and argue that we should exploit what LLMs do well while delegating responsibility to special purpose models and solvers for decision making. I will present some of our work in this space, focusing on supporting reasoning and information access in a range of quantitative, visual, and spatial reasoning tasks. Dan Roth 0001 |
KDD (2) | 1 |
| 2025 | H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on TablesabstractNikhil Abhyankar, Vivek Gupta, Dan Roth, Chandan K. Reddy. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nikhil Abhyankar, Vivek Gupta 0001, Dan Roth 0001, Chandan K. Reddy |
NAACL (Long Papers) | 3 |
| 2025 | Leveraging LLM For Synchronizing Information Across Multilingual TablesabstractSiddharth Khincha, Tushar Kataria, Ankita Anand, Dan Roth, Vivek Gupta. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Siddharth Khincha, Tushar Kataria, Dan Roth 0001, Vivek Gupta 0001 |
NAACL (Long Papers) | 4 |
| 2025 | MAPWise: Evaluating Vision-Language Models for Advanced Map QueriesabstractSrija Mukhopadhyay, Abhishek Rajgaria, Prerana Khatiwada, Manish Shrivastava, Dan Roth, Vivek Gupta. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Srija Mukhopadhyay, Abhishek Rajgaria, Prerana Khatiwada, Manish Shrivastava 0001, Dan Roth 0001, Vivek Gupta 0001 |
NAACL (Long Papers) | 5 |
| 2025 | TRANSIENTTABLES: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured TablesabstractAbhilash Shankarampeta, Harsh Mahajan, Tushar Kataria, Dan Roth, Vivek Gupta. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Abhilash Reddy Shankarampeta, Harsh Mahajan, Tushar Kataria, Dan Roth 0001, Vivek Gupta 0001 |
NAACL (Long Papers) | 4 |
| 2025 | Imbalances in Neurosymbolic Learning: Characterization and Mitigating StrategiesabstractWe study one of the most popular problems in **neurosymbolic learning** (NSL), that of learning neural classifiers given only the result of applying a symbolic component $\sigma$ to the gold labels of the elements of a vector $\mathbf x$. The gold labels of the elements in $\mathbf x$ are unknown to the learner. We make multiple contributions, theoretical and practical, to address a problem that has not been studied so far in this context, that of characterizing and mitigating *learning imbalances*, i.e., major differences in the errors that occur when classifying instances of different classes (aka **class-specific risks**). Our theoretical reveals a unique phenomenon: that $\sigma$ can greatly impact learning imbalances. This result sharply contrasts with previous research on supervised and weakly supervised learning, which only studies learning imbalances under data imbalances. On the practical side, we introduce a technique for estimating the marginal of the hidden gold labels using weakly supervised data. Then, we introduce algorithms that mitigate imbalances at training and testing time by treating the marginal of the hidden labels as a constraint. We demonstrate the effectiveness of our techniques using strong baselines from NSL and long-tailed learning, suggesting performance improvements of up to 14\%. Efthymia Tsamoura, Kaifu Wang, Dan Roth 0001 |
NeurIPS | 3 |
| 2025 | Contextualized Evaluations: Judging Language Model Responses to Underspecified QueriesabstractAbstract Language model users often issue queries that lack specification, where the context under which a query was issued—such as the user’s identity, the query’s intent, and the criteria for a response to be useful—is not explicit. For instance, a good response to a subjective query like “What book should I read next?” would depend on the user’s preferences, and a good response to an open-ended query like “How do antibiotics work against bacteria?” would depend on the user’s expertise. This makes evaluation of responses to such queries an ill-posed task, as evaluators may make arbitrary judgments about the response quality. To remedy this, we present contextualized evaluations, a protocol that synthetically constructs context surrounding an underspecified query and provides it during evaluation. We find that the presence of context can 1) alter conclusions drawn from evaluation, even flipping benchmark rankings between model pairs, 2) nudge evaluators to make fewer judgments based on surface-level criteria, like style, and 3) provide new insights about model behavior across diverse contexts. Specifically, our procedure suggests a potential bias towards WEIRD (Western, Educated, Industrialized, Rich and Democratic) contexts in models’ “default” responses and we find that models are not equally sensitive to following different contexts, even when they are provided in prompts.1 Chaitanya Malaviya, Joseph Chee Chang, Dan Roth 0001, Mohit Iyyer, Mark Yatskar, Kyle Lo |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table RetrievalabstractRetrieving relevant tables containing the necessary information to accurately answer a given question over tables is critical to open-domain question-answering (QA) systems.Previous methods assume the answer to such a question can be found either in a single table or multiple tables identified through question decomposition or rewriting.However, neither of these approaches is sufficient, as many questions require retrieving multiple tables and joining them through a join plan that cannot be discerned from the user query itself.If the join plan is not considered in the retrieval stage, the subsequent steps of reasoning and answering based on those retrieved tables are likely to be incorrect.To address this problem, we introduce a method that uncovers useful join relations for any query and database during table retrieval.We use a novel re-ranking method formulated as a mixed-integer program that considers not only table-query relevance but also table-table relevance that requires inferring join relationships.Our method outperforms the state-of-the-art approaches for table retrieval by up to 9.3% in F1 score and for end-to-end QA by up to 5.4% in accuracy. Peter Baile Chen, Yi Zhang 0001, Dan Roth 0001 |
ACL (1) | 3 |
| 2024 | ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsabstractIn this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon insights from disciplines such as user experience research and human behavioral psychology to ensure that the experimental design and results are reliable.The conclusions from these evaluations, thus, must consider factors such as usability, aesthetics, and cognitive biases.We highlight how cognitive biases can conflate fluent information and truthfulness, and how cognitive uncertainty affects the reliability of rating scores such as Likert.Furthermore, the evaluation should differentiate the capabilities and weaknesses of increasingly powerful large language models -which requires effective test sets.The scalability of human evaluation is also crucial to wider adoption.Hence, to design an effective human evaluation system in the age of generative NLP, we propose the ConSiDERS-The-Human evaluation framework consisting of 6 pillars -Consistency, Scoring Critera, Differentiating, User Experience, Responsible, and Scalability. Aparna Elangovan, Lei Xu 0040, Sravan Babu Bodapati, Dan Roth 0001 |
ACL (1) | 5 |
| 2024 | CoCoMIC: Code Completion by Jointly Modeling In-file and Cross-file ContextabstractWhile pre-trained language models (LM) for code have achieved great success in code completion, they generate code conditioned only on the contents within the file, i.e., in-file context, but ignore the rich semantics in other files within the same project, i.e., project-level cross-file context, a critical source of information that is especially useful in modern modular software development. Such overlooking constrains code LMs’ capacity in code completion, leading to unexpected behaviors such as generating hallucinated class member functions or function calls with unexpected arguments. In this work, we propose CoCoMIC, a novel framework that jointly learns the in-file and cross-file context on top of code LMs. To empower CoCoMIC, we develop CCFinder, a static-analysis-based tool that locates and retrieves the most relevant project-level cross-file context for code completion. CoCoMIC successfully improves the existing code LM with a 33.94% relative increase in exact match and 28.69% in identifier matching for code completion when the cross-file context is provided. Finally, we perform a series of ablation studies and share valuable insights for future research on integrating cross-file context into code LMs. Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang |
LREC/COLING | 7 |
| 2024 | BLINK: Multimodal Large Language Models Can See but Not Perceive
Yushi Hu, Bangzheng Li, Yu Feng 0013, Haoyu Wang 0005, Xudong Lin 0003, Dan Roth 0001, Noah A. Smith, Wei-Chiu Ma, Ranjay Krishna |
ECCV (23) | 7 |
| 2024 | Evaluating Concurrent Robustness of Language Models Across Diverse Challenge SetsabstractLanguage models, characterized by their blackbox nature, often hallucinate and display sensitivity to input perturbations, causing concerns about trust.To enhance trust, it is imperative to gain a comprehensive understanding of the model's failure modes and develop effective strategies to improve their performance.In this study, we introduce a methodology designed to examine how input perturbations affect language models across various scales, including pre-trained models and large language models (LLMs).Utilizing fine-tuning, we enhance the model's robustness to input perturbations.Additionally, we investigate whether exposure to one perturbation enhances or diminishes the model's performance with respect to other perturbations.To address robustness against multiple perturbations, we present three distinct fine-tuning strategies.Furthermore, we broaden the scope of our methodology to encompass large language models (LLMs) by leveraging a chain of thought (CoT) prompting approach augmented with exemplars.We employ the Tabular-NLI task to showcase how our proposed strategies adeptly train a robust model, enabling it to address diverse perturbations while maintaining accuracy on the original dataset. Pranshu Pandya, Tushar Kataria, Vivek Gupta 0001, Dan Roth 0001 |
EMNLP | 5 |
| 2024 | Event Causality Identification with Synthetic ControlabstractEvent causality identification (ECI), a process that extracts causal relations between events from text, is crucial for distinguishing causation from correlation.Traditional approaches to ECI have primarily utilized linguistic patterns and multi-hop relational inference, risking false causality identification due to informal usage of causality and specious graphical inference.In this paper, we adopt the Rubin Causal Model to identify event causality: given two temporally ordered events, we see the first event as the treatment and the second one as the observed outcome.Determining their causality involves manipulating the treatment and estimating the resultant change in the likelihood of the outcome.Given that it is only possible to implement manipulation conceptually in the text domain, as a work-around, we try to find a 'twin' for the protagonist from existing corpora.This 'twin' should have identical life experiences with the protagonist before the treatment but undergoes an intervention of treatment.However, the practical difficulty of locating such a match limits its feasibility.Addressing this issue, we use the synthetic control method to generate such a 'twin' from relevant historical data, leveraging text embedding synthesis and inversion techniques.This approach allows us to identify causal relations more robustly than previous methods, including GPT-4, which is demonstrated on a causality benchmark, COPES-hard. Haoyu Wang 0005, Fengze Liu, Jiayao Zhang 0001, Dan Roth 0001, Kyle Richardson 0001 |
EMNLP | 4 |
| 2024 | A Peek into Token Bias: Large Language Models Are Not Yet Genuine ReasonersabstractBowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie J Su, Camillo Jose Taylor, Dan Roth. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang 0002, Tanwi Mallick, Weijie J. Su, Camillo J. Taylor, Dan Roth 0001 |
EMNLP | 8 |
| 2024 | Code Representation Learning at ScaleabstractRecent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size. Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang |
ICLR | 6 |
| 2024 | Fewer Truncations Improve Language ModelingabstractIn large language model training, input documents are typically concatenated together and then split into sequences of equal length to avoid padding tokens. Despite its efficiency, the concatenation approach compromises data integrity—it inevitably breaks many documents into incomplete pieces, leading to excessive truncations that hinder the model from learning to compose logically coherent and factually consistent content that is grounded on the complete context. To address the issue, we propose Best-fit Packing, a scalable and efficient method that packs documents into training sequences through length-aware combinatorial optimization. Our method completely eliminates unnecessary truncations while retaining the same training efficiency as concatenation. Empirical results from both text and code pre-training show that our method achieves superior performance (e.g., +4.7% on reading comprehension; +16.8% in context following; and +9.2% on program synthesis), and reduces closed-domain hallucination effectively by up to 58.3%. Hantian Ding, Zijian Wang 0002, Giovanni Paolini, Anoop Deoras, Dan Roth 0001, Stefano Soatto |
ICML | 6 |
| 2024 | Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic RepresentationsabstractSihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hongming Zhang 0009, Ben Zhou, Wenhao Yu 0002, Dian Yu 0001, Baolin Peng, Hongwei Wang 0010, Dan Roth 0001, Dong Yu 0001 |
NAACL-HLT | 9 |
| 2024 | Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?abstractBangzheng Li, Ben Zhou, Fei Wang, Xingyu Fu, Dan Roth, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Bangzheng Li, Ben Zhou, Fei Wang 0060, Dan Roth 0001, Muhao Chen 0001 |
NAACL-HLT | 5 |
| 2024 | ExpertQA: Expert-Curated Questions and Attributed AnswersabstractChaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, Dan Roth. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chaitanya Malaviya, Elizabeth Sieber, Mark Yatskar, Dan Roth 0001 |
NAACL-HLT | 6 |
| 2024 | What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User PerceptionabstractChaitanya Malaviya, Subin Lee, Dan Roth, Mark Yatskar. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chaitanya Malaviya, Dan Roth 0001, Mark Yatskar |
NAACL-HLT | 3 |
| 2024 | Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsabstractHumans draw to facilitate reasoning: we draw auxiliary lines when solving geometry problems; we mark and circle when reasoning on maps; we use sketches to amplify our ideas and relieve our limited-capacity working memory. However, such actions are missing in current multimodal language models (LMs). Current chain-of-thought and tool-use paradigms only use text as intermediate reasoning steps. In this work, we introduce Sketchpad, a framework that gives multimodal LMs a visual sketchpad and tools to draw on the sketchpad. The LM conducts planning and reasoning according to the visual artifacts it has drawn. Different from prior work, which uses text-to-image models to enable LMs to draw, Sketchpad enables LMs to draw with lines, boxes, marks, etc., which is closer to human sketching and better facilitates reasoning. \name can also use specialist vision models during the sketching process (e.g., draw bounding boxes with object detection models, draw masks with segmentation models), to further enhance visual perception and reasoning. We experiment on a wide range of math tasks (including geometry, functions, graph, chess) and complex visual reasoning tasks. Sketchpad substantially improves performance on all tasks over strong base models with no sketching, yielding an average gain of 12.7% on math tasks, and 8.6% on vision tasks. GPT-4o with Sketchpad sets a new state of the art on all tasks, including V*Bench (80.3%), BLINK spatial reasoning (83.9%), and visual correspondence (80.8%). We will release all code and data. Yushi Hu, Dan Roth 0001, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Ranjay Krishna |
NeurIPS | 4 |
| 2024 | Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at ScaleabstractLLMs can now act as autonomous agents that interact with digital environments and complete specific objectives (e.g., arranging an online meeting). However, accuracy is still far from satisfactory, partly due to a lack of large-scale, direct demonstrations for digital tasks. Obtaining supervised data from humans is costly, and automatic data collection through exploration or reinforcement learning relies on complex environmental and content setup, resulting in datasets that lack comprehensive coverage of various scenarios. On the other hand, there is abundant knowledge that may indirectly assist task completion, such as online tutorials that were created for human consumption. In this work, we present Synatra, an approach that effectively transforms this indirect knowledge into direct supervision at scale. We define different types of indirect knowledge, and carefully study the available sources to obtain it, methods to encode the structure of direct demonstrations, and finally methods to transform indirect knowledge into direct demonstrations. We use 100k such synthetically-created demonstrations to finetune a 7B CodeLlama, and demonstrate that the resulting agent surpasses all comparably sized models on three web-based task benchmarks Mind2Web, MiniWoB++ and WebArena, as well as surpassing GPT-3.5 on WebArena and Mind2Web. In addition, while synthetic demonstrations prove to be only 3% the cost of human demonstrations (at $0.031 each), we show that the synthetic demonstrations can be more effective than an identical number of human demonstrations collected from limited domains. Tianyue Ou, Frank F. Xu, Aman Madaan, Jiarui Liu 0004, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth 0001, Graham Neubig, Shuyan Zhou |
NeurIPS | 8 |
| 2024 | Disparities in seizure outcomes revealed by large language modelsabstractOBJECTIVE: Large-language models (LLMs) can potentially revolutionize health care delivery and research, but risk propagating existing biases or introducing new ones. In epilepsy, social determinants of health are associated with disparities in care access, but their impact on seizure outcomes among those with access remains unclear. Here we (1) evaluated our validated, epilepsy-specific LLM for intrinsic bias, and (2) used LLM-extracted seizure outcomes to determine if different demographic groups have different seizure outcomes. MATERIALS AND METHODS: We tested our LLM for differences and equivalences in prediction accuracy and confidence across demographic groups defined by race, ethnicity, sex, income, and health insurance, using manually annotated notes. Next, we used LLM-classified seizure freedom at each office visit to test for demographic outcome disparities, using univariable and multivariable analyses. RESULTS: We analyzed 84 675 clinic visits from 25 612 unique patients seen at our epilepsy center. We found little evidence of bias in the prediction accuracy or confidence of outcome classifications across demographic groups. Multivariable analysis indicated worse seizure outcomes for female patients (OR 1.33, P ≤ .001), those with public insurance (OR 1.53, P ≤ .001), and those from lower-income zip codes (OR ≥1.22, P ≤ .007). Black patients had worse outcomes than White patients in univariable but not multivariable analysis (OR 1.03, P = .66). CONCLUSION: We found little evidence that our LLM was intrinsically biased against any demographic group. Seizure freedom extracted by LLM revealed disparities in seizure outcomes across several demographic groups. These findings quantify the critical need to reduce disparities in the care of people with epilepsy. Kevin Xie, William K. S. Ojemann, Ryan S. Gallagher, Russell T. Shinohara, Alfredo Lucas, Chloe E. Hill, Roy H. Hamilton, Kevin B. Johnson, Dan Roth 0001, Brian Litt, Colin A. Ellis |
J. Am. Medical Informatics Assoc. | 9 |
| 2023 | GLUECons: A Generic Benchmark for Learning under ConstraintsabstractRecent research has shown that integrating domain knowledge into deep learning architectures is effective; It helps reduce the amount of required data, improves the accuracy of the models' decisions, and improves the interpretability of models. However, the research community lacks a convened benchmark for systematically evaluating knowledge integration methods. In this work, we create a benchmark that is a collection of nine tasks in the domains of natural language processing and computer vision. In all cases, we model external knowledge as constraints, specify the sources of the constraints for each task, and implement various models that use these constraints. We report the results of these models using a new set of extended evaluation criteria in addition to the task performances for a more in-depth analysis. This effort provides a framework for a more comprehensive and systematic comparison of constraint integration techniques and for identifying related research challenges. It will facilitate further research for alleviating some problems of state-of-the-art neural models. Hossein Rajaby Faghihi, Aliakbar Nafar, Chen Zheng 0006, Roshanak Mirzaee, Yue Zhang 0004, Andrzej Uszok, Alexander Wan, Tanawan Premsri, Dan Roth 0001, Parisa Kordjamshidi |
AAAI | 9 |
| 2023 | ReCode: Robustness Evaluation of Code Generation ModelsabstractShiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shiqi Wang 0002, Haifeng Qian, Chenghao Yang 0001, Zijian Wang 0002, Mingyue Shang, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth 0001, Bing Xiang |
ACL (1) | 13 |
| 2023 | Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion ScaleabstractHritik Bansal, Karthik Gopalakrishnan, Saket Dingliwal, Sravan Bodapati, Katrin Kirchhoff, Dan Roth. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hritik Bansal, Karthik Gopalakrishnan 0001, Saket Dingliwal, Sravan Babu Bodapati, Katrin Kirchhoff, Dan Roth 0001 |
ACL (1) | 6 |
| 2023 | Characterizing and Measuring Linguistic Dataset DriftabstractTyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas, Yassine Benajiba, Miguel Ballesteros, Dan Roth. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas, Yassine Benajiba, Miguel Ballesteros, Dan Roth 0001 |
ACL (1) | 7 |
| 2023 | Generic Temporal Reasoning with Differential Analysis and ExplanationabstractTemporal reasoning is the task of predicting temporal relations of event pairs.While temporal reasoning models can perform reasonably well on in-domain benchmarks, we have little idea of these systems' generalizability due to existing datasets' limitations.In this work, we introduce a novel task named TODAY that bridges this gap with temporal differential analysis, which as the name suggests, evaluates whether systems can correctly understand the effect of incremental changes.Specifically, TODAY introduces slight contextual changes for given event pairs, and systems are asked to tell how this subtle contextual change would affect relevant temporal relation distributions.To facilitate learning, TODAY also annotates human explanations.We show that existing models, including GPT-3.5, drop to random guessing on TODAY, suggesting that they heavily rely on spurious information rather than proper reasoning for temporal predictions.On the other hand, we show that TODAY's supervision style and explanation annotations can be used in joint learning, encouraging models to use more appropriate signals during training and thus outperform across several benchmarks.TODAY can also be used to train models to solicit incidental supervision from noisy sources such as GPT-3.5, thus moving us more toward the goal of generic temporal reasoning systems. Yu Feng 0013, Ben Zhou, Haoyu Wang 0005, Helen Jin, Dan Roth 0001 |
ACL (1) | 5 |
| 2023 | Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source LearningabstractAlexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 0005, Patrick Ng, Zhiguo Wang 0006, Bonan Min, William Yang Wang, Kathy McKeown, Vittorio Castelli, Dan Roth 0001, Bing Xiang |
ACL (1) | 11 |
| 2023 | Incorporating Question Answering-Based Signals into Abstractive Summarization via Salient Span SelectionabstractIn this work, we propose a method for incorporating question-answering (QA) signals into a summarization model.Our method identifies salient noun phrases (NPs) in the input document by automatically generating wh-questions that are answered by the NPs and automatically determining whether those questions are answered in the gold summaries.This QA-based signal is incorporated into a two-stage summarization model which first marks salient NPs in the input document using a classification model, then conditionally generates a summary.Our experiments demonstrate that the models trained using QA-based supervision generate higher-quality summaries than baseline methods of identifying salient spans on benchmark summarization datasets.Further, we show that the content of the generated summaries can be controlled based on which NPs are marked in the input document.Finally, we propose a method of augmenting the training data so the gold summaries are more consistent with the marked input spans used during training and show how this results in models which learn to better exclude unmarked document content. 1 Daniel Deutsch, Dan Roth 0001 |
EACL | 2 |
| 2023 | Extracting or Guessing? Improving Faithfulness of Event Temporal Relation ExtractionabstractIn this paper, we seek to improve the faithfulness of TEMPREL extraction models from two perspectives.The first perspective is to extract genuinely based on contextual description.To achieve this, we propose to conduct counterfactual analysis to attenuate the effects of two significant types of training biases: the event trigger bias and the frequent label bias.We also add tense information into event representations to explicitly place an emphasis on the contextual description.The second perspective is to provide proper uncertainty estimation and abstain from extraction when no relation is described in the text.By parameterization of Dirichlet Prior over the model-predicted categorical distribution, we improve the model estimates of the correctness likelihood and make TEMPREL predictions more selective.We also employ temperature scaling to recalibrate the model confidence measure after bias mitigation.Through experimental analysis on MATRES, MATRES-DS, and TDDiscourse, we demonstrate that our model extracts TEMPREL and timelines more faithfully compared to SOTA methods, especially under distribution shifts. Haoyu Wang 0005, Hongming Zhang 0009, Yuqian Deng, Jacob R. Gardner, Dan Roth 0001, Muhao Chen 0001 |
EACL | 5 |
| 2023 | Event Linking: Grounding Event Mentions to WikipediaabstractComprehending an article requires understanding its constituent events.However, the context where an event is mentioned often lacks the details of this event.A question arises: how can the reader obtain more knowledge about this particular event in addition to what is provided by the local context in the article?This work defines Event Linking, a new natural language understanding task at the event level.Event linking tries to link an event mention appearing in an article to the most appropriate Wikipedia page.This page is expected to provide rich knowledge about what the event mention refers to.To standardize the research in this new direction, we contribute in fourfold.First, this is the first work in the community that formally defines the Event Linking task.Second, we collect a dataset for this new task.Specifically, we automatically gather the training set from Wikipedia, and then create two evaluation sets: one from the Wikipedia domain, reporting the in-domain performance, and a second from the real-world news domain, to evaluate out-of-domain performance.Third, we retrain and evaluate two state-of-theart (SOTA) entity linking models, showing the challenges of event linking, and we propose an event-specific linking system, EVELINK, to set a competitive result for the new task.Fourth, we conduct a detailed and insightful analysis to help understand the task and the limitations of the current model.Overall, as our analysis shows, Event Linking is a challenging and essential task requiring more effort from the community.1 Xiaodong Yu 0003, Wenpeng Yin 0001, Nitish Gupta, Dan Roth 0001 |
EACL | 4 |
| 2023 | Are All Steps Equally Important? Benchmarking Essentiality Detection in Event ProcessesabstractNatural language expresses events with varying granularities, where coarse-grained events (goals) can be broken down into finer-grained event sequences (steps).A critical yet overlooked aspect of understanding event processes is recognizing that not all step events hold equal importance toward the completion of a goal.In this paper, we address this gap by examining the extent to which current models comprehend the essentiality of step events in relation to a goal event.Cognitive studies suggest that such capability enables machines to emulate human commonsense reasoning about preconditions and necessary efforts of everyday tasks.We contribute a high-quality corpus of (goal, step) pairs gathered from the community guideline website WikiHow, with steps manually annotated for their essentiality concerning the goal by experts.The high inter-annotator agreement demonstrates that humans possess a consistent understanding of event essentiality.However, after evaluating multiple statistical and largescale pre-trained language models, we find that existing approaches considerably underperform compared to humans.This observation highlights the need for further exploration into this critical and challenging task 1 . Haoyu Wang 0005, Hongming Zhang 0009, Yueguan Wang, Yuqian Deng, Muhao Chen 0001, Dan Roth 0001 |
EMNLP | 6 |
| 2023 | Taxonomy Expansion for Named Entity RecognitionabstractKarthikeyan K, Yogarshi Vyas, Jie Ma, Giovanni Paolini, Neha John, Shuai Wang, Yassine Benajiba, Vittorio Castelli, Dan Roth, Miguel Ballesteros. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Karthikeyan K, Yogarshi Vyas, Jie Ma 0005, Giovanni Paolini, Neha Anna John, Yassine Benajiba, Vittorio Castelli, Dan Roth 0001, Miguel Ballesteros |
EMNLP | 9 |
| 2023 | Comparing Biases and the Impact of Multilingual Training across Multiple LanguagesabstractSharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sharon Levy, Neha Anna John, Yogarshi Vyas, Jie Ma 0005, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth 0001 |
EMNLP | 9 |
| 2023 | Bootstrapping Small & High Performance Language Models with Unmasking-Removal Training PolicyabstractBabyBERTa, a language model trained on small-scale child-directed speech while none of the words are unmasked during training, has been shown to achieve a level of grammaticality comparable to that of RoBERTa-base, which is trained on 6,000 times more words and 15 times more parameters (Huebner et al., 2021).Relying on this promising result, we explore in this paper the performance of BabyBERTabased models in downstream tasks, focusing on Semantic Role Labeling (SRL) and two Extractive Question Answering tasks, with the aim of building more efficient systems that rely on less data and smaller models.We investigate the influence of these models both alone and as a starting point to larger pre-trained models, separately examining the contribution of the pre-training data, the vocabulary, and the masking policy on the downstream task performance.Our results show that BabyBERTa trained with unmasking-removal policy is a much stronger starting point for downstream tasks compared to the use of RoBERTa masking policy when 10M words are used for training and that this tendency persists, although to a lesser extent, when adding more training data. 1 Yahan Yang, Elior Sulem, Insup Lee 0001, Dan Roth 0001 |
EMNLP | 4 |
| 2023 | STREET: A Multi-Task Structured Reasoning and Explanation Benchmark
Danilo Neves Ribeiro, Shen Wang 0005, Xiaofei Ma 0001, Henghui Zhu, Deguang Kong, Juliette Burger, Anjelica Ramos, Zhiheng Huang, William Yang Wang, George Karypis, Bing Xiang, Dan Roth 0001 |
ICLR | 13 |
| 2023 | On Regularization and Inference with Label ConstraintsabstractPrior knowledge and symbolic rules in machine learning are often expressed in the form of label constraints, especially in structured prediction problems. In this work, we compare two common strategies for encoding label constraints in a machine learning pipeline, regularization with constraints and constrained inference, by quantifying their impact on model performance. For regularization, we show that it narrows the generalization gap by precluding models that are inconsistent with the constraints. However, its preference for small violations introduces a bias toward a suboptimal model. For constrained inference, we show that it reduces the population risk by correcting a model’s violation, and hence turns the violation into an advantage. Given these differences, we further explore the use of two approaches together and propose conditions for constrained inference to compensate for the bias introduced by regularization, aiming to improve both the model complexity and optimal risk. Kaifu Wang, Hangfeng He 0001, Tin D. Nguyen, Dan Roth 0001 |
ICML | 5 |
| 2023 | Conversation Style Transfer using Few-Shot LearningabstractShamik Roy, Raphael Shu, Nikolaos Pappas, Elman Mansimov, Yi Zhang, Saab Mansour, Dan Roth. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shamik Roy, Raphael Shu, Nikolaos Pappas 0004, Elman Mansimov, Yi Zhang 0001, Saab Mansour, Dan Roth 0001 |
IJCNLP (1) | 7 |
| 2023 | CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code CompletionabstractCode completion models have made significant progress in recent years, yet current popular evaluation datasets, such as HumanEval and MBPP, predominantly focus on code completion tasks within a single file. This over-simplified setting falls short of representing the real-world software development scenario where repositories span multiple files with numerous cross-file dependencies, and accessing and understanding cross-file context is often required to complete the code correctly. To fill in this gap, we propose CrossCodeEval, a diverse and multilingual code completion benchmark that necessitates an in-depth cross-file contextual understanding to complete the code accurately. CrossCodeEval is built on a diverse set of real-world, open-sourced, permissively-licensed repositories in four popular programming languages: Python, Java, TypeScript, and C#. To create examples that strictly require cross-file context for accurate completion, we propose a straightforward yet efficient static-analysis-based approach to pinpoint the use of cross-file context within the current file. Extensive experiments on state-of-the-art code language models like CodeGen and StarCoder demonstrate that CrossCodeEval is extremely challenging when the relevant cross-file context is absent, and we see clear improvements when adding these context into the prompt. However, despite such improvements, the pinnacle of performance remains notably unattained even with the highest-performing model, indicating that CrossCodeEval is also capable of assessing model's capability in leveraging extensive context to make better code completion. Finally, we benchmarked various methods in retrieving cross-file context, and show that CrossCodeEval can also be used to measure the capability of code retrievers. Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Hantian Ding, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang |
NeurIPS | 10 |
| 2023 | On Learning Latent Models with Multi-Instance Weak SupervisionabstractWe consider a weakly supervised learning scenario where the supervision signal is generated by a transition function $\sigma$ of labels associated with multiple input instances. We formulate this problem as *multi-instance Partial Label Learning (multi-instance PLL)*, which is an extension to the standard PLL problem. Our problem is met in different fields, including latent structural learning and neuro-symbolic integration. Despite the existence of many learning techniques, limited theoretical analysis has been dedicated to this problem. In this paper, we provide the first theoretical study of multi-instance PLL with possibly an unknown transition $\sigma$. Our main contributions are as follows: First, we proposed a necessary and sufficient condition for the learnability of the problem. This condition nontrivially generalizes and relaxes the existing *small ambiguity degree* in PLL literature since we allow the transition to be deterministic. Second, we derived Rademacher-style error bounds based on the top-$k$ surrogate loss that is widely used in the neuro-symbolic literature. Furthermore, we conclude with empirical experiments for learning with an unknown transition. The empirical results align with our theoretical findings; however, they also expose the issue of scalability in the weak supervision literature. Kaifu Wang, Efthymia Tsamoura, Dan Roth 0001 |
NeurIPS | 3 |
| 2022 | There's a Time and Place for Reasoning Beyond the ImageabstractImages are often more significant than only the pixels to human eyes, as we can infer, associate, and reason with contextual information from other sources to establish a more complete picture. For example, in Figure This reasoning could provide the time and place the image was taken, which will help us in subsequent tasks, such as automatic storyline construction, correction of image source in intended effect photographs, and upper-stream processing such as image clustering for certain location or time. Ben Zhou, Ishaan Preetam Chandratreya, Carl Vondrick, Dan Roth 0001 |
ACL (1) | 5 |
| 2022 | Label Semantic Aware Pre-training for Few-shot Text ClassificationabstractAaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang, Dan Roth. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Aaron Mueller, Jason Krone, Salvatore Romeo, Saab Mansour, Elman Mansimov, Yi Zhang 0001, Dan Roth 0001 |
ACL (1) | 7 |
| 2022 | Cross-modal Map Learning for Vision and Language NavigationabstractWe consider the problem of Vision-and-Language Navigation (VLN). The majority of current methods for VLN are trained end-to-end using either unstructured memory such as LSTM, or using cross-modal attention over the egocentric observations of the agent. In contrast to other works, our key insight is that the association between language and vision is stronger when it occurs in explicit spatial representations. In this work, we propose a cross-modal map learning model for vision-and-language navigation that first learns to predict the top-down semantics on an egocentric map for both observed and unobserved regions, and then predicts a path towards the goal as a set of way-points. In both cases, the prediction is informed by the language through cross-modal attention mechanisms. We experimentally test the basic hypothesis that language-driven navigation can be solved given a map, and then show competitive results on the full VLN-CE benchmark. Georgios Georgakis, Karl Schmeckpeper, Karan Wanchoo, Soham Dan, Eleni Miltsakaki, Dan Roth 0001, Kostas Daniilidis |
CVPR | 6 |
| 2022 | On the Limitations of Reference-Free Evaluations of Generated TextabstractThere is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications.However, in this work, we demonstrate that these reference-free metrics are inherently biased and limited in their ability to evaluate generated text, and we argue that they should not be used to measure progress on tasks like machine translation or summarization.We show how reference-free metrics are equivalent to using one generation model to evaluate another, which has several limitations: (1) the metrics can be optimized at test time to find the approximate best-possible output, (2) they are inherently biased toward models which are more similar to their own, and (3) they can be biased against higher-quality outputs, including those written by humans.Therefore, we recommend that reference-free metrics should be used as diagnostic tools for analyzing and understanding model behavior instead of measures of how well models perform a task, in which the goal is to achieve as high of a score as possible.1 Daniel Deutsch, Rotem Dror, Dan Roth 0001 |
EMNLP | 3 |
| 2022 | Learning to Decompose: Hypothetical Question Decomposition Based on Comparable TextsabstractExplicit decomposition modeling, which involves breaking down complex tasks into more straightforward and often more interpretable sub-tasks, has long been a central theme in developing robust and interpretable NLU systems.However, despite the many datasets and resources built as part of this effort, the majority have small-scale annotations and limited scope, which is insufficient to solve general decomposition tasks.In this paper, we look at large-scale intermediate pre-training of decomposition-based transformers using distant supervision from comparable texts, particularly large-scale parallel news.We show that with such intermediate pre-training, developing robust decomposition-based models for a diverse range of tasks becomes more feasible.For example, on semantic parsing, our model, DECOMPT5, improves 20% to 30% on two datasets, Overnight and TORQUE, over the baseline language model.We further use DECOMPT5 to build a novel decompositionbased QA system named DECOMPENTAIL, improving over state-of-the-art models, including GPT-3, on both HotpotQA and StrategyQA by 8% and 4%, respectively. Ben Zhou, Kyle Richardson 0001, Xiaodong Yu 0003, Dan Roth 0001 |
EMNLP | 4 |
| 2022 | Weighted Training for Cross-Task Learning
Shuxiao Chen, Koby Crammer, Hangfeng He 0001, Dan Roth 0001, Weijie J. Su |
ICLR | 4 |
| 2022 | ROCK: Causal Inference Principles for Reasoning about Commonsense CausalityabstractCommonsense causality reasoning (CCR) aims at identifying plausible causes and effects in natural language descriptions that are deemed reasonable by an average person. Although being of great academic and practical interest, this problem is still shadowed by the lack of a well-posed theoretical framework; existing work usually relies on deep language models wholeheartedly, and is potentially susceptible to confounding co-occurrences. Motivated by classical causal principles, we articulate the central question of CCR and draw parallels between human subjects in observational studies and natural languages to adopt CCR to the potential-outcomes framework, which is the first such attempt for commonsense tasks. We propose a novel framework, ROCK, to Reason O(A)bout Commonsense K(C)ausality, which utilizes temporal signals as incidental supervision, and balances confounding effects using temporal propensities that are analogous to propensity scores. The ROCK implementation is modular and zero-shot, and demonstrates good CCR capabilities. Jiayao Zhang 0001, Hongming Zhang 0009, Weijie J. Su, Dan Roth 0001 |
ICML | 4 |
| 2022 | Neuro-Symbolic Language Modeling with Automaton-augmented RetrievalabstractRetrieval-based language models (R-LM) model the probability of natural language text by combining a standard language model (LM) with examples retrieved from an external datastore at test time. While effective, a major bottleneck of using these models in practice is the computationally costly datastore search, which can be performed as frequently as every time step. In this paper, we present RetoMaton - retrieval automaton - which approximates the datastore search, based on (1) saving pointers between consecutive datastore entries, and (2) clustering of entries into "states". This effectively results in a weighted finite automaton built on top of the datastore, instead of representing the datastore as a flat list. The creation of the automaton is unsupervised, and a RetoMaton can be constructed from any text collection: either the original training corpus or from another domain. Traversing this automaton at inference time, in parallel to the LM inference, reduces its perplexity by up to 1.85, or alternatively saves up to 83% of the nearest neighbor searches over $k$NN-LM (Khandelwal et al., 2020) without hurting perplexity. Our code and trained models are available at https://github.com/neulab/retomaton . Uri Alon 0002, Frank F. Xu, Junxian He, Sudipta Sengupta, Dan Roth 0001, Graham Neubig |
ICML | 5 |
| 2022 | Understanding Robust Generalization in Learning Regular LanguagesabstractA key feature of human intelligence is the ability to generalize beyond the training distribution, for instance, parsing longer sentences than seen in the past. Currently, deep neural networks struggle to generalize robustly to such shifts in the data distribution. We study robust generalization in the context of using recurrent neural networks (RNNs) to learn regular languages. We hypothesize that standard end-to-end modeling strategies cannot generalize well to systematic distribution shifts and propose a compositional strategy to address this. We compare an end-to-end strategy that maps strings to labels with a compositional strategy that predicts the structure of the deterministic finite state automaton (DFA) that accepts the regular language. We theoretically prove that the compositional strategy generalizes significantly better than the end-to-end strategy. In our experiments, we implement the compositional strategy via an auxiliary task where the goal is to predict the intermediate states visited by the DFA when parsing a string. Our empirical results support our hypothesis, showing that auxiliary tasks can enable robust generalization. Interestingly, the end-to-end RNN generalizes significantly better than the theoretical lower bound, suggesting that it is able to achieve atleast some degree of robust generalization. Soham Dan, Osbert Bastani, Dan Roth 0001 |
ICML | 3 |
| 2022 | Re-Examining System-Level Correlations of Automatic Summarization Evaluation MetricsabstractHow reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by systemlevel correlations.We identify two ways in which the definition of the system-level correlation is inconsistent with how metrics are used to evaluate systems in practice and propose changes to rectify this disconnect.First, we calculate the system score for an automatic metric using the full test set instead of the subset of summaries judged by humans, which is currently standard practice.We demonstrate how this small change leads to more precise estimates of system-level correlations.Second, we propose to calculate correlations only on pairs of systems that are separated by small differences in automatic scores which are commonly observed in practice.This allows us to demonstrate that our best estimate of the correlation of ROUGE to human judgments is near 0 in realistic scenarios.The results from the analyses point to the need to collect more high-quality human judgments and to improve automatic metrics when differences in system scores are small.1 Daniel Deutsch, Rotem Dror, Dan Roth 0001 |
NAACL-HLT | 3 |
| 2022 | Yes, No or IDK: The Challenge of Unanswerable Yes/No QuestionsabstractThe Yes/No QA task (Clark et al., 2019) consists of "Yes" or "No" questions about a given context.However, in realistic scenarios, the information provided in the context is not always sufficient in order to answer the question.For example, given the context "She married a lawyer from New-York.",we don't know whether the answer to the question "Did she marry in New York?" is "Yes" or "No".In this paper, we extend the Yes/No QA task, adding questions with an IDK answer, and show its considerable difficulty compared to the original 2-label task.For this purpose, we (i) enrich the BoolQ dataset (Clark et al., 2019) to include unanswerable questions and (ii) create out-ofdomain test sets for the Yes/No/IDK QA task.We study the contribution of training on other Natural Language Understanding tasks.We focus in particular on Extractive QA (Rajpurkar et al., 2018) and Recognizing Textual Entailments (RTE, Dagan et al., 2013), analyzing the differences between 2 and 3 labels using the new data. 1 Elior Sulem, Jamaal Hay, Dan Roth 0001 |
NAACL-HLT | 3 |
| 2022 | Extracting seizure frequency from epilepsy clinic notes: a machine reading approach to natural language processingabstractOBJECTIVE: Seizure frequency and seizure freedom are among the most important outcome measures for patients with epilepsy. In this study, we aimed to automatically extract this clinical information from unstructured text in clinical notes. If successful, this could improve clinical decision-making in epilepsy patients and allow for rapid, large-scale retrospective research. MATERIALS AND METHODS: We developed a finetuning pipeline for pretrained neural models to classify patients as being seizure-free and to extract text containing their seizure frequency and date of last seizure from clinical notes. We annotated 1000 notes for use as training and testing data and determined how well 3 pretrained neural models, BERT, RoBERTa, and Bio_ClinicalBERT, could identify and extract the desired information after finetuning. RESULTS: The finetuned models (BERTFT, Bio_ClinicalBERTFT, and RoBERTaFT) achieved near-human performance when classifying patients as seizure free, with BERTFT and Bio_ClinicalBERTFT achieving accuracy scores over 80%. All 3 models also achieved human performance when extracting seizure frequency and date of last seizure, with overall F1 scores over 0.80. The best combination of models was Bio_ClinicalBERTFT for classification, and RoBERTaFT for text extraction. Most of the gains in performance due to finetuning required roughly 70 annotated notes. DISCUSSION AND CONCLUSION: Our novel machine reading approach to extracting important clinical outcomes performed at or near human performance on several tasks. This approach opens new possibilities to support clinical practice and conduct large-scale retrospective clinical research. Future studies can use our finetuning pipeline with minimal training annotations to answer new clinical questions. Kevin Xie, Ryan S. Gallagher, Erin C. Conrad, Chadric O. Garrick, Steven Baldassano, John M. Bernabei, Peter D. Galer, Nina J. Ghosn, Adam S. Greenblatt, Tara Jennings, Alana Kornspun, Catherine V. Kulick-Soper, Jal M. Panchal, Akash R. Pattnaik, Brittany Scheid, Danmeng Wei, Micah Weitzman, Ramya Muthukrishnan, Joongwon Kim, Brian Litt, Colin A. Ellis, Dan Roth 0001 |
J. Am. Medical Informatics Assoc. | 22 |
| 2021 | Visual Pivoting for (Unsupervised) Entity AlignmentabstractThis work studies the use of visual semantic representations to align entities in heterogeneous knowledge graphs (KGs). Images are natural components of many existing KGs. By combining visual knowledge with other auxiliary information, we show that the proposed new approach, EVA, creates a holistic entity representation that provides strong signals for cross-graph entity alignment. Besides, previous entity alignment methods require human labelled seed alignment, restricting availability. EVA provides a completely unsupervised solution by leveraging the visual similarity of entities to create an initial seed dictionary (visual pivots). Experiments on benchmark data sets DBP15k and DWY15k show that EVA offers state-of-the-art performance on both monolingual and cross-lingual entity alignment tasks. Furthermore, we discover that images are particularly useful to align long-tail KG entities, which inherently lack the structural contexts necessary for capturing the correspondences. Code release: https://github.com/cambridgeltl/eva; project page: http://cogcomp.org/page/publication view/927. Fangyu Liu 0001, Muhao Chen 0001, Dan Roth 0001, Nigel Collier |
AAAI | 3 |
| 2021 | Coreference Reasoning in Machine Reading ComprehensionabstractMingzhu Wu, Nafise Sadat Moosavi, Dan Roth, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mingzhu Wu, Nafise Sadat Moosavi, Dan Roth 0001, Iryna Gurevych |
ACL/IJCNLP (1) | 3 |
| 2021 | What is Your Article Based On? Inferring Fine-grained ProvenanceabstractYi Zhang, Zachary Ives, Dan Roth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yi Zhang 0001, Zachary G. Ives, Dan Roth 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Understanding the Extent to which Content Quality Metrics Measure the Information Quality of SummariesabstractReference-based metrics such as ROUGE or BERTScore evaluate the content quality of a summary by comparing the summary to a reference.Ideally, this comparison should measure the summary's information quality by calculating how much information the summaries have in common.In this work, we analyze the token alignments used by ROUGE and BERTScore to compare summaries and argue that their scores largely cannot be interpreted as measuring information overlap.Rather, they are better estimates of the extent to which the summaries discuss the same topics.Further, we provide evidence that this result holds true for many other summarization evaluation metrics.The consequence of this result is that the most frequently used summarization evaluation metrics do not align with the community's research goal, to generate summaries with high-quality information.However, we conclude by demonstrating that a recently proposed metric, QAEval, which scores summaries using question-answering, appears to better capture information quality than current evaluations, highlighting a direction for future research. Daniel Deutsch, Dan Roth 0001 |
CoNLL | 2 |
| 2021 | BabyBERTa: Learning More Grammar With Small-Scale Child-Directed LanguageabstractTransformer-based language models have taken the NLP world by storm.However, their potential for addressing important questions in language acquisition research has been largely ignored.In this work, we examined the grammatical knowledge of RoBERTa (Liu et al., 2019) when trained on a 5M word corpus of language acquisition data to simulate the input available to children between the ages 1 and 6.Using the behavioral probing paradigm, we found that a smaller version of RoBERTa-base that never predicts unmasked tokens, which we term BabyBERTa, acquires grammatical knowledge comparable to that of pre-trained RoBERTa-base -and does so with approximately 15X fewer parameters and 6,000X fewer words.We discuss implications for building more efficient models and the learnability of grammar from input available to children.Lastly, to support research on this front, we release our novel grammar test suite that is compatible with the small vocabulary of child-directed input. Philip A. Huebner, Elior Sulem, Cynthia Fisher, Dan Roth 0001 |
CoNLL | 4 |
| 2021 | Cross-lingual Entity Alignment with Incidental SupervisionabstractMuch research effort has been put to multilingual knowledge graph (KG) embedding methods to address the entity alignment task, which seeks to match entities in different languagespecific KGs that refer to the same real-world object.Such methods are often hindered by the insufficiency of seed alignment provided between KGs.Therefore, we propose an incidentally supervised model, JEANS , which jointly represents multilingual KGs and text corpora in a shared embedding scheme, and seeks to improve entity alignment with incidental supervision signals from text.JEANS first deploys an entity grounding process to combine each KG with the monolingual text corpus.Then, two learning processes are conducted: (i) an embedding learning process to encode the KG and text of each language in one embedding space, and (ii) a selflearning based alignment learning process to iteratively induce the matching of entities and that of lexemes between embeddings.Experiments on benchmark datasets show that JEANS leads to promising improvement on entity alignment with incidental supervision, and significantly outperforms state-of-the-art methods that solely rely on internal information of KGs. 1 * Indicating equal contributions. Muhao Chen 0001, Ben Zhou, Dan Roth 0001 |
EACL | 4 |
| 2021 | How Good (really) are Grammatical Error Correction Systems?abstractStandard evaluations of Grammatical Error Correction (GEC) systems make use of a fixed reference text generated relative to the original text; they show, even when using multiple references, that we have a long way to go.This analysis paper studies the performance of GEC systems relative to closest-gold -a gold reference text created relative to the output of a system.Surprisingly, we show that the real performance is 20-40 points better than standard evaluations show.Moreover, the performance remains high even when considering any of the top-10 hypotheses produced by a system.Importantly, the type of mistakes corrected by lower-ranked hypotheses differs in interesting ways from the top one, providing an opportunity to focus on a range of errors -local spelling and grammar edits vs. more complex lexical improvements.Our study shows these results in English and Russian, and thus provides a preliminary proposal for a more realistic evaluation of GEC systems. Alla Rozovskaya, Dan Roth 0001 |
EACL | 2 |
| 2021 | Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd SchemaabstractThe Winograd Schema (WS) has been proposed as a test for measuring commonsense capabilities of models.Recently, pre-trained language model-based approaches have boosted performance on some WS benchmarks but the source of improvement is still not clear.This paper suggests that the apparent progress on WS may not necessarily reflect progress in commonsense reasoning.To support this claim, we first show that the current evaluation method of WS is sub-optimal and propose a modification that uses twin sentences for evaluation.We also propose two new baselines that indicate the existence of artifacts in WS benchmarks.We then develop a method for evaluating WS-like sentences in a zero-shot setting to account for the commonsense reasoning abilities acquired during the pretraining and observe that popular language models perform randomly in this setting when using our more strict evaluation.We conclude that the observed progress is mostly due to the use of supervision in training WS models, which is not likely to successfully support all the required commonsense reasoning skills and knowledge.1 Yanai Elazar, Hongming Zhang 0009, Yoav Goldberg, Dan Roth 0001 |
EMNLP (1) | 4 |
| 2021 | Paired Examples as Indirect Supervision in Latent Decision ModelsabstractCompositional, structured models are appealing because they explicitly decompose problems and provide interpretable intermediate outputs that give confidence that the model is not simply latching onto data artifacts.Learning these models is challenging, however, because end-task supervision only provides a weak indirect signal on what values the latent decisions should take.This often results in the model failing to learn to perform the intermediate tasks correctly.In this work, we introduce a way to leverage paired examples that provide stronger cues for learning latent decisions.When two related training examples share internal substructure, we add an additional training objective to encourage consistency between their latent decisions.Such an objective does not require external supervision for the values of the latent output, or even the end task, yet provides an additional training signal to that provided by individual training examples themselves.We apply our method to improve compositional question answering using neural module networks on the DROP dataset.We explore three ways to acquire paired questions in DROP: (a) discovering naturally occurring paired examples within the dataset, (b) constructing paired examples using templates, and (c) generating paired examples using a question generation model.We empirically demonstrate that our proposed approach improves both in-and outof-distribution generalization and leads to correct latent decision predictions. Nitish Gupta, Sameer Singh 0001, Matt Gardner 0001, Dan Roth 0001 |
EMNLP (1) | 4 |
| 2021 | ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic RelationsabstractUnderstanding how events are semantically related to each other is the essence of reading comprehension.Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations.While these tasks partially evaluate machines' ability of narrative understanding, human-like reading comprehension requires the capability to process event-based information beyond arguments and temporal reasoning.For example, to understand causality between events, we need to infer motivation or purpose; to establish event hierarchy, we need to understand the composition of events.To facilitate these tasks, we introduce ESTER, a comprehensive machine reading comprehension (MRC) dataset for Event Semantic Relation Reasoning.The dataset leverages natural language queries to reason about the five most common event semantic relations, provides more than 6K questions, and captures 10.1K event relation pairs.Experimental results show that the current SOTA systems achieve 22.1%, 63.3% and 83.5% for token-based exact-match (EM), F 1 and event-based HIT@1 scores, which are all significantly below human performances (36.0%, 79.6%, 100% respectively), highlighting our dataset as a challenging benchmark.1 Rujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon, Qiang Ning, Dan Roth 0001, Nanyun Peng 0001 |
EMNLP (1) | 6 |
| 2021 | Foreseeing the Benefits of Incidental SupervisionabstractReal-world applications often require improved models by leveraging a range of cheap incidental supervision signals.These could include partial labels, noisy labels, knowledgebased constraints, and cross-domain or crosstask annotations -all having statistical associations with gold annotations but not exactly the same.However, we currently lack a principled way to measure the benefits of these signals to a given target task, and the common practice of evaluating these benefits is through exhaustive experiments with various models and hyperparameters.This paper studies whether we can, in a single framework, quantify the benefits of various types of incidental signals for a given target task without going through combinatorial experiments.We propose a unified PAC-Bayesian motivated informativeness measure, PABI, that characterizes the uncertainty reduction provided by incidental supervision signals.We demonstrate PABI's effectiveness by quantifying the value added by various types of incidental signals to sequence tagging tasks.Experiments on named entity recognition (NER) and question answering (QA) show that PABI's predictions correlate well with learning performance, providing a promising way to determine, ahead of learning, which supervision signals would be beneficial. 1 Hangfeng He 0001, Qiang Ning, Dan Roth 0001 |
EMNLP (1) | 4 |
| 2021 | Learning Constraints and Descriptive Segmentation for Subevent DetectionabstractEvent mentions in text correspond to realworld events of varying degrees of granularity.The task of subevent detection aims to resolve this granularity issue, recognizing the membership of multi-granular events in event complexes.Since knowing the span of descriptive contexts of event complexes helps infer the membership of events, we propose the task of event-based text segmentation (EVENTSEG) as an auxiliary task to improve the learning for subevent detection.To bridge the two tasks together, we propose an approach to learning and enforcing constraints that capture dependencies between subevent detection and EVENTSEG prediction, as well as guiding the model to make globally consistent inference.Specifically, we adopt Rectifier Networks for constraint learning and then convert the learned constraints to a regularization term in the loss function of the neural model.Experimental results show that the proposed method outperforms baseline methods by 2.3% and 2.5% on benchmark datasets for subevent detection, HiEve and IC, respectively, while achieving a decent performance on EVENTSEG prediction 1 . Haoyu Wang 0005, Hongming Zhang 0009, Muhao Chen 0001, Dan Roth 0001 |
EMNLP (1) | 4 |
| 2021 | Context-Based Quotation Recommendation
Ansel MacLaughlin, Burcu Karagol Ayan, Dan Roth 0001 |
ICWSM | 4 |
| 2021 | Improving Faithfulness in Abstractive Summarization with Contrast Candidate Generation and SelectionabstractDespite significant progress in neural abstractive summarization, recent studies have shown that the current models are prone to generating summaries that are unfaithful to the original context.To address the issue, we study contrast candidate generation and selection as a model-agnostic post-processing technique to correct the extrinsic hallucinations (i.e.information not present in the source text) in unfaithful summaries.We learn a discriminative correction model by generating alternative candidate summaries where named entities and quantities in the generated summary are replaced with ones with compatible semantic types from the source document.This model is then used to select the best candidate as the final output summary.Our experiments and analysis across a number of neural summarization systems show that our proposed method is effective in identifying and correcting extrinsic hallucinations.We analyze the typical hallucination phenomenon by different types of neural summarization systems, in hope to provide insights for future work on the direction. Fan Zhang 0095, Kazoo Sone, Dan Roth 0001 |
NAACL-HLT | 4 |
| 2021 | Generalization in Instruction Following SystemsabstractUnderstanding and executing natural language instructions in a grounded domain is one of the hallmarks of artificial intelligence.In this paper, we focus on instruction understanding in the blocks world domain and investigate the language understanding abilities of two topperforming systems for the task.We aim to understand if the test performance of these models indicates an understanding of the spatial domain and of the natural language instructions relative to it, or whether they merely over-fit spurious signals in the dataset.We formulate a set of expectations one might have from an instruction following model and concretely characterize the different dimensions of robustness such a model should possess.Despite decent test performance, we find that state-of-the-art models fall short of these expectations and are extremely brittle.We then propose a learning strategy that involves data augmentation and show through extensive experiments that the proposed learning strategy yields models that are competitive on the original test set while satisfying our expectations much better.1 . Soham Dan, Michael Zhou, Dan Roth 0001 |
NAACL-HLT | 3 |
| 2021 | MultiOpEd: A Corpus of Multi-Perspective News EditorialsabstractSiyi Liu, Sihao Chen, Xander Uyttendaele, Dan Roth. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xander Uyttendaele, Dan Roth 0001 |
NAACL-HLT | 4 |
| 2021 | Event Time Extraction and Propagation via Graph Attention NetworksabstractHaoyang Wen, Yanru Qu, Heng Ji, Qiang Ning, Jiawei Han, Avi Sil, Hanghang Tong, Dan Roth. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Haoyang Wen, Yanru Qu, Heng Ji 0001, Qiang Ning, Jiawei Han 0001, Avirup Sil, Hanghang Tong, Dan Roth 0001 |
NAACL-HLT | 8 |
| 2021 | Learning to Decompose and Organize Complex TasksabstractYi Zhang, Sujay Kumar Jauhar, Julia Kiseleva, Ryen White, Dan Roth. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yi Zhang 0001, Sujay Kumar Jauhar, Julia Kiseleva, Ryen W. White, Dan Roth 0001 |
NAACL-HLT | 5 |
| 2021 | Temporal Reasoning on Implicit Events from Distant SupervisionabstractBen Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ben Zhou, Kyle Richardson 0001, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
NAACL-HLT | 6 |
| 2021 | Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a SummaryabstractAbstract A desirable property of a reference-based evaluation metric that measures the content quality of a summary is that it should estimate how much information that summary has in common with a reference. Traditional text overlap based metrics such as ROUGE fail to achieve this because they are limited to matching tokens, either lexically or via embeddings. In this work, we propose a metric to evaluate the content quality of a summary using question-answering (QA). QA-based methods directly measure a summary’s information overlap with a reference, making them fundamentally different than text overlap metrics. We demonstrate the experimental benefits of QA-based metrics through an analysis of our proposed metric, QAEval. QAEval outperforms current state-of-the-art metrics on most evaluations using benchmark datasets, while being competitive on others due to limitations of state-of-the-art models. Through a careful analysis of each component of QAEval, we identify its performance bottlenecks and estimate that its potential upper-bound performance surpasses all other automatic metrics, approaching that of the gold-standard Pyramid Method.1 Daniel Deutsch, Tania Bedrax-Weiss, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | A Statistical Analysis of Summarization Evaluation Metrics Using Resampling MethodsabstractAbstract The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently, it is unclear how precise these correlation estimates are, nor whether differences between two metrics’ correlations reflect a true difference or if it is due to mere chance. In this work, we address these two problems by proposing methods for calculating confidence intervals and running hypothesis tests for correlations using two resampling methods, bootstrapping and permutation. After evaluating which of the proposed methods is most appropriate for summarization through two simulation experiments, we analyze the results of applying these methods to several different automatic evaluation metrics across three sets of human annotations. We find that the confidence intervals are rather wide, demonstrating high uncertainty in the reliability of automatic metrics. Further, although many metrics fail to show statistical improvements over ROUGE, two recent works, QAEval and BERTScore, do so in some evaluation settings.1 Daniel Deutsch, Rotem Dror, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning StrategiesabstractAbstract A key limitation in current datasets for multi-hop reasoning is that the required steps for answering the question are mentioned in it explicitly. In this work, we introduce StrategyQA, a question answering (QA) benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy. A fundamental challenge in this setup is how to elicit such creative questions from crowdsourcing workers, while covering a broad range of potential strategies. We propose a data collection procedure that combines term-based priming to inspire annotators, careful control over the annotator population, and adversarial filtering for eliminating reasoning shortcuts. Moreover, we annotate each question with (1) a decomposition into reasoning steps for answering it, and (2) Wikipedia paragraphs that contain the answers to each step. Overall, StrategyQA includes 2,780 examples, each consisting of a strategy question, its decomposition, and evidence paragraphs. Analysis shows that questions in StrategyQA are short, topic-diverse, and cover a wide range of strategies. Empirically, we show that humans perform well (87%) on this task, while our best baseline reaches an accuracy of ∼ 66%. Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth 0001, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 5 |
| 2020 | Robust Named Entity Recognition with Truecasing PretrainingabstractAlthough modern named entity recognition (NER) systems show impressive performance on standard datasets, they perform poorly when presented with noisy data. In particular, capitalization is a strong signal for entities in many languages, and even state of the art models overfit to this feature, with drastically lower performance on uncapitalized text. In this work, we address the problem of robustness of NER systems in data with noisy or uncertain casing, using a pretraining objective that predicts casing in text, or a truecaser, leveraging unlabeled data. The pretrained truecaser is combined with a standard BiLSTM-CRF model for NER by appending output distributions to character embeddings. In experiments over several datasets of varying domain and casing quality, we show that our new model improves performance in uncased text, even adding value to uncased BERT embeddings. Our method achieves a new state of the art on the WNUT17 shared task dataset. Stephen Mayhew 0001, Nitish Gupta, Dan Roth 0001 |
AAAI | 3 |
| 2020 | Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess HypothesesabstractEmpirical research in Natural Language Processing (NLP) has adopted a narrow set of principles for assessing hypotheses, relying mainly on p-value computation, which suffers from several known issues.While alternative proposals have been well-debated and adopted in other fields, they remain rarely discussed or used within the NLP community.We address this gap by contrasting various hypothesis assessment techniques, especially those not commonly used in the field (such as evaluations based on Bayesian inference).Since these statistical techniques differ in the hypotheses they can support, we argue that practitioners should first decide their target hypothesis before choosing an assessment method.This is crucial because common fallacies, misconceptions, and misinterpretation surrounding hypothesis assessment methods often stem from a discrepancy between what one would like to claim versus what the method used actually assesses.Our survey reveals that these issues are omnipresent in the NLP research community.As a step forward, we provide best practices and guidelines tailored towards NLP research, as well as an easy-to-use package called HyBayes for Bayesian assessment of hypotheses, 1 complementing existing tools. Erfan Sadeqi Azer, Daniel Khashabi, Ashish Sabharwal, Dan Roth 0001 |
ACL | 4 |
| 2020 | QuASE: Question-Answer Driven Sentence EncodingabstractQuestion-answering (QA) data often encodes essential information in many facets. This paper studies a natural question: Can we get supervision from QA data for other tasks (typically, non-QA ones)? For example, can we use QAMR (Michael et al., 2017) to improve named entity recognition? We suggest that simply further pre-training BERT is often not the best option, and propose the question-answer driven sentence encoding (QuASE) framework. QuASE learns representations from QA data, using BERT or other state-of-the-art contextual language models. In particular, we observe the need to distinguish between two types of sentence encodings, depending on whether the target task is a single- or multi-sentence input; in both cases, the resulting encoding is shown to be an easy-to-use plugin for many downstream tasks. This work may point out an alternative way to supervise NLP tasks. Hangfeng He 0001, Qiang Ning, Dan Roth 0001 |
ACL | 3 |
| 2020 | "Who said it, and Why?" Provenance for Natural Language ClaimsabstractIn an era where generating content and publishing it is so easy, we are bombarded with information and are exposed to all kinds of claims, some of which do not always rank high on the truth scale. This paper suggests that the key to a longer-term, holistic, and systematic approach to navigating this information pollution is capturing the provenance of claims. To do that, we develop a formal definition of provenance graph for a given natural language claim, aiming to understand where the claim may come from and how it has evolved. To construct the graph, we model provenance inference, formulated mainly as an information extraction task and addressed via a textual entailment model. We evaluate our approach using two benchmark datasets, showing initial success in capturing the notion of provenance and its effectiveness on the application of claim verification. Yi Zhang 0001, Zachary G. Ives, Dan Roth 0001 |
ACL | 3 |
| 2020 | Temporal Common Sense Acquisition with Minimal SupervisionabstractTemporal common sense (e.g., duration and frequency of events) is crucial for understanding natural language.However, its acquisition is challenging, partly because such information is often not expressed explicitly in text, and human annotation on such concepts is costly.This work proposes a novel sequence modeling approach that exploits explicit and implicit mentions of temporal common sense, extracted from a large corpus, to build TACOLM, 1 a temporal common sense language model.Our method is shown to give quality predictions of various dimensions of temporal common sense (on UDST and a newly collected dataset from Real-News).It also produces representations of events for relevant tasks such as duration comparison, parent-child relations, event coreference and temporal QA (on TimeBank, HiEVE and MCTACO) that are better than using the standard BERT.Thus, it will be an important component of temporal NLP. Ben Zhou, Qiang Ning, Daniel Khashabi, Dan Roth 0001 |
ACL | 4 |
| 2020 | Is Killed More Significant than Fled? A Contextual Model for Salient Event DetectionabstractIdentifying the key events in a document is critical to holistically understanding its important information.Although measuring the salience of events is highly contextual, most previous work has used a limited representation of events that omits essential information.In this work, we propose a highly contextual model of event salience that uses a rich representation of events, incorporates document-level information and allows for interactions between latent event encodings.Our experimental results on an event salience dataset (Liu et al., 2018) demonstrate that our model improves over previous work by an absolute 2-4% on standard metrics, establishing a new state-of-the-art performance for the task.We also propose a new evaluation metric which addresses flaws in previous evaluation methodologies.Finally, we discuss the importance of salient event detection for the downstream task of summarization. 1 Disha Jindal, Daniel Deutsch, Dan Roth 0001 |
COLING | 3 |
| 2020 | QANom: Question-Answer driven SRL for NominalizationsabstractAyal Klein, Jonathan Mamou, Valentina Pyatkin, Daniela Stepanov, Hangfeng He, Dan Roth, Luke Zettlemoyer, Ido Dagan. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Ayal Klein, Jonathan Mamou, Valentina Pyatkin, Daniela Stepanov, Hangfeng He 0001, Dan Roth 0001, Luke Zettlemoyer, Ido Dagan |
COLING | 6 |
| 2020 | What Are You Trying to Do? Semantic Typing of Event ProcessesabstractThis paper studies a new cognitively motivated semantic typing task, multi-axis event process typing, that, given an event process, attempts to infer free-form type labels describing (i) the type of action made by the process and (ii) the type of object the process seeks to affect.This task is inspired by computational and cognitive studies of event understanding, which suggest that understanding processes of events is often directed by recognizing the goals, plans or intentions of the protagonist(s).We develop a large dataset containing over 60k event processes, featuring ultra fine-grained typing on both the action and object type axes with very large (10 3 ∼ 10 4 ) label vocabularies.We then propose a hybrid learning framework, P2GT, which addresses the challenging typing problem with indirect supervision from glosses 1 and a joint learning-to-rank framework.As our experiments indicate, P2GT supports identifying the intent of processes, as well as the fine semantic type of the affected object.It also demonstrates the capability of handling fewshot cases, and strong generalizability on outof-domain processes.2 * This work was done when the author was visiting the University of Pennsylvania.1 A gloss provides a sense definition for a lexeme. 2 The contributed learning resources, software and a system demonstration are available at http://cogcomp.org/page/publication_view/915. Muhao Chen 0001, Hongming Zhang 0009, Haoyu Wang 0005, Dan Roth 0001 |
CoNLL | 4 |
| 2020 | Design Challenges in Low-resource Cross-lingual Entity LinkingabstractCross-lingual Entity Linking (XEL), the problem of grounding mentions of entities in a foreign language text into an English knowledge base such as Wikipedia, has seen a lot of research in recent years, with a range of promising techniques.However, current techniques do not rise to the challenges introduced by text in low-resource languages (LRL) and, surprisingly, fail to generalize to text not taken from Wikipedia, on which they are usually trained.This paper provides a thorough analysis of low-resource XEL techniques, focusing on the key step of identifying candidate English Wikipedia titles that correspond to a given foreign language mention.Our analysis indicates that current methods are limited by their reliance on Wikipedia's interlanguage links and thus suffer when the foreign language's Wikipedia is small.We conclude that the LRL setting requires the use of outside-Wikipedia cross-lingual resources and present a simple yet effective zero-shot XEL system, QuEL, that utilizes search engines query logs.With experiments on 25 languages, QuEL shows an average increase of 25% in gold candidate recall and of 13% in end-to-end linking accuracy over state-of-the-art baselines.1 Xiaodong Yu 0003, Zian Zhao, Dan Roth 0001 |
EMNLP (1) | 5 |
| 2020 | "I'd rather just go to bed": Understanding Indirect AnswersabstractWe revisit a pragmatic inference problem in dialog: Understanding indirect responses to questions.Humans can interpret 'I'm starving.' in response to 'Hungry?', even without direct cue words such as 'yes' and 'no'.In dialog systems, allowing natural responses rather than closed vocabularies would be similarly beneficial.However, today's systems are only as sensitive to these pragmatic moves as their language model allows.We create and release 1 the first large-scale English language corpus 'Circa' with 34,268 (polar question, indirect answer) pairs to enable progress on this task.The data was collected via elaborate crowdsourcing, and contains utterances with yes/no meaning, as well as uncertain, middle-ground, and conditional responses.We also present BERT-based neural models to predict such categories for a question-answer pair.We find that while transfer learning from entailment works reasonably, performance is not yet sufficient for robust dialog.Our models reach 82-88% accuracy for a 4-class distinction, and 74-85% for 6 classes.* * Work done at Google "Want to get some dinner together?""I know a restaurant we could get a reservation at." "I have already eaten recently.""I hope to make it home by supper but I'm not sure I can." "Dinner would be lovely.""I'd rather just go to bed." "There's a few new restaurants we could go to.""I would like that.""We could do dinner this weekend.""I would like to go somewhere casual.""I'd like to try the new Italian place." Annie Louis, Dan Roth 0001, Filip Radlinski |
EMNLP (1) | 2 |
| 2020 | TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsabstractA critical part of reading is being able to understand the temporal relationships between events described in a passage of text, even when those relationships are not explicitly stated.However, current machine reading comprehension benchmarks have practically no questions that test temporal phenomena, so systems trained on these benchmarks have no capacity to answer questions such as "what happened before/after [some event]?"We introduce TORQUE, a new English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships.Results show that RoBERTa-large achieves an exact-match score of 51% on the test set of TORQUE, about 30% behind human performance.1 1 https://allennlp.org/torque.htmlHeavy snow is causing disruption to transport across the UK, with heavy rainfall bringing flooding to the south-west of England.Rescuers searching for a woman trapped in a landslide at her home in Looe, Cornwall, said they had found a body.Q1: What events have already finished?A: searching trapped landslide said found Q2: What events have begun but has not finished?A: snow causing disruption rainfall bringing flooding Q3: What will happen in the future?A: No answers.Q4: What happened before a woman was trapped?A: landslide Q5: What had started before a woman was trapped?A: snow rainfall landslide Q6: What happened while a woman was trapped?A: searching Q7: What happened after a woman was trapped?A: searching said found Q8: What happened at about the same time as the snow?A: rainfall Q9: What happened after the snow started?A: causing disruption bringing flooding searching trapped landslide said found Q10: What happened before the snow started?A: No answers.warm Qiang Ning, Hao Wu 0034, Rujun Han, Nanyun Peng 0001, Matt Gardner 0001, Dan Roth 0001 |
EMNLP (1) | 6 |
| 2020 | Joint Constrained Learning for Event-Event Relation ExtractionabstractUnderstanding natural language involves recognizing how multiple event mentions structurally and temporally interact with each other.In this process, one can induce event complexes that organize multi-granular events with temporal order and membership relations interweaving among them.Due to the lack of jointly labeled data for these relational phenomena and the restriction on the structures they articulate, we propose a joint constrained learning framework for modeling event-event relations.Specifically, the framework enforces logical constraints within and across multiple temporal and subevent relations by converting these constraints into differentiable learning objectives.We show that our joint constrained learning approach effectively compensates for the lack of jointly labeled data, and outperforms SOTA methods on benchmarks for both temporal relation extraction and event hierarchy construction, replacing a commonly used but more expensive global inference process.We also present a promising case study showing the effectiveness of our approach in inducing event complexes on an external corpus. 1 Haoyu Wang 0005, Muhao Chen 0001, Hongming Zhang 0009, Dan Roth 0001 |
EMNLP (1) | 4 |
| 2020 | Analogous Process Structure Induction for Sub-event Sequence PredictionabstractComputational and cognitive studies of event understanding suggest that identifying, comprehending, and predicting events depend on having structured representations of a sequence of events and on conceptualizing (abstracting) its components into (soft) event categories.Thus, knowledge about a known process such as "buying a car" can be used in the context of a new but analogous process such as "buying a house".Nevertheless, most event understanding work in NLP is still at the ground level and does not consider abstraction.In this paper, we propose an Analogous Process Structure Induction (APSI) framework, which leverages analogies among processes and conceptualization of sub-event instances to predict the whole sub-event sequence of previously unseen open-domain processes.As our experiments and analysis indicate, APSI 1 supports the generation of meaningful sub-event sequences for unseen processes and can help predict missing events. Hongming Zhang 0009, Muhao Chen 0001, Haoyu Wang 0005, Yangqiu Song, Dan Roth 0001 |
EMNLP (1) | 5 |
| 2020 | Neural Module Networks for Reasoning over Text
Nitish Gupta, Dan Roth 0001, Sameer Singh 0001, Matt Gardner 0001 |
ICLR | 3 |
| 2020 | Cross-Lingual Ability of Multilingual BERT: An Empirical Study
Karthikeyan K, Zihan Wang 0001, Stephen Mayhew 0001, Dan Roth 0001 |
ICLR | 4 |
| 2020 | TransOMCS: From Linguistic Graphs to Commonsense KnowledgeabstractCommonsense knowledge acquisition is a key problem for artificial intelligence. Conventional methods of acquiring commonsense knowledge generally require laborious and costly human annotations, which are not feasible on a large scale. In this paper, we explore a practical way of mining commonsense knowledge from linguistic graphs, with the goal of transferring cheap knowledge obtained with linguistic patterns into expensive commonsense knowledge. The result is a conversion of ASER [Zhang et al., 2020], a large-scale selectional preference knowledge resource, into TransOMCS, of the same representation as ConceptNet [Liu and Singh, 2004] but two orders of magnitude larger. Experimental results demonstrate the transferability of linguistic knowledge to commonsense knowledge and the effectiveness of the proposed approach in terms of quantity, novelty, and quality. TransOMCS is publicly available at: https://github.com/HKUST-KnowComp/TransOMCS. Hongming Zhang 0009, Daniel Khashabi, Yangqiu Song, Dan Roth 0001 |
IJCAI | 4 |
| 2020 | Understanding Spatial Relations through Multiple ModalitiesabstractRecognizing spatial relations and reasoning about them is essential in multiple applications including navigation, direction giving and human-computer interaction in general. Spatial relations between objects can either be explicit – expressed as spatial prepositions, or implicit – expressed by spatial verbs such as moving, walking, shifting, etc. Both these, but implicit relations in particular, require significant common sense understanding. In this paper, we introduce the task of inferring implicit and explicit spatial relations between two entities in an image. We design a model that uses both textual and visual information to predict the spatial relations, making use of both positional and size information of objects and image embeddings. We contrast our spatial model with powerful language models and show how our modeling complements the power of these, improving prediction accuracy and coverage and facilitates dealing with unseen subjects, objects and relations. Soham Dan, Hangfeng He 0001, Dan Roth 0001 |
LREC | 3 |
| 2020 | From Spatial Relations to Spatial ConfigurationsabstractSpatial Reasoning from language is essential for natural language understanding. Supporting it requires a representation scheme that can capture spatial phenomena encountered in language as well as in images and videos. Existing spatial representations are not sufficient for describing spatial configurations used in complex tasks. This paper extends the capabilities of existing spatial representation languages and increases coverage of the semantic aspects that are needed to ground spatial meaning of natural language text in the world. Our spatial relation language is able to represent a large, comprehensive set of spatial concepts crucial for reasoning and is designed to support composition of static and dynamic spatial configurations. We integrate this language with the Abstract Meaning Representation (AMR) annotation schema and present a corpus annotated by this extended AMR. To exhibit the applicability of our representation scheme, we annotate text taken from diverse datasets and show how we extend the capabilities of existing spatial representation languages with fine-grained decomposition of semantics and blend it seamlessly with AMRs of sentences and discourse representations as a whole. Soham Dan, Parisa Kordjamshidi, Julia Bonn, Archna Bhatia, Zheng Cai, Martha Palmer, Dan Roth 0001 |
LREC | 7 |
| 2020 | Learnability with Indirect Supervision SignalsabstractLearning from indirect supervision signals is important in real-world AI applications when, often, gold labels are missing or too costly. In this paper, we develop a unified theoretical framework for multi-class classification when the supervision is provided by a variable that contains nonzero mutual information with the gold label. The nature of this problem is determined by (i) the transition probability from the gold labels to the indirect supervision variables and (ii) the learner's prior knowledge about the transition. Our framework relaxes assumptions made in the literature, and supports learning with unknown, non-invertible and instance-dependent transitions. Our theory introduces a novel concept called \emph{separation}, which characterizes the learnability and generalization bounds. We also demonstrate the application of our framework via concrete novel results in a variety of learning scenarios such as learning with superset annotations and joint supervision signals. Kaifu Wang, Qiang Ning, Dan Roth 0001 |
NeurIPS | 3 |
| 2020 | Task-Oriented Dialogue as Dataflow SynthesisabstractWe describe an approach to task-oriented dialogue in which dialogue state is represented as a dataflow graph. A dialogue agent maps each user utterance to a program that extends this graph. Programs include metacomputation operators for reference and revision that reuse dataflow fragments from previous turns. Our graph-based state enables the expression and manipulation of complex user intents, and explicit metacomputation makes these intents easier for learned models to predict. We introduce a new dataset, SMCalFlow, featuring complex dialogues about events, weather, places, and people. Experiments show that dataflow graphs and metacomputation substantially improve representability and predictability in these natural dialogues. Additional experiments on the MultiWOZ dataset show that our dataflow representation enables an otherwise off-the-shelf sequence-to-sequence model to match the best existing task-specific state tracking model. The SMCalFlow dataset, code for replicating experiments, and a public leaderboard are available at https://www.microsoft.com/en-us/research/project/dataflow-based-dialogue-semantic-machines . Jacob Andreas, John Bufe, David Burkett, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, Hao Fang 0002, Alan Guo, David Hall 0006, Kristin Hayes, Kellie Hill, Diana Ho, Wendy Iwaszuk, Smriti Jha, Daniel Klein 0001, Jayant Krishnamurthy, Theo Lanman, Percy Liang, Christopher H. Lin, Ilya Lintsbakh, Andy McGovern, Aleksandr Nisnevich, Adam Pauls, Dmitrij Petters, Brent Read, Dan Roth 0001, Subhro Roy, Jesse Rusak, Beth Short, Div Slomin, Ben Snyder, Stephon Striplin, Yu Su 0001, Zachary Tellman, Sam Thomson, Andrei Vorobev, Izabela Witoszko, Jason Andrew Wolfe, Abby Wray, Yuchen Zhang 0002, Alexander Zotov |
Trans. Assoc. Comput. Linguistics | 30 |
| 2019 | How Large Are Lions? Inducing Distributions over Quantitative AttributesabstractMost current NLP systems have little knowledge about quantitative attributes of objects and events.We propose an unsupervised method for collecting quantitative information from large amounts of web data, and use it to create a new, very large resource consisting of distributions over physical quantities associated with objects, adjectives, and verbs which we call Distribution over Quantities (DOQ) 1 .This contrasts with recent work in this area which has focused on making only relative comparisons such as "Is a lion bigger than a wolf?".Our evaluation shows that DOQ compares favorably with state of the art results on existing datasets for relative comparisons of nouns and adjectives, and on a new dataset we introduce.* Work carried out during an internship at Google.† Work carried out during employment at Google. 1 The resource is available at https:// github.com/google-research-datasets/ distribution-over-quantities Yanai Elazar, Abhijit Mahabal, Deepak Ramachandran, Tania Bedrax-Weiss, Dan Roth 0001 |
ACL (1) | 5 |
| 2019 | Evidence-based TrustworthinessabstractThe information revolution brought with it information pollution.Information retrieval and extraction help us cope with abundant information from diverse sources.But some sources are of anonymous authorship, and some are of uncertain accuracy, so how can we determine what we should actually believe?Not all information sources are equally trustworthy, and simply accepting the majority view is often wrong.This paper develops a general framework for estimating the trustworthiness of information sources in an environment where multiple sources provide claims and supporting evidence, and each claim can potentially be produced by multiple sources.We consider two settings: one in which information sources directly assert claims, and a more realistic and challenging one, in which claims are inferred from evidence provided by sources, via (possibly noisy) NLP techniques.Our key contribution is to develop a family of probabilistic models that jointly estimate the trustworthiness of sources, and the credibility of claims they assert.This is done while accounting for the (possibly noisy) NLP needed to infer claims from evidence supplied by sources.We evaluate our framework on several datasets, showing strong results and significant improvement over baselines. Yi Zhang 0001, Zachary G. Ives, Dan Roth 0001 |
ACL (1) | 3 |
| 2019 | A General-Purpose Algorithm for Constrained Sequential InferenceabstractInference in structured prediction involves finding the best output structure for an input, subject to certain constraints.Many current approaches use sequential inference, which constructs the output in a left-to-right manner.However, there is no general framework to specify constraints in these approaches.We present a principled approach for incorporating constraints into sequential inference algorithms.Our approach expresses constraints using an automaton, which is traversed in lockstep during inference, guiding the search to valid outputs.We show that automata can express commonly used constraints and are easily incorporated into sequential inference.When it is more natural to represent constraints as a set of automata, our algorithm uses an active set method for demonstrably fast and efficient inference.We experimentally show the benefits of our algorithm on constituency parsing and semantic role labeling.For parsing, unlike unconstrained approaches, our algorithm always generates valid output, incurring only a small drop in performance.For semantic role labeling, imposing constraints using our algorithm corrects common errors, improving F 1 by 1.5 points.These benefits increase in low-resource settings.Our active set method achieves a 5.2x relative speedup over a naive approach.1 Daniel Deutsch, Shyam Upadhyay, Dan Roth 0001 |
CoNLL | 3 |
| 2019 | Named Entity Recognition with Partially Annotated Training DataabstractSupervised machine learning assumes the availability of fully-labeled data, but in many cases, such as low-resource languages, the only data available is partially annotated.We study the problem of Named Entity Recognition (NER) with partially annotated training data in which a fraction of the named entities are labeled, and all other tokens, entities or otherwise, are labeled as non-entity by default.In order to train on this noisy dataset, we need to distinguish between the true and false negatives.To this end, we introduce a constraintdriven iterative algorithm that learns to detect false negatives in the noisy set and downweigh them, resulting in a weighted training set.With this set, we train a weighted NER model.We evaluate our algorithm with weighted variants of neural and non-neural NER models on data in 8 languages from several language and script families, showing strong ability to learn from partial data.Finally, to show real-world efficacy, we evaluate on a Bengali NER corpus annotated by non-speakers, outperforming the prior state-of-the-art by over 5 points F1. Stephen Mayhew 0001, Snigdha Chaturvedi, Chen-Tse Tsai, Dan Roth 0001 |
CoNLL | 4 |
| 2019 | KnowSemLM: A Knowledge Infused Semantic Language ModelabstractStory understanding requires developing expectations of what events come next in text.Prior knowledge -both statistical and declarative -is essential in guiding such expectations.While existing semantic language models (SemLM) capture event co-occurrence information by modeling event sequences as semantic frames, entities, and other semantic units, this paper aims at augmenting them with causal knowledge (i.e., one event is likely to lead to another).Such knowledge is modeled at the frame and entity level, and can be obtained either statistically from text or stated declaratively.The proposed method, KnowSemLM 1 , infuses this knowledge into a semantic LM by joint training and inference, and is shown to be effective on both the event cloze test and story/referent prediction tasks. Haoruo Peng, Qiang Ning, Dan Roth 0001 |
CoNLL | 3 |
| 2019 | Evidence Sentence Extraction for Machine Reading ComprehensionabstractRemarkable success has been achieved in the last few years on some limited machine reading comprehension (MRC) tasks.However, it is still difficult to interpret the predictions of existing MRC models.In this paper, we focus on extracting evidence sentences that can explain or support the answers of multiplechoice MRC tasks, where the majority of answer options cannot be directly extracted from reference documents.Due to the lack of ground truth evidence sentence labels in most cases, we apply distant supervision to generate imperfect labels and then use them to train an evidence sentence extractor.To denoise the noisy labels, we apply a recently proposed deep probabilistic logic learning framework to incorporate both sentence-level and cross-sentence linguistic indicators for indirect supervision.We feed the extracted evidence sentences into existing MRC models and evaluate the end-to-end performance on three challenging multiplechoice MRC datasets: MultiRC, RACE, and DREAM, achieving comparable or better performance than the same models that take as input the full reference document.To the best of our knowledge, this is the first work extracting evidence sentences for multiple-choice MRC. Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001, David A. McAllester, Dan Roth 0001 |
CoNLL | 7 |
| 2019 | Summary Cloze: A New Task for Content Selection in Topic-Focused SummarizationabstractDaniel Deutsch, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Daniel Deutsch, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | ner and pos when nothing is capitalizedabstractStephen Mayhew, Tatiana Tsygankova, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Stephen Mayhew 0001, Tatiana Tsygankova, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | An Improved Neural Baseline for Temporal Relation ExtractionabstractQiang Ning, Sanjay Subramanian, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qiang Ning, Sanjay Subramanian, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Benchmarking Zero-shot Text Classification: Datasets, Evaluation and Entailment ApproachabstractWenpeng Yin, Jamaal Hay, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wenpeng Yin 0001, Jamaal Hay, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | "Going on a vacation" takes longer than "Going for a walk": A Study of Temporal Commonsense UnderstandingabstractBen Zhou, Daniel Khashabi, Qiang Ning, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ben Zhou, Daniel Khashabi, Qiang Ning, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Randomized Greedy Search for Structured Prediction: Amortized Inference and LearningabstractIn a structured prediction problem, we need to learn a predictor that can produce a structured output given a structured input (e.g., part-of-speech tagging). The key learning and inference challenge is due to the exponential size of the structured output space. This paper makes four contributions towards the goal of a computationally-efficient inference and training approach for structured prediction that allows to employ complex models and to optimize for non-decomposable loss functions. First, we define a simple class of randomized greedy search (RGS) based inference procedures that leverage classification algorithms for simple outputs. Second, we develop a RGS specific learning approach for amortized inference that can quickly produce high-quality outputs for a given set of structured inputs. Third, we plug our amortized RGS inference solver inside the inner loop of parameter-learning algorithms (e.g., structured SVM) to improve the speed of training. Fourth, we perform extensive experiments on diverse structured prediction tasks. Results show that our proposed approach is competitive or better than many state-of-the-art approaches in spite of its simplicity. Chao Ma 0001, F. A. Rezaur Rahman Chowdhury, Aryan Deshwal, Md. Rakibul Islam 0001, Janardhan Rao Doppa, Dan Roth 0001 |
IJCAI | 6 |
| 2019 | Learning and Inference for Structured Prediction: A Unifying PerspectiveabstractIn a structured prediction problem, one needs to learn a predictor that, given a structured input, produces a structured object, such as a sequence, tree, or clustering output. Prototypical structured prediction tasks include part-of-speech tagging (predicting POS tag sequence for an input sentence) and semantic segmentation of images (predicting semantic labels for pixels of an input image). Unlike simple classification problems, here there is a need to assign values to multiple output variables accounting for the dependencies between them. Consequently, the prediction step itself (aka ``inference" or ``decoding") is computationally-expensive, and so is the learning process, that typically requires making predictions as part of it. The key learning and inference challenge is due to the exponential size of the structured output space and depend on its complexity. In this paper, we present a unifying perspective of the different frameworks that address structured prediction problems and compare them in terms of their strengths and weaknesses. We also discuss important research directions including integration of deep learning advances into structured prediction, and learning from weakly supervised signals and active querying to overcome the challenges of building structured predictors from small amount of labeled data. Aryan Deshwal, Janardhan Rao Doppa, Dan Roth 0001 |
IJCAI | 3 |
| 2019 | DiAd: Domain Adaptation for Learning at ScaleabstractMassive online courses occupy an important place in the educational landscape of today. We study an approach to scale predictive analytic models derived from online course discussion fora--specifically that of confusion detection--onto other courses. The primary challenge here is the lack of labeled examples in a new course and this calls for unsupervised domain adaptation (DA). As a first step in exploring DA in the education domain, we propose a simple algorithm, DiAd, which adapts a classifier trained on a course with labeled data by selectively choosing instances from a new course (with no labeled data) that are most dissimilar to the course with labeled data and on which the classifier is very confident of classification. Our algorithm is empirically validated on the confusion detection task across multiple online courses. We find that DiAd outperforms other methods on the target domain, while showing a comparable performance to a popular method that uses labeled data from the target domain. Ziheng Zeng, Snigdha Chaturvedi, Suma Bhat, Dan Roth 0001 |
LAK | 4 |
| 2019 | Toward any-language zero-shot topic classification of textual documents
Yangqiu Song, Shyam Upadhyay, Haoruo Peng, Stephen Mayhew 0001, Dan Roth 0001 |
Artif. Intell. | 5 |
| 2019 | Discourse in Multimedia: A Case Study in Extracting Geometry Knowledge from TextbooksabstractTo ensure readability, text is often written and presented with due formatting. These text formatting devices help the writer to effectively convey the narrative. At the same time, these help the readers pick up the structure of the discourse and comprehend the conveyed information. There have been a number of linguistic theories on discourse structure of text. However, these theories only consider unformatted text. Multimedia text contains rich formatting features that can be leveraged for various NLP tasks. In this article, we study some of these discourse features in multimedia text and what communicative function they fulfill in the context. As a case study, we use these features to harvest structured subject knowledge of geometry from textbooks. We conclude that the discourse and text layout features provide information that is complementary to lexical semantic information. Finally, we show that the harvested structured knowledge can be used to improve an existing solver for geometry problems, making it more accurate as well as more explainable. Mrinmaya Sachan, Avinava Dubey, Eduard H. Hovy, Tom M. Mitchell, Dan Roth 0001, Eric P. Xing |
Comput. Linguistics | 5 |
| 2019 | Planning with actively eliciting preferences
Mayukh Das, Phillip Odom, Md. Rakibul Islam 0001, Janardhan Rao Doppa, Dan Roth 0001, Sriraam Natarajan |
Knowl. Based Syst. | 5 |
| 2019 | Grammar Error Correction in Morphologically-Rich Languages: The Case of RussianabstractAbstract Until now, most of the research in grammar error correction focused on English, and the problem has hardly been explored for other languages. We address the task of correcting writing mistakes in morphologically rich languages, with a focus on Russian. We present a corrected and error-tagged corpus of Russian learner writing and develop models that make use of existing state-of-the-art methods that have been well studied for English. Although impressive results have recently been achieved for grammar error correction of non-native English writing, these results are limited to domains where plentiful training data are available. Because annotation is extremely costly, these approaches are not suitable for the majority of domains and languages. We thus focus on methods that use “minimal supervision”; that is, those that do not rely on large amounts of annotated training data, and show how existing minimal-supervision approaches extend to a highly inflectional language such as Russian. The results demonstrate that these methods are particularly useful for correcting mistakes in grammatical phenomena that involve rich morphology. Alla Rozovskaya, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | Question Answering as Global Reasoning Over Semantic AbstractionsabstractWe propose a novel method for exploiting the semantic structure of text to answer multiple-choice questions. The approach is especially suitable for domains that require reasoning over a diverse set of linguistic constructs but have limited training data. To address these challenges, we present the first system, to the best of our knowledge, that reasons over a wide range of semantic abstractions of the text, which are derived using off-the-shelf, general-purpose, pre-trained natural language modules such as semantic role labelers, coreference resolvers, and dependency parsers. Representing multiple abstractions as a family of graphs, we translate question answering (QA) into a search for an optimal subgraph that satisfies certain global and local properties. This formulation generalizes several prior structured QA systems. Our system, SEMANTICILP, demonstrates strong performance on two domains simultaneously. In particular, on a collection of challenging science QA datasets, it outperforms various state-of-the-art approaches, including neural models, broad coverage information retrieval, and specialized techniques using structured knowledge bases, by 2%-6%. Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
AAAI | 4 |
| 2018 | Learning Better Name Translation for Cross-Lingual WikificationabstractA notable challenge in cross-lingual wikification is the problem of retrieving English Wikipedia title candidates given a non-English mention, a step that requires translating names written in a foreign language into English. Creating training data for name translation requires significant amount of human efforts. In order to cover as many languages as possible, we propose a probabilistic model that leverages indirect supervision signals in a knowledge base. More specifically, the model learns name translation from title pairs obtained from the inter-language links in Wikipedia. The model jointly considers word alignment and word transliteration. Comparing to 6 other approaches on 9 languages, we show that the proposed model outperforms others not only on the transliteration metric, but also on the ability to generate target English titles for a cross-lingual wikifier. Consequently, as we show, it improves the end-to-end performance of a cross-lingual wikifier on the TAC 2016 EDL dataset. Chen-Tse Tsai, Dan Roth 0001 |
AAAI | 2 |
| 2018 | A Distributional and Orthographic Aggregation Model for English Derivational MorphologyabstractModeling derivational morphology to generate words with particular semantics is useful in many text generation tasks, such as machine translation or abstractive question answering.In this work, we tackle the task of derived word generation.That is, given the word "run," we attempt to generate the word "runner" for "someone who runs."We identify two key problems in generating derived words from root words and transformations: suffix ambiguity and orthographic irregularity.We contribute a novel aggregation model of derived word generation that learns derivational transformations both as orthographic functions using sequence-to-sequence models and as functions in distributional word embedding space.Our best open-vocabulary model, which can generate novel words, and our best closed-vocabulary model, show 22% and 37% relative error reductions over current state-of-the-art systems on the same dataset. Daniel Deutsch, John Hewitt, Dan Roth 0001 |
ACL (1) | 3 |
| 2018 | A Multi-Axis Annotation Scheme for Event Temporal RelationsabstractExisting temporal relation (TempRel) annotation schemes often have low interannotator agreements (IAA) even between experts, suggesting that the current annotation task needs a better definition.This paper proposes a new multi-axis modeling to better capture the temporal structure of events.In addition, we identify that event end-points are a major source of confusion in annotation, so we also propose to annotate TempRels based on start-points only.A pilot expert annotation effort using the proposed scheme shows significant improvement in IAA from the conventional 60's to 80's (Cohen's Kappa).This better-defined annotation scheme further enables the use of crowdsourcing to alleviate the labor intensity for each annotator.We hope that this work can foster more interesting studies towards event understanding. 1 Qiang Ning, Hao Wu 0034, Dan Roth 0001 |
ACL (1) | 3 |
| 2018 | Joint Reasoning for Temporal and Causal RelationsabstractUnderstanding temporal and causal relations between events is a fundamental natural language understanding task.Because a cause must occur earlier than its effect, temporal and causal relations are closely related and one relation often dictates the value of the other.However, limited attention has been paid to studying these two relations jointly.This paper presents a joint inference framework for them using constrained conditional models (CCMs).Specifically, we formulate the joint problem as an integer linear programming (ILP) problem, enforcing constraints that are inherent in the nature of time and causality.We show that the joint inference framework results in statistically significant improvement in the extraction of both temporal and causal relations from text. 1 Qiang Ning, Zhili Feng, Hao Wu 0034, Dan Roth 0001 |
ACL (1) | 4 |
| 2018 | Gold Standard Annotations for Preposition and Verb Sense with Semantic Role Labels in Adult-Child InteractionsabstractThis paper describes the augmentation of an existing corpus of child-directed speech. The resulting corpus is a gold-standard labeled corpus for supervised learning of semantic role labels in adult-child dialogues. Semantic role labeling (SRL) models assign semantic roles to sentence constituents, thus indicating who has done what to whom (and in what way). The current corpus is derived from the Adam files in the Brown corpus (Brown 1973) of the CHILDES corpora, and augments the partial annotation described in Connor et al. (2010). It provides labels for both semantic arguments of verbs and semantic arguments of prepositions. The semantic role labels and senses of verbs follow Propbank guidelines Kingsbury and Palmer, 2002; Gildea and Palmer 2002; Palmer et al., 2005) and those for prepositions follow Srikumar and Roth (2011). The corpus was annotated by two annotators. Inter-annotator agreement is given separately for prepositions and verbs, and for adult speech and child speech. Overall, across child and adult samples, including verbs and prepositions, the kappa score for sense is 72.6, for the number of semantic-role-bearing arguments, the kappa score is 77.4, for identical semantic role labels on a given argument, the kappa score is 91.1, for the span of semantic role labels, and the kappa for agreement is 93.9. The sense and number of arguments was often open to multiple interpretations in child speech, due to the rapidly changing discourse and omission of constituents in production. Annotators used a discourse context window of ten sentences before and ten sentences after the target utterance to determine the annotation labels. The derived corpus is available for use in CHAT (MacWhinney, 2000) and XML format. Lori Moon, Christos Christodoulopoulos 0001, Cynthia Fisher, Sandra Franco, Dan Roth 0001 |
COLING | 5 |
| 2018 | TwoWingOS: A Two-Wing Optimization Strategy for Evidential Claim VerificationabstractDetermining whether a given claim is supported by evidence is a fundamental NLP problem that is best modeled as Textual Entailment.However, given a large collection of text, finding evidence that could support or refute a given claim is a challenge in itself, amplified by the fact that different evidence might be needed to support or refute a claim.Nevertheless, most prior work decouples evidence identification from determining the truth value of the claim given the evidence.We propose to consider these two aspects jointly.We develop TWOWINGOS (twowing optimization strategy), a system that, while identifying appropriate evidence for a claim, also determines whether or not the claim is supported by the evidence.Given the claim, TWOWINGOS attempts to identify a subset of the evidence candidates; given the predicted evidence, it then attempts to determine the truth value of the corresponding claim.We treat this challenge as coupled optimization problems, training a joint model for it.TWOWINGOS offers two advantages: (i) Unlike pipeline systems, it facilitates flexible-size evidence set, and (ii) Joint training improves both the claim verification and the evidence identification.Experiments on a benchmark dataset show state-of-the-art performance.1 Wenpeng Yin 0001, Dan Roth 0001 |
EMNLP | 2 |
| 2018 | Joint Multilingual Supervision for Cross-lingual Entity LinkingabstractCross-lingual Entity Linking (XEL) aims to ground entity mentions written in any language to an English Knowledge Base (KB), such as Wikipedia.XEL for most languages is challenging, owing to limited availability of resources as supervision.We address this challenge by developing the first XEL approach that combines supervision from multiple languages jointly.This enables our approach to: (a) augment the limited supervision in the target language with additional supervision from a high-resource language (like English), and (b) train a single entity linking model for multiple languages, improving upon individually trained models for each language.Extensive evaluation on three benchmark datasets across 8 languages shows that our approach significantly improves over the current state-of-theart.We also provide analyses in two limited resource settings: (a) zero-shot setting, when no supervision in the target language is available, and in (b) low-resource setting, when some supervision in the target language is available.Our analysis provides insights into the limitations of zero-shot XEL approaches in realistic scenarios, and shows the value of joint supervision in low-resource settings.1 Shyam Upadhyay, Nitish Gupta, Dan Roth 0001 |
EMNLP | 3 |
| 2018 | Bootstrapping Transliteration with Guided Discovery for Low-Resource LanguagesabstractGenerating the English transliteration of a name written in a foreign script is an important and challenging step in multilingual knowledge acquisition and information extraction.Existing approaches to transliteration generation require a large (>5000) number of training examples.This difficulty contrasts with transliteration discovery, a somewhat easier task that involves picking a plausible transliteration from a given list.In this work, we present a bootstrapping algorithm that uses constrained discovery to improve generation, and can be used with as few as 500 training examples, which we show can be sourced from annotators in a matter of hours.This opens the task to languages for which large number of training examples are unavailable.We evaluate transliteration generation performance itself, as well the improvement it brings to crosslingual candidate generation for entity linking, a typical downstream task.We present a comprehensive evaluation of our approach on nine languages, each written in a unique script. 1 Shyam Upadhyay, Jordan Kodner, Dan Roth 0001 |
EMNLP | 3 |
| 2018 | On the Strength of Character Language Models for Multilingual Named Entity RecognitionabstractCharacter-level patterns have been widely used as features in English Named Entity Recognition (NER) systems.However, to date there has been no direct investigation of the inherent differences between name and nonname tokens in text, nor whether this property holds across multiple languages.This paper analyzes the capabilities of corpus-agnostic Character-level Language Models (CLMs) in the binary task of distinguishing name tokens from non-name tokens.We demonstrate that CLMs provide a simple and powerful model for capturing these differences, identifying named entity tokens in a diverse set of languages at close to the performance of full NER systems.Moreover, by adding very simple CLM-based features we can significantly improve the performance of an off-the-shelf NER system for multiple languages.1 Xiaodong Yu 0003, Stephen Mayhew 0001, Mark Sammons, Dan Roth 0001 |
EMNLP | 4 |
| 2018 | Zero-Shot Open Entity Typing as Type-Compatible GroundingabstractThe problem of entity-typing has been studied predominantly in supervised learning fashion, mostly with task-specific annotations (for coarse types) and sometimes with distant supervision (for fine types).While such approaches have strong performance within datasets, they often lack the flexibility to transfer across text genres and to generalize to new type taxonomies.In this work we propose a zero-shot entity typing approach that requires no annotated data and can flexibly identify newly defined types.Given a type taxonomy defined as Boolean functions of FREEBASE "types", we ground a given mention to a set of type-compatible Wikipedia entries and then infer the target mention's types using an inference algorithm that makes use of the types of these entries.We evaluate our system on a broad range of datasets, including standard fine-grained and coarse-grained entity typing datasets, and also a dataset in the biological domain.Our system is shown to be competitive with state-of-theart supervised NER systems and outperforms them on out-of-domain datasets.We also show that our system significantly outperforms other zero-shot fine typing systems. Ben Zhou, Daniel Khashabi, Chen-Tse Tsai, Dan Roth 0001 |
EMNLP | 4 |
| 2018 | Systems AI: A Declarative Learning Based Programming PerspectiveabstractData-driven approaches are becoming dominant problem-solving techniques in many areas of research and industry. Unfortunately, current technologies do not make such techniques easy to use for application experts who are not fluent in machine learning nor for machine learning experts who aim at testing ideas on real-world data and need to evaluate those as a part of an end-to-end system. We review key efforts made by various AI communities to provide languages for high-level abstractions over learning and reasoning techniques needed for designing complex AI systems. We classify the existing frameworks based on the type of techniques as well as the data and knowledge representations they use, provide a comparative study of the way they address the challenges of programming real-world applications, and highlight some shortcomings and future directions. Parisa Kordjamshidi, Dan Roth 0001, Kristian Kersting |
IJCAI | 2 |
| 2018 | CogCompNLP: Your Swiss Army Knife for NLP
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos 0001, Vivek Srikumar, Nick Rizzolo, Lev-Arie Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew 0001, Zhili Feng, John Wieting, Xiaodong Yu 0003, Yangqiu Song, Shashank Gupta 0007, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth 0001 |
LREC | 23 |
| 2018 | Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple SentencesabstractDaniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, Dan Roth. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Daniel Khashabi, Snigdha Chaturvedi, Michael Roth 0001, Shyam Upadhyay, Dan Roth 0001 |
NAACL-HLT | 5 |
| 2018 | Improving Temporal Relation Extraction with a Globally Acquired Statistical ResourceabstractQiang Ning, Hao Wu, Haoruo Peng, Dan Roth. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Qiang Ning, Hao Wu 0034, Haoruo Peng, Dan Roth 0001 |
NAACL-HLT | 4 |
| 2018 | Robust Cross-Lingual Hypernymy Detection Using Dependency ContextabstractShyam Upadhyay, Yogarshi Vyas, Marine Carpuat, Dan Roth. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shyam Upadhyay, Yogarshi Vyas, Marine Carpuat, Dan Roth 0001 |
NAACL-HLT | 4 |
| 2018 | Learning Pipelines with Limited Data and Domain Knowledge: A Study in Parsing Physics ProblemsabstractAs machine learning becomes more widely used in practice, we need new methods to build complex intelligent systems that integrate learning with existing software, and with domain knowledge encoded as rules. As a case study, we present such a system that learns to parse Newtonian physics problems in textbooks. This system, Nuts&Bolts, learns a pipeline process that incorporates existing code, pre-learned machine learning models, and human engineered rules. It jointly trains the entire pipeline to prevent propagation of errors, using a combination of labelled and unlabelled data. Our approach achieves a good performance on the parsing task, outperforming the simple pipeline and its variants. Finally, we also show how Nuts&Bolts can be used to achieve improvements on a relation extraction task and on the end task of answering Newtonian physics problems. Mrinmaya Sachan, Avinava Dubey, Tom M. Mitchell, Dan Roth 0001, Eric P. Xing |
NeurIPS | 4 |
| 2018 | Illinois CCG LoReHLT 2016 named entity recognition and situation frame systems
Chen-Tse Tsai, Stephen Mayhew 0001, Yangqiu Song, Mark Sammons, Dan Roth 0001 |
Mach. Transl. | 5 |
| 2018 | Mapping to Declarative Knowledge for Word Problem SolvingabstractMath word problems form a natural abstraction to a range of quantitative reasoning problems, such as understanding financial news, sports results, and casualties of war. Solving such problems requires the understanding of several mathematical concepts such as dimensional analysis, subset relationships, etc. In this paper, we develop declarative rules which govern the translation of natural language description of these concepts to math expressions. We then present a framework for incorporating such declarative knowledge into word problem solving. Our method learns to map arithmetic word problem text to math expressions, by learning to select the relevant declarative knowledge for each operation of the solution expression. This provides a way to handle multiple concepts in the same problem while, at the same time, supporting interpretability of the answer expression. Our method models the mapping to declarative knowledge as a latent variable, thus removing the need for expensive annotations. Experimental evaluation suggests that our domain knowledge based solver outperforms all other systems, and that it generalizes better in the realistic case where the training data it is exposed to is biased in a different way than the test data. Subhro Roy, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | Incidental Supervision: Moving beyond Supervised LearningabstractMachine Learning and Inference methods have become ubiquitous in our attempt to induce more abstract representations of natural language text, visual scenes, and other messy, naturally occurring data, and support decisions that depend on it. However, learning models for these tasks is difficult partly because generating the necessary supervision signals for it is costly and does not scale. This paper describes several learning paradigms that are designed to alleviate the supervision bottleneck. It will illustrate their benefit in the context of multiple problems, all pertaining to inducing various levels of semantic representations from text. In particular, we discuss (i) esponse Driven Learning of models, a learning protocol that supports inducing meaning representations simply by observing the model's behavior in its environment, (ii) the exploitation of Incidental Supervision signals that exist in the data, independently of the task at hand, to learn models that identify and classify semantic predicates, and (iii) the use of weak supervision to combine simple models to support global decisions where joint supervision is not available. Dan Roth 0001 |
AAAI | 1 |
| 2017 | Unit Dependency Graph and Its Application to Arithmetic Word Problem SolvingabstractMath word problems provide a natural abstraction to a range of natural language understanding problems that involve reasoning about quantities, such as interpreting election results, news about casualties, and the financial section of a newspaper. Units associated with the quantities often provide information that is essential to support this reasoning. This paper proposes a principled way to capture and reason about units and shows how it can benefit an arithmetic word problem solver. This paper presents the concept of Unit Dependency Graphs (UDGs), which provides a compact representation of the dependencies between units of numbers mentioned in a given problem. Inducing the UDG alleviates the brittleness of the unit extraction system and allows for a natural way to leverage domain knowledge about unit compatibility, for word problem solving. We introduce a decomposed model for inducing UDGs with minimal additional annotations, and use it to augment the expressions used in the arithmetic word problem solver of (Roy and Roth 2015) via a constrained inference framework. We show that introduction of UDGs reduces the error of the solver by over 10 %, surpassing all existing systems for solving arithmetic word problems. In addition, it also makes the system more robust to adaptation to new vocabulary and equation forms . Subhro Roy, Dan Roth 0001 |
AAAI | 2 |
| 2017 | Learning What is Essential in QuestionsabstractQuestion answering (QA) systems are easily distracted by irrelevant or redundant words in questions, especially when faced with long or multi-sentence questions in difficult domains. This paper introduces and studies the notion of essential question terms with the goal of improving such QA solvers. We illustrate the importance of essential question terms by showing that humans' ability to answer questions drops significantly when essential terms are eliminated from questions.We then develop a classifier that reliably (90% mean average precision) identifies and ranks essential terms in questions. Finally, we use the classifier to demonstrate that the notion of question term essentiality allows state-of-the-art QA solver for elementary-level science questions to make better and more informed decisions,improving performance by up to 5%.We also introduce a new dataset of over 2,200 crowd-sourced essential terms annotated science questions. Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
CoNLL | 4 |
| 2017 | A Joint Model for Semantic Sequences: Frames, Entities, SentimentsabstractUnderstanding stories -sequences of events -is a crucial yet challenging natural language understanding task.These events typically carry multiple aspects of semantics including actions, entities and emotions.Not only does each individual aspect contribute to the meaning of the story, so does the interaction among these aspects.Building on this intuition, we propose to jointly model important aspects of semantic knowledge -frames, entities and sentiments -via a semantic language model.We achieve this by first representing these aspects' semantic units at an appropriate level of abstraction and then using the resulting vector representations for each semantic aspect to learn a joint representation via a neural language model.We show that the joint semantic language model is of high quality and can generate better semantic sequences than models that operate on the word level.We further demonstrate that our joint model can be applied to story cloze test and shallow discourse parsing tasks with improved performance and that each semantic aspect contributes to the model. Haoruo Peng, Snigdha Chaturvedi, Dan Roth 0001 |
CoNLL | 3 |
| 2017 | Story Comprehension for Predicting What Happens NextabstractAutomatic story comprehension is a fundamental challenge in Natural Language Understanding, and can enable computers to learn about social norms, human behavior and commonsense.In this paper, we present a story comprehension model that explores three distinct semantic aspects: (i) the sequence of events described in the story, (ii) its emotional trajectory, and (iii) its plot consistency.We judge the model's understanding of real-world stories by inquiring if, like humans, it can develop an expectation of what will happen next in a given story.Specifically, we use it to predict the correct ending of a given short story from possible alternatives.The model uses a hidden variable to weigh the semantic aspects in the context of the story.Our experiments demonstrate the potential of our approach to characterize these semantic aspects, and the strength of the hidden variable based approach.The model outperforms the stateof-the-art approaches and achieves best results on a publicly available dataset. Snigdha Chaturvedi, Haoruo Peng, Dan Roth 0001 |
EMNLP | 3 |
| 2017 | Entity Linking via Joint Encoding of Types, Descriptions, and ContextabstractFor accurate entity linking, we need to capture various information aspects of an entity, such as its description in a KB, contexts in which it is mentioned, and structured knowledge.Additionally, a linking system should work on texts from different domains without requiring domain-specific training data or hand-engineered features.In this work we present a neural, modular entity linking system that learns a unified dense representation for each entity using multiple sources of information, such as its description, contexts around its mentions, and its fine-grained types.We show that the resulting entity linking system is effective at combining these sources, and performs competitively, sometimes out-performing current state-of-theart systems across datasets, without requiring any domain-specific training data or hand-engineered features.We also show that our model can effectively "embed" entities that are new to the KB, and is able to link its mentions accurately. Nitish Gupta, Sameer Singh 0001, Dan Roth 0001 |
EMNLP | 3 |
| 2017 | Cheap Translation for Cross-Lingual Named Entity RecognitionabstractRecent work in NLP has attempted to deal with low-resource languages but still assumed a resource level that is not present for most languages, e.g., the availability of Wikipedia in the target language. We propose a simple method for cross-lingual named entity recognition (NER) that works well in settings with very minimal resources. Our approach makes use of a lexicon to "translate" annotated data available in one or several high resource language(s) into the target language, and learns a standard monolingual NER model there. Further, when Wikipedia is available in the target language, our method can enhance Wikipedia based methods to yield state-of-the-art NER results; we evaluate on 7 diverse languages, improving the state-of-the-art by an average of 5.5% F1 points. With the minimal resources required, this is an extremely portable cross-lingual NER approach, as illustrated using a truly low-resource language, Uyghur. Stephen Mayhew 0001, Chen-Tse Tsai, Dan Roth 0001 |
EMNLP | 3 |
| 2017 | A Structured Learning Approach to Temporal Relation ExtractionabstractIdentifying temporal relations between events is an essential step towards natural language understanding.However, the temporal relation between two events in a story depends on, and is often dictated by, relations among other events.Consequently, effectively identifying temporal relations between events is a challenging problem even for human annotators.This paper suggests that it is important to take these dependencies into account while learning to identify these relations and proposes a structured learning approach to address this challenge.As a byproduct, this provides a new perspective on handling missing relations, a known issue that hurts existing methods.As we show, the proposed approach results in significant improvements on the two commonly used data sets for this problem. Qiang Ning, Zhili Feng, Dan Roth 0001 |
EMNLP | 3 |
| 2017 | Adapting to Learner Errors with Minimal SupervisionabstractThis article considers the problem of correcting errors made by English as a Second Language writers from a machine learning perspective, and addresses an important issue of developing an appropriate training paradigm for the task, one that accounts for error patterns of non-native writers using minimal supervision. Existing training approaches present a trade-off between large amounts of cheap data offered by the native-trained models and additional knowledge of learner error patterns provided by the more expensive method of training on annotated learner data. We propose a novel training approach that draws on the strengths offered by the two standard training paradigms—of training either on native or on annotated learner data—and that outperforms both of these standard methods. Using the key observation that parameters relating to error regularities exhibited by non-native writers are relatively simple, we develop models that can incorporate knowledge about error regularities based on a small annotated sample but that are otherwise trained on native English data. The key contribution of this article is the introduction and analysis of two methods for adapting the learned models to error patterns of non-native writers; one method that applies to generative classifiers and a second that applies to discriminative classifiers. Both methods demonstrated state-of-the-art performance in several text correction competitions. In particular, the Illinois system that implements these methods ranked at the top in two recent CoNLL shared tasks on error correction.1We conduct further evaluation of the proposed approaches studying the effect of using error data from speakers of the same native language, languages that are closely related linguistically, and unrelated languages.2 Alla Rozovskaya, Dan Roth 0001, Mark Sammons |
Comput. Linguistics | 2 |
| 2016 | Labeling the Semantic Roles of CommasabstractCommas and the surrounding sentence structure often express relations that are essential to understanding the meaning of the sentence. This paper proposes a set of relations commas participate in, expanding on previous work in this area, and develops a new dataset annotated with this set of labels. We identify features that are important to achieve a good performance on comma labeling and then develop a machine learning method that achieves high accuracy on identifying comma relations, improving over previous work. Finally, we discuss a variety of possible uses, both as syntactic and discourse-oriented features and constraints for downstream tasks. Naveen Arivazhagan, Christos Christodoulopoulos 0001, Dan Roth 0001 |
AAAI | 3 |
| 2016 | Two Discourse Driven Language Models for SemanticsabstractNatural language understanding often requires deep semantic knowledge.Expanding on previous proposals, we suggest that some important aspects of semantic knowledge can be modeled as a language model if done at an appropriate level of abstraction.We develop two distinct models that capture semantic frame chains and discourse information while abstracting over the specific mentions of predicates and entities.For each model, we investigate four implementations: a "standard" N-gram language model and three discriminatively trained "neural" language models that generate embeddings for semantic frames.The quality of the semantic language models (SemLM) is evaluated both intrinsically, using perplexity and a narrative cloze test and extrinsically -we show that our SemLM helps improve performance on semantic natural language processing tasks such as co-reference resolution and discourse parsing. Haoruo Peng, Dan Roth 0001 |
ACL (1) | 2 |
| 2016 | Grammatical Error Correction: Machine Translation and ClassifiersabstractWe focus on two leading state-of-the-art approaches to grammatical error correction -machine learning classification and machine translation.Based on the comparative study of the two learning frameworks and through error analysis of the output of the state-of-the-art systems, we identify key strengths and weaknesses of each of these approaches and demonstrate their complementarity.In particular, the machine translation method learns from parallel data without requiring further linguistic input and is better at correcting complex mistakes.The classification approach possesses other desirable characteristics, such as the ability to easily generalize beyond what was seen in training, the ability to train without human-annotated data, and the flexibility to adjust knowledge sources for individual error types.Based on this analysis, we develop an algorithmic approach that combines the strengths of both methods.We present several systems based on resources used in previous work with a relative improvement of over 20% (and 7.4 F score points) over the previous state-of-the-art.System Method Performance P R F0.5 CoNLL-2014 top 3 MT 41.62 21.40 35.01 CoNLL-2014 top 2 Classif.41.78 24.88 36.79CoNLL-2014 top 1 MT, rules 39.71 30.10 37.33 Susanto et al. (2014) MT, classif.53.55 19.14 39.39 Miz.& Mats.(2016) MT 45.80 26.60 40.00This work MT, classif.60.17 25.64 47.40 Alla Rozovskaya, Dan Roth 0001 |
ACL (1) | 2 |
| 2016 | Cross-lingual Models of Word Embeddings: An Empirical ComparisonabstractDespite interest in using cross-lingual knowledge to learn word embeddings for various tasks, a systematic comparison of the possible approaches is lacking in the literature.We perform an extensive evaluation of four popular approaches of inducing cross-lingual embeddings, each requiring a different form of supervision, on four typologically different language pairs.Our evaluation setup spans four different tasks, including intrinsic evaluation on mono-lingual and cross-lingual similarity, and extrinsic evaluation on downstream semantic and syntactic applications.We show that models which require expensive cross-lingual knowledge almost always perform better, but cheaply supervised models often prove competitive on certain tasks. Shyam Upadhyay, Manaal Faruqui, Chris Dyer, Dan Roth 0001 |
ACL (1) | 4 |
| 2016 | Better call Saul: Flexible Programming for Learning and Inference in NLPabstractWe present a novel way for designing complex joint inference and learning models using Saul (Kordjamshidi et al., 2015), a recently-introduced declarative learning-based programming language (DeLBP). We enrich Saul with components that are necessary for a broad range of learning based Natural Language Processing tasks at various levels of granularity. We illustrate these advances using three different, well-known NLP problems, and show how these generic learning and inference modules can directly exploit Saul’s graph-based data representation. These properties allow the programmer to easily switch between different model formulations and configurations, and consider various kinds of dependencies and correlations among variables of interest with minimal programming effort. We argue that Saul provides an extremely useful paradigm both for the design of advanced NLP systems and for supporting advanced research in NLP. Parisa Kordjamshidi, Daniel Khashabi, Christos Christodoulopoulos 0001, Bhargav Mangipudi, Sameer Singh 0001, Dan Roth 0001 |
COLING | 6 |
| 2016 | Revisiting the Evaluation for Cross Document Event CoreferenceabstractCross document event coreference (CDEC) is an important task that aims at aggregating event-related information across multiple documents. We revisit the evaluation for CDEC, and discover that past works have adopted different, often inconsistent, evaluation settings, which either overlook certain mistakes in coreference decisions, or make assumptions that simplify the coreference task considerably. We suggest a new evaluation methodology which overcomes these limitations, and allows for an accurate assessment of CDEC systems. Our new evaluation setting better reflects the corpus-wide information aggregation ability of CDEC systems by separating event-coreference decisions made across documents from those made within a document. In addition, we suggest a better baseline for the task and semi-automatically identify several inconsistent annotations in the evaluation dataset. Shyam Upadhyay, Nitish Gupta, Christos Christodoulopoulos 0001, Dan Roth 0001 |
COLING | 4 |
| 2016 | Cross-Lingual Named Entity Recognition via WikificationabstractNamed Entity Recognition (NER) models for language L are typically trained using annotated data in that language.We study cross-lingual NER, where a model for NER in L is trained on another, source, language (or multiple source languages).We introduce a language independent method for NER, building on cross-lingual wikification, a technique that grounds words and phrases in non-English text into English Wikipedia entries.Thus, mentions in any language can be described using a set of categories and FreeBase types, yielding, as we show, strong language-independent features.With this insight, we propose an NER model that can be applied to all languages in Wikipedia.When trained on English, our model outperforms comparable approaches on the standard CoNLL datasets (Spanish, German, and Dutch) and also performs very well on lowresource languages (e.g., Turkish, Tagalog, Yoruba, Bengali, and Tamil) that have significantly smaller Wikipedia.Moreover, our method allows us to train on multiple source languages, typically improving NER results on the target languages.Finally, we show that our languageindependent features can be used also to enhance monolingual NER systems, yielding improved results for all 9 languages. Chen-Tse Tsai, Stephen Mayhew 0001, Dan Roth 0001 |
CoNLL | 3 |
| 2016 | Event Detection and Co-reference with Minimal SupervisionabstractAn important aspect of natural language understanding involves recognizing and categorizing events and the relations among them. However, these tasks are quite subtle and annotating training data for machine learning based approaches is an expensive task, resulting in supervised systems that attempt to learn complex models from small amounts of data, which they over-fit. This paper addresses this challenge by developing an event detection and co-reference system with minimal supervision, in the form of a few event examples. We view these tasks as semantic similarity problems between event mentions or event mentions and an ontology of types, thus facilitating the use of large amounts of out of domain text data. Notably, our semantic relatedness function exploits the structure of the text by making use of a semantic-role-labeling based representation of an event. We show that our approach to event detection is competitive with the top supervised methods. More significantly, we outperform state-of-the-art supervised methods for event co-reference on benchmark data sets, and support significantly better transfer across domains. Haoruo Peng, Yangqiu Song, Dan Roth 0001 |
EMNLP | 3 |
| 2016 | Equation Parsing : Mapping Sentences to Grounded EquationsabstractIdentifying mathematical relations expressed in text is essential to understanding a broad range of natural language text from election reports, to financial news, to sport commentaries to mathematical word problems.This paper focuses on identifying and understanding mathematical relations described within a single sentence.We introduce the problem of Equation Parsing -given a sentence, identify noun phrases which represent variables, and generate the mathematical equation expressing the relation described in the sentence.We introduce the notion of projective equation parsing and provide an efficient algorithm to parse text to projective equations.Our system makes use of a high precision lexicon of mathematical expressions and a pipeline of structured predictors, and generates correct equations in 70% of the cases.In 60% of the time, it also identifies the correct noun phrase → variables mapping, significantly outperforming baselines.We also release a new annotated dataset for task evaluation. Subhro Roy, Shyam Upadhyay, Dan Roth 0001 |
EMNLP | 3 |
| 2016 | Question Answering via Integer Programming over Semi-Structured Knowledge
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Peter Clark, Oren Etzioni, Dan Roth 0001 |
IJCAI | 6 |
| 2016 | Cross-Lingual Dataless Classification for Many Languages
Yangqiu Song, Shyam Upadhyay, Haoruo Peng, Dan Roth 0001 |
IJCAI | 4 |
| 2016 | EDISON: Feature Extraction for NLP, Simplified
Mark Sammons, Christos Christodoulopoulos 0001, Parisa Kordjamshidi, Daniel Khashabi, Vivek Srikumar, Dan Roth 0001 |
LREC | 6 |
| 2016 | Cross-lingual Wikification Using Multilingual EmbeddingsabstractCross-lingual Wikification is the task of grounding mentions written in non-English documents to entries in the English Wikipedia.This task involves the problem of comparing textual clues across languages, which requires developing a notion of similarity between text snippets across languages.In this paper, we address this problem by jointly training multilingual embeddings for words and Wikipedia titles.The proposed method can be applied to all languages represented in Wikipedia, including those for which no machine translation technology is available.We create a challenging dataset in 12 languages and show that our proposed approach outperforms various baselines.Moreover, our model compares favorably with the best systems on the TAC KBP2015 Entity Linking task including those that relied on the availability of translation from the target language to English. Chen-Tse Tsai, Dan Roth 0001 |
HLT-NAACL | 2 |
| 2016 | Learning invariants using decision trees and implication counterexamplesabstractInductive invariants can be robustly synthesized using a learning model where the teacher is a program verifier who instructs the learner through concrete program configurations, classified as positive, negative, and implications. We propose the first learning algorithms in this model with implication counter-examples that are based on machine learning techniques. In particular, we extend classical decision-tree learning algorithms in machine learning to handle implication samples, building new scalable ways to construct small decision trees using statistical measures. We also develop a decision-tree learning algorithm in this model that is guaranteed to converge to the right concept (invariant) if one exists. We implement the learners and an appropriate teacher, and show that the resulting invariant synthesis is efficient and convergent for a large suite of programs. Pranav Garg 0001, Daniel Neider, P. Madhusudan, Dan Roth 0001 |
POPL | 4 |
| 2016 | Entity Disambiguation with Linkless Knowledge BasesabstractNamed Entity Disambiguation is the task of disambiguating named entity mentions in natural language text and link them to their corresponding entries in a reference knowledge base (e.g. Wikipedia). Such disambiguation can help add semantics to plain text and distinguish homonymous entities. Previous research has tackled this problem by making use of two types of context-aware features derived from the reference knowledge base, namely, the context similarity and the semantic relatedness. Both features heavily rely on the cross-document hyperlinks within the knowledge base: the semantic relatedness feature is directly measured via those hyperlinks, while the context similarity feature implicitly makes use of those hyperlinks to expand entity candidates' descriptions and then compares them against the query context. Unfortunately, cross-document hyperlinks are rarely available in many closed domain knowledge bases and it is very expensive to manually add such links. Therefore few algorithms can work well on linkless knowledge bases. In this work, we propose the challenging Named Entity Disambiguation with Linkless Knowledge Bases (LNED) problem and tackle it by leveraging the useful disambiguation evidences scattered across the reference knowledge base. We propose a generative model to automatically mine such evidences out of noisy information. The mined evidences can mimic the role of the missing links and help boost the LNED performance. Experimental results show that our proposed method substantially improves the disambiguation accuracy over the baseline approaches. Yang Li 0150, Shulong Tan, Huan Sun 0001, Jiawei Han 0001, Dan Roth 0001, Xifeng Yan |
WWW | 5 |
| 2016 | litewi: A combined term extraction and entity linking method for eliciting educational ontologies from textbooksabstractMajor efforts have been conducted on ontology learning, that is, semiautomatic processes for the construction of domain ontologies from diverse sources of information. In the past few years, a research trend has focused on the construction of educational ontologies, that is, ontologies to be used for educational purposes. The identification of the terminology is crucial to build ontologies. Term extraction techniques allow the identification of the domain‐related terms from electronic resources. This paper presents LiTeWi, a novel method that combines current unsupervised term extraction approaches for creating educational ontologies for technology supported learning systems from electronic textbooks. LiTeWi uses Wikipedia as an additional information source. Wikipedia contains more than 30 million articles covering the terminology of nearly every domain in 288 languages, which makes it an appropriate generic corpus for term extraction. Furthermore, given that its content is available in several languages, it promotes both domain and language independence. LiTeWi is aimed at being used by teachers, who usually develop their didactic material from textbooks. To evaluate its performance, LiTeWi was tuned up using a textbook on object oriented programming and then tested with two textbooks of different domains—astronomy and molecular biology. Angel Conde, Mikel Larrañaga, Ana Arruarte Lasa, Jon A. Elorriaga, Dan Roth 0001 |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2016 | Concept Grounding to Multiple Knowledge Bases via Indirect SupervisionabstractWe consider the problem of disambiguating concept mentions appearing in documents and grounding them in multiple knowledge bases, where each knowledge base addresses some aspects of the domain. This problem poses a few additional challenges beyond those addressed in the popular Wikification problem. Key among them is that most knowledge bases do not contain the rich textual and structural information Wikipedia does; consequently, the main supervision signal used to train Wikification rankers does not exist anymore. In this work we develop an algorithmic approach that, by carefully examining the relations between various related knowledge bases, generates an indirect supervision signal it uses to train a ranking model that accurately chooses knowledge base entries for a given mention; moreover, it also induces prior knowledge that can be used to support a global coherent mapping of all the concepts in a given document to the knowledge bases. Using the biomedical domain as our application, we show that our indirectly supervised ranking model outperforms other unsupervised baselines and that the quality of this indirect supervision scheme is very close to a supervised model. We also show that considering multiple knowledge bases together has an advantage over grounding concepts to each knowledge base individually. Chen-Tse Tsai, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | World Knowledge as Indirect Supervision for Document ClusteringabstractOne of the key obstacles in making learning protocols realistic in applications is the need to supervise them, a costly process that often requires hiring domain experts. We consider the framework to use the world knowledge as indirect supervision. World knowledge is general-purpose knowledge, which is not designed for any specific domain. Then, the key challenges are how to adapt the world knowledge to domains and how to represent it for learning. In this article, we provide an example of using world knowledge for domain-dependent document clustering. We provide three ways to specify the world knowledge to domains by resolving the ambiguity of the entities and their types, and represent the data with world knowledge as a heterogeneous information network. Then, we propose a clustering algorithm that can cluster multiple types and incorporate the sub-type information as constraints. In the experiments, we use two existing knowledge bases as our sources of world knowledge. One is Freebase, which is collaboratively collected knowledge about entities and their organizations. The other is YAGO2, a knowledge base automatically extracted from Wikipedia and maps knowledge to the linguistic knowledge base, WordNet. Experimental results on two text benchmark datasets (20newsgroups and RCV1) show that incorporating world knowledge as indirect supervision can significantly outperform the state-of-the-art clustering algorithms as well as clustering algorithms enhanced with world knowledge features. A preliminary version of this work appeared in the proceedings of KDD 2015 [Wang et al. 2015a]. This journal version has made several major improvements. First, we have proposed a new and general learning framework for machine learning with world knowledge as indirect supervision, where document clustering is a special case in the original paper. Second, in order to make our unsupervised semantic parsing method more understandable, we add several real cases from the original sentences to the resulting logic forms with all the necessary information. Third, we add details of the three semantic filtering methods and conduct deep analysis of the three semantic filters, by using case studies to show why the conceptualization-based semantic filter can produce more accurate indirect supervision. Finally, in addition to the experiment on 20 newsgroup data and Freebase, we have extended the experiments on clustering results by using all the combinations of text (20 newsgroup, MCAT, CCAT, ECAT) and world knowledge sources (Freebase, YAGO2). Chenguang Wang 0001, Yangqiu Song, Dan Roth 0001, Ming Zhang 0004, Jiawei Han 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2015 | Structural Learning with Amortized InferenceabstractTraining a structured prediction model involves performing several loss-augmented inference steps. Over the lifetime of the training, many of these inference problems, although different, share the same solution. We propose AI-DCD, an Amortized Inference framework for Dual Coordinate Descent method, an approximate learning algorithm, that accelerates the training process by exploiting this redundancy of solutions, without compromising the performance of the model. We show the efficacy of our method by training a structured SVM using dual coordinate descent for an entityrelation extraction task. Our method learns the same model as an exact training algorithm would, but call the inference engine only in 10% – 24% of the inference problems encountered during training. We observe similar gains on a multi-label classification task and with a Structured Perceptron model for the entity-relation task. Kai-Wei Chang 0001, Shyam Upadhyay, Gourab Kundu, Dan Roth 0001 |
AAAI | 4 |
| 2015 | A Joint Framework for Coreference Resolution and Mention Head DetectionabstractIn coreference resolution, a fair amount of research treats mention detection as a preprocessed step and focuses on developing algorithms for clustering coreferred mentions. However, there are significant gaps between the performance on gold mentions and the performance on the real problem, when mentions are predicted from raw text via an imperfect Mention Detection (MD) module. Motivated by the goal of reducing such gaps, we develop an ILP-based joint coreference resolution and mention head formulation that is shown to yield significant improvements on coreference from raw text, outperforming existing state-ofart systems on both the ACE-2004 and the CoNLL-2012 datasets. At the same time, our joint approach is shown to improve mention detection by close to 15% F1. One key insight underlying our approach is that identifying and co-referring mention heads is not only sufficient but is more robust than working with complete mentions. Haoruo Peng, Kai-Wei Chang 0001, Dan Roth 0001 |
CoNLL | 3 |
| 2015 | Joint Mention Extraction and Classification with Mention HypergraphsabstractWe present a novel model for the task of joint mention extraction and classifi-cation. Unlike existing approaches, our model is able to effectively capture over-lapping mentions with unbounded lengths. The model is highly scalable, with a time complexity that is linear in the number of words in the input sentence and linear in the number of possible mention classes. Our model can be extended to additionally capture mention heads explicitly in a joint manner under the same time complexity. We demonstrate the effectiveness of our model through extensive experiments on standard datasets. 1 Wei Lu 0011, Dan Roth 0001 |
EMNLP | 2 |
| 2015 | Solving General Arithmetic Word ProblemsabstractThis paper presents a novel approach to automatically solving arithmetic word problems.This is the first algorithmic approach that can handle arithmetic problems with multiple steps and operations, without depending on additional annotations or predefined templates.We develop a theory for expression trees that can be used to represent and evaluate the target arithmetic expressions; we use it to uniquely decompose the target arithmetic problem to multiple classification problems; we then compose an expression tree, combining these with world knowledge through a constrained inference framework.Our classifiers gain from the use of quantity schemas that supports better extraction of features.Experimental results show that our method outperforms existing systems, achieving state of the art performance on benchmark datasets of arithmetic word problems. Subhro Roy, Dan Roth 0001 |
EMNLP | 2 |
| 2015 | Distributed Box-Constrained Quadratic Optimization for Dual Linear SVMabstractTraining machine learning models sometimes needs to be done on large amounts of data that exceed the capacity of a single machine, motivating recent works on developing algorithms that train in a distributed fashion. This paper proposes an efficient box-constrained quadratic optimization algorithm for distributedly training linear support vector machines (SVMs) with large data. Our key technical contribution is an analytical solution to the problem of computing the optimal step size at each iteration, using an efficient method that requires only O(1) communication cost to ensure fast convergence. With this optimal step size, our approach is superior to other methods by possessing global linear convergence, or, equivalently, O(\log(1/ε)) iteration complexity for an epsilon-accurate solution, for distributedly solving the non-strongly-convex linear SVM dual problem. Experiments also show that our method is significantly faster than state-of- the-art distributed linear SVM algorithms including DSVM-AVE, DisDCA and TRON. Ching-Pei Lee, Dan Roth 0001 |
ICML | 2 |
| 2015 | Saul: Towards Declarative Learning Based Programming
Parisa Kordjamshidi, Dan Roth 0001, Hao Wu 0034 |
IJCAI | 2 |
| 2015 | Constrained Information-Theoretic Tripartite Graph Clustering to Identify Semantically Similar Relations
Chenguang Wang 0001, Yangqiu Song, Dan Roth 0001, Chi Wang 0001, Jiawei Han 0001, Heng Ji 0001, Ming Zhang 0004 |
IJCAI | 3 |
| 2015 | Incorporating World Knowledge to Document Clustering via Heterogeneous Information NetworksabstractOne of the key obstacles in making learning protocols realistic in applications is the need to supervise them, a costly process that often requires hiring domain experts. We consider the framework to use the world knowledge as indirect supervision. World knowledge is general-purpose knowledge, which is not designed for any specific domain. Then the key challenges are how to adapt the world knowledge to domains and how to represent it for learning. In this paper, we provide an example of using world knowledge for domain dependent document clustering. We provide three ways to specify the world knowledge to domains by resolving the ambiguity of the entities and their types, and represent the data with world knowledge as a heterogeneous information network. Then we propose a clustering algorithm that can cluster multiple types and incorporate the sub-type information as constraints. In the experiments, we use two existing knowledge bases as our sources of world knowledge. One is Freebase, which is collaboratively collected knowledge about entities and their organizations. The other is YAGO2, a knowledge base automatically extracted from Wikipedia and maps knowledge to the linguistic knowledge base, Word-Net. Experimental results on two text benchmark datasets (20newsgroups and RCV1) show that incorporating world knowledge as indirect supervision can significantly outperform the state-of-the-art clustering algorithms as well as clustering algorithms enhanced with world knowledge features. Chenguang Wang 0001, Yangqiu Song, Ahmed El-Kishky, Dan Roth 0001, Ming Zhang 0004, Jiawei Han 0001 |
KDD | 4 |
| 2015 | Debiasing Crowdsourced BatchesabstractCrowdsourcing is the de-facto standard for gathering annotated data. While, in theory, data annotation tasks are assumed to be attempted by workers independently, in practice, data annotation tasks are often grouped into batches to be presented and annotated by workers together, in order to save on the time or cost overhead of providing instructions or necessary background. Thus, even though independence is usually assumed between annotations on data items within the same batch, in most cases, a worker's judgment on a data item can still be affected by other data items within the batch, leading to additional errors in collected labels. In this paper, we study the data annotation bias when data items are presented as batches to be judged by workers simultaneously. We propose a novel worker model to characterize the annotating behavior on data batches, and present how to train the worker model on annotation data sets. We also present a debiasing technique to remove the effect of such annotation bias from adversely affecting the accuracy of labels obtained. Our experimental results on both synthetic data and real-world data demonstrate the effectiveness of our proposed method. Honglei Zhuang, Aditya G. Parameswaran, Dan Roth 0001, Jiawei Han 0001 |
KDD | 3 |
| 2015 | Solving Hard Coreference ProblemsabstractCoreference resolution is a key problem in natural language understanding that still escapes reliable solutions.One fundamental difficulty has been that of resolving instances involving pronouns since they often require deep language understanding and use of background knowledge.In this paper we propose an algorithmic solution that involves a new representation for the knowledge required to address hard coreference problems, along with a constrained optimization framework that uses this knowledge in coreference decision making.Our representation, Predicate Schemas, is instantiated with knowledge acquired in an unsupervised way, and is compiled automatically into constraints that impact the coreference decision.We present a general coreference resolution system that significantly improves state-of-the-art performance on hard, Winograd-style, pronoun resolution cases, while still performing at the stateof-the-art level on standard coreference resolution datasets. Haoruo Peng, Daniel Khashabi, Dan Roth 0001 |
HLT-NAACL | 3 |
| 2015 | Unsupervised Sparse Vector Densification for Short Text SimilarityabstractSparse representations of text such as bag-ofwords models or extended explicit semantic analysis (ESA) representations are commonly used in many NLP applications.However, for short texts, the similarity between two such sparse vectors is not accurate due to the small term overlap.While there have been multiple proposals for dense representations of words, measuring similarity between short texts (sentences, snippets, paragraphs) requires combining these token level similarities.In this paper, we propose to combine ESA representations and word2vec representations as a way to generate denser representations and, consequently, a better similarity measure between short texts.We study three densification mechanisms that involve aligning sparse representation via many-to-many, many-to-one, and oneto-one mappings.We then show the effectiveness of these mechanisms on measuring similarity between short texts. Yangqiu Song, Dan Roth 0001 |
HLT-NAACL | 2 |
| 2015 | Structured learning for spatial information extraction from biomedical text: bacteria biotopesabstractBACKGROUND: We aim to automatically extract species names of bacteria and their locations from webpages. This task is important for exploiting the vast amount of biological knowledge which is expressed in diverse natural language texts and putting this knowledge in databases for easy access by biologists. The task is challenging and the previous results are far below an acceptable level of performance, particularly for extraction of localization relationships. Therefore, we aim to design a new system for such extractions, using the framework of structured machine learning techniques. RESULTS: We design a new model for joint extraction of biomedical entities and the localization relationship. Our model is based on a spatial role labeling (SpRL) model designed for spatial understanding of unrestricted text. We extend SpRL to extract discourse level spatial relations in the biomedical domain and apply it on the BioNLP-ST 2013, BB-shared task. We highlight the main differences between general spatial language understanding and spatial information extraction from the scientific text which is the focus of this work. We exploit the text's structure and discourse level global features. Our model and the designed features substantially improve on the previous systems, achieving an absolute improvement of approximately 57 percent over F1 measure of the best previous system for this task. CONCLUSIONS: Our experimental results indicate that a joint learning model over all entities and relationships in a document outperforms a model which extracts entities and relationships independently. Our global learning model significantly improves the state-of-the-art results on this task and has a high potential to be adopted in other natural language processing (NLP) tasks in the biomedical domain. Parisa Kordjamshidi, Dan Roth 0001, Marie-Francine Moens |
BMC Bioinform. | 2 |
| 2015 | Overcoming bias to learn about controversial topicsabstractDeciding whether a claim is true or false often requires a deeper understanding of the evidence supporting and contradicting the claim. However, when presented with many evidence documents, users do not necessarily read and trust them uniformly. Psychologists and other researchers have shown that users tend to follow and agree with articles and sources that hold viewpoints similar to their own, a phenomenon known as confirmation bias. This suggests that when learning about a controversial topic, human biases and viewpoints about the topic may affect what is considered “trustworthy” or credible. It is an interesting challenge to build systems that can help users overcome this bias and help them decide the truthfulness of claims. In this article, we study various factors that enable humans to acquire additional information about controversial claims in an unbiased fashion. Specifically, we designed a user study to understand how presenting evidence with contrasting viewpoints and source expertise ratings affect how users learn from the evidence documents. We find that users do not seek contrasting viewpoints by themselves, but explicitly presenting contrasting evidence helps them get a well‐rounded understanding of the topic. Furthermore, explicit knowledge of the credibility of the sources and the context in which the source provides the evidence document not only affects what users read but also whether they perceive the document to be credible. V. G. Vinod Vydiswaran, ChengXiang Zhai, Dan Roth 0001, Peter Pirolli |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2015 | Reasoning about Quantities in Natural LanguageabstractLittle work from the Natural Language Processing community has targeted the role of quantities in Natural Language Understanding. This paper takes some key steps towards facilitating reasoning about quantities expressed in natural language. We investigate two different tasks of numerical reasoning. First, we consider Quantity Entailment, a new task formulated to understand the role of quantities in general textual inference tasks. Second, we consider the problem of automatically understanding and solving elementary school math word problems. In order to address these quantitative reasoning problems we first develop a computational approach which we show to successfully recognize and normalize textual expressions of quantities. We then use these capabilities to further develop algorithms to assist reasoning in the context of the aforementioned tasks. Subhro Roy, Tim Vieira, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2014 | On Dataless Hierarchical Text ClassificationabstractIn this paper, we systematically study the problem of dataless hierarchical text classification. Unlike standard text classification schemes that rely on supervised training, dataless classification depends on understanding the labels of the sought after categories and requires no labeled data. Given a collection of text documents and a set of labels, we show that understanding the labels can be used to accurately categorize the documents. This is done by embedding both labels and documents in a semantic space that allows one to compute meaningful semantic similarity between a document and a potential label. We show that this scheme can be used to support accurate multiclass classification without any supervision. We study several semantic representations and show how to improve the classification using bootstrapping. Our results show that bootstrapped dataless classification is competitive with supervised classification with thousands of labeled examples. Yangqiu Song, Dan Roth 0001 |
AAAI | 2 |
| 2014 | Correcting Grammatical Verb ErrorsabstractVerb errors are some of the most common mistakes made by non-native writers of English but some of the least studied. The reason is that dealing with verb errors requires a new paradigm; essentially all research done on correcting grammatical errors assumes a closed set of triggers ‐ e.g., correcting the use of prepositions or articles ‐ but identifying mistakes in verbs necessitates identifying potentially ambiguous triggers first, and then determining the type of mistake made and correcting it. Moreover, once the verb is identified, modeling verb errors is challenging because verbs fulfill many grammatical functions, resulting in a variety of mistakes. Consequently, the little earlier work done on verb errors assumed that the error type is known in advance. We propose a linguistically-motivated approach to verb error correction that makes use of the notion of verb finiteness to identify triggers and types of mistakes, before using a statistical machine learning approach to correct these mistakes. We show that the linguistically-informed model significantly improves the accuracy of the verb correction approach. Alla Rozovskaya, Dan Roth 0001, Vivek Srikumar |
EACL | 2 |
| 2014 | A Discriminative Latent Variable Model for Online ClusteringabstractThis paper presents a latent variable structured prediction model for discriminative supervised clustering of items called the Latent Left-linking Model (L3M). We present an online clustering algorithm for L3M based on a feature-based item similarity function. We provide a learning framework for estimating the similarity function and present a fast stochastic gradient-based learning technique. In our experiments on coreference resolution and document clustering, L3 M outperforms several existing online as well as batch supervised clustering techniques. Rajhans Samdani, Kai-Wei Chang 0001, Dan Roth 0001 |
ICML | 3 |
| 2014 | ILLINOISCLOUDNLP: Text Analytics Services in the Cloud
Hao Wu 0034, Zhiye Fei, Aaron Dai, Mark Sammons, Dan Roth 0001, Stephen Mayhew 0001 |
LREC | 5 |
| 2014 | Soft-constrained inference for Named Entity Recognition
Elisabetta Fersini, Enza Messina, Giovanni Felici, Dan Roth 0001 |
Inf. Process. Manag. | 4 |
| 2014 | Introduction to the special issue on learning semantics
Antoine Bordes, Léon Bottou, Ronan Collobert, Dan Roth 0001, Jason Weston, Luke Zettlemoyer |
Mach. Learn. | 4 |
| 2014 | Learning from natural instructions
Dan Goldwasser, Dan Roth 0001 |
Mach. Learn. | 2 |
| 2014 | Building a State-of-the-Art Grammatical Error Correction SystemabstractThis paper identifies and examines the key principles underlying building a state-of-the-art grammatical error correction system. We do this by analyzing the Illinois system that placed first among seventeen teams in the recent CoNLL-2013 shared task on grammatical error correction. The system focuses on five different types of errors common among non-native English writers. We describe four design principles that are relevant for correcting all of these errors, analyze the system along these dimensions, and show how each of these dimensions contributes to the performance. Alla Rozovskaya, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Margin-based Decomposed Amortized Inference
Gourab Kundu, Vivek Srikumar, Dan Roth 0001 |
ACL (1) | 3 |
| 2013 | Concept-based analysis of scientific literatureabstractThis paper studies the importance of identifying and categorizing scientific concepts as a way to achieve a deeper understanding of the research literature of a scientific community. To reach this goal, we propose an unsupervised bootstrapping algorithm for identifying and categorizing mentions of concepts. We then propose a new clustering algorithm that uses citations' context as a way to cluster the extracted mentions into coherent concepts. Our evaluation of the algorithms against gold standards shows significant improvement over state-of-the-art results. More importantly, we analyze the computational linguistic literature using the proposed algorithms and show four different ways to summarize and understand the research community which are difficult to obtain using existing techniques. Chen-Tse Tsai, Gourab Kundu, Dan Roth 0001 |
CIKM | 3 |
| 2013 | A Constrained Latent Variable Model for Coreference ResolutionabstractCoreference resolution is a well known clustering task in Natural Language Processing.In this paper, we describe the Latent Left Linking model (L 3 M), a novel, principled, and linguistically motivated latent structured prediction approach to coreference resolution.We show that L 3 M admits efficient inference and can be augmented with knowledge-based constraints; we also present a fast stochastic gradient based learning.Experiments on ACE and Ontonotes data show that L 3 M and its constrained version, CL 3 M, are more accurate than several state-of-the-art approaches as well as some structured prediction models proposed in the literature. Kai-Wei Chang 0001, Rajhans Samdani, Dan Roth 0001 |
EMNLP | 3 |
| 2013 | Relational Inference for WikificationabstractWikification, commonly referred to as Disambiguation to Wikipedia (D2W), is the task of identifying concepts and entities in text and disambiguating them into the most specific corresponding Wikipedia pages.Previous approaches to D2W focused on the use of local and global statistics over the given text, Wikipedia articles and its link structures, to evaluate context compatibility among a list of probable candidates.However, these methods fail (often, embarrassingly), when some level of text understanding is needed to support Wikification.In this paper we introduce a novel approach to Wikification by incorporating, along with statistical methods, richer relational analysis of the text.We provide an extensible, efficient and modular Integer Linear Programming (ILP) formulation of Wikification that incorporates the entity-relation inference problem, and show that the ability to identify relations in text helps both candidate generation and ranking Wikipedia titles considerably.Our results show significant improvements in both Wikification and the TAC Entity Linking task. Dan Roth 0001 |
EMNLP | 2 |
| 2013 | Using Soft Constraints in Joint Inference for Clinical Concept RecognitionabstractThis paper introduces IQPs (Integer Quadratic Programs) as a way to model joint inference for the task of concept recognition in clinical domain.IQPs make it possible to easily incorporate soft constraints in the optimization framework and still support exact global inference.We show that soft constraints give statistically significant performance improvements when compared to hard constraints. Prateek Jindal, Dan Roth 0001 |
EMNLP | 2 |
| 2013 | Joint Learning and Inference for Grammatical Error CorrectionabstractState-of-the-art systems for grammatical error correction are based on a collection of independently-trained models for specific errors.Such models ignore linguistic interactions at the sentence level and thus do poorly on mistakes that involve grammatical dependencies among several words.In this paper, we identify linguistic structures with interacting grammatical properties and propose to address such dependencies via joint inference and joint learning.We show that it is possible to identify interactions well enough to facilitate a joint approach and, consequently, that joint methods correct incoherent predictions that independentlytrained classifiers tend to produce.Furthermore, because the joint learning model considers interacting phenomena during training, it is able to identify mistakes that require making multiple changes simultaneously and that standard approaches miss.Overall, our model significantly outperforms the Illinois system that placed first in the CoNLL-2013 shared task on grammatical error correction. Alla Rozovskaya, Dan Roth 0001 |
EMNLP | 2 |
| 2013 | End-to-End Coreference Resolution for Clinical Narratives
Prateek Jindal, Dan Roth 0001 |
IJCAI | 2 |
| 2013 | Mining evidences for named entity disambiguationabstractNamed entity disambiguation is the task of disambiguating named entity mentions in natural language text and link them to their corresponding entries in a knowledge base such as Wikipedia. Such disambiguation can help enhance readability and add semantics to plain text. It is also a central step in constructing high-quality information network or knowledge graph from unstructured text. Previous research has tackled this problem by making use of various textual and structural features from a knowledge base. Most of the proposed algorithms assume that a knowledge base can provide enough explicit and useful information to help disambiguate a mention to the right entity. However, the existing knowledge bases are rarely complete (likely will never be), thus leading to poor performance on short queries with not well-known contexts. In such cases, we need to collect additional evidences scattered in internal and external corpus to augment the knowledge bases and enhance their disambiguation power. In this work, we propose a generative model and an incremental algorithm to automatically mine useful evidences across documents. With a specific modeling of "background topic" and "unknown entities", our model is able to harvest useful evidences out of noisy information. Experimental results show that our proposed method outperforms the state-of-the-art approaches significantly: boosting the disambiguation accuracy from 43% (baseline) to 86% on short queries derived from tweets. Yang Li 0150, Chi Wang 0001, Fangqiu Han, Jiawei Han 0001, Dan Roth 0001, Xifeng Yan |
KDD | 5 |
| 2013 | Understanding evolution of research themes: a probabilistic generative model for citationsabstractUnderstanding how research themes evolve over time in a research community is useful in many ways (e.g., revealing important milestones and discovering emerging major research trends). In this paper, we propose a novel way of analyzing literature citation to explore the research topics and the theme evolution by modeling article citation relations with a probabilistic generative model. The key idea is to represent a research paper by a ``bag of citations'' and model such a ``citation document'' with a probabilistic topic model. We explore the extension of a particular topic model, i.e., Latent Dirichlet Allocation~(LDA), for citation analysis, and show that such a Citation-LDA can facilitate discovering of individual research topics as well as the theme evolution from multiple related topics, both of which in turn lead to the construction of evolution graphs for characterizing research themes. We test the proposed citation-LDA on two datasets: the ACL Anthology Network(AAN) of natural language research literatures and PubMed Central(PMC) archive of biomedical and life sciences literatures, and demonstrate that Citation-LDA can effectively discover the evolution of research themes, with better formed topics than (conventional) Content-LDA. Xiaolong Wang 0009, ChengXiang Zhai, Dan Roth 0001 |
KDD | 3 |
| 2013 | Multi-core Structural SVM Training
Kai-Wei Chang 0001, Vivek Srikumar, Dan Roth 0001 |
ECML/PKDD (2) | 3 |
| 2013 | Latent credibility analysisabstractA frequent problem when dealing with data gathered from multiple sources on the web (ranging from booksellers to Wikipedia pages to stock analyst predictions) is that these sources disagree, and we must decide which of their (often mutually exclusive) claims we should accept. Current state-of-the-art information credibility algorithms known as "fact-finders" are transitive voting systems with rules specifying how votes iteratively flow from sources to claims and then back to sources. While this is quite tractable and often effective, fact-finders also suffer from substantial limitations; in particular, a lack of transparency obfuscates their credibility decisions and makes them difficult to adapt and analyze: knowing the mechanics of how votes are calculated does not readily tell us what those votes mean, and finding, for example, that a source has a score of 6 is not informative. We introduce a new approach to information credibility, Latent Credibility Analysis (LCA), constructing strongly principled, probabilistic models where the truth of each claim is a latent variable and the credibility of a source is captured by a set of model parameters. This gives LCA models clear semantics and modularity that make extending them to capture additional observed and latent credibility factors straightforward. Experiments over four real-world datasets demonstrate that LCA models can outperform the best fact-finders in both unsupervised and semi-supervised settings. Jeff Pasternack, Dan Roth 0001 |
WWW | 2 |
| 2013 | Using domain knowledge and domain-inspired discourse model for coreference resolution for clinical narrativesabstractOBJECTIVE: This paper presents a coreference resolution system for clinical narratives. Coreference resolution aims at clustering all mentions in a single document to coherent entities. MATERIALS AND METHODS: A knowledge-intensive approach for coreference resolution is employed. The domain knowledge used includes several domain-specific lists, a knowledge intensive mention parsing, and task informed discourse model. Mention parsing allows us to abstract over the surface form of the mention and represent each mention using a higher-level representation, which we call the mention's semantic representation (SR). SR reduces the mention to a standard form and hence provides better support for comparing and matching. Existing coreference resolution systems tend to ignore discourse aspects and rely heavily on lexical and structural cues in the text. The authors break from this tradition and present a discourse model for "person" type mentions in clinical narratives, which greatly simplifies the coreference resolution. RESULTS: This system was evaluated on four different datasets which were made available in the 2011 i2b2/VA coreference challenge. The unweighted average of F1 scores (over B-cubed, MUC and CEAF) varied from 84.2% to 88.1%. These experiments show that domain knowledge is effective for different mention types for all the datasets. DISCUSSION: Error analysis shows that most of the recall errors made by the system can be handled by further addition of domain knowledge. The precision errors, on the other hand, are more subtle and indicate the need to understand the relations in which mentions participate for building a robust coreference system. CONCLUSION: This paper presents an approach that makes an extensive use of domain knowledge to significantly improve coreference resolution. The authors state that their system and the knowledge sources developed will be made publicly available. Prateek Jindal, Dan Roth 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Extraction of events and temporal expressions from clinical narratives
Prateek Jindal, Dan Roth 0001 |
J. Biomed. Informatics | 2 |
| 2013 | Modeling Semantic Relations Expressed by PrepositionsabstractThis paper introduces the problem of predicting semantic relations expressed by prepositions and develops statistical learning models for predicting the relations, their arguments and the semantic types of the arguments. We define an inventory of 32 relations, building on the word sense disambiguation task for prepositions and collapsing related senses across prepositions. Given a preposition in a sentence, our computational task to jointly model the preposition relation and its arguments along with their semantic types, as a way to support the relation prediction. The annotated data, however, only provides labels for the relation label, and not the arguments and types. We address this by presenting two models for preposition relation labeling. Our generalization of latent structure SVM gives close to 90% accuracy on relation labeling. Further, by jointly predicting the relation, arguments, and their types along with preposition sense, we show that we can not only improve the relation accuracy, but also significantly improve sense prediction accuracy. Vivek Srikumar, Dan Roth 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2012 | Automatic Event Extraction with Structured Preference Modeling
Wei Lu 0011, Dan Roth 0001 |
ACL (1) | 2 |
| 2012 | Unsupervised discovery of opposing opinion networks from forum discussionsabstractWith more and more people freely express opinions as well as actively interact with each other in discussion threads, online forums are becoming a gold mine with rich information about people's opinions and social behaviors. In this paper, we study an interesting new problem of automatically discovering opposing opinion networks of users from forum discussions, which are subset of users who are strongly against each other on some topic. Toward this goal, we propose to use signals from both textual content (e.g., who says what) and social interactions (e.g., who talks to whom) which are both abundant in online forums. We also design an optimization formulation to combine all the signals in an unsupervised way. We created a data set by manually annotating forum data on five controversial topics and our experimental results show that the proposed optimization method outperforms several baselines and existing approaches, demonstrating the power of combining both text analysis and social network analysis in analyzing and generating the opposing opinion networks. Yue Lu 0002, Hongning Wang, ChengXiang Zhai, Dan Roth 0001 |
CIKM | 4 |
| 2012 | BiasTrust: teaching biased users about controversial topicsabstractDeciding whether a claim is true or false often requires understanding the evidence supporting and contradicting the claim. However, when learning about a controversial claim, human biases and viewpoints may affect which evidence documents are considered "trustworthy" or credible. It is important to overcome this bias and know both viewpoints to get a balanced perspective. In this paper, we study various factors that affect learning about the truthfulness of controversial claims. We designed a user study to understand the impact of these factors. Specifically, we studied the impact of presenting evidence with contrasting viewpoints and source expertise rating on how users accessed the evidence documents. This would help us optimize how to teach users about controversial topics in the most effective way, and to design better claim verification systems. We find that users do not seek contrasting viewpoints by themselves, but explicitly presenting contrasting evidence helps them get a well-rounded understanding of the topic. Furthermore, explicit knowledge of the source credibility and the context not only affects what users read, but also how credible they perceive the document to be. V. G. Vinod Vydiswaran, ChengXiang Zhai, Dan Roth 0001, Peter Pirolli |
CIKM | 3 |
| 2012 | New Frontiers in Computational Models of Grammatical Development
Micah B. Goldwater, Scott Friedman 0001, Dedre Gentner, Kenneth D. Forbus, Cynthia Fisher, Michael Connor, Dan Roth 0001, Franklin Chang, Gary S. Dell |
CogSci | 7 |
| 2012 | Using Knowledge and Constraints To Find the Best Antecedent
Prateek Jindal, Dan Roth 0001 |
COLING | 2 |
| 2012 | Joint Inference for Event Timeline Construction
Quang Do, Wei Lu 0011, Dan Roth 0001 |
EMNLP-CoNLL | 3 |
| 2012 | A Discriminative Model for Query Spelling Correction with Latent Structural SVM
Huizhong Duan, Yanen Li, ChengXiang Zhai, Dan Roth 0001 |
EMNLP-CoNLL | 4 |
| 2012 | Learning-based Multi-Sieve Co-reference Resolution with Knowledge
Lev-Arie Ratinov, Dan Roth 0001 |
EMNLP-CoNLL | 2 |
| 2012 | On Amortizing Inference Cost for Structured Prediction
Vivek Srikumar, Gourab Kundu, Dan Roth 0001 |
EMNLP-CoNLL | 3 |
| 2012 | Efficient Pattern-Based Time Series Classification on GPUabstractTime series shapelet discovery algorithm finds subsequences from a set of time series for use as primitives for time series classification. This algorithm has drawn a lot of interest because of the interpretability of its results. However, computation requirements restrict the algorithm from dealing with large data sets and may limit its application in many domains. In this paper, we address this issue by redesigning the algorithm for implementation on highly parallel Graphics Process Units (GPUs). We investigate several concepts of GPU programming and propose a dynamic programming algorithm that is suitable for implementation on GPUs. Results show that the proposed GPU implementation significantly reduces the running time of the shapelet discovery algorithm. For example, on the largest sample dataset from the original authors, the running time is reduced from half a day to two minutes. Kai-Wei Chang 0001, Biplab Deka, Wen-Mei W. Hwu, Dan Roth 0001 |
ICDM | 4 |
| 2012 | Efficient Decomposed Learning for Structured Prediction
Rajhans Samdani, Dan Roth 0001 |
ICML | 2 |
| 2012 | An NLP Curator (or: How I Learned to Stop Worrying and Love NLP Pipelines)
James Clarke, Vivek Srikumar, Mark Sammons, Dan Roth 0001 |
LREC | 4 |
| 2012 | Predicting Structures in NLP: Constrained Conditional Models and Integer Linear Programming in NLP
Dan Goldwasser, Vivek Srikumar, Dan Roth 0001 |
HLT-NAACL | 3 |
| 2012 | Unified Expectation Maximization
Rajhans Samdani, Ming-Wei Chang, Dan Roth 0001 |
HLT-NAACL | 3 |
| 2012 | A Robust Shallow Temporal Reasoning System
Quang Do, Dan Roth 0001 |
HLT-NAACL | 3 |
| 2012 | Structured learning with constrained conditional models
Ming-Wei Chang, Lev-Arie Ratinov, Dan Roth 0001 |
Mach. Learn. | 3 |
| 2012 | Exploiting the Wikipedia structure in local and global classification of taxonomic relationsabstractAbstract Determining whether two terms have an ancestor relation (e.g. Toyota Camry and car) or a sibling relation (e.g. Toyota and Honda) is an essential component of textual inference in Natural Language Processing applications such as Question Answering, Summarization, and Textual Entailment. Significant work has been done on developing knowledge sources that could support these tasks, but these resources usually suffer from low coverage, noise, and are inflexible when dealing with ambiguous and general terms that may not appear in any stationary resource, making their use as general purpose background knowledge resources difficult. In this paper, rather than building a hierarchical structure of concepts and relations, we describe an algorithmic approach that, given two terms, determines the taxonomic relation between them using a machine learning-based approach that makes use of existing resources. Moreover, we develop a global constraint-based inference process that leverages an existing knowledge base to enforce relational constraints among terms and thus improves the classifier predictions. Our experimental evaluation shows that our approach significantly outperforms other systems built upon the existing well-known knowledge sources. Quang Xuan Do, Dan Roth 0001 |
Nat. Lang. Eng. | 2 |
| 2011 | Exploiting Syntactico-Semantic Structures for Relation Extraction
Yee Seng Chan, Dan Roth 0001 |
ACL | 2 |
| 2011 | Confidence Driven Unsupervised Semantic Parsing
Dan Goldwasser, Roi Reichart, James Clarke, Dan Roth 0001 |
ACL | 4 |
| 2011 | Local and Global Algorithms for Disambiguation to Wikipedia
Lev-Arie Ratinov, Dan Roth 0001, Doug Downey |
ACL | 2 |
| 2011 | Algorithm Selection and Model Adaptation for ESL Correction Tasks
Alla Rozovskaya, Dan Roth 0001 |
ACL | 2 |
| 2011 | Adapting Text instead of the Model: An Open Domain Approach
Gourab Kundu, Dan Roth 0001 |
CoNLL | 2 |
| 2011 | Minimally Supervised Event Causality Identification
Quang Do, Yee Seng Chan, Dan Roth 0001 |
EMNLP | 3 |
| 2011 | A Joint Model for Extended Semantic Role Labeling
Vivek Srikumar, Dan Roth 0001 |
EMNLP | 2 |
| 2011 | Constrained conditional models for information fusion
Gourab Kundu, Dan Roth 0001, Rajhans Samdani |
FUSION | 2 |
| 2011 | On Bayesian interpretation of fact-finding in information networks
Dong Wang 0002, Tarek F. Abdelzaher, Hossein Ahmadi 0001, Jeff Pasternack, Dan Roth 0001, Manish Gupta 0001, Jiawei Han 0001, Omid Fatemieh, Hieu Khac Le, Charu C. Aggarwal |
FUSION | 5 |
| 2011 | Learning from Negative Examples in Set-ExpansionabstractThis paper addresses the task of set-expansion on free text. Set-expansion has been viewed as a problem of generating an extensive list of instances of a concept of interest, given a few examples of the concept as input. Our key contribution is that we show that the concept definition can be significantly improved by specifying some negative examples in the input, along with the positive examples. The state-of-the art centroid-based approach to set-expansion doesn't readily admit the negative examples. We develop an inference-based approach to set-expansion which naturally allows for negative examples and show that it performs significantly better than a strong baseline. Prateek Jindal, Dan Roth 0001 |
ICDM | 2 |
| 2011 | Online Latent Structure Training for Language AcquisitionabstractA fundamental step in sentence comprehension involves assigning semantic roles to sentence constituents. To accomplish this, the listener must parse the sentence, find constituents that are candidate arguments, and assign semantic roles to those constituents. Where do children learning their first languages begin in solving this problem? Even assuming children can derive a rough meaning for the sentence from the situation, how do they begin to map this meaning to the structure and the structure to the form of the sentence? In this paper we use feedback from a semantic role labeling (SRL) task to improve the intermediate syntactic representations that feed the SRL. We accomplish this by training an intermediate classifier using signals derived from latent structure optimization techniques. By using a separate classifier to predict internal structure we see benefits due to knowledge embedded in the classifier's feature representation. This extra structure allows the system to begin to learn using weaker, more plausible semantic feedback. Michael Connor, Cynthia Fisher, Dan Roth 0001 |
IJCAI | 3 |
| 2011 | Learning from Natural InstructionsabstractMachine learning is traditionally formalized and researched as the study of learning concepts and decision functions from labeled examples, requiring a representation that encodes information about the domain of the decision function to be learned. We are interested in providing a way for a human teacher to interact with an automated learner using natural instructions, thus allowing the teacher to communicate the relevant domain expertise to the learner without necessarily knowing anything about the internal representations used in the learning process. In this paper we suggest to view the process of learning a decision function as a natural language lesson interpretation problem instead of learning from labeled examples. This interpretation of machine learning is motivated by human learning processes, in which the learner is given a lesson describing the target concept directly, and a few instances exemplifying it. We introduce a learning algorithm for the lesson interpretation problem that gets feedback from its performance on the final task, while learning jointly (1) how to interpret the lesson and (2) how to use this interpretation to do well on the final task. This approach alleviates the supervision burden of traditional machine learning by focusing on supplying the learner with only human-level task expertise for learning. We evaluate our approach by applying it to the rules of the Freecell solitaire card game. We show that our learning approach can eventually use natural language instructions to learn the target concept and play the game legally. Furthermore, we show that the learned semantic interpreter also generalizes to previously unseen instructions. 1 Dan Goldwasser, Dan Roth 0001 |
IJCAI | 2 |
| 2011 | Making Better Informed Trust Decisions with Generalized Fact-Finding
Jeff Pasternack, Dan Roth 0001 |
IJCAI | 2 |
| 2011 | Apollo: Towards factfinding in participatory sensing
Hieu Khac Le, Jeff Pasternack, Hossein Ahmadi 0001, Manish Gupta 0001, Yizhou Sun, Tarek F. Abdelzaher, Jiawei Han 0001, Dan Roth 0001, Boleslaw K. Szymanski, Sibel Adali |
IPSN | 8 |
| 2011 | Selective block minimization for faster convergence of limited memory large-scale linear modelsabstractAs the size of data sets used to build classifiers steadily increases, training a linear model efficiently with limited memory becomes essential. Several techniques deal with this problem by loading blocks of data from disk one at a time, but usually take a considerable number of iterations to converge to a reasonable model. Even the best block minimization techniques [1] require many block loads since they treat all training examples uniformly. As disk I/O is expensive, reducing the amount of disk access can dramatically decrease the training time. This paper introduces a selective block minimization (SBM) algorithm, a block minimization method that makes use of selective sampling. At each step, SBM updates the model using data consisting of two parts: (1) new data loaded from disk and (2) a set of informative samples already in memory from previous steps. We prove that, by updating the linear model in the dual form, the proposed method fully utilizes the data in memory and converges to a globally optimal solution on the entire data. Experiments show that the SBM algorithm dramatically reduces the number of blocks loaded from disk and consequently obtains an accurate and stable model quickly on both binary and multi-class classification. Kai-Wei Chang 0001, Dan Roth 0001 |
KDD | 2 |
| 2011 | Content-driven trust propagation frameworkabstractExisting fact-finding models assume availability of structured data or accurate information extraction. However, as online data gets more unstructured, these assumptions are no longer valid. To overcome this, we propose a novel, content-based, trust propagation framework that relies on signals from the textual content to ascertain veracity of free-text claims and compute trustworthiness of their sources. We incorporate the quality of relevant content into the framework and present an iterative algorithm for propagation of trust scores. We show that existing fact finders on structured data can be modeled as specific instances of this framework. Using a retrieval-based approach to find relevant articles, we instantiate the framework to compute trustworthiness of news sources and articles. We show that the proposed framework helps ascertain trustworthiness of sources better. We also show that ranking news articles based on trustworthiness learned from the content-driven framework is significantly better than baselines that ignore either the content quality or the trust framework. V. G. Vinod Vydiswaran, ChengXiang Zhai, Dan Roth 0001 |
KDD | 3 |
| 2011 | Portfolios in Stochastic Local Search: Efficiently Computing Most Probable Explanations in Bayesian Networks
Ole J. Mengshoel, Dan Roth 0001, David C. Wilkins |
J. Autom. Reason. | 2 |
| 2011 | Initialization and Restart in Stochastic Local Search: Computing a Most Probable Explanation in Bayesian NetworksabstractFor hard computational problems, stochastic local search has proven to be a competitive approach to finding optimal or approximately optimal problem solutions. Two key research questions for stochastic local search algorithms are: Which algorithms are effective for initialization? When should the search process be restarted? In the present work, we investigate these research questions in the context of approximate computation of most probable explanations (MPEs) in Bayesian networks (BNs). We introduce a novel approach, based on the Viterbi algorithm, to explanation initialization in BNs. While the Viterbi algorithm works on sequences and trees, our approach works on BNs with arbitrary topologies. We also give a novel formalization of stochastic local search, with focus on initialization and restart, using probability theory and mixture models. Experimentally, we apply our methods to the problem of MPE computation, using a stochastic local search algorithm known as Stochastic Greedy Search. By carefully optimizing both initialization and restart, we reduce the MPE search time for application BNs by several orders of magnitude compared to using uniform at random initialization without restart. On several BNs from applications, the performance of Stochastic Greedy Search is competitive with clique tree clustering, a state-of-the-art exact algorithm used for MPE computation in BNs. Ole J. Mengshoel, David C. Wilkins, Dan Roth 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | Starting from Scratch in Semantic Role Labeling
Michael Connor, Yael Gertner, Cynthia Fisher, Dan Roth 0001 |
ACL | 4 |
| 2010 | "Ask Not What Textual Entailment Can Do for You..."
Mark Sammons, V. G. Vinod Vydiswaran, Dan Roth 0001 |
ACL | 3 |
| 2010 | Exploiting Background Knowledge for Relation Extraction
Yee Seng Chan, Dan Roth 0001 |
COLING | 2 |
| 2010 | Knowing What to Believe (when you already know something)
Jeff Pasternack, Dan Roth 0001 |
COLING | 2 |
| 2010 | Driving Semantic Parsing from the World's Response
James Clarke, Dan Goldwasser, Ming-Wei Chang, Dan Roth 0001 |
CoNLL | 4 |
| 2010 | The Necessity of Combining Adaptation Methods
Ming-Wei Chang, Michael Connor, Dan Roth 0001 |
EMNLP | 3 |
| 2010 | Constraints Based Taxonomic Relation Classification
Quang Do, Dan Roth 0001 |
EMNLP | 2 |
| 2010 | Generating Confusion Sets for Context-Sensitive Error Correction
Alla Rozovskaya, Dan Roth 0001 |
EMNLP | 2 |
| 2010 | Structured Output Learning with Indirect Supervision
Ming-Wei Chang, Vivek Srikumar, Dan Goldwasser, Dan Roth 0001 |
ICML | 4 |
| 2010 | Learning Based Java for Rapid Development of NLP Systems
Nick Rizzolo, Dan Roth 0001 |
LREC | 2 |
| 2010 | Discriminative Learning over Constrained Latent Representations
Ming-Wei Chang, Dan Goldwasser, Dan Roth 0001, Vivek Srikumar |
HLT-NAACL | 3 |
| 2010 | Training Paradigms for Correcting Errors in Grammar and Usage
Alla Rozovskaya, Dan Roth 0001 |
HLT-NAACL | 2 |
| 2010 | Automatic Model Adaptation for Complex Structured Domains
Geoffrey Levine, Gerald DeJong, Li-Lun Wang, Rajhans Samdani, Shankar Vembu, Dan Roth 0001 |
ECML/PKDD (2) | 6 |
| 2010 | Recognizing textual entailment: Rational, evaluation and approaches - ErratumabstractDue to publisher error, this article was omitted from the printed issue ofNatural Language Engineeringvolume 15 issue 4. It is published online in the correct volume ( journals.cambridge.org/nle ) and also printed here in volume 16 issue 1. Sincere apologies are extended to the authors for this error. Ido Dagan, William B. Dolan, Bernardo Magnini, Dan Roth 0001 |
Nat. Lang. Eng. | 4 |
| 2009 | Learning better transliterationsabstractWe introduce a new probabilistic model for transliteration that performs significantly better than previous approaches, is language-agnostic, requiring no knowledge of the source or target languages, and is capable of both generation (creating the most likely transliteration of a source word) and discovery (selecting the most likely transliteration from a list of candidate words). Our experimental results demonstrate improved accuracy over the existing state-of-the-art by more than 10% in Chinese, Hebrew and Russian. While past work has commonly made use of fixed-size n-gram features along with more traditional models such as HMM or Perceptron, we utilize an intuitive notion of "productions", where each source word can be segmented into a series of contiguous, non-overlapping substrings of any size, each of which independently transliterates to a substring in the target language with a given probability. To learn these parameters, we employ Expectation-Maximization (EM), with the alignment between substrings in the source and target word training pairs as our latent data. Despite the size of the parameter space and the 2(|w|-1) possible segmentations to consider for each word, by using dynamic programming each iteration of EM takes O(m^6 * n) time, where m is the length of the longest word in the data and n is the number of word pairs, and is very fast in practice. Furthermore, discovering transliterations takes only O(m^4 * w) time, where w is the number of candidate words to choose from, and generating a transliteration takes O(m2 * k2) time, where k is a pruning constant (we used a value of 100). Additionally, we are able to obtain training examples in an unsupervised fashion from Wikipedia by using a relatively simple algorithm to filter potential word pairs. Jeff Pasternack, Dan Roth 0001 |
CIKM | 2 |
| 2009 | Minimally Supervised Model of Early Language Acquisition
Michael Connor, Yael Gertner, Cynthia Fisher, Dan Roth 0001 |
CoNLL | 4 |
| 2009 | Design Challenges and Misconceptions in Named Entity Recognition
Lev-Arie Ratinov, Dan Roth 0001 |
CoNLL | 2 |
| 2009 | Interactive Feature Space Construction using Semantic Information
Dan Roth 0001, Kevin Small |
CoNLL | 1 |
| 2009 | Reading to Learn: Constructing Features from Semantic Abstracts
Jacob Eisenstein, James Clarke, Dan Goldwasser, Dan Roth 0001 |
EMNLP | 4 |
| 2009 | Aspect Guided Text Categorization with Unobserved LabelsabstractThis paper proposes a novel multiclass classification method and exhibits its advantage in the domain of text categorization with a large label space and, most importantly, when some of the labels were not observed in the training data. The key insight is the introduction of intermediate aspect variables that encode properties of the labels. Aspect variables serve as a joint representation for observed and unobserved labels. This way the classification problem can be viewed as a structure learning problem with natural constraints on assignments to the aspect variables. We solve the problem as a constrained optimization problem over multiple learners and show significant improvement in classifying short sentences into a large label space of categories, including previously unobserved categories. Dan Roth 0001, Yuancheng Tu |
ICDM | 1 |
| 2009 | Unsupervised Rank Aggregation with Domain-Specific Expertise
Alexandre Klementiev, Dan Roth 0001, Kevin Small, Ivan Titov 0001 |
IJCAI | 2 |
| 2009 | Unsupervised Constraint Driven Learning For Transliteration Discovery
Ming-Wei Chang, Dan Goldwasser, Dan Roth 0001, Yuancheng Tu |
HLT-NAACL | 3 |
| 2009 | Learning Multi-linear Representations of Distributions for Efficient Inference
Dan Roth 0001, Rajhans Samdani |
ECML/PKDD (1) | 1 |
| 2009 | Extracting article text from the web with maximum subsequence segmentationabstractMuch of the information on the Web is found in articles from online news outlets, magazines, encyclopedias, review collections, and other sources. However, extracting this content from the original HTML document is complicated by the large amount of less informative and typically unrelated material such as navigation menus, forms, user comments, and ads. Existing approaches tend to be either brittle and demand significant expert knowledge and time (manual or tool-assisted generation of rules or code), necessitate labeled examples for every different page structure to be processed (wrapper induction), require relatively uniform layout (template detection), or, as with Visual Page Segmentation (VIPS), are computationally expensive. We introduce maximum subsequence segmentation, a method of global optimization over token-level local classifiers, and apply it to the domain of news websites. Training examples are easy to obtain, both learning and prediction are linear time, and results are excellent (our semi-supervised algorithm yields an overall F1-score of 97.947%), surpassing even those produced by VIPS with a hypothetical perfect block-selection heuristic. We also evaluate against the recent CleanEval shared task with surprisingly good cross-task performance cleaning general web pages, exceeding the top "text-only" score (based on Levenshtein distance), 87.8% versus 84.1%. Jeff Pasternack, Dan Roth 0001 |
WWW | 2 |
| 2009 | Learning multi-linear representations of distributions for efficient inference
Dan Roth 0001, Rajhans Samdani |
Mach. Learn. | 1 |
| 2008 | Learning and Inference with Constraints
Ming-Wei Chang, Lev-Arie Ratinov, Nick Rizzolo, Dan Roth 0001 |
AAAI | 4 |
| 2008 | Importance of Semantic Representation: Dataless Classification
Ming-Wei Chang, Lev-Arie Ratinov, Dan Roth 0001, Vivek Srikumar |
AAAI | 3 |
| 2008 | Proactive Intrusion Detection
Benjamin Liebald, Dan Roth 0001, Neelay Shah, Vivek Srikumar |
AAAI | 2 |
| 2008 | Active Learning for Pipeline Models
Dan Roth 0001, Kevin Small |
AAAI | 1 |
| 2008 | Extraction of Entailed Semantic Relations Through Syntax-Based Comma Resolution
Vivek Srikumar, Roi Reichart, Mark Sammons, Ari Rappoport, Dan Roth 0001 |
ACL | 5 |
| 2008 | Baby SRL: Modeling Early Language Acquisition
Michael Connor, Yael Gertner, Cynthia Fisher, Dan Roth 0001 |
CoNLL | 4 |
| 2008 | Understanding the Value of Features for Coreference Resolution
Eric Bengtson, Dan Roth 0001 |
EMNLP | 2 |
| 2008 | Transliteration as Constrained Optimization
Dan Goldwasser, Dan Roth 0001 |
EMNLP | 2 |
| 2008 | Unsupervised rank aggregation with distance-based modelsabstractThe need to meaningfully combine sets of rankings often comes up when one deals with ranked data. Although a number of heuristic and supervised learning approaches to rank aggregation exist, they require domain knowledge or supervised ranked data, both of which are expensive to acquire. In order to address these limitations, we propose a mathematical and algorithmic framework for learning to aggregate (partial) rankings without supervision. We instantiate the framework for the cases of combining permutations and combining top-k lists, and propose a novel metric for the latter. Experiments in both scenarios demonstrate the effectiveness of the proposed formalism. Alexandre Klementiev, Dan Roth 0001, Kevin Small |
ICML | 2 |
| 2008 | Which "Apple" are you talking about ?abstractIn a higher level task such as clustering of web results or Mandar Rahurkar, Dan Roth 0001, Thomas S. Huang |
WWW | 2 |
| 2008 | Identifying Semitic Roots: Machine Learning with Linguistic ConstraintsabstractWords in Semitic languages are formed by combining two morphemes: a root and a pattern. The root consists of consonants only, by default three, and the pattern is a combination of vowels and consonants, with non-consecutive “slots” into which the root consonants are inserted. Identifying the root of a given word is an important task, considered to be an essential part of the morphological analysis of Semitic languages, and information on roots is important for linguistics research as well as for practical applications. We present a machine learning approach, augmented by limited linguistic knowledge, to the problem of identifying the roots of Semitic words. Although programs exist which can extract the root of words in Arabic and Hebrew, they are all dependent on labor-intensive construction of large-scale lexicons which are components of full-scale morphological analyzers. The advantage of our method is an automation of this process, avoiding the bottleneck of having to laboriously list the root and pattern of each lexeme in the language. To the best of our knowledge, this is the first application of machine learning to this problem, and one of the few attempts to directly address non-concatenative morphology using machine learning. More generally, our results shed light on the problem of combining classifiers under (linguistically motivated) constraints. Ezra Daya, Dan Roth 0001, Shuly Wintner |
Comput. Linguistics | 2 |
| 2008 | The Importance of Syntactic Parsing and Inference in Semantic Role LabelingabstractWe present a general framework for semantic role labeling. The framework combines a machine-learning technique with an integer linear programming-based inference procedure, which incorporates linguistic and structural constraints into a global decision process. Within this framework, we study the role of syntactic parsing information in semantic role labeling. We show that full syntactic parsing information is, by far, most relevant in identifying the argument, especially, in the very first stage—the pruning stage. Surprisingly, the quality of the pruning stage cannot be solely determined based on its recall and precision. Instead, it depends on the characteristics of the output candidates that determine the difficulty of the downstream problems. Motivated by this observation, we propose an effective and simple approach of combining different semantic role labeling systems through joint inference, which significantly improves its performance. Our system has been evaluated in the CoNLL-2005 shared task on semantic role labeling, and achieves the highest F1 score among 19 participants. Vasin Punyakanok, Dan Roth 0001, Scott Yih |
Comput. Linguistics | 2 |
| 2007 | Guiding Semi-Supervision with Constraint-Driven Learning
Ming-Wei Chang, Lev-Arie Ratinov, Dan Roth 0001 |
ACL | 3 |
| 2007 | Context Sensitive Paraphrasing with a Global Unsupervised Classifier
Michael Connor, Dan Roth 0001 |
ECML | 2 |
| 2007 | An Unsupervised Learning Algorithm for Rank Aggregation
Alexandre Klementiev, Dan Roth 0001, Kevin Small |
ECML | 2 |
| 2007 | Maximum Margin Coresets for Active and Noise Tolerant Learning
Sariel Har-Peled, Dan Roth 0001, Dav Zimak |
IJCAI | 2 |
| 2007 | Audio-Visual Affect RecognitionabstractThe ability of a computer to detect and appropriately respond to changes in a user's affective state has significant implications to human-computer interaction (HCI). In this paper, we present our efforts toward audio-visual affect recognition on 11 affective states customized for HCI application (four cognitive/motivational and seven basic affective states) of 20 nonactor subjects. A smoothing method is proposed to reduce the detrimental influence of speech on facial expression recognition. The feature selection analysis shows that subjects are prone to use brow movement in face, pitch and energy in prosody to express their affects while speaking. For person-dependent recognition, we apply the voting method to combine the frame-based classification results from both audio and visual channels. The result shows 7.5% improvement over the best unimodal performance. For person-independent test, we apply multistream HMM to combine the information from multiple component streams. This test shows 6.1% improvement over the best component performance Zhihong Zeng, Jilin Tu, Ming Liu 0009, Thomas S. Huang, Brian Pianfetti, Dan Roth 0001, Stephen E. Levinson |
IEEE Trans. Multim. | 6 |
| 2006 | MPE and Partial Inversion in Lifted Probabilistic Variable Elimination
Rodrigo de Salvo Braz, Eyal Amir, Dan Roth 0001 |
AAAI | 3 |
| 2006 | A Pipeline Framework for Dependency Parsing
Ming-Wei Chang, Quang Do, Dan Roth 0001 |
ACL | 3 |
| 2006 | Weakly Supervised Named Entity Transliteration and Discovery from Multilingual Comparable CorporaabstractNamed Entity recognition (NER) is an important part of many natural language processing tasks. Current approaches often employ machine learning techniques and require supervised data. However, many languages lack such resources. This paper presents an (almost) unsupervised learning algorithm for automatic discovery of Named Entities (NEs) in a resource free language, given a bilingual corpora in which it is weakly temporally aligned with a resource rich language. NEs have similar time distributions across such corpora, and often some of the tokens in a multi-word NE are transliterated. We develop an algorithm that exploits both observations iteratively. The algorithm makes use of a new, frequency based, metric for time distributions and a resource free discriminative approach to transliteration. Seeded with a small number of transliteration pairs, our algorithm discovers multi-word NEs, and takes advantage of a dictionary (if one exists) to account for translated or partially translated NEs. We evaluate the algorithm on an English-Russian corpus, and show high level of NEs discovery in Russian. Alexandre Klementiev, Dan Roth 0001 |
ACL | 2 |
| 2006 | A Pipeline Model for Bottom-Up Dependency Parsing
Ming-Wei Chang, Quang Do, Dan Roth 0001 |
CoNLL | 3 |
| 2006 | Margin-Based Active Learning for Structured Output Spaces
Dan Roth 0001, Kevin Small |
ECML | 1 |
| 2006 | Named Entity Transliteration and Discovery from Multilingual Comparable Corpora
Alexandre Klementiev, Dan Roth 0001 |
HLT-NAACL | 2 |
| 2006 | Controlled generation of hard and easy Bayesian networks: Impact on maximal clique size in tree clustering
Ole J. Mengshoel, David C. Wilkins, Dan Roth 0001 |
Artif. Intell. | 3 |
| 2006 | Learning question classifiers: the role of semantic informationabstractTo respond correctly to a free form factual question given a large collection of text data, one needs to understand the question to a level that allows determining some of the constraints the question imposes on a possible answer. These constraints may include a semantic classification of the sought after answer and may even suggest using different strategies when looking for and verifying a candidate answer. This work presents a machine learning approach to question classification. Guided by a layered semantic hierarchy of answer types, we develop a hierarchical classifier that classifies questions into fine-grained classes. This work also performs a systematic study of the use of semantic information sources in natural language classification tasks. It is shown that, in the context of question classification, augmenting the input of the classifier with appropriate semantic category information results in significant improvements to classification accuracy. We show accurate results on a large collection of free-form questions used in TREC 10 and 11. Xin Li 0018, Dan Roth 0001 |
Nat. Lang. Eng. | 2 |
| 2005 | An Inference Model for Semantic Entailment in Natural Language
Rodrigo de Salvo Braz, Roxana Girju, Vasin Punyakanok, Dan Roth 0001, Mark Sammons |
AAAI | 4 |
| 2005 | Learnability of Bipartite Ranking Functions
Shivani Agarwal 0001, Dan Roth 0001 |
COLT | 2 |
| 2005 | Generalized Inference with Multiple Semantic Role Labeling Systems
Peter Koomen, Vasin Punyakanok, Dan Roth 0001, Scott Yih |
CoNLL | 3 |
| 2005 | Discriminative Training of Clustering Functions: Theory and Experiments with Entity Identification
Xin Li 0018, Dan Roth 0001 |
CoNLL | 2 |
| 2005 | Integer linear programming inference for conditional random fieldsabstractInference in Conditional Random Fields and Hidden Markov Models is done using the Viterbi algorithm, an efficient dynamic programming algorithm. In many cases, general (non-local and non-sequential) constraints may exist over the output sequence, but cannot be incorporated and exploited in a natural way by this inference procedure. This paper proposes a novel inference procedure based on integer linear programming (ILP) and extends CRF models to naturally and efficiently support general constraint structures. For sequential constraints, this procedure reduces to simple linear programming as the inference process. Experimental evidence is supplied in the context of an important NLP problem, semantic role labeling. Dan Roth 0001, Scott Yih |
ICML | 1 |
| 2005 | Lifted First-Order Probabilistic Inference
Rodrigo de Salvo Braz, Eyal Amir, Dan Roth 0001 |
IJCAI | 3 |
| 2005 | An Inference Model for Semantic Entailment in Natural Language
Rodrigo de Salvo Braz, Roxana Girju, Vasin Punyakanok, Dan Roth 0001, Mark Sammons |
IJCAI | 4 |
| 2005 | The Necessity of Syntactic Parsing for Semantic Role Labeling
Vasin Punyakanok, Dan Roth 0001, Scott Yih |
IJCAI | 2 |
| 2005 | Learning and Inference over Constrained Output
Vasin Punyakanok, Dan Roth 0001, Scott Yih, Dav Zimak |
IJCAI | 2 |
| 2005 | Efficiency versus Convergence of Boolean Kernels for On-Line Learning AlgorithmsabstractThe paper studies machine learning problems where each example is described using a set of Boolean features and where hypotheses are represented by linear threshold elements. One method of increasing the expressiveness of learned hypotheses in this context is to expand the feature set to include conjunctions of basic features. This can be done explicitly or where possible by using a kernel function. Focusing on the well known Perceptron and Winnow algorithms, the paper demonstrates a tradeoff between the computational efficiency with which the algorithm can be run over the expanded feature space and the generalization ability of the corresponding learning algorithm. We first describe several kernel functions which capture either limited forms of conjunctions or all conjunctions. We show that these kernels can be used to efficiently run the Perceptron algorithm over a feature space of exponentially many conjunctions; however we also show that using such kernels, the Perceptron algorithm can provably make an exponential number of mistakes even when learning simple functions. We then consider the question of whether kernel functions can analogously be used to run the multiplicative-update Winnow algorithm over an expanded feature space of exponentially many conjunctions. Known upper bounds imply that the Winnow algorithm can learn Disjunctive Normal Form (DNF) formulae with a polynomial mistake bound in this setting. However, we prove that it is computationally hard to simulate Winnow's behavior for learning DNF over such a feature set. This implies that the kernel functions which correspond to running Winnow for this problem are not efficiently computable, and that there is no general construction that can run Winnow with kernels. Roni Khardon, Dan Roth 0001, Rocco A. Servedio |
J. Artif. Intell. Res. | 2 |
| 2005 | Generalization Bounds for the Area Under the ROC CurveabstractWe study generalization properties of the area under the ROC curve (AUC), a quantity that has been advocated as an evaluation criterion for the bipartite ranking problem. The AUC is a different term than the error rate used for evaluation in classification problems; consequently, existing generalization bounds for the classification error rate cannot be used to draw conclusions about the AUC. In this paper, we define the expected accuracy of a ranking function (analogous to the expected error rate of a classification function), and derive distribution-free probabilistic bounds on the deviation of the empirical AUC of a ranking function (observed on a finite data sequence) from its expected accuracy. We derive both a large deviation bound, which serves to bound the expected accuracy of a ranking function in terms of its empirical AUC on a test sequence, and a uniform convergence bound, which serves to bound the expected accuracy of a learned ranking function in terms of its empirical AUC on a training sequence. Our uniform convergence bound is expressed in terms of a new set of combinatorial parameters that we term the bipartite rank-shatter coefficients; these play the same role in our result as do the standard VC-dimension related shatter coefficients (also known as the growth function) in uniform convergence results for the classification error rate. A comparison of our result with a recent uniform convergence result derived by Freund et al. (2003) for a quantity closely related to the AUC shows that the bound provided by our result can be considerably tighter. Shivani Agarwal 0001, Thore Graepel, Ralf Herbrich, Sariel Har-Peled, Dan Roth 0001 |
J. Mach. Learn. Res. | 5 |
| 2005 | Guest Editors Introduction: Machine Learning in Speech and Language Technologies
Pascale Fung, Dan Roth 0001 |
Mach. Learn. | 2 |
| 2004 | Identification and Tracing of Ambiguous Names: Discriminative and Generative Approaches
Xin Li 0018, Paul Morie, Dan Roth 0001 |
AAAI | 3 |
| 2004 | Semantic Role Labeling Via Integer Linear Programming Inference
Vasin Punyakanok, Dan Roth 0001, Scott Yih, Dav Zimak |
COLING | 2 |
| 2004 | Semantic Role Labeling Via Generalized Inference Over Classifiers
Vasin Punyakanok, Dan Roth 0001, Scott Yih, Dav Zimak, Yuancheng Tu |
CoNLL | 2 |
| 2004 | A Linear Programming Formulation for Global Inference in Natural Language Tasks
Dan Roth 0001, Scott Yih |
CoNLL | 1 |
| 2004 | Learning Hebrew Roots: Machine Learning with Linguistic Constraints
Ezra Daya, Dan Roth 0001, Shuly Wintner |
EMNLP | 2 |
| 2004 | Bimodal HCI-related affect recognitionabstractPerhaps the most fundamental application of affective computing will be Human-Computer Interaction (HCI) in which the computer should have the ability to detect and track the user's affective states, and make corresponding feedback. The human multi-sensor affect system defines the expectation of multimodal affect analyzer. In this paper, we present our efforts toward audio-visual HCI-related affect recognition. With HCI applications in mind, we take into account some special affective states which indicate users' cognitive/motivational states. Facing the fact that a facial expression is influenced by both an affective state and speech content, we apply a smoothing method to extract the information of the affective state from facial features. In our fusion stage, a voting method is applied to combine audio and visual modalities so that the final affect recognition accuracy is greatly improved. We test our bimodal affect recognition approach on 38 subjects with 11 HCI-related affect states. The extensive experimental results show that the average person-dependent affect recognition accuracy is almost 90% for our bimodal fusion. Zhihong Zeng, Jilin Tu, Ming Liu 0009, Tong Zhang 0005, Nick Rizzolo, ZhenQiu Zhang, Thomas S. Huang, Dan Roth 0001, Stephen E. Levinson |
ICMI | 8 |
| 2004 | Robust Reading: Identification and Tracing of Ambiguous Names
Xin Li 0018, Paul Morie, Dan Roth 0001 |
HLT-NAACL | 3 |
| 2004 | A Large Deviation Bound for the Area Under the ROC CurveabstractThe area under the ROC curve (AUC) has been advocated as an evalu- ation criterion for the bipartite ranking problem. We study large devi- ation properties of the AUC; in particular, we derive a distribution-free large deviation bound for the AUC which serves to bound the expected accuracy of a ranking function in terms of its empirical AUC on an inde- pendent test sequence. A comparison of our result with a corresponding large deviation result for the classification error rate suggests that the test sample size required to obtain an -accurate estimate of the expected ac- curacy of a ranking function with δ-confidence is larger than that required to obtain an -accurate estimate of the expected error rate of a classifi- cation function with the same confidence. A simple application of the union bound allows the large deviation bound to be extended to learned ranking functions chosen from finite function classes. Shivani Agarwal 0001, Thore Graepel, Ralf Herbrich, Dan Roth 0001 |
NIPS | 4 |
| 2004 | Learning to Detect Objects in Images via a Sparse, Part-Based RepresentationabstractWe study the problem of detecting objects in still, gray-scale images. Our primary focus is the development of a learning-based approach to the problem that makes use of a sparse, part-based representation. A vocabulary of distinctive object parts is automatically constructed from a set of sample images of the object class of interest; images are then represented using parts from this vocabulary, together with spatial relations observed among the parts. Based on this representation, a learning algorithm is used to automatically learn to detect instances of the object class in new images. The approach can be applied to any object with distinguishable parts in a relatively fixed spatial configuration; it is evaluated here on difficult sets of real-world images containing side views of cars, and is seen to successfully detect objects in varying conditions amidst background clutter and mild occlusion. In evaluating object detection approaches, several important methodological issues arise that have not been satisfactorily addressed in previous work. A secondary focus of this paper is to highlight these issues and to develop rigorous evaluation standards for the object detection problem. A critical evaluation of our approach under the proposed standards is presented. Shivani Agarwal 0001, Aatif Awan, Dan Roth 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Phrasenet: towards context sensitive lexical semantics
Xin Li 0018, Dan Roth 0001, Yuancheng Tu |
CoNLL | 2 |
| 2003 | On Kernel Methods for Relational Learning
Chad M. Cumby, Dan Roth 0001 |
ICML | 2 |
| 2003 | Margin Distribution and Learning
Ashutosh Garg 0001, Dan Roth 0001 |
ICML | 2 |
| 2002 | Constraint Classification: A New Approach to Multiclass Classification
Sariel Har-Peled, Dan Roth 0001, Dav Zimak |
ALT | 2 |
| 2002 | Learning Question Classifiers
Xin Li 0018, Dan Roth 0001 |
COLING | 2 |
| 2002 | Probabilistic Reasoning for Entity & Relation Recognition
Dan Roth 0001, Scott Yih |
COLING | 1 |
| 2002 | Learning a Sparse Representation for Object Detection
Shivani Agarwal 0001, Dan Roth 0001 |
ECCV (4) | 2 |
| 2002 | A Tale of Two Classifiers: SNoW vs. SVM in Visual Recognition
Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja |
ECCV (4) | 2 |
| 2002 | Learning and Inference for Clause Identification
Xavier Carreras, Lluís Màrquez, Vasin Punyakanok, Dan Roth 0001 |
ECML | 4 |
| 2002 | Reasoning with Classifiers
Dan Roth 0001 |
ECML | 1 |
| 2002 | On generalization bounds, projection profile, and margin distribution
Ashutosh Garg 0001, Sariel Har-Peled, Dan Roth 0001 |
ICML | 3 |
| 2002 | Learning with Feature Description Logics
Chad M. Cumby, Dan Roth 0001 |
ILP | 2 |
| 2002 | Constraint Classification for Multiclass Classification and RankingabstractThe constraint classification framework captures many flavors of mul- ticlass classification including winner-take-all multiclass classification, multilabel classification and ranking. We present a meta-algorithm for learning in this framework that learns via a single linear classifier in high dimension. We discuss distribution independent as well as margin-based generalization bounds and present empirical and theoretical evidence showing that constraint classification benefits over existing methods of multiclass classification. Sariel Har-Peled, Dan Roth 0001, Dav Zimak |
NIPS | 2 |
| 2002 | Reasoning with Classifiers
Dan Roth 0001 |
PKDD | 1 |
| 2002 | Learning cost-sensitive active classifiers
Russell Greiner, Adam J. Grove, Dan Roth 0001 |
Artif. Intell. | 3 |
| 2002 | Learning to Recognize Three-Dimensional ObjectsabstractA learning account for the problem of object recognition is developed within the probably approximately correct (PAC) model of learnability. The key assumption underlying this work is that objects can be recognized (or discriminated) using simple representations in terms of syntactically simple relations over the raw image. Although the potential number of these simple relations could be huge, only a few of them are actually present in each observed image, and a fairly small number of those observed are relevant to discriminating an object. We show that these properties can be exploited to yield an efficient learning approach in terms of sample and computational complexity within the PAC model. No assumptions are needed on the distribution of the observed objects, and the learning performance is quantified relative to its experience. Most important, the success of learning an object representation is naturally tied to the ability to represent it as a function of some intermediate representations extracted from the image. We evaluate this approach in a large-scale experimental study in which the SNoW learning architecture is used to learn representations for the 100 objects in the Columbia Object Image Library. Experimental results exhibit good generalization and robustness properties of the SNoW-based method relative to other approaches. SNoW's recognition rate degrades more gracefully when the training data contains fewer views, and it shows similar behavior in some preliminary experiments with partially occluded objects. Dan Roth 0001, Ming-Hsuan Yang 0001, Narendra Ahuja |
Neural Comput. | 1 |
| 2001 | Learning Coherent Concepts
Ashutosh Garg 0001, Dan Roth 0001 |
ALT | 2 |
| 2001 | Understanding Probabilistic Classifiers
Ashutosh Garg 0001, Dan Roth 0001 |
ECML | 2 |
| 2001 | A Sequential Model for Multi-Class Classification
Yair Even-Zohar, Dan Roth 0001 |
EMNLP | 2 |
| 2001 | Scaling Up Context-Sensitive Text Correction
Andrew J. Carlson, Jeffrey Rosen, Dan Roth 0001 |
IAAI | 3 |
| 2001 | Face detection using large margin classifiersabstractLarge margin classifiers have demonstrated their advantages in many visual learning tasks, and have attracted much attention in vision and image processing communities. We apply and compare two large margin classifiers, support vector machines and sparse network of winnows, to detect faces in still gray scale images Furthermore, we study the theoretical frameworks of these classifiers and analyze the empirical results. Experiments on a test set of 24,045 images exhibit good generalization and robustness, and conform to theoretical analysis. Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja |
ICIP (2) | 2 |
| 2001 | Relational Learning via Propositional Algorithms: An Information Extraction Case Study
Dan Roth 0001, Scott Yih |
IJCAI | 1 |
| 2001 | Efficiency versus Convergence of Boolean Kernels for On-Line Learning AlgorithmsabstractWe study online learning in Boolean domains using kernels which cap- ture feature expansions equivalent to using conjunctions over basic fea- tures. We demonstrate a tradeoff between the computational efficiency with which these kernels can be computed and the generalization abil- ity of the resulting classifier. We first describe several kernel functions which capture either limited forms of conjunctions or all conjunctions. We show that these kernels can be used to efficiently run the Percep- tron algorithm over an exponential number of conjunctions; however we also prove that using such kernels the Perceptron algorithm can make an exponential number of mistakes even when learning simple func- tions. We also consider an analogous use of kernel functions to run the multiplicative-update Winnow algorithm over an expanded feature space of exponentially many conjunctions. While known upper bounds imply that Winnow can learn DNF formulae with a polynomial mistake bound in this setting, we prove that it is computationally hard to simulate Win- now’s behavior for learning DNF over such a feature set, and thus that such kernel functions for Winnow are not efficiently computable. Roni Khardon, Dan Roth 0001, Rocco A. Servedio |
NIPS | 2 |
| 2001 | Linear Concepts and Hidden Variables
Adam J. Grove, Dan Roth 0001 |
Mach. Learn. | 2 |
| 2000 | Applying System Combination to Base Noun Phrase Identification
Erik F. Tjong Kim Sang, Walter Daelemans, Hervé Déjean, Rob Koeling, Yuval Krymolowski, Vasin Punyakanok, Dan Roth 0001 |
COLING | 7 |
| 2000 | Learning to Recognize ObjectsabstractA learning account for the problem of object recognition is developed within the PAC (Probably Approximately Correct) model of learnability. The proposed approach makes no assumptions on the distribution of the observed objects, but quantifies success relative to its past experience. Most importantly, the success of learning an object representation is naturally tied to the ability to represent at as a function of some intermediate representations extracted from the image. We evaluate this approach an a large scale experimental study in which the SNoW learning architecture is used to learn representations for the 100 objects in the Columbia Object Image Database (COIL-100). The SNoW-based method is shown to outperform other methods in terms of recognition rates; its performance degrades gracefully when the training data contains fewer views and in the presence of occlusion noise. Dan Roth 0001, Ming-Hsuan Yang 0001, Narendra Ahuja |
CVPR | 1 |
| 2000 | Learning to Recognize 3D Objects with SNoW
Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja |
ECCV (1) | 2 |
| 2000 | Relational Representations that Facilitate Learning
Chad M. Cumby, Dan Roth 0001 |
KR | 2 |
| 2000 | The Use of Classifiers in Sequential InferenceabstractWe study the problem of combining the outcomes of several different classifiers in a way that provides a coherent inference that satisfies some constraints. In particular, we develop two general approaches for an im(cid:173) portant subproblem - identifying phrase structure. The first is a Marko(cid:173) vian approach that extends standard HMMs to allow the use of a rich ob(cid:173) servation structure and of general classifiers to model state-observation dependencies. The second is an extension of constraint satisfaction for(cid:173) malisms. We develop efficient combination algorithms under both mod(cid:173) els and study them experimentally in the context of shallow parsing. Vasin Punyakanok, Dan Roth 0001 |
NIPS | 2 |
| 1999 | A Learning Approach to Shallow Parsing
Marcia Muñoz, Vasin Punyakanok, Dan Roth 0001, Dav Zimak |
EMNLP | 3 |
| 1999 | Relational Learning for NLP using Linear Threshold Elements
Roni Khardon, Dan Roth 0001, Leslie G. Valiant |
IJCAI | 2 |
| 1999 | Learning in Natural Language
Dan Roth 0001 |
IJCAI | 1 |
| 1999 | A SNoW-Based Face Detector
Ming-Hsuan Yang 0001, Dan Roth 0001, Narendra Ahuja |
NIPS | 2 |
| 1999 | Coherent Concepts, Robust Learning
Dan Roth 0001, Dmitry Zelenko |
SOFSEM | 1 |
| 1999 | Reasoning with Examples: Propositional Formulae and Database Dependencies
Roni Khardon, Heikki Mannila, Dan Roth 0001 |
Acta Informatica | 3 |
| 1999 | A Winnow-Based Approach to Context-Sensitive Spelling Correction
Andrew R. Golding, Dan Roth 0001 |
Mach. Learn. | 2 |
| 1999 | Learning to Reason with a Restricted View
Roni Khardon, Dan Roth 0001 |
Mach. Learn. | 2 |
| 1999 | Linearizable Read/Write Objects
Marios Mavronicolas, Dan Roth 0001 |
Theor. Comput. Sci. | 2 |
| 1998 | Clustering Appearances of 3D ObjectsabstractWe introduce a method for unsupervised clustering of images of 3D objects. Our method examines the space of all images and partitions the images into sets that form smooth and parallel surfaces in this space. It further uses sequences of images to obtain more reliable clustering. Finally, since our method relies on a non-Euclidean similarity measure we introduce algebraic techniques for estimating local properties of these surfaces without first embedding the images in a Euclidean space. We demonstrate our method by applying it to a large database of images. Ronen Basri, Dan Roth 0001, David Jacobs 0001 |
CVPR | 2 |
| 1998 | On Learning Read-k-Satisfy-j DNFabstractWe study the learnability of read-k-satisfy-j (RkSj) DNF formulas. These are boolean formulas in disjunctive normal form (DNF), in which the maximum number of occurrences of a variable is bounded by k, and the number of terms satisfied by any assignment is at most j. After motivating the investigation of this class of DNF formulas, we present an algorithm that for any unknown RkSj DNF formula to be learned, with high probability finds a logically equivalent DNF formula using the well-studied protocol of equivalence and membership queries. The algorithm runs in polynomial time for $k\cdot j=O({\log n\over\log\log n})$, where n is the number of input variables. Howard Aizenstein, Avrim Blum, Roni Khardon, Eyal Kushilevitz, Leonard Pitt, Dan Roth 0001 |
SIAM J. Comput. | 6 |
| 1997 | Mistake-Driven Learning in Text Categorization
Ido Dagan, Yael Karov, Dan Roth 0001 |
EMNLP | 3 |
| 1997 | Learning to Perform Knowledge-Intensive Inferences
Dan Roth 0001 |
MFCS | 1 |
| 1997 | Linear Concepts and Hidden Variables: An Empirical Study
Adam J. Grove, Dan Roth 0001 |
NIPS | 2 |
| 1997 | Defaults and Relevance in Model-Based Reasoning
Roni Khardon, Dan Roth 0001 |
Artif. Intell. | 2 |
| 1997 | Finding the Largest Area Axis-parallel Rectangle in a Polygon
Karen L. Daniels, Victor J. Milenkovic, Dan Roth 0001 |
Comput. Geom. | 3 |
| 1997 | Learning to reasonabstractWe introduce a new framework for the study of reasoning. The Learning (in order) to Reason approach developed here views learning as an integral part of the inference process, and suggests that learning and reasoning should be studied together. The Learning to Reason framework combines the interfaces to the world used by known learning models with the reasoning task and a performance criterion suitable for it. In this framework, the intelligent agent is given access to its favorite learning interface, and is also given a grace period in with it can interact with this interface and construct a representation KB of the world W . The reasoning performance is measured only after this period, when the agent is presented with queries α from some query language, relevant to the world, and has to answer whether W implies α. The approach is meant to overcome the main computational difficulties in the traditional treatment of reasoning which stem from its separation from the “world”. Since the agent interacts with the world when construction its knowledge representation it can choose a representation that is useful for the task at hand. Moreover, we can now make explicit the dependence of the reasoning performance on the environment the agent interacts with. We show how previous results from learning theory and reasoning fit into this framwork and illustrate the usefulness of the Learning to Reason approach by exhibiting new results that are not possible in the traditional setting. First, we give Learning to Reason algorithms for classes of propositional languages for which there are no efficient reasoning algorithms, when represented as a traditional (formula-based) knowledge base. Second, we exhibit a Learning to Reason algorithm for a class of propositional languages that is not know to be learnable in the traditional sense. Roni Khardon, Dan Roth 0001 |
J. ACM | 2 |
| 1996 | Applying Winnow to Context-Sensitive Spelling Correction
Andrew R. Golding, Dan Roth 0001 |
ICML | 2 |
| 1996 | Learning Active Classifiers
Russell Greiner, Adam J. Grove, Dan Roth 0001 |
ICML | 3 |
| 1996 | Learning in Order to Reason: The Approach
Dan Roth 0001 |
SOFSEM | 1 |
| 1996 | Reasoning with Models
Roni Khardon, Dan Roth 0001 |
Artif. Intell. | 2 |
| 1996 | On the Hardness of Approximate Reasoning
Dan Roth 0001 |
Artif. Intell. | 1 |
| 1996 | On Learning Visual Concepts and DNF Formulae
Eyal Kushilevitz, Dan Roth 0001 |
Mach. Learn. | 2 |
| 1995 | Learning to Reason with a Restricted ViewabstractThe Learning to Reason framework combines the study of Learning and Reasoning into a single task. Within it, learning is done specifically for the purpose of reasoning with the learned knowledge. Computational considerations show that this is a useful paradigm; in some cases learning and reasoning problems that are intractable when studied separately become tractable when performed as a task of Learning to Reason. In this paper we study Learning to Reason problems where the interaction with the world supplies the learner only partial information in the form of partial assignments. Several natural interpretations of partial assignments are considered and learning and reasoning algorithms using these are developed. The results presented exhibit a tradeo between learnability, the strength of the oracles used in the interface, and the range of reasoning queries the learner is guaranteed to answer correctly. Roni Khardon, Dan Roth 0001 |
COLT | 2 |
| 1995 | Default-Reasoning with Models
Roni Khardon, Dan Roth 0001 |
IJCAI | 2 |
| 1995 | Learning to Reason: The Non-Monotonic Case
Dan Roth 0001 |
IJCAI | 1 |
| 1994 | Learning to Reason
Roni Khardon, Dan Roth 0001 |
AAAI | 2 |
| 1994 | Reasoning with Models
Roni Khardon, Dan Roth 0001 |
AAAI | 2 |
| 1994 | On Learning Read-k-Satisfy-j DNFabstractWe study the learnability of Read-k-Satisfy-j (RkSj) DNF formulae. These are DNF formulae in which the maximal number of occurrences of a variable is bounded by k, and the number of terms satisfied by any assignment is at most j. We show that this class of functions is learnable in polynomial time, using Equivalence and Membership Queries, as long as k•j=O(logn/loglogn). Learnability was previously known only in case that both k and j are constants. We also present a family of boolean functions that have short (poly(n)) Read-2-Satisfy-1 DNF formulae but require CNF formulae of size > 2W(n). Therefore, our result does not seem to follow from the recent learnability result of [Bsh93]. Avrim Blum, Roni Khardon, Eyal Kushilevitz, Leonard Pitt, Dan Roth 0001 |
COLT | 5 |
| 1993 | On Learning Visual Concepts and DNF FormulaeabstractWe consider the problem of learning visual concepts in the mistake-bound and the PAC models.We develop an approach that is shown to be useful also for the problem of learning DNF formulae in these models.As a result, we extend the class of learnable DNF (and CNF) formulae.The classes shown to be learnable are not limited in the number of terms or in the number of variables per term, and they contain the classes of k-DNF and kterm-DNF (and the corresponding classes of CNF) as special cases.We discuss some of the limitations of this approach and a number of applications. Eyal Kushilevitz, Dan Roth 0001 |
COLT | 2 |
| 1993 | On the Hardness of Approximate Reasoning
Dan Roth 0001 |
IJCAI | 1 |