EDBT 2026 Demo / reviewers in the wild / expert
Swarnadeep Saha
dblp:203/9296
· DBLP profile ↗
21ranked-venue papers
14as first author
13since 2021 · last 2025
0000-0002-6972-3448ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 13 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for ReasoningabstractLarge language model (LLM) reasoning can be improved by scaling test-time compute with aggregation, i.e., generating multiple samples and aggregating over them.While improving performance, this strategy often reaches a saturation point beyond which additional compute provides no return.Refinement offers an alternative by using model-generated feedback to improve answer quality.However, refinement faces three key challenges: (1) Excessive refinement: Uniformly refining all instances can cause over-correction and reduce overall performance.(2) Inability to localize and address errors: LLMs struggle to identify and correct their own mistakes.(3) Insufficient refinement: Stopping refinement too soon could leave errors unaddressed.To tackle these issues, we propose MAGICORE, a framework for Multi-Agent Iteration for Coarse-to-fine Refinement.MAGICORE mitigates excessive refinement by categorizing problems as easy or hard, solving easy problems with coarsegrained aggregation, and solving the hard ones with fine-grained multi-agent refinement.To better localize errors, we incorporate external step-wise reward model scores, and to ensure sufficient refinement, we iteratively refine the solutions using a multi-agent setup.We evaluate MAGICORE on Llama-3-8B and GPT-3.5 and show its effectiveness across seven reasoning datasets.One iteration of MAGI-CORE beats Self-Consistency by 3.4%, Bestof-k by 3.2%, and Self-Refine by 4.0% even when these baselines use k = 120, and MAGI-CORE uses less than 50% of the compute. 1 Justin Chih-Yao Chen, Archiki Prasad, Swarnadeep Saha, Elias Stengel-Eskin, Mohit Bansal |
EMNLP | 3 |
| 2025 | System 1.x: Learning to Balance Fast and Slow Planning with Language ModelsabstractLanguage models can be used to solve long-horizon planning problems in two distinct modes. In a fast 'System-1' mode, models directly generate plans without any explicit search or backtracking, and in a slow 'System-2' mode, they plan step-by-step by explicitly searching over possible actions. System-2 planning, while typically more effective, is also computationally more expensive and often infeasible for long plans or large action spaces. Moreover, isolated System-1 or System-2 planning ignores the user's end goals and constraints (e.g., token budget), failing to provide ways for the user to control the model's behavior. To this end, we propose the System-1.x Planner, a framework for controllable planning with language models that is capable of generating hybrid plans and balancing between the two planning modes based on the difficulty of the problem at hand. System-1.x consists of (i) a controller, (ii) a System-1 Planner, and (iii) a System-2 Planner. Based on a user-specified hybridization factor x governing the degree to which the system uses System-1 vs. System-2, the controller decomposes a planning problem into subgoals, and classifies them as easy or hard to be solved by either System-1 or System-2, respectively. We fine-tune all three components on top of a single base LLM, requiring only search traces as supervision. Experiments with two diverse planning tasks -- Maze Navigation and Blocksworld -- show that our System-1.x Planner outperforms a System-1 Planner, a System-2 Planner trained to approximate A* search, and also a symbolic planner (A* search), given a state exploration budget. We also demonstrate the following key properties of our planner: (1) controllability: by adjusting the hybridization factor x (e.g., System-1.75 vs. System-1.5) we can perform more (or less) search, improving performance, (2) flexibility: by building a neuro-symbolic variant composed of a neural System-1 planner and a symbolic System-2 planner, we can take advantage of existing symbolic methods, and (3) generalizability: by learning from different search algorithms (BFS, DFS, A*), we show that our method is robust to the choice of search algorithm used for training. Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen, Peter Hase, Elias Stengel-Eskin, Mohit Bansal |
ICLR | 1 |
| 2025 | Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-JudgeabstractLLM-as-a-Judge models generate chain-of-thought (CoT) sequences intended to capture the step-by-step reasoning process that underlies the final evaluation of a response. However, due to the lack of human-annotated CoTs for evaluation, the required components and structure of effective reasoning traces remain understudied. Consequently, previous approaches often (1) constrain reasoning traces to hand-designed components, such as a list of criteria, reference answers, or verification questions and (2) structure them such that planning is intertwined with the reasoning for evaluation. In this work, we propose EvalPlanner, a preference optimization algorithm for Thinking-LLM-as-a-Judge that first generates an unconstrained evaluation plan, followed by its execution, and then the final judgment. In a self-training loop, EvalPlanner iteratively optimizes over synthetically constructed evaluation plans and executions, leading to better final verdicts. Our method achieves a new state-of-the-art performance for generative reward models on RewardBench and PPE, despite being trained on fewer amount of, and synthetically generated, preference pairs. Additional experiments on other benchmarks like RM-Bench, JudgeBench, and FollowBenchEval further highlight the utility of both planning and reasoning for building robust LLM-as-a-Judge reasoning models. Swarnadeep Saha, Xian Li 0003, Marjan Ghazvininejad, Jason Weston |
ICML | 1 |
| 2024 | ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMsabstractLarge Language Models (LLMs) still struggle with natural language reasoning tasks.Motivated by the society of minds (Minsky, 1988), we propose RECONCILE, a multi-model multiagent framework designed as a round table conference among diverse LLM agents.RECON-CILE enhances collaborative reasoning between LLM agents via multiple rounds of discussion, learning to convince other agents to improve their answers, and employing a confidenceweighted voting mechanism that leads to a better consensus.In each round, RECONCILE initiates discussion between agents via a 'discussion prompt' that consists of (a) grouped answers and explanations generated by each agent in the previous round, (b) their confidence scores, and (c) demonstrations of answerrectifying human explanations, used for convincing other agents.Experiments on seven benchmarks demonstrate that RECONCILE significantly improves LLMs' reasoning -both individually and as a team -surpassing prior single-agent and multi-agent baselines by up to 11.4% and even outperforming GPT-4 on three datasets.RECONCILE also flexibly incorporates different combinations of agents, including API-based, open-source, and domainspecific models, leading to an 8% improvement on MATH.Finally, we analyze the individual components of RECONCILE, demonstrating that the diversity originating from different models is critical to its superior performance.1 Self-Refine MAD+Judge Multi-Agent Debate (MAD) ReConcile (Group-Discuss-and-Convince) Yes, with 95% confidence No, with 50% confidence No, with 40% confidence yes no no yes no no yes no no Question (Q): Is an ammonia fighting cleaner good for pet owners?Human Explanation (Exp): Ammonia is a component in pet urine.It has an unpleasant odor. Justin Chih-Yao Chen, Swarnadeep Saha, Mohit Bansal |
ACL (1) | 2 |
| 2024 | MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language ModelsabstractMulti-agent interactions between Large Language Model (LLM) agents have shown major improvements on diverse reasoning tasks. However, these involve long generations from multiple models across several rounds, making them expensive. Moreover, these multi-agent approaches fail to provide a final, single model for efficient inference. To address this, we introduce MAGDi, a new method for structured distillation of the reasoning interactions between multiple LLMs into smaller LMs. MAGDi teaches smaller models by representing multi-agent interactions as graphs, augmenting a base student model with a graph encoder, and distilling knowledge using three objective functions: next-token prediction, a contrastive loss between correct and incorrect reasoning, and a graph-based objective to model the interaction structure. Experiments on seven widely used commonsense and math reasoning benchmarks show that MAGDi improves the reasoning capabilities of smaller models, outperforming several methods that distill from a single teacher and multiple teachers. Moreover, MAGDi also demonstrates an order of magnitude higher efficiency over its teachers. We conduct extensive analyses to show that MAGDi (1) enhances the generalizability to out-of-domain tasks, (2) scales positively with the size and strength of the base student model, and (3) obtains larger improvements (via our multi-teacher training) when applying self-consistency – an inference technique that relies on model diversity. Justin Chih-Yao Chen, Swarnadeep Saha, Elias Stengel-Eskin, Mohit Bansal |
ICML | 2 |
| 2024 | Branch-Solve-Merge Improves Large Language Model Evaluation and GenerationabstractSwarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li 0003 |
NAACL-HLT | 1 |
| 2023 | ReCEval: Evaluating Reasoning Chains via Correctness and InformativenessabstractMulti-step reasoning ability is fundamental to many natural language tasks, yet it is unclear what constitutes a good reasoning chain and how to evaluate them.Most existing methods focus solely on whether the reasoning chain leads to the correct conclusion, but this answeroriented view may confound reasoning quality with other spurious shortcuts to predict the answer.To bridge this gap, we evaluate reasoning chains by viewing them as informal proofs that derive the final answer.Specifically, we propose RECEVAL (Reasoning Chain Evaluation), a framework that evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, i.e., each step provides new information that is helpful towards deriving the generated answer.We evaluate these properties by developing metrics using natural language inference models and V-Information.On multiple datasets, we show that RECEVAL effectively identifies various error types and yields notable improvements compared to prior methods.We analyze the impact of step boundaries, and previous steps on evaluating correctness and demonstrate that our informativeness metric captures the expected flow of information in high-quality reasoning chains.Finally, we show that scoring reasoning chains based on RECEVAL improves downstream task performance.1 Archiki Prasad, Swarnadeep Saha, Mohit Bansal |
EMNLP | 2 |
| 2023 | Summarization Programs: Interpretable Abstractive Summarization with Neural Modular Trees
Swarnadeep Saha, Shiyue Zhang 0001, Peter Hase, Mohit Bansal |
ICLR | 1 |
| 2023 | Can Language Models Teach? Teacher Explanations Improve Student Performance via PersonalizationabstractA hallmark property of explainable AI models is the ability to teach other agents, communicating knowledge of how to perform a task. While Large Language Models (LLMs) perform complex reasoning by generating explanations for their predictions, it is unclear whether they also make good teachers for weaker agents. To address this, we consider a student-teacher framework between two LLM agents and study if, when, and how the teacher should intervene with natural language explanations to improve the student’s performance. Since communication is expensive, we define a budget such that the teacher only communicates explanations for a fraction of the data, after which the student should perform well on its own. We decompose the teaching problem along four axes: (1) if teacher’s test time in- tervention improve student predictions, (2) when it is worth explaining a data point, (3) how the teacher should personalize explanations to better teach the student, and (4) if teacher explanations also improve student performance on future unexplained data. We first show that teacher LLMs can indeed intervene on student reasoning to improve their performance. Next, inspired by the Theory of Mind abilities of effective teachers, we propose building two few-shot mental models of the student. The first model defines an Intervention Function that simulates the utility of an intervention, allowing the teacher to intervene when this utility is the highest and improving student performance at lower budgets. The second model enables the teacher to personalize explanations for a particular student and outperform unpersonalized teachers. We also demonstrate that in multi-turn interactions, teacher explanations generalize and learning from explained data improves student performance on future unexplained data. Finally, we also verify that misaligned teachers can lower student performance to random chance by intentionally misleading them. Swarnadeep Saha, Peter Hase, Mohit Bansal |
NeurIPS | 1 |
| 2022 | Explanation Graph Generation via Pre-trained Language Models: An Empirical Study with Contrastive LearningabstractPre-trained sequence-to-sequence language models have led to widespread success in many natural language generation tasks.However, there has been relatively less work on analyzing their ability to generate structured outputs such as graphs.Unlike natural language, graphs have distinct structural and semantic properties in the context of a downstream NLP task, e.g., generating a graph that is connected and acyclic can be attributed to its structural constraints, while the semantics of a graph can refer to how meaningfully an edge represents the relation between two node concepts.In this work, we study pre-trained language models that generate explanation graphs in an end-to-end manner and analyze their ability to learn the structural constraints and semantics of such graphs.We first show that with limited supervision, pre-trained language models often generate graphs that either violate these constraints or are semantically incoherent.Since curating large amount of humanannotated graphs is expensive and tedious, we propose simple yet effective ways of graph perturbations via node and edge edit operations that lead to structurally and semantically positive and negative graphs.Next, we leverage these graphs in different contrastive learning models with Max-Margin and InfoNCE losses.Our methods lead to significant improvements in both structural and semantic accuracy of explanation graphs and also generalize to other similar graph generation tasks.Lastly, we show that human errors are the best negatives for contrastive learning and also that automatically generating more such human-like negative graphs can lead to further improvements.1 Swarnadeep Saha, Prateek Yadav, Mohit Bansal |
ACL (1) | 1 |
| 2022 | Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated ExplanationsabstractRecent work on explainable NLP has shown that few-shot prompting can enable large pretrained language models (LLMs) to generate grammatical and factual natural language explanations for data labels.In this work, we study the connection between explainability and sample hardness by investigating the following research question -"Are LLMs and humans equally good at explaining data labels for both easy and hard samples?"We answer this question by first collecting humanwritten explanations in the form of generalizable commonsense rules on the task of Winograd Schema Challenge (Winogrande dataset).We compare these explanations with those generated by GPT-3 while varying the hardness of the test samples as well as the in-context samples.We observe that (1) GPT-3 explanations are as grammatical as human explanations regardless of the hardness of the test samples, (2) for easy examples, GPT-3 generates highly supportive explanations but human explanations are more generalizable, and (3) for hard examples, human explanations are significantly better than GPT-3 explanations both in terms of label-supportiveness and generalizability judgements.We also find that hardness of the in-context examples impacts the quality of GPT-3 explanations.Finally, we show that the supportiveness and generalizability aspects of human explanations are also impacted by sample hardness, although by a much smaller margin than models. 1 Swarnadeep Saha, Peter Hase, Nazneen Fatema Rajani, Mohit Bansal |
EMNLP | 1 |
| 2021 | ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense ReasoningabstractRecent commonsense-reasoning tasks are typically discriminative in nature, where a model answers a multiple-choice question for a certain context.Discriminative tasks are limiting because they fail to adequately evaluate the model's ability to reason and explain predictions with underlying commonsense knowledge.They also allow such models to use reasoning shortcuts and not be "right for the right reasons".In this work, we present EX-PLAGRAPHS, a new generative and structured commonsense-reasoning task (and an associated dataset) of explanation graph generation for stance prediction.Specifically, given a belief and an argument, a model has to predict if the argument supports or counters the belief and also generate a commonsense-augmented graph that serves as non-trivial, complete, and unambiguous explanation for the predicted stance.We collect explanation graphs through a novel Create-Verify-And-Refine graph collection framework that improves the graph quality (up to 90%) via multiple rounds of verification and refinement.A significant 79% of our graphs contain external commonsense nodes with diverse structures and reasoning depths.Next, we propose a multi-level evaluation framework, consisting of automatic metrics and human evaluation, that check for the structural and semantic correctness of the generated graphs and their degree of match with ground-truth graphs.Finally, we present several structured, commonsense-augmented, and text generation models as strong starting points for this explanation graph generation task, and observe that there is a large gap with human performance, thereby encouraging future work for this new challenging task. 1 Swarnadeep Saha, Prateek Yadav, Lisa Bauer, Mohit Bansal |
EMNLP (1) | 1 |
| 2021 | multiPRover: Generating Multiple Proofs for Improved Interpretability in Rule ReasoningabstractWe focus on a type of linguistic formal reasoning where the goal is to reason over explicit knowledge in the form of natural language facts and rules (Clark et al., 2020).A recent work, named PROVER (Saha et al., 2020), performs such reasoning by answering a question and also generating a proof graph that explains the answer.However, compositional reasoning is not always unique and there may be multiple ways of reaching the correct answer.Thus, in our work, we address a new and challenging problem of generating multiple proof graphs for reasoning over natural language rule-bases.Each proof provides a different rationale for the answer, thereby improving the interpretability of such reasoning systems.In order to jointly learn from all proof graphs and exploit the correlations between multiple proofs for a question, we pose this task as a set generation problem over structured output spaces where each proof is represented as a directed graph.We propose two variants of a proof-set generation model, MULTIPROVER.Our first model, Multilabel-MULTIPROVER, generates a set of proofs via multi-label classification and implicit conditioning between the proofs; while the second model, Iterative-MULTIPROVER, generates proofs iteratively by explicitly conditioning on the previously generated proofs.Experiments on multiple synthetic, zero-shot, and human-paraphrased datasets reveal that both MULTIPROVER models significantly outperform PROVER on datasets containing multiple gold proofs.Iterative-MULTIPROVER obtains state-of-the-art proof F1 in zero-shot scenarios where all examples have single correct proofs.It also generalizes better to questions requiring higher depths of reasoning where multiple proofs are more frequent. Swarnadeep Saha, Prateek Yadav, Mohit Bansal |
NAACL-HLT | 1 |
| 2020 | PRover: Proof Generation for Interpretable Reasoning over RulesabstractRecent work by Clark et al. (2020) shows that transformers can act as "soft theorem provers" by answering questions over explicitly provided knowledge in natural language.In our work, we take a step closer to emulating formal theorem provers, by proposing PROVER, an interpretable transformer-based model that jointly answers binary questions over rule-bases and generates the corresponding proofs.Our model learns to predict nodes and edges corresponding to proof graphs in an efficient constrained training paradigm.During inference, a valid proof, satisfying a set of global constraints is generated.We conduct experiments on synthetic, hand-authored, and human-paraphrased rule-bases to show promising results for QA and proof generation, with strong generalization performance.First, PROVER generates proofs with an accuracy of 87%, while retaining or improving performance on the QA task, compared to RuleTakers (up to 6% improvement on zero-shot evaluation).Second, when trained on questions requiring lower depths of reasoning, it generalizes significantly better to higher depths (up to 15% improvement).Third, PROVER obtains near perfect QA accuracy of 98% using only 40% of the training data.However, generating proofs for questions requiring higher depths of reasoning becomes challenging, and the accuracy drops to 65% for "depth 5", indicating significant scope for future work. 1Facts : F 1 : The bald eagle eats the lion.F2: The bald eagle sees the tiger.F3: The lion chases the bald eagle.F 4 : The lion eats the mouse.F5: The mouse eats the tiger.F6: The tiger eats the bald eagle.F 7 : The tiger is red.Rules : R1: If the lion is green and the lion is not kind then the lion sees the bald eagle.R2: If someone sees the lion then they eat the mouse.R 3 : If someone is kind and not green then they see the bald eagle.R4: If someone is rough then they see the lion.R5: If someone sees the lion and they do not eat the tiger then the tiger is rough.R 6 : If someone eats the bald eagle and the bald eagle is not kind then the bald eagle is rough.R7: If someone does not eat the lion then the lion is big.R8: If someone is kind then they do not eat the mouse. Q4:The bald eagle eats the mouse. Swarnadeep Saha, Mohit Bansal |
EMNLP (1) | 1 |
| 2020 | ConjNLI: Natural Language Inference Over Conjunctive SentencesabstractReasoning about conjuncts in conjunctive sentences is important for a deeper understanding of conjunctions in English and also how their usages and semantics differ from conjunctive and disjunctive boolean logic.Existing NLI stress tests do not consider non-boolean usages of conjunctions and use templates for testing such model knowledge.Hence, we introduce CONJNLI, a challenge stress-test for natural language inference over conjunctive sentences, where the premise differs from the hypothesis by conjuncts removed, added, or replaced.These sentences contain single and multiple instances of coordinating conjunctions ("and", "or", "but", "nor") with quantifiers, negations, and requiring diverse boolean and non-boolean inferences over conjuncts.We find that large-scale pre-trained language models like RoBERTa do not understand conjunctive semantics well and resort to shallow heuristics to make inferences over such sentences.As some initial solutions, we first present an iterative adversarial fine-tuning method that uses synthetically created training data based on boolean and non-boolean heuristics.We also propose a direct model advancement by making RoBERTa aware of predicate semantic roles.While we observe some performance gains, CONJNLI is still challenging for current methods, thus encouraging interesting future work for better understanding of conjunctions. 1 Swarnadeep Saha, Yixin Nie, Mohit Bansal |
EMNLP (1) | 1 |
| 2019 | Pre-Training BERT on Domain Resources for Short Answer GradingabstractChul Sung, Tejas Dhamecha, Swarnadeep Saha, Tengfei Ma, Vinay Reddy, Rishi Arora. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chul Sung, Tejas I. Dhamecha, Swarnadeep Saha, Tengfei Ma 0001, Vinay Reddy, Rishi Arora |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Aligning Learning Outcomes to Learning Resources: A Lexico-Semantic Spatial ApproachabstractAligning Learning Outcomes (LO) to relevant portions of Learning Resources (LR) is necessary to help students quickly navigate within the recommended learning material. In general, the problem can be viewed as finding the relevant sections of a document (LR) that is pertinent to a broad question (LO). In this paper, we introduce the novel problem of aligning LOs (LO is usually a sentence long text) to relevant pages of LRs (LRs are in the form of slide decks). We observe that the set of relevant pages can be composed of multiple chunks (a chunk is a contiguous set of pages) and the same page of an LR might be relevant to multiple LOs. To this end, we develop a novel Lexico-Semantic Spatial approach that captures the lexical, semantic, and spatial aspects of the task, and also alleviates the limited availability of training data. Our approach first identifies the relevancy of a page to an LO by using lexical and semantic features from each page independently. The spatial model at a later stage exploits the dependencies between the sequence of pages in the LR to further improve the alignment task. We empirically establish the importance of the lexical, semantic, and spatial models within the proposed approach. We show that, on average, a student can navigate to a relevant page from the first predicted page by about four clicks within a 38 page slide deck, as compared to two clicks by human experts. Swarnadeep Saha, Malolan Chetlur, Tejas I. Dhamecha, K. Gayathri Wijayarathna, Red Mendoza, Paul Gagnon, Nabil Zary, Shantanu Godbole |
IJCAI | 1 |
| 2018 | Balancing Human Efforts and Performance of Student Response Analyzer in Dialog-Based Tutors
Tejas I. Dhamecha, Smit Marvaniya, Swarnadeep Saha, Renuka Sindhgatta, Bikram Sengupta |
AIED (1) | 3 |
| 2018 | Sentence Level or Token Level Features for Automatic Short Answer Grading?: Use Both
Swarnadeep Saha, Tejas I. Dhamecha, Smit Marvaniya, Renuka Sindhgatta, Bikram Sengupta |
AIED (1) | 1 |
| 2018 | Creating Scoring Rubric from Representative Student Answers for Improved Short Answer GradingabstractAutomatic short answer grading remains one of the key challenges of any dialog-based tutoring system due to the variability in the student answers. Typically, each question may have no or few expert authored exemplary answers which make it difficult to (1) generalize to all correct ways of answering the question, or (2) represent answers which are either partially correct or incorrect. In this paper, we propose an affinity propagation based clustering technique to obtain class-specific representative answers from the graded student answers. Our novelty lies in formulating the Scoring Rubric by incorporating class-specific representatives obtained after proposed clustering, selecting, and ranking of graded student answers. We experiment with baseline as well as stateof-the-art sentence-embedding based features to demonstrate the feature-agnostic utility of class-specific representative answers. Experimental evaluations on our large-scale industry dataset and a benchmarking dataset show that the Scoring Rubric significantly improves the classification performance of short answer grading. Smit Marvaniya, Swarnadeep Saha, Tejas I. Dhamecha, Peter W. Foltz, Renuka Sindhgatta, Bikram Sengupta |
CIKM | 2 |
| 2018 | Open Information Extraction from Conjunctive SentencesabstractWe develop CALM, a coordination analyzer that improves upon the conjuncts identified from dependency parses. It uses a language model based scoring and several linguistic constraints to search over hierarchical conjunct boundaries (for nested coordination). By splitting a conjunctive sentence around these conjuncts, CALM outputs several simple sentences. We demonstrate the value of our coordination analyzer in the end task of Open Information Extraction (Open IE). State-of-the-art Open IE systems lose substantial yield due to ineffective processing of conjunctive sentences. Our Open IE system, CALMIE, performs extraction over the simple sentences identified by CALM to obtain up to 1.8x yield with a moderate increase in precision compared to extractions from original sentences. Swarnadeep Saha, Mausam |
COLING | 1 |