Mirella Lapata

dblp:59/6701 · DBLP profile ↗
← Back
220ranked-venue papers
13as first author
60since 2021 · last 2026
0000-0002-2107-1516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 217 · 12 first-author · 60 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
abstract
Recent large language models support inputs of up to 10 million tokens, yet they perform poorly on long-context tasks that require complex reasoning.Such tasks can be solved using only a subset of the input -a proxy contextrather than the full sequence.Despite sharing the same underlying reasoning process, models exhibit a significant performance disparity between proxy and full contexts.To improve long-context reasoning, we propose ProxyCoT, a novel training framework that transfers reasoning capabilities from short proxy contexts to full long contexts.Specifically, we first obtain high-quality chain-of-thought reasoning traces on proxy contexts through reinforcement learning or distillation from a larger teacher model, and then ground the generated traces in full long contexts with supervised fine-tuning.Experiments across different datasets demonstrate that ProxyCoT consistently outperforms strong baselines with reduced computational overhead.Furthermore, models trained with ProxyCoT generalize their long-context reasoning capabilities to out-of-domain tasks. 1
Irina Saparina, Alexander Gurung, Mirella Lapata
ACL (1)4
2026 Generating Visual Stories with Grounded and Coreferent Characters
abstract
Abstract Characters are important in narratives. They move the plot forward, create emotional connections, and embody the story’s themes. Visual storytelling methods focus more on the plot and events relating to it, without building the narrative around specific characters. As a result, the generated stories feel generic, with character mentions being absent, vague, or incorrect. To mitigate these issues, we introduce a new character-centric approach to visual story generation. We present the first model capable of predicting visual stories with consistently grounded and coreferent character mentions. Our model is finetuned on a new dataset which we build on top of the widely used VIST (Huang et al., 2016) benchmark. Specifically, we develop an automated pipeline to enrich VIST with visual and textual character coreference chains. We also propose new evaluation metrics to measure the richness of characters and coreference in stories. Experimental results show that our model generates stories with recurring characters which are consistent and coreferent to larger extent compared to baselines and state-of-the-art systems.1 Our code and dataset are available at https://github.com/iz2late/character-centric-vist.
Mirella Lapata, Frank Keller
Trans. Assoc. Comput. Linguistics2
2025 What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
abstract
Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains. VISTA contains 18,599 recorded AI conference presentations paired with their corresponding paper abstracts. We benchmark the performance of state-of-the-art large models and apply a plan-based framework to better capture the structured nature of abstracts. Both human and automated evaluations confirm that explicit planning enhances summary quality and factual consistency. However, a considerable gap remains between models and human performance, highlighting the challenges of our dataset. This study aims to pave the way for future research on scientific video-to-text summarization.
Dongqi Liu 0001, Chenxi Whitehouse, Louis Mahon, Rohit Saxena, Zheng Zhao 0005, Yifu Qiu, Mirella Lapata, Vera Demberg
ACL (1)8
2025 Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
abstract
Multi-hop Question Answering (MHQA) adds layers of complexity to question answering, making it more challenging.When Language Models (LMs) are prompted with multiple search results, they are tasked not only with retrieving relevant information but also employing multi-hop reasoning across the information sources.Although LMs perform well on traditional question-answering tasks, the causal mask can hinder their capacity to reason across complex contexts.In this paper, we explore how LMs respond to multi-hop questions by permuting search results (retrieved documents) under various configurations.Our study reveals interesting findings as follows: 1) Encoderdecoder models, such as the ones in the Flan-T5 family, generally outperform causal decoderonly LMs in MHQA tasks, despite being significantly smaller in size; 2) altering the order of gold documents reveals distinct trends in both Flan T5 models and fine-tuned decoderonly models, with optimal performance observed when the document order aligns with the reasoning chain order; 3) enhancing causal decoder-only models with bi-directional attention by modifying the causal mask can effectively boost their end performance.In addition to the above, we conduct a thorough investigation of the distribution of LM attention weights in the context of MHQA.Our experiments reveal that attention weights tend to peak at higher values when the resulting answer is correct.We leverage this finding to heuristically improve LMs' performance on this task.Our code is publicly available at https://github. com/hwy9855/MultiHopQA-Reasoning.
Wenyu Huang, Pavlos Vougiouklis, Mirella Lapata, Jeff Z. Pan
ACL (1)3
2025 Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
abstract
Can LLMs provide support to creative writers by giving meaningful writing feedback?In this paper, we explore the challenges and limitations of model-generated writing feedback by defining a new task, dataset, and evaluation frameworks.To study model performance in a controlled manner, we present a novel test set of 1,300 stories that we corrupted to intentionally introduce writing issues.We study the performance of commonly used LLMs in this task with both automatic and human evaluation metrics.Our analysis shows that current models have strong out-of-the-box behavior in many respects-providing specific and mostly accurate writing feedback.However, models often fail to identify the biggest writing issue in the story and to correctly decide when to offer critical vs. positive feedback.
Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata
ACL (1)4
2025 Compositional Generalisation for Explainable Hate Speech Detection
abstract
Hate speech detection is key to online content moderation, but current models struggle to generalise beyond their training data.This has been linked to dataset biases and the use of sentence-level labels, which fail to teach models the underlying structure of hate speech.In this work, we show that even when models are trained with more fine-grained, spanlevel annotations (e.g., "artists" is labeled as target and "are parasites" as dehumanising comparison), they struggle to disentangle the meaning of these labels from the surrounding context.As a result, combinations of expressions that deviate from those seen during training remain particularly difficult for models to detect.We investigate whether training on a dataset where expressions occur with equal frequency across all contexts can improve generalisation.To this end, we create Unseen-PLEAD (U-PLEAD), a dataset of ∼364,000 synthetic posts, along with a novel compositional generalisation benchmark of ∼8,000 posts.Training on a combination of U-PLEAD and real data improves compositional generalisation while achieving state-of-the-art performance on the human-sourced PLEAD.
Agostina Calabrese, Tom Sherborne, Björn Ross, Mirella Lapata
EMNLP4
2025 Long-Form Information Alignment Evaluation Beyond Atomic Facts
abstract
Information alignment evaluators are vital for various NLG evaluation tasks and trustworthy LLM deployment, reducing hallucinations and enhancing user trust.Current fine-grained methods, like FactScore, verify facts individually but neglect inter-fact dependencies, enabling subtle vulnerabilities.In this work, we introduce MONTAGELIE, a challenging benchmark that constructs deceptive narratives by "montaging" truthful statements without introducing explicit hallucinations.We demonstrate that both coarse-grained LLM-based evaluators and current fine-grained frameworks are susceptible to this attack, with AUC-ROC scores falling below 65%.To enable more robust fine-grained evaluation, we propose DOVESCORE, a novel framework that jointly verifies factual accuracy and event-order consistency.By modeling inter-fact relationships, DOVESCORE outperforms existing finegrained methods by over 8%, providing a more robust solution for long-form text alignment evaluation.Our code and datasets are available at https://github.com/dannalily/DoveScore.
Danna Zheng, Mirella Lapata, Jeff Z. Pan
EMNLP2
2025 Faithfulness and Content Selection in Long-Input Multi-Document Summarisation of U.S. Civil Rights Litigation
abstract
Automatic summarisation is a promising tool for condensing lengthy legal texts. However, abstractive methods often hallucinate, undermining trust in legal settings, where accuracy and reliability are crucial. To address these challenges, we propose a content selection strategy prior to summarisation. We investigate (1) whether content selection improves the quality and faithfulness of summaries, and (2) whether domain-specific pretraining further enhances results. We show that salient content selection improves summary faithfulness by +0.2614 and achieves gains of up to +5.56 ROUGE-1 F1, +5.46 ROUGE-2 F1, +2.70 ROUGE-L F1, and +2.15 BERTScore. Furthermore, we demonstrate that legal-pretrained models outperform general ones, and we identify persistent hallucination sources, highlighting areas for future work.
Isabel Sebire, Claire Barale, Mirella Lapata
ICAIL3
2025 Agents' Room: Narrative Generation through Multi-step Collaboration
abstract
Writing compelling fiction is a multifaceted process combining elements such as crafting a plot, developing interesting characters, and using evocative language. While large language models (LLMs) show promise for story writing, they currently rely heavily on intricate prompting, which limits their use. We propose Agents' Room, a generation framework inspired by narrative theory, that decomposes narrative writing into subtasks tackled by specialized agents. To illustrate our method, we introduce Tell Me A Story, a high-quality dataset of complex writing prompts and human-written stories, and a novel evaluation framework designed specifically for assessing long narratives. We show that Agents' Room generates stories that are preferred by expert evaluators over those produced by baseline systems by leveraging collaboration and specialization to decompose the complex story writing task into tractable components. We provide extensive analysis with automated and human-based metrics of the generated output.
Fantine Huot, Reinald Kim Amplayo, Jennimaria Palomaki, Alice Shoshana Jakobovits, Elizabeth Clark, Mirella Lapata
ICLR6
2025 Prompting large language models with knowledge graphs for question answering involving long-tail facts
Wenyu Huang, Guancheng Zhou, Mirella Lapata, Pavlos Vougiouklis, Sébastien Montella, Jeff Z. Pan
Knowl. Based Syst.3
2025 Explanatory Summarization with Discourse-Driven Planning
abstract
Abstract Lay summaries for scientific documents typically include explanations to help readers grasp sophisticated concepts or arguments. However, current automatic summarization methods do not explicitly model explanations, which makes it difficult to align the proportion of explanatory content with human-written summaries. In this paper, we present a plan-based approach that leverages discourse frameworks to organize summary generation and guide explanatory sentences by prompting responses to the plan. Specifically, we propose two discourse-driven planning strategies, where the plan is conditioned as part of the input or part of the output prefix, respectively. Empirical experiments on three lay summarization datasets show that our approach outperforms existing state-of-the-art methods in terms of summary quality, and it enhances model robustness, controllability, and mitigates hallucination. The project information is available at https://dongqi.me/projects/ExpSum.
Dongqi Liu 0001, Vera Demberg, Mirella Lapata
Trans. Assoc. Comput. Linguistics4
2025 Dolomites: Domain-Specific Long-Form Methodical Tasks
abstract
Abstract Experts in various fields routinely perform methodical writing tasks to plan, organize, and report their work. From a clinician writing a differential diagnosis for a patient, to a teacher writing a lesson plan for students, these tasks are pervasive, requiring to methodically generate structured long-form output for a given input. We develop a typology of methodical tasks structured in the form of a task objective, procedure, input, and output, and introduce DoLoMiTes, a novel benchmark with specifications for 519 such tasks elicited from hundreds of experts from across 25 fields. Our benchmark further contains specific instantiations of methodical tasks with concrete input and output examples (1,857 in total) which we obtain by collecting expert revisions of up to 10 model-generated examples of each task. We use these examples to evaluate contemporary language models, highlighting that automating methodical tasks is a challenging long-form generation problem, as it requires performing complex inferences, while drawing upon the given context as well as domain knowledge. Our dataset is available at https://dolomites-benchmark.github.io/.
Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das 0001, Mirella Lapata, Christopher Alberti
Trans. Assoc. Comput. Linguistics9
2024 PixT3: Pixel-based Table-To-Text Generation
abstract
Table -to-text generation involves generating appropriate textual descriptions given structured tabular data.It has attracted increasing attention in recent years thanks to the popularity of neural network models and the availability of large-scale datasets.A common feature across existing methods is their treatment of the input as a string, i.e., by employing linearization techniques that do not always preserve information in the table, are verbose, and lack space efficiency.We propose to rethink data-to-text generation as a visual recognition task, removing the need for rendering the input in a string format.We present PixT3, a multimodal tableto-text model that overcomes the challenges of linearization and input size limitations encountered by existing models.PixT3 is trained with a new self-supervised learning objective to reinforce table structure awareness and is applicable to open-ended and controlled generation settings.Experiments on the ToTTo (Parikh et al., 2020a) and Logic2Text (Chen et al., 2020c) benchmarks show that PixT3 is competitive and, in some settings, superior to generators that operate solely on text. 1
Iñigo Alonso 0001, Eneko Agirre, Mirella Lapata
ACL (1)3
2024 Learning to Plan and Generate Text with Citations
abstract
Constanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Constanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata
ACL (1)7
2024 A Modular Approach for Multimodal Summarization of TV Shows
abstract
In this paper we address the task of summarizing television shows, which touches key areas in AI research: complex reasoning, multiple modalities, and long narratives.We present a modular approach where separate components perform specialized sub-tasks which we argue affords greater flexibility compared to end-toend methods.Our modules involve detecting scene boundaries, reordering scenes so as to minimize the number of cuts between different events, converting visual information to text, summarizing the dialogue in each scene, and fusing the scene summaries into a final summary for the entire episode.We also present a new metric, PRISMA (Precision and Recall Evaluation of Summary Facts), to measure both precision and recall of generated summaries, which we decompose into atomic facts.Tested on the recently released SummScreen3D dataset (Papalampidi and Lapata, 2023), our method produces higher quality summaries than comparison models, as measured with ROUGE and our new fact-based metric.
Louis Mahon, Mirella Lapata
ACL (1)2
2024 Little Red Riding Hood Goes around the Globe: Crosslingual Story Planning and Generation with Large Language Models
abstract
Previous work has demonstrated the effectiveness of planning for story generation exclusively in a monolingual setting focusing primarily on English. We consider whether planning brings advantages to automatic story generation across languages. We propose a new task of crosslingual story generation with planning and present a new dataset for this task. We conduct a comprehensive study of different plans and generate stories in several languages, by leveraging the creative and reasoning capabilities of large pretrained language models. Our results demonstrate that plans which structure stories into three acts lead to more coherent and interesting narratives, while allowing to explicitly control their content and structure.
Evgeniia Razumovskaia, Joshua Maynez, Annie Louis, Mirella Lapata, Shashi Narayan
LREC/COLING4
2024 μPLAN: Summarizing using a Content Plan as Cross-Lingual Bridge
abstract
Fantine Huot, Joshua Maynez, Chris Alberti, Reinald Kim Amplayo, Priyanka Agrawal, Constanza Fierro, Shashi Narayan, Mirella Lapata. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Fantine Huot, Joshua Maynez, Christopher Alberti, Reinald Kim Amplayo, Priyanka Agrawal, Constanza Fierro, Shashi Narayan, Mirella Lapata
EACL (1)8
2024 Improving Generalization in Semantic Parsing by Increasing Natural Language Variation
abstract
Text-to-SQL semantic parsing has made significant progress in recent years, with various models demonstrating impressive performance on the challenging Spider benchmark.However, it has also been shown that these models often struggle to generalize even when faced with small perturbations of previously (accurately) parsed expressions.This is mainly due to the linguistic form of questions in Spider which are overly specific, unnatural, and display limited variation.In this work, we use data augmentation to enhance the robustness of text-to-SQL parsers against natural language variations.Existing approaches generate question reformulations either via models trained on Spider or only introduce local changes.In contrast, we leverage the capabilities of large language models to generate more realistic and diverse questions.Using only a few prompts, we achieve a two-fold increase in the number of questions in Spider.Training on this augmented dataset yields substantial improvements on a range of evaluation sets, including robustness benchmarks and out-of-domain data. 1
Irina Saparina, Mirella Lapata
EACL (1)2
2024 Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning
abstract
We present Archer, a challenging bilingual textto-SQL dataset specific to complex reasoning, including arithmetic, commonsense and hypothetical reasoning.It contains 1,042 English questions and 1,042 Chinese questions, along with 521 unique SQL queries, covering 20 English databases across 20 domains.Notably, this dataset demonstrates a significantly higher level of complexity compared to existing publicly available datasets.Our evaluation shows that Archer challenges the capabilities of current state-of-the-art models, with a high-ranked model on the Spider leaderboard achieving only 6.73% execution accuracy on Archer test set.Thus, Archer presents a significant challenge for future research in this field.Which 4-cylinder car needs the most fuel to drive 300 miles?List how many gallons it needs, and its make and model. 开300英⾥耗油最多的四缸⻋的品牌和型号分别是什么,它 需要多少加仑的油?Commonsense Knowledge: Fuel used is calculated by divding distance driven by fuel consumption.SELECT B. Make, B.Model, 1.0 * 300 / mpg AS n_gallon FROM cars_data A JOIN car_names B ON A.Id=B.MakeId WHERE cylinders="4" ORDER BY mpg ASC LIMIT 1 Commonsense ReasoningIf all cars produced by the Daimler Benz company have 4cylinders, then in all 4-cylinder cars, which one needs the most fuel to drive 300 miles?Please list how many gallons it needs, along with its make and model.假如⽣产⾃奔驰公司的⻋都是四缸,开300英⾥耗油最多的 四缸⻋的品牌和型号分别是什么,它需要多少加仑的油? SELECT B.Make, B.Model, 1.0 * 300 / mpg AS n_gallon FROM cars_data A JOIN car_names B ON A.id=B.makeid JOIN model_list C ON B.model=C.modelJOIN car_makers D on C.maker=D.idWHERE D.fullname="Daimler Benz" or A.cylinders="4" ORDER BY mpg ASC LIMIT 1 Hypothetical Reasoning How much higher is the maximum power of a BMW car than the maximum power of a Fiat car? 宝⻢汽⻋的最⾼功率⽐⻜雅特汽⻋的最⾼功率⾼多少? SELECT MAX(horsepower) -(SELECT MAX (horsepower) FROM cars_data A JOIN car_names B ON A.id=B.makeid WHERE B.model="fiat") AS diff FROM cars_data A JOIN car_names B ON A.id=B.makeid WHERE B.model="bmw"
Danna Zheng, Mirella Lapata, Jeff Z. Pan
EACL (1)2
2024 Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts
abstract
One useful application of NLP models is to support people in reading complex text from unfamiliar domains (e.g., scientific articles).Simplifying the entire text makes it understandable but sometimes removes important details.On the contrary, helping adult readers understand difficult concepts in context can enhance their vocabulary and knowledge.In a preliminary human study, we first identify that lack of context and unfamiliarity with difficult concepts is a major reason for adult readers' difficulty with domain-specific text.We then introduce targeted concept simplification, a simplification task for rewriting text to help readers comprehend text containing unfamiliar concepts.We also introduce WIKIDOMAINS 1 , a new dataset of 22k definitions from 13 academic domains paired with a difficult concept within each definition.We benchmark the performance of open-source and commercial LLMs, and a simple dictionary baseline on this task across human judgments of ease of understanding and meaning preservation.Interestingly, our human judges preferred explanations about the difficult concept more than simplification of the concept phrase.Further, no single model achieved superior performance across all quality dimensions, and automated metrics also show low correlations with human evaluations of concept simplification (∼ 0.2), opening up rich avenues for research on personalized human reading comprehension support.* Work done as student researcher at Google DeepMind. 1 https://github.com/google-deepmind/wikidomains
Sumit Asthana, Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata
EMNLP5
2024 AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
abstract
Practical semantic parsers are expected to understand user utterances and map them to executable programs, even when these are ambiguous. We introduce a new benchmark, AMBROSIA, which we hope will inform and inspire the development of text-to-SQL parsers capable of recognizing and interpreting ambiguous requests. Our dataset contains questions showcasing three different types of ambiguity (scope ambiguity, attachment ambiguity, and vagueness), their interpretations, and corresponding SQL queries. In each case, the ambiguity persists even when the database context is provided. This is achieved through a novel approach that involves controlled generation of databases from scratch. We benchmark various LLMs on AMBROSIA, revealing that even the most advanced models struggle to identify and interpret ambiguity in questions.
Irina Saparina, Mirella Lapata
NeurIPS2
2024 Finding the Right Moment: Human-Assisted Trailer Creation via Task Composition
abstract
Movie trailers perform multiple functions: they introduce viewers to the story, convey the mood and artistic style of the film, and encourage audiences to see the movie. These diverse functions make trailer creation a challenging endeavor. In this work, we focus on finding trailer moments in a movie, i.e., shots that could be potentially included in a trailer. We decompose this task into two subtasks: narrative structure identification and sentiment prediction. We model movies as graphs, where nodes are shots and edges denote semantic relations between them. We learn these relations using joint contrastive training which distills rich textual information (e.g., characters, actions, situations) from screenplays. An unsupervised algorithm then traverses the graph and selects trailer moments from the movie that human judges prefer to ones selected by competitive supervised approaches. A main advantage of our algorithm is that it uses interpretable criteria, which allows us to deploy it in an interactive tool for trailer creation with a human in the loop. Our tool allows users to select trailer shots in under 30 minutes that are superior to fully automatic methods and comparable to (exclusive) manual selection by experts.
Pinelopi Papalampidi, Frank Keller, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Hierarchical Indexing for Retrieval-Augmented Opinion Summarization
abstract
Abstract We propose a method for unsupervised abstractive opinion summarization, that combines the attributability and scalability of extractive approaches with the coherence and fluency of Large Language Models (LLMs). Our method, HIRO, learns an index structure that maps sentences to a path through a semantically organized discrete hierarchy. At inference time, we populate the index and use it to identify and retrieve clusters of sentences containing popular opinions from input reviews. Then, we use a pretrained LLM to generate a readable summary that is grounded in these extracted evidential clusters. The modularity of our approach allows us to evaluate its efficacy at each stage. We show that HIRO learns an encoding space that is more semantically structured than prior work, and generates summaries that are more representative of the opinions in the input reviews. Human evaluation confirms that HIRO generates significantly more coherent, detailed, and accurate summaries.
Tom Hosking, Mirella Lapata
Trans. Assoc. Comput. Linguistics3
2023 Attributable and Scalable Opinion Summarization
abstract
We propose a method for unsupervised opinion summarization that encodes sentences from customer reviews into a hierarchical discrete latent space, then identifies common opinions based on the frequency of their encodings.We are able to generate both abstractive summaries by decoding these frequent encodings, and extractive summaries by selecting the sentences assigned to the same frequent encodings.Our method is attributable, because the model identifies sentences used to generate the summary as part of the summarization process.It scales easily to many hundreds of input reviews, because aggregation is performed in the latent space rather than over long sequences of tokens.We also demonstrate that our appraoch enables a degree of control, generating aspectspecific summaries by restricting the model to parts of the encoding space that correspond to desired aspects (e.g., location or food).Automatic and human evaluation on two datasets from different domains demonstrates that our method generates summaries that are more informative than prior work and better grounded in the input reviews.
Tom Hosking, Mirella Lapata
ACL (1)3
2023 Semantic Parsing for Conversational Question Answering over Knowledge Graphs
abstract
In this paper, we are interested in developing semantic parsers which understand natural language questions embedded in a conversation with a user and ground them to formal queries over definitions in a general purpose knowledge graph (KG) with very large vocabularies (covering thousands of concept names and relations, and millions of entities).To this end, we develop a dataset where user questions are annotated with SPARQL parses and system answers correspond to execution results thereof.We present two different semantic parsing approaches and highlight the challenges of the task: dealing with large vocabularies, modelling conversation context, predicting queries with multiple entities, and generalising to new questions at test time.We hope our dataset will serve as useful testbed for the development of conversational semantic parsers. 1
Laura Perez-Beltrachini, Parag Jain, Emilio Monti, Mirella Lapata
EACL4
2023 Conversational Semantic Parsing using Dynamic Context Graphs
abstract
In this paper we consider the task of conversational semantic parsing over general purpose knowledge graphs (KGs) with millions of entities, and thousands of relation-types.We focus on models which are capable of interactively mapping user utterances into executable logical forms (e.g., SPARQL) in the context of the conversational history.Our key idea is to represent information about an utterance and its context via a subgraph which is created dynamically, i.e., the number of nodes varies per utterance.Rather than treating the subgraph as a sequence, we exploit its underlying structure and encode it with a graph neural network which further allows us to represent a large number of (unseen) nodes.Experimental results show that dynamic context modeling is superior to static approaches, delivering performance improvements across the board (i.e., for simple and complex questions).Our results further confirm that modeling the structure of context is better at processing discourse information, (i.e., at handling ellipsis and resolving coreference) and longer interactions.
Parag Jain, Mirella Lapata
EMNLP2
2023 Text Summarization with Oracle Expectation
Yumo Xu, Mirella Lapata
ICLR2
2023 Retrieval Augmented Generation with Rich Answer Encoding
abstract
Wenyu Huang, Mirella Lapata, Pavlos Vougiouklis, Nikos Papasarantopoulos, Jeff Pan. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Wenyu Huang, Mirella Lapata, Pavlos Vougiouklis, Nikos Papasarantopoulos, Jeff Z. Pan
IJCNLP (1)2
2023 QAmeleon: Multilingual QA with Only 5 Examples
abstract
Abstract The availability of large, high-quality datasets has been a major driver of recent progress in question answering (QA). Such annotated datasets, however, are difficult and costly to collect, and rarely exist in languages other than English, rendering QA technology inaccessible to underrepresented languages. An alternative to building large monolingual training datasets is to leverage pre-trained language models (PLMs) under a few-shot learning setting. Our approach, QAmeleon, uses a PLM to automatically generate multilingual data upon which QA models are fine-tuned, thus avoiding costly annotation. Prompt tuning the PLM with only five examples per language delivers accuracy superior to translation-based baselines; it bridges nearly 60% of the gap between an English-only baseline and a fully-supervised upper bound fine-tuned on almost 50,000 hand-labeled examples; and consistently leads to improvements compared to directly fine-tuning a QA model on labeled examples in low resource settings. Experiments on the TyDiqa-GoldP and MLQA benchmarks show that few-shot prompt tuning for data synthesis scales across languages and is a viable alternative to large-scale annotation.1
Priyanka Agrawal, Christopher Alberti, Fantine Huot, Joshua Maynez, Ji Ma 0004, Sebastian Ruder, Kuzman Ganchev, Dipanjan Das 0001, Mirella Lapata
Trans. Assoc. Comput. Linguistics9
2023 Conditional Generation with a Question-Answering Blueprint
abstract
Abstract The ability to convey relevant and faithful information is critical for many tasks in conditional generation and yet remains elusive for neural seq-to-seq models whose outputs often reveal hallucinations and fail to correctly cover important details. In this work, we advocate planning as a useful intermediate representation for rendering conditional generation less opaque and more grounded. We propose a new conceptualization of text plans as a sequence of question-answer (QA) pairs and enhance existing datasets (e.g., for summarization) with a QA blueprint operating as a proxy for content selection (i.e., what to say) and planning (i.e., in what order). We obtain blueprints automatically by exploiting state-of-the-art question generation technology and convert input-output pairs into input-blueprint-output tuples. We develop Transformer-based models, each varying in how they incorporate the blueprint in the generated output (e.g., as a global plan or iteratively). Evaluation across metrics and datasets demonstrates that blueprint models are more factual than alternatives which do not resort to planning and allow tighter control of the generation output.
Shashi Narayan, Joshua Maynez, Reinald Kim Amplayo, Kuzman Ganchev, Annie Louis, Fantine Huot, Anders Sandholm 0001, Dipanjan Das 0001, Mirella Lapata
Trans. Assoc. Comput. Linguistics9
2023 Optimal Transport Posterior Alignment for Cross-lingual Semantic Parsing
abstract
Abstract Cross-lingual semantic parsing transfers parsing capability from a high-resource language (e.g., English) to low-resource languages with scarce training data. Previous work has primarily considered silver-standard data augmentation or zero-shot methods; exploiting few-shot gold data is comparatively unexplored. We propose a new approach to cross-lingual semantic parsing by explicitly minimizing cross-lingual divergence between probabilistic latent variables using Optimal Transport. We demonstrate how this direct guidance improves parsing from natural languages using fewer examples and less training. We evaluate our method on two datasets, MTOP and MultiATIS++SQL, establishing state-of-the-art results under a few-shot cross-lingual regime. Ablation studies further reveal that our method improves performance even without parallel input translations. In addition, we show that our model better captures cross-lingual structure in the latent space to improve semantic representation similarity.1
Tom Sherborne, Tom Hosking, Mirella Lapata
Trans. Assoc. Comput. Linguistics3
2023 Meta-Learning a Cross-lingual Manifold for Semantic Parsing
abstract
Abstract Localizing a semantic parser to support new languages requires effective cross-lingual generalization. Recent work has found success with machine-translation or zero-shot methods, although these approaches can struggle to model how native speakers ask questions. We consider how to effectively leverage minimal annotated examples in new languages for few-shot cross-lingual semantic parsing. We introduce a first-order meta-learning algorithm to train a semantic parser with maximal sample efficiency during cross-lingual transfer. Our algorithm uses high-resource languages to train the parser and simultaneously optimizes for cross-lingual generalization to lower-resource languages. Results across six languages on ATIS demonstrate that our combination of generalization steps yields accurate semantic parsers sampling ≤10% of source training data in each new language. Our approach also trains a competitive model on Spider using English with generalization to Chinese similarly sampling ≤10% of training data.1
Tom Sherborne, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2022 Hierarchical Sketch Induction for Paraphrase Generation
abstract
We propose a generative model of paraphrase generation, that encourages syntactic diversity by conditioning on an explicit syntactic sketch.We introduce Hierarchical Refinement Quantized Variational Autoencoders (HRQ-VAE), a method for learning decompositions of dense encodings as a sequence of discrete latent variables that make iterative refinements of increasing granularity.This hierarchy of codes is learned through end-to-end training, and represents fine-to-coarse grained information about the input.We use HRQ-VAE to encode the syntactic form of an input sentence as a path through the hierarchy, allowing us to more easily predict syntactic sketches at test time.Extensive experiments, including a human evaluation, confirm that HRQ-VAE learns a hierarchical representation of the input space, and generates paraphrases of higher quality than previous systems.
Tom Hosking, Mirella Lapata
ACL (1)3
2022 A Well-Composed Text is Half Done! Composition Sampling for Diverse Conditional Generation
abstract
Shashi Narayan, Gonçalo Simões, Yao Zhao, Joshua Maynez, Dipanjan Das, Michael Collins, Mirella Lapata. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shashi Narayan, Gonçalo Simões, Joshua Maynez, Dipanjan Das 0001, Michael Collins 0001, Mirella Lapata
ACL (1)7
2022 Zero-Shot Cross-lingual Semantic Parsing
abstract
Recent work in cross-lingual semantic parsing has successfully applied machine translation to localize parsers to new languages.However, these advances assume access to highquality machine translation systems and word alignment tools.We remove these assumptions and study cross-lingual semantic parsing as a zero-shot problem, without parallel data (i.e., utterance-logical form pairs) for new languages.We propose a multi-task encoderdecoder model to transfer parsing knowledge to additional languages using only Englishlogical form paired data and in-domain natural language corpora in each new language.Our model encourages language-agnostic encodings by jointly optimizing for logical-form generation with auxiliary objectives designed for cross-lingual latent representation alignment.Our parser performs significantly above translation-based baselines and, in some cases, competes with the supervised upper-bound. 1
Tom Sherborne, Mirella Lapata
ACL (1)2
2022 Disentangled Sequence to Sequence Learning for Compositional Generalization
abstract
There is mounting evidence that existing neural network models, in particular the very popular sequence-to-sequence architecture, struggle to systematically generalize to unseen compositions of seen components.We demonstrate that one of the reasons hindering compositional generalization relates to representations being entangled.We propose an extension to sequenceto-sequence models which encourages disentanglement by adaptively re-encoding (at each time step) the source input.Specifically, we condition the source representations on the newly decoded target context which makes it easier for the encoder to exploit specialized information for each prediction rather than capturing it all in a single forward pass.Experimental results on semantic parsing and machine translation empirically show that our proposal delivers more disentangled representations and better generalization.1 1 Our code is available at https://github.com/ mswellhao/Dangle.Training Set A boy ate the cake on the table in a house.*cake(x4); *table(x7); boy(x1) AND eat.agent(x2, x1) AND eat.theme(x2, x4) AND cake.nmod.on(x4,x7) AND table.nmod.in(x7,x10) AND house(x10) Test Set (Lexical Generalization) A boy likes the cake on the table in a house.*cake(x4); *table(x7); boy(x1) AND like.agent(x2,x1) AND like.theme(x2,x4) AND cake.nmod.on(x4,x7) AND table.nmod.in(x7,x10) AND house(x10) Test Set (Structural Generalization) A boy ate the cake on the table in a house beside the tree.*cake(x4); *table(x7); *tree(x13); boy(x1) AND eat.agent(x2, x1) AND eat.theme(x2, x4) AND cake.nmod.on(x4,x7) AND table.nmod.in(x7,x10) AND house(x10) AND house.nmod.beside(x10,x13)
Mirella Lapata
ACL (1)2
2022 Scaling Structured Inference with Randomization
abstract
Deep discrete structured models have seen considerable progress recently, but traditional inference using dynamic programming (DP) typically works with a small number of states (less than hundreds), which severely limits model capacity. At the same time, across machine learning, there is a recent trend of using randomized truncation techniques to accelerate computations involving large sums. Here, we propose a family of randomized dynamic programming (RDP) algorithms for scaling structured models to tens of thousands of latent states. Our method is widely applicable to classical DP-based inference (partition, marginal, reparameterization, entropy) and different graph structures (chains, trees, and more general hypergraphs). It is also compatible with automatic differentiation: it can be integrated with neural networks seamlessly and learned with gradient-based optimizers. Our core technique approximates the sum-product by restricting and reweighting DP on a small subset of nodes, which reduces computation by orders of magnitude. We further achieve low bias and variance via Rao-Blackwellization and importance sampling. Experiments over different graphs demonstrate the accuracy and efficiency of our approach. Furthermore, when using RDP for training a structured variational autoencoder with a scaled inference network, we achieve better test likelihood than baselines and successfully prevent posterior collapse.
John P. Cunningham, Mirella Lapata
ICML3
2022 Explainable Abuse Detection as Intent Classification and Slot Filling
abstract
Abstract To proactively offer social media users a safe online experience, there is a need for systems that can detect harmful posts and promptly alert platform moderators. In order to guarantee the enforcement of a consistent policy, moderators are provided with detailed guidelines. In contrast, most state-of-the-art models learn what abuse is from labeled examples and as a result base their predictions on spurious cues, such as the presence of group identifiers, which can be unreliable. In this work we introduce the concept of policy-aware abuse detection, abandoning the unrealistic expectation that systems can reliably learn which phenomena constitute abuse from inspecting the data alone. We propose a machine-friendly representation of the policy that moderators wish to enforce, by breaking it down into a collection of intents and slots. We collect and annotate a dataset of 3,535 English posts with such slots, and show how architectures for intent classification and slot filling can be used for abuse detection, while providing a rationale for model decisions.1
Agostina Calabrese, Björn Ross, Mirella Lapata
Trans. Assoc. Comput. Linguistics3
2022 Data-to-text Generation with Variational Sequential Planning
abstract
Abstract We consider the task of data-to-text generation, which aims to create textual output from non-linguistic input. We focus on generating long-form text, that is, documents with multiple paragraphs, and propose a neural model enhanced with a planning component responsible for organizing high-level information in a coherent and meaningful way. We infer latent plans sequentially with a structured variational model, while interleaving the steps of planning and generation. Text is generated by conditioning on previous variational decisions and previously generated text. Experiments on two data-to-text benchmarks (RotoWire and MLB) show that our model outperforms strong baselines and is sample-efficient in the face of limited training data (e.g., a few hundred instances).
Ratish Puduppully, Mirella Lapata
Trans. Assoc. Comput. Linguistics3
2022 Document Summarization with Latent Queries
abstract
Abstract The availability of large-scale datasets has driven the development of neural models that create generic summaries for single or multiple documents. For query-focused summarization (QFS), labeled training data in the form of queries, documents, and summaries is not readily available. We provide a unified modeling framework for any kind of summarization, under the assumption that all summaries are a response to a query, which is observed in the case of QFS and latent in the case of generic summarization. We model queries as discrete latent variables over document tokens, and learn representations compatible with observed and unobserved query verbalizations. Our framework formulates summarization as a generative process, and jointly optimizes a latent query model and a conditional language model. Despite learning from generic summarization data only, our approach outperforms strong comparison systems across benchmarks, query types, document settings, and target domains.1
Yumo Xu, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2021 Unsupervised Opinion Summarization with Content Planning
abstract
The recent success of deep learning techniques for abstractive summarization is predicated on the availability of large-scale datasets. When summarizing reviews (e.g., for products or movies), such training data is neither available nor can be easily sourced, motivating the development of methods which rely on synthetic datasets for supervised training. We show that explicitly incorporating content planning in a summarization model not only yields output of higher quality, but also allows the creation of synthetic datasets which are more natural, resembling real world document-summary pairs. Our content plans take the form of aspect and sentiment distributions which we induce from data without access to expensive annotations. Synthetic datasets are created by sampling pseudo-reviews from a Dirichlet distribution parametrized by our content planner, while our model generates summaries based on input reviews and induced content plans. Experimental results on three domains show that our approach outperforms competitive models in generating informative, coherent, and fluent summaries that capture opinion consensus.
Reinald Kim Amplayo, Stefanos Angelidis, Mirella Lapata
AAAI3
2021 Movie Summarization via Sparse Graph Construction
abstract
We summarize full-length movies by creating shorter videos containing their most informative scenes. We explore the hypothesis that a summary can be created by assembling scenes which are turning points (TPs), i.e., key events in a movie that describe its storyline. We propose a model that identifies TP scenes by building a sparse movie graph that represents relations between scenes and is constructed using multimodal information. According to human judges, the summaries created by our approach are more informative and complete, and receive higher ratings, than the outputs of sequence-based models and general-purpose summarization algorithms. The induced graphs are interpretable, displaying different topology for different movie genres.
Pinelopi Papalampidi, Frank Keller, Mirella Lapata
AAAI3
2021 Exploring Explainable Selection to Control Abstractive Summarization
abstract
Like humans, document summarization models can interpret a document’s contents in a number of ways. Unfortunately, the neural models of today are largely black boxes that provide little explanation of how or why they generated a summary in the way they did. Therefore, to begin prying open the black box and to inject a level of control into the substance of the final summary, we developed a novel select-and-generate framework that focuses on explainability. By revealing the latent centrality and interactions between sentences, along with scores for novelty and relevance, users are given a window into the choices a model is making and an opportunity to guide those choices in a more desirable direction. A novel pair-wise matrix captures the sentence interactions, centrality and attribute scores, and a mask with tunable attribute thresholds allows the user to control which sentences are likely to be included in the extraction. A sentence-deployed attention mechanism in the abstractor ensures the final summary emphasizes the desired content. Additionally, the encoder is adaptable, supporting both Transformer- and BERT-based configurations. In a series of experiments assessed with ROUGE metrics and two human evaluations, ESCA outperformed eight state-of-the-art models on the CNN/DailyMail and NYT50 benchmark datasets.
Yang Gao 0016, Yu Bai 0018, Mirella Lapata, Heyan Huang
AAAI4
2021 Factorising Meaning and Form for Intent-Preserving Paraphrasing
abstract
Tom Hosking, Mirella Lapata. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tom Hosking, Mirella Lapata
ACL/IJCNLP (1)2
2021 Generating Query Focused Summaries from Query-Free Resources
abstract
Yumo Xu, Mirella Lapata. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yumo Xu, Mirella Lapata
ACL/IJCNLP (1)2
2021 Categorization in the Wild: Category and Feature Learning across Languages
Lea Frermann, Mirella Lapata
CogSci2
2021 Informative and Controllable Opinion Summarization
abstract
Opinion summarization is the task of automatically generating summaries for a set of reviews about a specific target (e.g., a movie or a product).Since the number of reviews for each target can be prohibitively large, neural network-based methods follow a two-stage approach where an extractive step first pre-selects a subset of salient opinions and an abstractive step creates the summary while conditioning on the extracted subset.However, the extractive model leads to loss of information which may be useful depending on user needs.In this paper we propose a summarization framework that eliminates the need to rely only on pre-selected content and waste possibly useful information, especially when customizing summaries.The framework enables the use of all input reviews by first condensing them into multiple dense vectors which serve as input to an abstractive model.We showcase an effective instantiation of our framework which produces more informative summaries and also allows to take user preferences into account using our zero-shot customization technique.Experimental results demonstrate that our model improves the state of the art on the Rotten Tomatoes dataset and generates customized summaries effectively."Coach Carter" Reviews • Samuel L. Jackson plays the real-life coach of a high school basketball team in this solid sports drama ...• Great performance by Samuel Jackson but predictable as a slam dunk ...• ... excellent basketball choreography, Coach Carter is fun, hopeful, occasionally silly and, what can I say, inspiring.Consensus Summary Even though it's based on a true story, Coach Carter is pretty formulaic stuff, but it's effective and energetic, thanks to a strong central performance from Samuel L. Jackson.EXTRACT-ABSTRACT Framework Coach Carter is a preposterously plotted thriller that
Reinald Kim Amplayo, Mirella Lapata
EACL2
2021 Aspect-Controllable Opinion Summarization
abstract
Recent work on opinion summarization produces general summaries based on a set of input reviews and the popularity of opinions expressed in them.In this paper, we propose an approach that allows the generation of customized summaries based on aspect queries (e.g., describing the location and room of a hotel).Using a review corpus, we create a synthetic training dataset of (review, summary) pairs enriched with aspect controllers which are induced by a multi-instance learning model that predicts the aspects of a document at different levels of granularity.We fine-tune a pretrained model using our synthetic dataset and generate aspect-specific summaries by modifying the aspect controllers.Experiments on two benchmarks show that our model outperforms the previous state of the art and generates personalized summaries by controlling the number of aspects discussed in them.
Reinald Kim Amplayo, Stefanos Angelidis, Mirella Lapata
EMNLP (1)3
2021 Learning Opinion Summarizers by Selecting Informative Reviews
abstract
Opinion summarization has been traditionally approached with unsupervised, weaklysupervised and few-shot learning techniques.In this work, we collect a large dataset of summaries paired with user reviews for over 31,000 products, enabling supervised training.However, the number of reviews per product is large (320 on average), making summarization -and especially training a summarizerimpractical.Moreover, the content of many reviews is not reflected in the human-written summaries, and, thus, the summarizer trained on random review subsets hallucinates.In order to deal with both of these challenges, we formulate the task as jointly learning to select informative subsets of reviews and summarizing the opinions expressed in these subsets.The choice of the review subset is treated as a latent variable, predicted by a small and simple selector.The subset is then fed into a more powerful summarizer.For joint training, we use amortized variational inference and policy gradient methods.Our experiments demonstrate the importance of selecting informative reviews resulting in improved quality of summaries and reduced hallucinations.
Arthur Brazinskas, Mirella Lapata, Ivan Titov 0001
EMNLP (1)2
2021 Models and Datasets for Cross-Lingual Summarisation
abstract
We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language.The corpus covers twelve language pairs and directions for four European languages, namely Czech, English, French and German, and the methodology for its creation can be applied to several other languages.We derive cross-lingual document-summary instances from Wikipedia by combining lead paragraphs and articles' bodies from language aligned Wikipedia titles.We analyse the proposed cross-lingual summarisation task with automatic metrics and validate it with a human study.To illustrate the utility of our dataset we report experiments with multi-lingual pretrained models in supervised, zero-and fewshot, and out-of-domain scenarios.
Laura Perez-Beltrachini, Mirella Lapata
EMNLP (1)2
2021 Text Generation from Discourse Representation Structures
abstract
We propose neural models to generate text from formal meaning representations based on Discourse Representation Structures (DRSs).DRSs are document-level representations which encode rich semantic detail pertaining to rhetorical relations, presupposition, and co-reference within and across sentences.We formalize the task of neural DRS-to-text generation and provide modeling solutions for the problems of condition ordering and variable naming which render generation from DRSs non-trivial.Our generator relies on a novel sibling treeLSTM model which is able to accurately represent DRS structures and is more generally suited to trees with wide branches.We achieve competitive performance (59.48 BLEU) on the GMB benchmark against several strong baselines.
Jiangming Liu, Shay B. Cohen, Mirella Lapata
NAACL-HLT3
2021 Noisy Self-Knowledge Distillation for Text Summarization
abstract
In this paper we apply self-knowledge distillation to text summarization which we argue can alleviate problems with maximumlikelihood training on single reference and noisy datasets.Instead of relying on one-hot annotation labels, our student summarization model is trained with guidance from a teacher which generates smoothed labels to help regularize training.Furthermore, to better model uncertainty during training, we introduce multiple noise signals for both teacher and student models.We demonstrate experimentally on three benchmarks that our framework boosts the performance of both pretrained and nonpretrained summarizers achieving state-of-theart results. 1
Yang Liu 0124, Sheng Shen 0001, Mirella Lapata
NAACL-HLT3
2021 Meta-Learning for Domain Generalization in Semantic Parsing
abstract
The importance of building semantic parsers which can be applied to new domains and generate programs unseen at training has long been acknowledged, and datasets testing outof-domain performance are becoming increasingly available.However, little or no attention has been devoted to learning algorithms or objectives which promote domain generalization, with virtually all existing approaches relying on standard supervised learning.In this work, we use a meta-learning framework which targets zero-shot domain generalization for semantic parsing.We apply a modelagnostic training algorithm that simulates zeroshot parsing by constructing virtual train and test sets from disjoint domains.The learning objective capitalizes on the intuition that gradient steps that improve source-domain performance should also improve target-domain performance, thus encouraging a parser to generalize to unseen target domains.Experimental results on the (English) Spider and Chinese Spider datasets show that the meta-learning objective significantly boosts the performance of a baseline parser.
Bailin Wang, Mirella Lapata, Ivan Titov 0001
NAACL-HLT2
2021 Learning from Executions for Semantic Parsing
abstract
Semantic parsing aims at translating natural language (NL) utterances onto machineinterpretable programs, which can be executed against a real-world environment.The expensive annotation of utterance-program pairs has long been acknowledged as a major bottleneck for the deployment of contemporary neural models to real-life applications.In this work, we focus on the task of semi-supervised learning where a limited amount of annotated data is available together with many unlabeled NL utterances.Based on the observation that programs which correspond to NL utterances must be always executable, we propose to encourage a parser to generate executable programs for unlabeled utterances.Due to the large search space of executable programs, conventional methods that use approximations based on beam-search such as self-training and top-k marginal likelihood training, do not perform as well.Instead, we view the problem of learning from executions from the perspective of posterior regularization and propose a set of new training objectives.Experimental results on OVERNIGHT and GEOQUERY show that our new objectives outperform conventional methods, bridging the gap between semi-supervised and supervised learning.
Bailin Wang, Mirella Lapata, Ivan Titov 0001
NAACL-HLT2
2021 Structured Reordering for Modeling Latent Alignments in Sequence Transduction
abstract
Despite success in many domains, neural models struggle in settings where train and test examples are drawn from different distributions. In particular, in contrast to humans, conventional sequence-to-sequence (seq2seq) models fail to generalize systematically, i.e., interpret sentences representing novel combinations of concepts (e.g., text segments) seen in training. Traditional grammar formalisms excel in such settings by implicitly encoding alignments between input and output segments, but are hard to scale and maintain. Instead of engineering a grammar, we directly model segment-to-segment alignments as discrete structured latent variables within a neural seq2seq model. To efficiently explore the large space of alignments, we introduce a reorder-first align-later framework whose central component is a neural reordering module producing separable permutations. We present an efficient dynamic programming algorithm performing exact marginal inference of separable permutations, and, thus, enabling end-to-end differentiable training of our model. The resulting seq2seq model exhibits better systematic generalization than standard models on synthetic problems and NLP tasks (i.e., semantic parsing and machine translation).
Bailin Wang, Mirella Lapata, Ivan Titov 0001
NeurIPS2
2021 Universal Discourse Representation Structure Parsing
abstract
Abstract We consider the task of crosslingual semantic parsing in the style of Discourse Representation Theory (DRT) where knowledge from annotated corpora in a resource-rich language is transferred via bitext to guide learning in other languages. We introduce Universal Discourse Representation Theory (UDRT), a variant of DRT that explicitly anchors semantic representations to tokens in the linguistic input. We develop a semantic parsing framework based on the Transformer architecture and utilize it to obtain semantic resources in multiple languages following two learning schemes. The Many-to-One approach translates non-English text to English, and then runs a relatively accurate English parser on the translated text, while the One-to-Many approach translates gold standard English to non-English text and trains multiple parsers (one per language) on the translations. Experimental results on the Parallel Meaning Bank show that our proposal outperforms strong baselines by a wide margin and can be used to construct (silver-standard) meaning banks for 99 languages.
Jiangming Liu, Shay B. Cohen, Mirella Lapata, Johan Bos
Comput. Linguistics3
2021 Multi-Document Summarization with Determinantal Point Process Attention
abstract
The ability to convey relevant and diverse information is critical in multi-document summarization and yet remains elusive for neural seq-to-seq models whose outputs are often redundant and fail to correctly cover important details. In this work, we propose an attention mechanism which encourages greater focus on relevance and diversity. Attention weights are computed based on (proportional) probabilities given by Determinantal Point Processes (DPPs) defined on the set of content units to be summarized. DPPs have been successfully used in extractive summarisation, here we use them to select relevant and diverse content for neural abstractive summarisation. We integrate DPP-based attention with various seq-to-seq architectures ranging from CNNs to LSTMs, and Transformers. Experimental evaluation shows that our attention mechanism consistently improves summarization and delivers performance comparable with the state-of-the-art on the MultiNews dataset
Laura Perez-Beltrachini, Mirella Lapata
J. Artif. Intell. Res.2
2021 Extractive Opinion Summarization in Quantized Transformer Spaces
abstract
Abstract We present the Quantized Transformer (QT), an unsupervised system for extractive opinion summarization. QT is inspired by Vector- Quantized Variational Autoencoders, which we repurpose for popularity-driven summarization. It uses a clustering interpretation of the quantized space and a novel extraction algorithm to discover popular opinions among hundreds of reviews, a significant step towards opinion summarization of practical scope. In addition, QT enables controllable summarization without further training, by utilizing properties of the quantized space to extract aspect-specific summaries. We also make publicly available Space, a large-scale evaluation benchmark for opinion summarizers, comprising general and aspect-specific summaries for 50 hotels. Experiments demonstrate the promise of our approach, which is validated by human studies where judges showed clear preference for our method over competitive baselines.
Stefanos Angelidis, Reinald Kim Amplayo, Yoshihiko Suhara, Xiaolan Wang 0001, Mirella Lapata
Trans. Assoc. Comput. Linguistics5
2021 Memory-Based Semantic Parsing
abstract
Abstract We present a memory-based model for context- dependent semantic parsing. Previous approaches focus on enabling the decoder to copy or modify the parse from the previous utterance, assuming there is a dependency between the current and previous parses. In this work, we propose to represent contextual information using an external memory. We learn a context memory controller that manages the memory by maintaining the cumulative meaning of sequential user utterances. We evaluate our approach on three semantic parsing benchmarks. Experimental results show that our model can better process context-dependent information and demonstrates improved performance without using task-specific decoders.
Parag Jain, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2021 Data-to-text Generation with Macro Planning
abstract
Abstract Recent approaches to data-to-text generation have adopted the very successful encoder-decoder architecture or variants thereof. These models generate text that is fluent (but often imprecise) and perform quite poorly at selecting appropriate content and ordering it coherently. To overcome some of these issues, we propose a neural model with a macro planning stage followed by a generation stage reminiscent of traditional methods which embrace separate modules for planning and surface realization. Macro plans represent high level organization of important content such as entities, events, and their interactions; they are learned from data and given as input to the generator. Extensive experiments on two data-to-text benchmarks (RotoWire and MLB) show that our approach outperforms competitive baselines in terms of automatic and human evaluation.
Ratish Puduppully, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2020 Unsupervised Opinion Summarization with Noising and Denoising
abstract
The supervised training of high-capacity models on large datasets containing hundreds of thousands of document-summary pairs is critical to the recent success of deep learning techniques for abstractive summarization.Unfortunately, in most domains (other than news) such training data is not available and cannot be easily sourced.In this paper we enable the use of supervised learning for the setting where there are only documents available (e.g., product or business reviews) without ground truth summaries.We create a synthetic dataset from a corpus of user reviews by sampling a review, pretending it is a summary, and generating noisy versions thereof which we treat as pseudo-review input.We introduce several linguistically motivated noise generation functions and a summarization model which learns to denoise the input and generate the original review.At test time, the model accepts genuine reviews and generates a summary containing salient opinions, treating those that do not reach consensus as noise.Extensive automatic and human evaluation shows that our model brings substantial improvements over both abstractive and extractive baselines.
Reinald Kim Amplayo, Mirella Lapata
ACL2
2020 Unsupervised Opinion Summarization as Copycat-Review Generation
abstract
Opinion summarization is the task of automatically creating summaries that reflect subjective information expressed in multiple documents, such as product reviews.While the majority of previous work has focused on the extractive setting, i.e., selecting fragments from input reviews to produce a summary, we let the model generate novel sentences and hence produce abstractive summaries.Recent progress in summarization has seen the development of supervised models which rely on large quantities of document-summary pairs.Since such training data is expensive to acquire, we instead consider the unsupervised setting, in other words, we do not use any summaries in training.We define a generative model for a review collection which capitalizes on the intuition that when generating a new review given a set of other reviews of a product, we should be able to control the "amount of novelty" going into the new review or, equivalently, vary the extent to which it deviates from the input.At test time, when generating summaries, we force the novelty to be minimal, and produce a text reflecting consensus opinions.We capture this intuition by defining a hierarchical variational autoencoder model.Both individual reviews and the products they correspond to are associated with stochastic latent codes, and the review generator ("decoder") has direct access to the text of input reviews through the pointergenerator mechanism.Experiments on Amazon and Yelp datasets, show that setting at test time the review's latent code to its mean, allows the model to produce fluent and coherent summaries reflecting common opinions.
Arthur Brazinskas, Mirella Lapata, Ivan Titov 0001
ACL2
2020 Dscorer: A Fast Evaluation Metric for Discourse Representation Structure Parsing
abstract
Discourse representation structures (DRSs) are scoped semantic representations for texts of arbitrary length.Evaluation of the accuracy of predicted DRSs plays a key role in developing semantic parsers and improving their performance.DRSs are typically visualized as nested boxes, in a way that is not straightforward to process automatically.COUNTER, an evaluation algorithm for DRSs, transforms them to clauses and measures clause overlap by searching for variable mappings between two DRSs.Unfortunately, COUNTER is computationally costly (with respect to memory and CPU time) and does not scale with longer texts.We introduce DSCORER, an efficient new metric which converts box-style DRSs to graphs and then measures the overlap of n-grams in the graphs.Experiments show that DSCORER computes accuracy scores that correlate with scores from COUNTER at a fraction of the time.
Jiangming Liu, Shay B. Cohen, Mirella Lapata
ACL3
2020 Screenplay Summarization Using Latent Narrative Structure
abstract
Most general-purpose extractive summarization models are trained on news articles, which are short and present all important information upfront.As a result, such models are biased by position and often perform a smart selection of sentences from the beginning of the document.When summarizing long narratives, which have complex structure and present information piecemeal, simple position heuristics are not sufficient.In this paper, we propose to explicitly incorporate the underlying structure of narratives into general unsupervised and supervised extractive summarization models.We formalize narrative structure in terms of key narrative events (turning points) and treat it as latent in order to summarize screenplays (i.e., extract an optimal sequence of scenes).Experimental results on the CSI corpus of TV screenplays, which we augment with scene-level summarization labels, show that latent turning points correlate with important aspects of a CSI episode and improve summarization performance over general extractive algorithms, leading to more complete and diverse summaries.Victim: Mike Kimble, found in a Body Farm.Died 6 hours ago, unknown cause of death.CSI discover cow tissue in Mike's body.Cross-contamination is suggested.Probable cause of death: Mike's house has been set on fire.CSI finds blood: Mike was murdered, fire was a cover up.First suspects: Mike's fiance, Jane and her ex-husband, Russ.CSI finds photos in Mike's house of Jane's daughter, Jodie, posing naked.Mike is now a suspect of abusing Jodie.Russ allows CSI to examine his gun.CSI discovers that the bullet that killed Mike was made of frozen beef that melt inside him.They also find beef in Russ' gun.Russ confesses that he knew that Mike was abusing Jody, so he confronted and killed him.CSI discovers that the naked photos were taken on a boat, which belongs to Russ.CSI discovers that it was Russ who was abusing his daughter based on fluids found in his sleeping bag and later killed Mike who tried to help Jodie.Russ is given bail, since no jury would convict a protective father.Russ receives a mandatory life sentence.
Pinelopi Papalampidi, Frank Keller, Lea Frermann, Mirella Lapata
ACL4
2020 Experience Grounds Language
abstract
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph Turian. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Y. Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph P. Turian
EMNLP (1)7
2020 Few-Shot Learning for Opinion Summarization
abstract
Opinion summarization is the automatic creation of text reflecting subjective information expressed in multiple documents, such as user reviews of a product.The task is practically important and has attracted a lot of attention.However, due to the high cost of summary production, datasets large enough for training supervised models are lacking.Instead, the task has been traditionally approached with extractive methods that learn to select text fragments in an unsupervised or weakly-supervised way.Recently, it has been shown that abstractive summaries, potentially more fluent and better at reflecting conflicting information, can also be produced in an unsupervised fashion.However, these models, not being exposed to actual summaries, fail to capture their essential properties.In this work, we show that even a handful of summaries is sufficient to bootstrap generation of the summary text with all expected properties, such as writing style, informativeness, fluency, and sentiment preservation.We start by training a conditional Transformer language model to generate a new product review given other available reviews of the product.The model is also conditioned on review properties that are directly related to summaries; the properties are derived from reviews with no manual effort.In the second stage, we fine-tune a plug-in module that learns to predict property values on a handful of summaries.This lets us switch the generator to the summarization mode.We show on Amazon and Yelp datasets that our approach substantially outperforms previous extractive and abstractive methods in automatic and human evaluation.
Arthur Brazinskas, Mirella Lapata, Ivan Titov 0001
EMNLP (1)2
2020 Alignment-free Cross-lingual Semantic Role Labeling
abstract
Cross-lingual semantic role labeling (SRL) aims at leveraging resources in a source language to minimize the effort required to construct annotations or models for a new target language.Recent approaches rely on word alignments, machine translation engines, or preprocessing tools such as parsers or taggers.We propose a cross-lingual SRL model which only requires annotations in a source language and access to raw text in the form of a parallel corpus.The backbone of our model is an LSTM-based semantic role labeler jointly trained with a semantic role compressor and multilingual word embeddings.The compressor collects useful information from the output of the semantic role labeler, filtering noisy and conflicting evidence.It lives in a multilingual embedding space and provides direct supervision for predicting semantic roles in the target language.Results on the Universal Proposition Bank and manually annotated datasets show that our method is highly effective, even against systems utilizing supervised features.1
Rui Cai 0005, Mirella Lapata
EMNLP (1)2
2020 Multi-view Story Characterization from Movie Plot Synopses and Reviews
abstract
This paper considers the problem of characterizing stories by inferring properties such as theme and style using written synopses and reviews of movies. We experiment with a multi-label dataset of movie synopses and a tagset representing various attributes of stories (e.g., genre, type of events). Our proposed multi-view model encodes the synopses and reviews using hierarchical attention and shows improvement over methods that only use synopses. Finally, we demonstrate how can we take advantage of such a model to extract a complementary set of story-attributes from reviews without direct supervision. We have made our dataset and source code publicly available at https://ritual.uh.edu/ multiview-tag-2020.
Sudipta Kar, Gustavo Aguilar, Mirella Lapata, Thamar Solorio
EMNLP (1)3
2020 Multi-Step Inference for Reasoning Over Paragraphs
abstract
Complex reasoning over text requires understanding and chaining together free-form predicates and logical connectives.Prior work has largely tried to do this either symbolically or with black-box transformers.We present a middle ground between these two extremes: a compositional model reminiscent of neural module networks that can perform chained logical reasoning.This model first finds relevant sentences in the context and then chains them together using neural modules.Our model gives significant performance improvements (up to 29% relative error reduction when combined with a reranker) on ROPES, a recentlyintroduced complex reasoning dataset.
Jiangming Liu, Matt Gardner 0001, Shay B. Cohen, Mirella Lapata
EMNLP (1)4
2020 Zero-Shot Crosslingual Sentence Simplification
abstract
Sentence simplification aims to make sentences easier to read and understand.Recent approaches have shown promising results with encoder-decoder models trained on large amounts of parallel data which often only exists in English.We propose a zero-shot modeling framework which transfers simplification knowledge from English to another language (for which no parallel simplification corpus exists) while generalizing across languages and tasks.A shared transformer encoder constructs language-agnostic representations, with a combination of task-specific encoder layers added on top (e.g., for translation and simplification).Empirical results using both human and automatic metrics show that our approach produces better simplifications than unsupervised and pivot-based methods.
Jonathan Mallinson, Rico Sennrich, Mirella Lapata
EMNLP (1)3
2020 Coarse-to-Fine Query Focused Multi-Document Summarization
abstract
We consider the problem of better modeling query-cluster interactions to facilitate query focused multi-document summarization.Due to the lack of training data, existing work relies heavily on retrieval-style methods for assembling query relevant summaries.We propose a coarse-to-fine modeling framework which employs progressively more accurate modules for estimating whether text segments are relevant, likely to contain an answer, and central.The modules can be independently developed and leverage training data if available.We present an instantiation of this framework with a trained evidence estimator which relies on distant supervision from question answering (where various resources exist) to identify segments which are likely to answer the query and should be included in the summary.Our framework 1 is robust across domains and query types (i.e., long vs short) and outperforms strong comparison systems on benchmark datasets.
Yumo Xu, Mirella Lapata
EMNLP (1)2
2019 Data-to-Text Generation with Content Selection and Planning
abstract
Recent advances in data-to-text generation have led to the use of large-scale datasets and neural network models which are trained end-to-end, without explicitly modeling what to say and in what order. In this work, we present a neural network architecture which incorporates content selection and planning without sacrificing end-to-end training. We decompose the generation task into two stages. Given a corpus of data records (paired with descriptive documents), we first generate a content plan highlighting which information should be mentioned and in which order and then generate the document while taking the content plan into account. Automatic and human-based evaluation experiments show that our model1 outperforms strong baselines improving the state-of-the-art on the recently released RotoWIRE dataset.
Ratish Puduppully, Li Dong 0004, Mirella Lapata
AAAI3
2019 Discourse Representation Parsing for Sentences and Documents
abstract
We introduce a novel semantic parsing task based on Discourse Representation Theory (DRT; Kamp and Reyle 1993).Our model operates over Discourse Representation Tree Structures which we formally define for sentences and documents.We present a general framework for parsing discourse structures of arbitrary length and granularity.We achieve this with a neural model equipped with a supervised hierarchical attention mechanism and a linguistically-motivated copy strategy.Experimental results on sentence-and documentlevel benchmarks show that our model outperforms competitive baselines by a wide margin.
Jiangming Liu, Shay B. Cohen, Mirella Lapata
ACL (1)3
2019 Hierarchical Transformers for Multi-Document Summarization
abstract
In this paper, we develop a neural summarization model which can effectively process multiple input documents and distill Transformer architecture with the ability to encode documents in a hierarchical manner. We represent cross-document relationships via an attention mechanism which allows to share information as opposed to simply concatenating text spans and processing them as a flat sequence. Our model learns latent dependencies among textual units, but can also take advantage of explicit graph representations focusing on similarity or discourse relations. Empirical results on the WikiSum dataset demonstrate that the proposed architecture brings substantial improvements over several strong baselines.
Yang Liu 0124, Mirella Lapata
ACL (1)2
2019 Generating Summaries with Topic Templates and Structured Convolutional Decoders
abstract
Existing neural generation approaches create multi-sentence text as a single sequence.In this paper we propose a structured convolutional decoder that is guided by the content structure of target summaries.We compare our model with existing sequential decoders on three data sets representing different domains.Automatic and human evaluation demonstrate that our summaries have better content coverage.
Laura Perez-Beltrachini, Yang Liu 0124, Mirella Lapata
ACL (1)3
2019 Data-to-text Generation with Entity Modeling
abstract
Recent approaches to data-to-text generation have shown great promise thanks to the use of large-scale datasets and the application of neural network architectures which are trained end-to-end.These models rely on representation learning to select content appropriately, structure it coherently, and verbalize it grammatically, treating entities as nothing more than vocabulary tokens.In this work we propose an entity-centric neural architecture for data-to-text generation.Our model creates entity-specific representations which are dynamically updated.Text is generated conditioned on the data input and entity memory representations using hierarchical attention at each time step.We present experiments on the ROTOWIRE benchmark and a (five times larger) new dataset on the baseball domain which we create.Our results show that the proposed model outperforms competitive baselines in automatic and human evaluation. 1TEAM Inn1 Inn2 Inn3 Inn4 . . .R H E . . .Orioles 1 0 0 0 . . . 2 4 0 . . .Royals 1 0 0 3 . . .9 14 1 . . .
Ratish Puduppully, Li Dong 0004, Mirella Lapata
ACL (1)3
2019 Sentence Centrality Revisited for Unsupervised Summarization
abstract
Single document summarization has enjoyed renewed interest in recent years thanks to the popularity of neural network models and the availability of large-scale datasets.In this paper we develop an unsupervised approach arguing that it is unrealistic to expect large-scale and high-quality training data to be available or created for different types of summaries, domains, or languages.We revisit a popular graph-based ranking algorithm and modify how node (aka sentence) centrality is computed in two ways: (a) we employ BERT, a state-of-the-art neural representation learning model to better capture sentential meaning and (b) we build graphs with directed edges arguing that the contribution of any two nodes to their respective centrality is influenced by their relative position in a document.Experimental results on three news summarization datasets representative of different languages and writing styles show that our approach outperforms strong baselines by a wide margin. 1
Mirella Lapata
ACL (1)2
2019 Semi-Supervised Semantic Role Labeling with Cross-View Training
abstract
Rui Cai, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Rui Cai 0005, Mirella Lapata
EMNLP/IJCNLP (1)2
2019 Semantic graph parsing with recurrent neural network DAG grammars
abstract
Federico Fancellu, Sorcha Gilroy, Adam Lopez, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Federico Fancellu, Sorcha Gilroy, Adam Lopez, Mirella Lapata
EMNLP/IJCNLP (1)4
2019 Text Summarization with Pretrained Encoders
abstract
Yang Liu, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yang Liu 0124, Mirella Lapata
EMNLP/IJCNLP (1)2
2019 Movie Plot Analysis via Turning Point Identification
abstract
Pinelopi Papalampidi, Frank Keller, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Pinelopi Papalampidi, Frank Keller, Mirella Lapata
EMNLP/IJCNLP (1)3
2019 Partners in Crime: Multi-view Sequential Inference for Movie Understanding
abstract
Nikos Papasarantopoulos, Lea Frermann, Mirella Lapata, Shay B. Cohen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Nikos Papasarantopoulos, Lea Frermann, Mirella Lapata, Shay B. Cohen
EMNLP/IJCNLP (1)3
2019 Learning Semantic Parsers from Denotations with Latent Structured Alignments and Abstract Programs
abstract
Bailin Wang, Ivan Titov, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Bailin Wang, Ivan Titov 0001, Mirella Lapata
EMNLP/IJCNLP (1)3
2019 Learning Natural Language Interfaces with Neural Models
Mirella Lapata
INTERSPEECH1
2019 Learning an Executable Neural Semantic Parser
abstract
This article describes a neural semantic parser that maps natural language utterances onto logical forms that can be executed against a task-specific environment, such as a knowledge base or a database, to produce a response. The parser generates tree-structured logical forms with a transition-based approach, combining a generic tree-generation algorithm with domain-general grammar defined by the logical language. The generation process is modeled by structured recurrent neural networks, which provide a rich encoding of the sentential context and generation history for making predictions. To tackle mismatches between natural language and logical form tokens, various attention mechanisms are explored. Finally, we consider different training settings for the neural semantic parser, including fully supervised training where annotated logical forms are given, weakly supervised training where denotations are provided, and distant supervision where only unlabeled sentences and a knowledge base are available. Experiments across a wide range of data sets demonstrate the effectiveness of our parser.
Jianpeng Cheng 0001, Siva Reddy, Vijay A. Saraswat, Mirella Lapata
Comput. Linguistics4
2019 Word n-gram attention models for sentence similarity and inference
Iñigo Lopez-Gazpio, Montse Maritxalar, Mirella Lapata, Eneko Agirre
Expert Syst. Appl.3
2019 What is this Article about? Extreme Summarization with Topic-aware Convolutional Neural Networks
abstract
We introduce "extreme summarization," a new single-document summarization task which aims at creating a short, one-sentence news summary answering the question "What is the article about?". We argue that extreme summarization, by nature, is not amenable to extractive strategies and requires an abstractive modeling approach. In the hope of driving research on this task further: (a) we collect a real-world, large scale dataset by harvesting online articles from the British Broadcasting Corporation (BBC); and (b) propose a novel abstractive model which is conditioned on the article's topics and based entirely on convolutional neural networks. We demonstrate experimentally that this architecture captures long-range dependencies in a document and recognizes pertinent content, outperforming an oracle extractive system and state-of-the-art abstractive approaches when evaluated automatically and by humans on the extreme summarization dataset.
Shashi Narayan, Shay B. Cohen, Mirella Lapata
J. Artif. Intell. Res.3
2019 Disambiguating Visual Verbs
abstract
In this article, we introduce a new task, visual sense disambiguation for verbs: given an image and a verb, assign the correct sense of the verb, i.e., the one that describes the action depicted in the image. Just as textual word sense disambiguation is useful for a wide range of NLP tasks, visual sense disambiguation can be useful for multimodal tasks such as image retrieval, image description, and text illustration. We introduce a new dataset, which we call VerSe (short for Verb Sense) that augments existing multimodal datasets (COCO and TUHOI) with verb and sense labels. We explore supervised and unsupervised models for the sense disambiguation task using textual, visual, and multimodal embeddings. We also consider a scenario in which we must detect the verb depicted in an image prior to predicting its sense (i.e., there is no verbal information associated with the image). We find that textual embeddings perform well when gold-standard annotations (object labels and image descriptions) are available, while multimodal embeddings perform well on unannotated images. VerSe is publicly available at https://github.com/spandanagella/verse.
Spandana Gella, Frank Keller, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Syntax-aware Semantic Role Labeling without Parsing
abstract
In this paper we focus on learning dependency aware representations for semantic role labeling without recourse to an external parser. The backbone of our model is an LSTM-based semantic role labeler jointly trained with two auxiliary tasks: predicting the dependency label of a word and whether there exists an arc linking it to the predicate. The auxiliary tasks provide syntactic information that is specific to semantic role labeling and are learned from training data (dependency annotations) without relying on existing dependency parsers, which can be noisy (e.g., on out-of-domain data or infrequent constructions). Experimental results on the CoNLL-2009 benchmark dataset show that our model outperforms the state of the art in English, and consistently improves performance in other languages, including Chinese, German, and Spanish.
Rui Cai 0005, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2019 Weakly Supervised Domain Detection
abstract
In this paper we introduce domain detection as a new natural language processing task. We argue that the ability to detect textual segments that are domain-heavy (i.e., sentences or phrases that are representative of and provide evidence for a given domain) could enhance the robustness and portability of various text classification applications. We propose an encoder-detector framework for domain detection and bootstrap classifiers with multiple instance learning. The model is hierarchically organized and suited to multilabel classification. We demonstrate that despite learning with minimal supervision, our model can be applied to text spans of different granularities, languages, and genres. We also showcase the potential of domain detection for text summarization.
Yumo Xu, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2018 Discourse Representation Structure Parsing
abstract
We introduce an open-domain neural semantic parser which generates formal meaning representations in the style of Discourse Representation Theory (DRT; Kamp and Reyle 1993).We propose a method which transforms Discourse Representation Structures (DRSs) to trees and develop a structure-aware model which decomposes the decoding process into three stages: basic DRS structure prediction, condition prediction (i.e., predicates and relations), and referent prediction (i.e., variables).Experimental results on the Groningen Meaning Bank (GMB) show that our model outperforms competitive baselines by a wide margin.
Jiangming Liu, Shay B. Cohen, Mirella Lapata
ACL (1)3
2018 Coarse-to-Fine Decoding for Neural Semantic Parsing
abstract
Semantic parsing aims at mapping natural language utterances into structured meaning representations.In this work, we propose a structure-aware neural architecture which decomposes the semantic parsing process into two stages.Given an input utterance, we first generate a rough sketch of its meaning, where low-level information (such as variable names and arguments) is glossed over.Then, we fill in missing details by taking into account the natural language input and the sketch itself.Experimental results on four datasets characteristic of different domains and meaning representations show that our approach consistently improves performance, achieving competitive results despite the use of relatively simple decoders.
Li Dong 0004, Mirella Lapata
ACL (1)2
2018 Confidence Modeling for Neural Semantic Parsing
abstract
In this work we focus on confidence modeling for neural semantic parsers which are built upon sequence-to-sequence models.We outline three major causes of uncertainty, and design various metrics to quantify these factors.These metrics are then used to estimate confidence scores that indicate whether model predictions are likely to be correct.Beyond confidence estimation, we identify which parts of the input contribute to uncertain predictions allowing users to interpret their model, and verify or refine its input.Experimental results show that our confidence model significantly outperforms a widely used method that relies on posterior probability, and improves the quality of interpretation compared to simply relying on attention scores.
Li Dong 0004, Chris Quirk, Mirella Lapata
ACL (1)3
2018 Document Modeling with External Attention for Sentence Extraction
abstract
Shashi Narayan, Ronald Cardenas, Nikos Papasarantopoulos, Shay B. Cohen, Mirella Lapata, Jiangsheng Yu, Yi Chang. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Shashi Narayan, Ronald Cardenas, Nikos Papasarantopoulos, Shay B. Cohen, Mirella Lapata, Jiangsheng Yu, Yi Chang 0001
ACL (1)5
2018 Weakly-Supervised Neural Semantic Parsing with a Generative Ranker
abstract
Weakly-supervised semantic parsers are trained on utterance-denotation pairs, treating logical forms as latent.The task is challenging due to the large search space and spuriousness of logical forms.In this paper we introduce a neural parser-ranker system for weakly-supervised semantic parsing.The parser generates candidate tree-structured logical forms from utterances using clues of denotations.These candidates are then ranked based on two criterion: their likelihood of executing to the correct denotation, and their agreement with the utterance semantics.We present a scheduled training procedure to balance the contribution of the two objectives.Furthermore, we propose to use a neurally encoded lexicon to inject prior domain knowledge to the model.Experiments on three Freebase datasets demonstrate the effectiveness of our semantic parser, achieving results within the state-of-the-art range.
Jianpeng Cheng 0001, Mirella Lapata
CoNLL2
2018 Summarizing Opinions: Aspect Extraction Meets Sentiment Prediction and They Are Both Weakly Supervised
abstract
We present a neural framework for opinion summarization from online product reviews which is knowledge-lean and only requires light supervision (e.g., in the form of product domain labels and user-provided ratings).Our method combines two weakly supervised components to identify salient opinions and form extractive summaries from multiple reviews: an aspect extractor trained under a multi-task objective, and a sentiment predictor based on multiple instance learning.We introduce an opinion summarization dataset that includes a training set of product reviews from six diverse domains and human-annotated development and test sets with gold standard aspect annotations, salience labels, and opinion summaries.Automatic evaluation shows significant improvements over baselines, and a largescale study indicates that our opinion summaries are preferred by human judges according to multiple criteria.1
Stefanos Angelidis, Mirella Lapata
EMNLP2
2018 Structured Alignment Networks for Matching Sentences
abstract
Many tasks in natural language processing involve comparing two sentences to compute some notion of relevance, entailment, or similarity.Typically, this comparison is done either at the word level or at the sentence level, with no attempt to leverage the inherent structure of the sentence.When sentence structure is used for comparison, it is obtained during a non-differentiable pre-processing step, leading to propagation of errors.We introduce a model of structured alignments between sentences, showing how to compare two sentences by matching their latent structures.Using a structured attention mechanism, our model matches candidate spans in the first sentence to candidate spans in the second sentence, simultaneously discovering the tree structure of each sentence.Our model is fully differentiable and trained only on the matching objective.We evaluate this model on two tasks, entailment detection and answer sentence selection, and find that modeling latent tree structures results in superior performance.Analysis of the learned sentence structures shows they can reflect some syntactic phenomena.
Yang Liu 0124, Matt Gardner 0001, Mirella Lapata
EMNLP3
2018 Sentence Compression for Arbitrary Languages via Multilingual Pivoting
abstract
In this paper we advocate the use of bilingual corpora which are abundantly available for training sentence compression models.Our approach borrows much of its machinery from neural machine translation and leverages bilingual pivoting: compressions are obtained by translating a source string into a foreign language and then back-translating it into the source while controlling the translation length.Our model can be trained for any language as long as a bilingual corpus is available and performs arbitrary rewrites without access to compression specific data.We release 1 MOSS, a new parallel Multilingual Compression dataset for English, German, and French which can be used to evaluate compression models across languages and genres.
Jonathan Mallinson, Rico Sennrich, Mirella Lapata
EMNLP3
2018 Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
abstract
We introduce extreme summarization, a new single-document summarization task which does not favor extractive strategies and calls for an abstractive modeling approach.The idea is to create a short, one-sentence news summary answering the question "What is the article about?".We collect a real-world, large scale dataset for this task by harvesting online articles from the British Broadcasting Corporation (BBC).We propose a novel abstractive model which is conditioned on the article's topics and based entirely on convolutional neural networks.We demonstrate experimentally that this architecture captures longrange dependencies in a document and recognizes pertinent content, outperforming an oracle extractive system and state-of-the-art abstractive approaches when evaluated automatically and by humans. 1
Shashi Narayan, Shay B. Cohen, Mirella Lapata
EMNLP3
2018 Neural Latent Extractive Document Summarization
abstract
Extractive summarization models require sentence-level labels, which are usually created heuristically (e.g., with rule-based methods) given that most summarization datasets only have document-summary pairs.Since these labels might be suboptimal, we propose a latent variable extractive model where sentences are viewed as latent variables and sentences with activated variables are used to infer gold summaries.During training the loss comes directly from gold summaries.Experiments on the CNN/Dailymail dataset show that our model improves over a strong extractive baseline trained on heuristically approximated labels and also performs competitively to several recent models.
Xingxing Zhang 0002, Mirella Lapata, Furu Wei, Ming Zhou 0001
EMNLP2
2018 What's This Movie About? A Joint Neural Network Architecture for Movie Content Analysis
abstract
Philip John Gorinski, Mirella Lapata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Philip John Gorinski, Mirella Lapata
NAACL-HLT2
2018 Ranking Sentences for Extractive Summarization with Reinforcement Learning
abstract
Shashi Narayan, Shay B. Cohen, Mirella Lapata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Shashi Narayan, Shay B. Cohen, Mirella Lapata
NAACL-HLT3
2018 Bootstrapping Generators from Noisy Data
abstract
Laura Perez-Beltrachini, Mirella Lapata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Laura Perez-Beltrachini, Mirella Lapata
NAACL-HLT2
2018 Understanding visual scenes
abstract
Abstract A growing body of recent work focuses on the challenging problem of scene understanding using a variety of cross-modal methods which fuse techniques from image and text processing. In this paper, we develop representations for the semantics of scenes by explicitly encoding the objects detected in them and their spatial relations. We represent image content via two well-known types of tree representations, namely constituents and dependencies. Our representations are created deterministically, can be applied to any image dataset irrespective of the task at hand, and are amenable to standard NLP tools developed for tree-based structures. We show that we can apply syntax-based SMT and tree kernel methods in order to build models for image description generation and image-based retrieval. Experimental results on real-world images demonstrate the effectiveness of the framework.
Carina Silberer, Jasper R. R. Uijlings, Mirella Lapata
Nat. Lang. Eng.3
2018 Multiple Instance Learning Networks for Fine-Grained Sentiment Analysis
abstract
We consider the task of fine-grained sentiment analysis from the perspective of multiple instance learning (MIL). Our neural model is trained on document sentiment labels, and learns to predict the sentiment of text segments, i.e. sentences or elementary discourse units (EDUs), without segment-level supervision. We introduce an attention-based polarity scoring method for identifying positive and negative text snippets and a new dataset which we call SpoT (as shorthand for Segment-level POlariTy annotations) for evaluating MIL-style sentiment models like ours. Experimental results demonstrate superior performance against multiple baselines, whereas a judgement elicitation study shows that EDU-level opinion extraction produces more informative summaries than sentence-based alternatives.
Stefanos Angelidis, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2018 Whodunnit? Crime Drama as a Case for Natural Language Understanding
abstract
In this paper we argue that crime drama exemplified in television programs such as CSI: Crime Scene Investigation is an ideal testbed for approximating real-world natural language understanding and the complex inferences associated with it. We propose to treat crime drama as a new inference task, capitalizing on the fact that each episode poses the same basic question (i.e., who committed the crime) and naturally provides the answer when the perpetrator is revealed. We develop a new dataset based on CSI episodes, formalize perpetrator identification as a sequence labeling problem, and develop an LSTM-based model which learns from multi-modal data. Experimental results show that an incremental inference strategy is key to making accurate guesses as well as learning from representations fusing textual, visual, and acoustic input.
Lea Frermann, Shay B. Cohen, Mirella Lapata
Trans. Assoc. Comput. Linguistics3
2018 Learning Structured Text Representations
abstract
In this paper, we focus on learning structure-aware document representations from data without recourse to a discourse parser or additional annotations. Drawing inspiration from recent efforts to empower neural networks with a structural bias (Cheng et al., 2016; Kim et al., 2017), we propose a model that can encode a document while automatically inducing rich structural dependencies. Specifically, we embed a differentiable non-projective parsing algorithm into a neural model and use attention mechanisms to incorporate the structural biases. Experimental evaluations across different tasks and datasets show that the proposed model achieves state-of-the-art results on document modeling tasks while inducing intermediate structures which are both interpretable and meaningful.
Yang Liu 0124, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2017 Learning Structured Natural Language Representations for Semantic Parsing
abstract
We introduce a neural semantic parser which is interpretable and scalable.Our model converts natural language utterances to intermediate, domain-general natural language representations in the form of predicate-argument structures, which are induced with a transition system and subsequently mapped to target domains.The semantic parser is trained end-to-end using annotated logical forms or their denotations.We achieve the state of the art on SPADES and GRAPHQUESTIONS and obtain competitive results on GEO-QUERY and WEBQUESTIONS.The induced predicate-argument structures shed light on the types of representations useful for semantic parsing and how these are different from linguistically motivated ones. 1
Jianpeng Cheng 0001, Siva Reddy, Vijay A. Saraswat, Mirella Lapata
ACL (1)4
2017 Paraphrasing Revisited with Neural Machine Translation
abstract
Recognizing and generating paraphrases is an important component in many natural language processing applications.A wellestablished technique for automatically extracting paraphrases leverages bilingual corpora to find meaning-equivalent phrases in a single language by "pivoting" over a shared translation in another language.In this paper we revisit bilingual pivoting in the context of neural machine translation and present a paraphrasing model based purely on neural networks.Our model represents paraphrases in a continuous space, estimates the degree of semantic relatedness between text segments of arbitrary length, or generates candidate paraphrases for any source input.Experimental results across tasks and datasets show that neural paraphrases outperform those obtained with conventional phrase-based pivoting approaches.
Jonathan Mallinson, Rico Sennrich, Mirella Lapata
EACL (1)3
2017 Dependency Parsing as Head Selection
abstract
Conventional graph-based dependency parsers guarantee a tree structure both during training and inference.Instead, we formalize dependency parsing as the problem of independently selecting the head of each word in a sentence.Our model which we call DENSE (as shorthand for Dependency Neural Selection) produces a distribution over possible heads for each word using features obtained from a bidirectional recurrent neural network.Without enforcing structural constraints during training, DENSE generates (at inference time) trees for the overwhelming majority of sentences, while non-tree outputs can be adjusted with a maximum spanning tree algorithm.We evaluate DENSE on four languages (English, Chinese, Czech, and German) with varying degrees of non-projectivity.Despite the simplicity of the approach, our parsers are on par with the state of the art. 1
Xingxing Zhang 0002, Jianpeng Cheng 0001, Mirella Lapata
EACL (1)3
2017 Learning to Generate Product Reviews from Attributes
abstract
Li Dong, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou, Ke Xu. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Li Dong 0004, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou 0001, Ke Xu 0001
EACL (1)4
2017 Learning to Paraphrase for Question Answering
abstract
Question answering (QA) systems are sensitive to the many different ways natural language expresses the same information need.In this paper we turn to paraphrases as a means of capturing this knowledge and present a general framework which learns felicitous paraphrases for various QA tasks.Our method is trained end-toend using question-answer pairs as a supervision signal.A question and its paraphrases serve as input to a neural scoring model which assigns higher weights to linguistic expressions most likely to yield correct answers.We evaluate our approach on QA over Freebase and answer sentence selection.Experimental results on three datasets show that our framework consistently improves performance, achieving competitive results despite the use of simple QA models.
Li Dong 0004, Jonathan Mallinson, Siva Reddy, Mirella Lapata
EMNLP4
2017 Image Pivoting for Learning Multilingual Multimodal Representations
abstract
In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image understanding.Our model learns a common representation for images and their descriptions in two different languages (which need not be parallel) by considering the image as a pivot between two languages.We introduce a new pairwise ranking loss function which can handle both symmetric and asymmetric similarity between the two modalities.We evaluate our models on image-description ranking for German and English, and on semantic textual similarity of image descriptions in English.In both cases we achieve state-of-the-art performance.
Spandana Gella, Rico Sennrich, Frank Keller, Mirella Lapata
EMNLP4
2017 Learning Contextually Informed Representations for Linear-Time Discourse Parsing
abstract
Recent advances in RST discourse parsing have focused on two modeling paradigms: (a) high order parsers which jointly predict the tree structure of the discourse and the relations it encodes; or (b) lineartime parsers which are efficient but mostly based on local features.In this work, we propose a linear-time parser with a novel way of representing discourse constituents based on neural networks which takes into account global contextual information and is able to capture long-distance dependencies.Experimental results show that our parser obtains state-of-the art performance on benchmark datasets, while being efficient (with time complexity linear in the number of sentences in the document) and requiring minimal feature engineering.
Yang Liu 0124, Mirella Lapata
EMNLP2
2017 Universal Semantic Parsing
abstract
Universal Dependencies (UD) offer a uniform cross-lingual syntactic representation, with the aim of advancing multilingual applications.Recent work shows that semantic parsing can be accomplished by transforming syntactic dependencies to logical forms.However, this work is limited to English, and cannot process dependency graphs, which allow handling complex phenomena such as control.In this work, we introduce UDEPLAMBDA, a semantic interface for UD, which maps natural language to logical forms in an almost language-independent fashion and can process dependency graphs.We perform experiments on question answering against Freebase and provide German and Spanish translations of the WebQuestions and GraphQuestions datasets to facilitate multilingual evaluation.Results show that UDEPLAMBDA outperforms strong baselines across languages and datasets.For English, it achieves a 4.9 F 1 point improvement over the state-of-the-art on Graph-Questions.ENTITY ⇒ λx.word(x a ); e.g.Oscar ⇒ λx.Oscar(x a ) EVENT ⇒ λx.word(x e ); e.g. won ⇒ λx.won(x e ) FUNCTIONAL ⇒ λx.TRUE; e.g. an ⇒ λx.TRUE COPY ⇒ λ f gx.∃y.f (x) ∧ g(y) ∧ rel(x, y) e.g.nsubj, dobj, nmod, advmod INVERT ⇒ λ f gx.∃y.f (x) ∧ g(y) ∧ rel i (y, x) e.g.amod, acl MERGE ⇒ λ f gx.f (x) ∧ g(x) e.g.compound, appos, amod, acl HEAD ⇒ λ f gx.f (x) e.g.case, punct, aux, mark .
Siva Reddy, Oscar Täckström, Slav Petrov, Mark Steedman, Mirella Lapata
EMNLP5
2017 Sentence Simplification with Deep Reinforcement Learning
abstract
Sentence simplification aims to make sentences easier to read and understand.Most recent approaches draw on insights from machine translation to learn simplification rewrites from monolingual corpora of complex and simple sentences.We address the simplification problem with an encoder-decoder model coupled with a deep reinforcement learning framework.Our model, which we call DRESS (as shorthand for Deep REinforcement Sentence Simplification), explores the space of possible simplifications while learning to optimize a reward function that encourages outputs which are simple, fluent, and preserve the meaning of the input.Experiments on three datasets demonstrate that our model outperforms competitive simplification systems. 1
Xingxing Zhang 0002, Mirella Lapata
EMNLP2
2017 Text Rewriting Improves Semantic Role Labeling (Extended Abstract)
abstract
Large-scale annotated corpora are a prerequisite to developing high-performance NLP systems. Such corpora are expensive to produce, limited in size, often demanding linguistic expertise. In this paper we use text rewriting as a means of increasing the amount of labeled data available for model training. Our method uses automatically extracted rewrite rules from comparable corpora and bitexts to generate multiple versions of sentences annotated with gold standard labels. We apply this idea to semantic role labeling and show that a model trained on rewritten data outperforms the state of the art on the CoNLL-2009 benchmark dataset.
Kristian Woodsend, Mirella Lapata
IJCAI2
2017 Visually Grounded Meaning Representations
abstract
In this paper we address the problem of grounding distributional representations of lexical meaning. We introduce a new model which uses stacked autoencoders to learn higher-level representations from textual and visual input. The visual modality is encoded via vectors of attributes obtained automatically from images. We create a new large-scale taxonomy of 600 visual attributes representing more than 500 concepts and 700 K images. We use this dataset to train attribute classifiers and integrate their predictions with text-based distributional models of word meaning. We evaluate our model on its ability to simulate word similarity judgments and concept categorization. On both tasks, our model yields a better fit to behavioral data compared to baselines and related models which either rely on a single modality or do not make use of attribute-based input.
Carina Silberer, Vittorio Ferrari, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Autofolding for Source Code Summarization
abstract
Developers spend much of their time reading and browsing source code, raising new opportunities for summarization methods. Indeed, modern code editors provide code folding, which allows one to selectively hide blocks of code. However this is impractical to use as folding decisions must be made manually or based on simple rules. We introduce the autofolding problem, which is to automatically create a code summary by folding less informative code regions. We present a novel solution by formulating the problem as a sequence of AST folding decisions, leveraging a scoped topic model for code tokens. On an annotated set of popular open source projects, we show that our summarizer outperforms simpler baselines, yielding a 28 percent error reduction. Furthermore, we find through a case study that our summarizer is strongly preferred by experienced developers. More broadly, we hope this work will aid program comprehension by turning code folding into a usable and valuable tool.
Jaroslav M. Fowkes, Pankajan Chanthirasegaran, Razvan Ranca, Miltiadis Allamanis, Mirella Lapata, Charles Sutton
IEEE Trans. Software Eng.5
2016 Neural Summarization by Extracting Sentences and Words
abstract
Traditional approaches to extractive summarization rely heavily on humanengineered features.In this work we propose a data-driven approach based on neural networks and continuous sentence features.We develop a general framework for single-document summarization composed of a hierarchical document encoder and an attention-based extractor.This architecture allows us to develop different classes of summarization models which can extract sentences or words.We train our models on large scale corpora containing hundreds of thousands of document-summary pairs 1 .Experimental results on two summarization datasets demonstrate that our models obtain results comparable to the state of the art without any access to linguistic annotation.
Jianpeng Cheng 0001, Mirella Lapata
ACL (1)2
2016 Language to Logical Form with Neural Attention
abstract
Semantic parsing aims at mapping natural language to machine interpretable meaning representations.Traditional approaches rely on high-quality lexicons, manually-built templates, and linguistic features which are either domainor representation-specific.In this paper we present a general method based on an attention-enhanced encoder-decoder model.We encode input utterances into vector representations, and generate their logical forms by conditioning the output sequences or trees on the encoding vectors.Experimental results on four datasets show that our approach performs competitively without using hand-engineered features and is easy to adapt across domains and meaning representations.
Li Dong 0004, Mirella Lapata
ACL (1)2
2016 Neural Semantic Role Labeling with Dependency Path Embeddings
abstract
This paper introduces a novel model for semantic role labeling that makes use of neural sequence modeling techniques.Our approach is motivated by the observation that complex syntactic structures and related phenomena, such as nested subordinations and nominal predicates, are not handled well by existing models.Our model treats such instances as subsequences of lexicalized dependency paths and learns suitable embedding representations.We experimentally demonstrate that such embeddings can improve results over previous state-of-the-art semantic role labelers, and showcase qualitative improvements obtained by our method.
Michael Roth 0001, Mirella Lapata
ACL (1)2
2016 Long Short-Term Memory-Networks for Machine Reading
abstract
In this paper we address the question of how to render sequence-level networks better at handling structured input. We propose a machine reading simulator which processes text incrementally from left to right and performs shallow reasoning with memory and attention. The reader extends the Long Short-Term Memory architecture with a memory network in place of a single memory cell. This enables adaptive memory usage during recurrence with neural attention, offering a way to weakly induce relations among tokens. The system is initially designed to process a single sequence but we also demonstrate how to integrate it with an encoder-decoder architecture. Experiments on language modeling, sentiment analysis, and natural language inference show that our model matches or outperforms the state of the art.
Jianpeng Cheng 0001, Li Dong 0004, Mirella Lapata
EMNLP3
2016 Unsupervised Visual Sense Disambiguation for Verbs using Multimodal Embeddings
abstract
We introduce a new task, visual sense disambiguation for verbs: given an image and a verb, assign the correct sense of the verb, i.e., the one that describes the action depicted in the image.Just as textual word sense disambiguation is useful for a wide range of NLP tasks, visual sense disambiguation can be useful for multimodal tasks such as image retrieval, image description, and text illustration.We introduce VerSe, a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.We propose an unsupervised algorithm based on Lesk which performs visual sense disambiguation using textual, visual, or multimodal embeddings.We find that textual embeddings perform well when goldstandard textual annotations (object labels and image descriptions) are available, while multimodal embeddings perform well on unannotated images.We also verify our findings by using the textual and multimodal embeddings as features in a supervised setting and analyse the performance of visual sense disambiguation task.VerSe is made publicly available and can be downloaded at: https://github.com/spandanagella/verse.
Spandana Gella, Mirella Lapata, Frank Keller
HLT-NAACL2
2016 Top-down Tree Long Short-Term Memory Networks
abstract
Long Short-Term Memory (LSTM) networks, a type of recurrent neural network with a more complex computational unit, have been successfully applied to a variety of sequence modeling tasks.In this paper we develop Tree Long Short-Term Memory (TREELSTM), a neural network model based on LSTM, which is designed to predict a tree rather than a linear sequence.TREELSTM defines the probability of a sentence by estimating the generation probability of its dependency tree.At each time step, a node is generated based on the representation of the generated subtree.We further enhance the modeling power of TREELSTM by explicitly representing the correlations between left and right dependents.Application of our model to the MSR sentence completion challenge achieves results beyond the current state of the art.We also report results on dependency parsing reranking achieving competitive performance.
Xingxing Zhang 0002, Liang Lu 0001, Mirella Lapata
HLT-NAACL3
2016 A Bayesian Model of Diachronic Meaning Change
abstract
Word meanings change over time and an automated procedure for extracting this information from text would be useful for historical exploratory studies, information retrieval or question answering. We present a dynamic Bayesian model of diachronic meaning change, which infers temporal word representations as a set of senses and their prevalence. Unlike previous work, we explicitly model language change as a smooth, gradual process. We experimentally show that this modeling decision is beneficial: our model performs competitively on meaning change detection tasks whilst inducing discernible word senses and their development over time. Application of our model to the SemEval-2015 temporal classification benchmark datasets further reveals that it performs on par with highly optimized task-specific systems.
Lea Frermann, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2016 Transforming Dependency Structures to Logical Forms for Semantic Parsing
abstract
The strongly typed syntax of grammar formalisms such as CCG, TAG, LFG and HPSG offers a synchronous framework for deriving syntactic structures and semantic logical forms. In contrast—partly due to the lack of a strong type system—dependency structures are easy to annotate and have become a widely used form of syntactic analysis for many languages. However, the lack of a type system makes a formal mechanism for deriving logical forms from dependency structures challenging. We address this by introducing a robust system based on the lambda calculus for deriving neo-Davidsonian logical forms from dependency trees. These logical forms are then used for semantic parsing of natural language to Freebase. Experiments on the Free917 and Web-Questions datasets show that our representation is superior to the original dependency trees and that it outperforms a CCG-based representation on this task. Compared to prior work, we obtain the strongest result to date on Free917 and competitive results on WebQuestions.
Siva Reddy, Oscar Täckström, Michael Collins 0001, Tom Kwiatkowski, Dipanjan Das 0001, Mark Steedman, Mirella Lapata
Trans. Assoc. Comput. Linguistics7
2015 Distributed Representations for Unsupervised Semantic Role Labeling
abstract
We present a new approach for unsupervised semantic role labeling that leverages distributed representations.We induce embeddings to represent a predicate, its arguments and their complex interdependence.Argument embeddings are learned from surrounding contexts involving the predicate and neighboring arguments, while predicate embeddings are learned from argument contexts.The induced representations are clustered into roles using a linear programming formulation of hierarchical clustering, where we can model task-specific knowledge.Experiments show improved performance over previous unsupervised semantic role labeling approaches and other distributed word representation models.
Kristian Woodsend, Mirella Lapata
EMNLP2
2015 A Bayesian Model for Joint Learning of Categories and their Features
abstract
Categories such as ANIMAL or FURNITURE are acquired at an early age and play an important role in processing, organizing, and conveying world knowledge.Theories of categorization largely agree that categories are characterized by features such as function or appearance and that feature and category acquisition go hand-in-hand, however previous work has considered these problems in isolation.We present the first model that jointly learns categories and their features.The set of features is shared across categories, and strength of association is inferred in a Bayesian framework.We approximate the learning environment with natural language text which allows us to evaluate performance on a large scale.Compared to highly engineered pattern-based approaches, our model is cognitively motivated, knowledge-lean, and learns categories and features which are perceived by humans as more meaningful.
Lea Frermann, Mirella Lapata
HLT-NAACL2
2015 Movie Script Summarization as Graph-based Scene Extraction
abstract
In this paper we study the task of movie script summarization, which we argue could enhance script browsing, give readers a rough idea of the script's plotline, and speed up reading time.We formalize the process of generating a shorter version of a screenplay as the task of finding an optimal chain of scenes.We develop a graph-based model that selects a chain by jointly optimizing its logical progression, diversity, and importance.Human evaluation based on a question-answering task shows that our model produces summaries which are more informative compared to competitive baselines.
Philip John Gorinski, Mirella Lapata
HLT-NAACL2
2015 Learning to Interpret and Describe Abstract Scenes
abstract
Luis Gilberto Mateos Ortiz, Clemens Wolff, Mirella Lapata. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Luis Gilberto Mateos Ortiz, Clemens Wolff, Mirella Lapata
HLT-NAACL3
2015 Which Step Do I Take First? Troubleshooting with Bayesian Models
abstract
Online discussion forums and community question-answering websites provide one of the primary avenues for online users to share information. In this paper, we propose text mining techniques which aid users navigate troubleshooting-oriented data such as questions asked on forums and their suggested solutions. We introduce Bayesian generative models of the troubleshooting data and apply them to two interrelated tasks: (a) predicting the complexity of the solutions (e.g., plugging a keyboard in the computer is easier compared to installing a special driver) and (b) presenting them in a ranked order from least to most complex. Experimental results show that our models are on par with human performance on these tasks, while outperforming baselines based on solution length or readability.
Annie Louis, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2015 Context-aware Frame-Semantic Role Labeling
abstract
Frame semantic representations have been useful in several applications ranging from text-to-scene generation, to question answering and social network analysis. Predicting such representations from raw text is, however, a challenging task and corresponding models are typically only trained on a small set of sentence-level annotations. In this paper, we present a semantic role labeling system that takes into account sentence and discourse context. We introduce several new features which we motivate based on linguistic insights and experimentally demonstrate that they lead to significant improvements over the current state-of-the-art in FrameNet-based semantic role labeling.
Michael Roth 0001, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2014 Learning Grounded Meaning Representations with Autoencoders
abstract
In this paper we address the problem of grounding distributional representations of lexical meaning. We introduce a new model which uses stacked autoencoders to learn higher-level embeddings from tex-tual and visual input. The two modali-ties are encoded as vectors of attributes and are obtained automatically from text and images, respectively. We evaluate our model on its ability to simulate similar-ity judgments and concept categorization. On both tasks, our approach outperforms baselines and related models. 1
Carina Silberer, Mirella Lapata
ACL (1)2
2014 Unsupervised Interpretation of Eventive Propositions
Anselmo Peñas, Bernardo Cabaleiro, Mirella Lapata
CICLing (1)3
2014 Incremental Bayesian Learning of Semantic Categories
abstract
Models of category learning have been extensively studied in cognitive science and primarily tested on perceptual abstractions or artificial stimuli. In this paper we focus on categories acquired from natural language stimuli, that is words (e.g., chair is a member of the FURNITURE category). We present a Bayesian model which, unlike previous work, learns both categories and their features in a single process. Our model employs particle filters, a sequential Monte Carlo method commonly used for approximate probabilistic inference in an incremental setting. Comparison against a state-of-the-art graph-based approach reveals that our model learns qualitatively better categories and demonstrates cognitive plausibility during learning.
Lea Frermann, Mirella Lapata
EACL2
2014 Incremental Semantic Role Labeling with Tree Adjoining Grammar
abstract
We introduce the task of incremental semantic role labeling (iSRL), in which semantic roles are assigned to incomplete input (sentence prefixes).iSRL is the semantic equivalent of incremental parsing, and is useful for language modeling, sentence completion, machine translation, and psycholinguistic modeling.We propose an iSRL system that combines an incremental TAG parser with a semantically enriched lexicon, a role propagation algorithm, and a cascade of classifiers.Our approach achieves an SRL Fscore of 78.38% on the standard CoNLL 2009 dataset.It substantially outperforms a strong baseline that combines gold-standard syntactic dependencies with heuristic role assignment, as well as a baseline based on Nivre's incremental dependency parser.
Ioannis Konstas, Frank Keller, Vera Demberg, Mirella Lapata
EMNLP4
2014 Chinese Poetry Generation with Recurrent Neural Networks
abstract
We propose a model for Chinese poem generation based on recurrent neural networks which we argue is ideally suited to capturing poetic content and form.Our generator jointly performs content selection ("what to say") and surface realization ("how to say") by learning representations of individual characters, and their combinations into one or more lines as well as how these mutually reinforce and constrain each other.Poem lines are generated incrementally by taking into account the entire history of what has been generated so far rather than the limited horizon imposed by the previous line or lexical n-grams.Experimental results show that our model outperforms competitive Chinese poetry generation systems using both automatic and manual evaluation methods.
Xingxing Zhang 0002, Mirella Lapata
EMNLP2
2014 Similarity-Driven Semantic Role Induction via Graph Partitioning
abstract
As in many natural language processing tasks, data-driven models based on supervised learning have become the method of choice for semantic role labeling. These models are guaranteed to perform well when given sufficient amount of labeled training data. Producing this data is costly and time-consuming, however, thus raising the question of whether unsupervised methods offer a viable alternative. The working hypothesis of this article is that semantic roles can be induced without human supervision from a corpus of syntactically parsed sentences based on three linguistic principles: (1) arguments in the same syntactic position (within a specific linking) bear the same semantic role, (2) arguments within a clause bear a unique role, and (3) clusters representing the same semantic role should be more or less lexically and distributionally equivalent. We present a method that implements these principles and formalizes the task as a graph partitioning problem, whereby argument instances of a verb are represented as vertices in a graph whose edges express similarities between these instances. The graph consists of multiple edge layers, each one capturing a different aspect of argument-instance similarity, and we develop extensions of standard clustering algorithms for partitioning such multi-layer graphs. Experiments for English and German demonstrate that our approach is able to induce semantic role clusters that are consistently better than a strong baseline and are competitive with the state of the art.
Joel Lang, Mirella Lapata
Comput. Linguistics2
2014 Text Rewriting Improves Semantic Role Labeling
abstract
Large-scale annotated corpora are a prerequisite to developing high-performance NLP systems. Such corpora are expensive to produce, limited in size, often demanding linguistic expertise. In this paper we use text rewriting as a means of increasing the amount of labeled data available for model training. Our method uses automatically extracted rewrite rules from comparable corpora and bitexts to generate multiple versions of sentences annotated with gold standard labels. We apply this idea to semantic role labeling and show that a model trained on rewritten data outperforms the state of the art on the CoNLL-2009 benchmark dataset.
Kristian Woodsend, Mirella Lapata
J. Artif. Intell. Res.2
2014 Large-scale Semantic Parsing without Question-Answer Pairs
abstract
In this paper we introduce a novel semantic parsing approach to query Freebase in natural language without requiring manual annotations or question-answer pairs. Our key insight is to represent natural language via semantic graphs whose topology shares many commonalities with Freebase. Given this representation, we conceptualize semantic parsing as a graph matching problem. Our model converts sentences to semantic graphs using CCG and subsequently grounds them to Freebase guided by denotations as a form of weak supervision. Evaluation experiments on a subset of the Free917 and WebQuestions benchmark datasets show our semantic parser improves over the state of the art.
Siva Reddy, Mirella Lapata, Mark Steedman
Trans. Assoc. Comput. Linguistics2
2013 Models of Semantic Representation with Visual Attributes
Carina Silberer, Vittorio Ferrari, Mirella Lapata
ACL (1)3
2013 Inducing Document Plans for Concept-to-Text Generation
abstract
In a language generation system, a content planner selects which elements must be included in the output text and the ordering between them.Recent empirical approaches perform content selection without any ordering and have thus no means to ensure that the output is coherent.In this paper we focus on the problem of generating text from a database and present a trainable end-to-end generation system that includes both content selection and ordering.Content plans are represented intuitively by a set of grammar rules that operate on the document level and are acquired automatically from training data.We develop two approaches: the first one is inspired from Rhetorical Structure Theory and represents the document as a tree of discourse relations between database records; the second one requires little linguistic sophistication and uses tree structures to represent global patterns of database record sequences within a document.Experimental evaluation on two domains yields considerable improvements over the state of the art for both approaches.
Ioannis Konstas, Mirella Lapata
EMNLP2
2013 Unsupervised Relation Extraction with General Domain Knowledge
abstract
In this paper we present an unsupervised approach to relational information extraction.Our model partitions tuples representing an observed syntactic relationship between two named entities (e.g., "X was born in Y" and "X is from Y") into clusters corresponding to underlying semantic relation types (e.g., BornIn, Located).Our approach incorporates general domain knowledge which we encode as First Order Logic rules and automatically combine with a topic model developed specifically for the relation extraction task.Evaluation results on the ACE 2007 English Relation Detection and Categorization (RDC) task show that our model outperforms competitive unsupervised approaches by a wide margin and is able to produce clusters shaped by both the data and the rules.
Oier Lopez de Lacalle, Mirella Lapata
EMNLP2
2013 i, Poet: Automatic Chinese Poetry Composition through a Generative Summarization Framework under Constrained Optimization
Rui Yan 0001, Mirella Lapata, Shou-De Lin, Xueqiang Lv, Xiaoming Li 0001
IJCAI3
2013 Semantic v.s. Positions: Utilizing Balanced Proximity in Language Model Smoothing for Information Retrieval
Rui Yan 0001, Mirella Lapata, Shou-De Lin, Xueqiang Lv, Xiaoming Li 0001
IJCNLP3
2013 A Quantum-Theoretic Approach to Distributional Semantics
William Blacoe, Elham Kashefi, Mirella Lapata
HLT-NAACL3
2013 A Global Model for Concept-to-Text Generation
abstract
Concept-to-text generation refers to the task of automatically producing textual output from non-linguistic input. We present a joint model that captures content selection ("what to say") and surface realization ("how to say") in an unsupervised domain-independent fashion. Rather than breaking up the generation process into a sequence of local decisions, we define a probabilistic context-free grammar that globally describes the inherent structure of the input (a corpus of database records and text describing some of them). We recast generation as the task of finding the best derivation tree for a set of database records and describe an algorithm for decoding in this framework that allows to intersect the grammar with additional information capturing fluency and syntactic well-formedness constraints. Experimental evaluation on several domains achieves results competitive with state-of-the-art systems that use domain specific constraints, explicit feature engineering or labeled data.
Ioannis Konstas, Mirella Lapata
J. Artif. Intell. Res.2
2013 Automatic Caption Generation for News Images
abstract
This paper is concerned with the task of automatically generating captions for images, which is important for many image-related applications. Examples include video and image retrieval as well as the development of tools that aid visually impaired individuals to access pictorial information. Our approach leverages the vast resource of pictures available on the web and the fact that many of them are captioned and colocated with thematically related documents. Our model learns to create captions from a database of news articles, the pictures embedded in them, and their captions, and consists of two stages. Content selection identifies what the image and accompanying article are about, whereas surface realization determines how to verbalize the chosen content. We approximate content selection with a probabilistic image annotation model that suggests keywords for an image. The model postulates that images and their textual descriptions are generated by a shared set of latent variables (topics) and is trained on a weakly labeled dataset (which treats the captions and associated news articles as image labels). Inspired by recent work in summarization, we propose extractive and abstractive surface realization models. Experimental results show that it is viable to generate captions that are pertinent to the specific content of an image and its associated article, while permitting creativity in the description. Indeed, the output of our abstractive model compares favorably to handwritten captions and is often superior to extractive methods.
Yansong Feng 0002, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 An abstractive approach to sentence compression
abstract
In this article we generalize the sentence compression task. Rather than simply shorten a sentence by deleting words or constituents, as in previous work, we rewrite it using additional operations such as substitution, reordering, and insertion. We present an experimental study showing that humans can naturally create abstractive sentences using a variety of rewrite operations, not just deletion. We next create a new corpus that is suited to the abstractive compression task and formulate a discriminative tree-to-tree transduction model that can account for structural and lexical mismatches. The model incorporates a grammar extraction method, uses a language model for coherent output, and can be easily tuned to a wide range of compression-specific loss functions.
Trevor Cohn, Mirella Lapata
ACM Trans. Intell. Syst. Technol.2
2012 Concept-to-text Generation via Discriminative Reranking
Ioannis Konstas, Mirella Lapata
ACL (1)2
2012 Tweet Recommendation with Graph Co-Ranking
Rui Yan 0001, Mirella Lapata, Xiaoming Li 0001
ACL (1)2
2012 Visualizing timelines: evolutionary summarization via iterative reinforcement between text and image streams
abstract
We present a novel graph-based framework for timeline summarization, the task of creating different summaries for different timestamps but for the same topic. Our work extends timeline summarization to a multimodal setting and creates timelines that are both textual and visual. Our approach exploits the fact that news documents are often accompanied by pictures and the two share some common content. Our model optimizes local summary creation and global timeline generation jointly following an iterative approach based on mutual reinforcement and co-ranking. In our algorithm, individual summaries are generated by taking into account the mutual dependencies between sentences and images, and are iteratively refined by considering how they contribute to the global timeline and its coherence. Experiments on real-world datasets show that the timelines produced by our model outperform several competitive baselines both in terms of ROUGE and when assessed by human evaluators.
Rui Yan 0001, Xiaojun Wan 0001, Mirella Lapata, Wayne Xin Zhao, Pu-Jen Cheng, Xiaoming Li 0001
CIKM3
2012 A Comparison of Vector-based Representations for Semantic Composition
William Blacoe, Mirella Lapata
EMNLP-CoNLL2
2012 Grounded Models of Semantic Representation
Carina Silberer, Mirella Lapata
EMNLP-CoNLL2
2012 Multiple Aspect Summarization Using Integer Linear Programming
Kristian Woodsend, Mirella Lapata
EMNLP-CoNLL2
2012 Taxonomy Induction Using Hierarchical Random Graphs
Trevor Fountain, Mirella Lapata
HLT-NAACL2
2012 Unsupervised Concept-to-text Generation with Hypergraphs
Ioannis Konstas, Mirella Lapata
HLT-NAACL2
2012 Semi-Supervised Semantic Role Labeling via Structural Alignment
abstract
Large-scale annotated corpora are a prerequisite to developing high-performance semantic role labeling systems. Unfortunately, such corpora are expensive to produce, limited in size, and may not be representative. Our work aims to reduce the annotation effort involved in creating resources for semantic role labeling via semi-supervised learning. The key idea of our approach is to find novel instances for classifier training based on their similarity to manually labeled seed instances. The underlying assumption is that sentences that are similar in their lexical material and syntactic structure are likely to share a frame semantic analysis. We formalize the detection of similar sentences and the projection of role annotations as a graph alignment problem, which we solve exactly using integer linear programming. Experimental results on semantic role labeling show that the automatic annotations produced by our method improve performance over using hand-labeled instances alone.
Hagen Fürstenau, Mirella Lapata
Comput. Linguistics2
2011 WikiSimple: Automatic Simplification of Wikipedia Articles
abstract
Text simplification aims to rewrite text into simpler versions and thus make information accessible to a broader audience (e.g., non-native speakers, children, and individuals with language impairments). In this paper, we propose a model that simplifies documents automatically while selecting their most important content and rewriting them in a simpler style. We learn content selection rules from same-topic Wikipedia articles written in the main encyclopedia and its Simple English variant. We also use the revision histories of Simple Wikipedia articles to learn a quasi-synchronous grammar of simplification rewrite rules. Based on an integer linear programming formulation, we develop a joint model where preferences based on content and style are optimized simultaneously. Experiments on simplifying main Wikipedia articles show that our method significantly reduces the reading difficulty, while still capturing the important content.
Kristian Woodsend, Mirella Lapata
AAAI2
2011 Unsupervised Semantic Role Induction via Split-Merge Clustering
Joel Lang, Mirella Lapata
ACL2
2011 Incremental Models of Natural Language Category Acquisition
Trevor Fountain, Mirella Lapata
CogSci2
2011 Unsupervised Semantic Role Induction with Graph Partitioning
Joel Lang, Mirella Lapata
EMNLP2
2011 Learning to Simplify Sentences with Quasi-Synchronous Grammar and Integer Programming
Kristian Woodsend, Mirella Lapata
EMNLP2
2010 How Many Words Is a Picture Worth? Automatic Caption Generation for News Images
Yansong Feng 0002, Mirella Lapata
ACL2
2010 Plot Induction and Evolutionary Search for Story Generation
Neil Duncan McIntyre, Mirella Lapata
ACL2
2010 Syntactic and Semantic Factors in Processing Difficulty: An Integrated Measure
Jeff Mitchell 0001, Mirella Lapata, Vera Demberg, Frank Keller
ACL2
2010 Automatic Generation of Story Highlights
Kristian Woodsend, Mirella Lapata
ACL2
2010 Image and Natural Language Processing for Multimedia Information Retrieval
Mirella Lapata
ECIR1
2010 Measuring Distributional Similarity in Context
Georgiana Dinu, Mirella Lapata
EMNLP2
2010 Title Generation with Quasi-Synchronous Grammar
Kristian Woodsend, Yansong Feng 0002, Mirella Lapata
EMNLP3
2010 Visual Information in Semantic Representation
Yansong Feng 0002, Mirella Lapata
HLT-NAACL2
2010 Topic Models for Image Annotation and Text Illustration
Yansong Feng 0002, Mirella Lapata
HLT-NAACL2
2010 Unsupervised Induction of Semantic Roles
Joel Lang, Mirella Lapata
HLT-NAACL2
2010 Discourse Constraints for Document Compression
abstract
Sentence compression holds promise for many applications ranging from summarization to subtitle generation. The task is typically performed on isolated sentences without taking the surrounding context into account, even though most applications would operate over entire documents. In this article we present a discourse-informed model which is capable of producing document compressions that are coherent and informative. Our model is inspired by theories of local coherence and formulated within the framework of integer linear programming. Experimental results show significant improvements over a state-of-the-art discourse agnostic approach.
James Clarke, Mirella Lapata
Comput. Linguistics2
2010 An Experimental Study of Graph Connectivity for Unsupervised Word Sense Disambiguation
abstract
Word sense disambiguation (WSD), the task of identifying the intended meanings (senses) of words in context, has been a long-standing research objective for natural language processing. In this paper, we are concerned with graph-based algorithms for large-scale WSD. Under this framework, finding the right sense for a given word amounts to identifying the most "important" node among the set of graph nodes representing its senses. We introduce a graph-based WSD algorithm which has few parameters and does not require sense-annotated data for training. Using this algorithm, we investigate several measures of graph connectivity with the aim of identifying those best suited for WSD. We also examine how the chosen lexicon and its connectivity influences WSD performance. We report results on standard data sets and show that our graph-based approach performs comparably to the state of the art.
Roberto Navigli, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Learning to Tell Tales: A Data-driven Approach to Story Generation
Neil Duncan McIntyre, Mirella Lapata
ACL/IJCNLP2
2009 Bayesian Word Sense Induction
Samuel Brody, Mirella Lapata
EACL2
2009 Semi-Supervised Semantic Role Labeling
Hagen Fürstenau, Mirella Lapata
EACL2
2009 Graph Alignment for Semi-Supervised Semantic Role Labeling
Hagen Fürstenau, Mirella Lapata
EMNLP2
2009 Language Models Based on Semantic Composition
Jeff Mitchell 0001, Mirella Lapata
EMNLP2
2009 Sentence Compression as Tree Transduction
abstract
This paper presents a tree-to-tree transduction method for sentence compression. Our model is based on synchronous tree substitution grammar, a formalism that allows local distortion of the tree topology and can thus naturally capture structural mismatches. We describe an algorithm for decoding in this framework and show how the model can be trained discriminatively within a large margin framework. Experimental results on sentence compression bring significant improvements over a state-of-the-art model.
Trevor Cohn, Mirella Lapata
J. Artif. Intell. Res.2
2009 Cross-lingual Annotation Projection for Semantic Roles
abstract
This article considers the task of automatically inducing role-semantic annotations in the FrameNet paradigm for new languages. We propose a general framework that is based on annotation projection, phrased as a graph optimization problem. It is relatively inexpensive and has the potential to reduce the human effort involved in creating role-semantic resources. Within this framework, we present projection models that exploit lexical and syntactic information. We provide an experimental evaluation on an English-German parallel corpus which demonstrates the feasibility of inducing high-precision German semantic role annotation both for manually and automatically annotated English data.
Sebastian Padó, Mirella Lapata
J. Artif. Intell. Res.2
2008 Automatic Image Annotation Using Auxiliary Text Information
Yansong Feng 0002, Mirella Lapata
ACL2
2008 Vector-based Models of Semantic Composition
Jeff Mitchell 0001, Mirella Lapata
ACL2
2008 Good Neighbors Make Good Senses: Exploiting Distributional Similarity for Unsupervised WSD
Samuel Brody, Mirella Lapata
COLING2
2008 ParaMetric: An Automatic Evaluation Metric for Paraphrasing
Chris Callison-Burch, Trevor Cohn, Mirella Lapata
COLING3
2008 Sentence Compression Beyond Word Deletion
Trevor Cohn, Mirella Lapata
COLING2
2008 Modeling Local Coherence: An Entity-Based Approach
abstract
This article proposes a novel framework for representing and measuring local coherence. Central to this approach is the entity-grid representation of discourse, which captures patterns of entity distribution in a text. The algorithm introduced in the article automatically abstracts a text into a set of entity transition sequences and records distributional, syntactic, and referential information about discourse entities. We re-conceptualize coherence assessment as a learning task and show that our entity-based representation is well-suited for ranking-based generation and text classification tasks. Using the proposed representation, we achieve good performance on text ordering, summary coherence evaluation, and readability assessment.
Regina Barzilay, Mirella Lapata
Comput. Linguistics2
2008 Constructing Corpora for the Development and Evaluation of Paraphrase Systems
abstract
Automatic paraphrasing is an important component in many natural language processing tasks. In this article we present a new parallel corpus with paraphrase annotations. We adopt a definition of paraphrase based on word alignments and show that it yields high inter-annotator agreement. As Kappa is suited to nominal data, we employ an alternative agreement statistic which is appropriate for structured alignment tasks. We discuss how the corpus can be usefully employed in evaluating paraphrase systems automatically (e.g., by measuring precision, recall, and F1) and also in developing linguistically rich paraphrase models based on syntactic structure.
Trevor Cohn, Chris Callison-Burch, Mirella Lapata
Comput. Linguistics3
2008 Global Inference for Sentence Compression: An Integer Linear Programming Approach
abstract
Sentence compression holds promise for many applications ranging from summarization to subtitle generation. Our work views sentence compression as an optimization problem and uses integer linear programming (ILP) to infer globally optimal compressions in the presence of linguistically motivated constraints. We show how previous formulations of sentence compression can be recast as ILPs and extend these models with novel global constraints. Experimental results on written and spoken texts demonstrate improvements over state-of-the-art models.
James Clarke, Mirella Lapata
J. Artif. Intell. Res.2
2007 Machine Translation by Triangulation: Making Effective Use of Multi-Parallel Corpora
Trevor Cohn, Mirella Lapata
ACL2
2007 Modelling Compression with Discourse Constraints
James Clarke, Mirella Lapata
EMNLP-CoNLL2
2007 Large Margin Synchronous Generation and its Application to Sentence Compression
Trevor Cohn, Mirella Lapata
EMNLP-CoNLL2
2007 Using Semantic Roles to Improve Question Answering
Mirella Lapata
EMNLP-CoNLL2
2007 Graph Connectivity Measures for Unsupervised Word Sense Disambiguation
Roberto Navigli, Mirella Lapata
IJCAI2
2007 An Information Retrieval Approach to Sense Ranking
Mirella Lapata, Frank Keller
HLT-NAACL1
2007 Dependency-Based Construction of Semantic Space Models
abstract
Traditionally, vector-based semantic space models use word co-occurrence counts from large corpora to represent lexical meaning. In this article we present a novel framework for constructing semantic spaces that takes syntactic relations into account. We introduce a formalization for this class of models, which allows linguistic knowledge to guide the construction process. We evaluate our framework on a range of tasks relevant for cognitive science and natural language processing: semantic priming, synonymy detection, and word sense disambiguation. In all cases, our framework obtains results that are comparable or superior to the state of the art.
Sebastian Padó, Mirella Lapata
Comput. Linguistics2
2006 Ensemble Methods for Unsupervised WSD
abstract
Combination methods are an effective way of improving system performance. This paper examines the benefits of system combination for unsupervised WSD. We investigate several voting- and arbiter-based combination strategies over a diverse pool of unsupervised WSD systems. Our combination methods rely on predominant senses which are derived automatically from raw text. Experiments using the SemCor and Senseval-3 data sets demonstrate that our ensembles yield significantly better results when compared with state-of-the-art.
Samuel Brody, Roberto Navigli, Mirella Lapata
ACL3
2006 Models for Sentence Compression: A Comparison across Domains, Training Requirements and Evaluation Measures
abstract
Sentence compression is the task of producing a summary at the sentence level. This paper focuses on three aspects of this task which have not received detailed treatment in the literature: training requirements, scalability, and automatic evaluation. We provide a novel comparison between a supervised constituent-based and an weakly supervised word-based compression algorithm and examine how these models port to different domains (written vs. spoken text). To achieve this, a human-authored compression corpus has been created and our study highlights potential problems with the automatically gathered compression corpora currently used. Finally, we assess whether automatic evaluation measures can be used to determine compression quality.
James Clarke, Mirella Lapata
ACL2
2006 Constraint-Based Sentence Compression: An Integer Programming Approach
James Clarke, Mirella Lapata
ACL2
2006 Optimal Constituent Alignment with Edge Covers for Semantic Projection
abstract
Given a parallel corpus, semantic projection attempts to transfer semantic role annotations from one language to another, typically by exploiting word alignments. In this paper, we present an improved method for obtaining constituent alignments between parallel sentences to guide the role projection task. Our extensions are twofold: (a) we model constituent alignment as minimum weight edge covers in a bipartite graph, which allows us to find a globally optimal solution efficiently; (b) we propose tree pruning as a promising strategy for reducing alignment noise. Experimental results on an English-German parallel corpus demonstrate improvements over state-of-the-art models.
Sebastian Padó, Mirella Lapata
ACL2
2006 Aggregation via Set Partitioning for Natural Language Generation
Regina Barzilay, Mirella Lapata
HLT-NAACL2
2006 Automatic Evaluation of Information Ordering: Kendall's Tau
abstract
This article considers the automatic evaluation of information ordering, a task underlying many text-based applications such as concept-to-text generation and multidocument summarization. We propose an evaluation method based on Kendall's τ, a metric of rank correlation. The method is inexpensive, robust, and representation independent. We show that Kendall's τ correlates reliably with human ratings and reading times.
Mirella Lapata
Comput. Linguistics1
2006 Learning Sentence-internal Temporal Relations
abstract
In this paper we propose a data intensive approach for inferring sentence-internal temporal relations. Temporal inference is relevant for practical NLP applications which either extract or synthesize temporal information (e.g., summarisation, question answering). Our method bypasses the need for manual coding by exploiting the presence of markers like ``after", which overtly signal a temporal relation. We first show that models trained on main and subordinate clauses connected with a temporal marker achieve good performance on a pseudo-disambiguation task simulating temporal inference (during testing the temporal marker is treated as unseen and the models must select the right marker from a set of possible candidates). Secondly, we assess whether the proposed approach holds promise for the semi-automatic creation of temporal annotations. Specifically, we use a model trained on noisy and approximate data (i.e., main and subordinate clauses) to predict intra-sentential relations present in TimeBank, a corpus annotated rich temporal information. Our experiments compare and contrast several probabilistic models differing in their feature space, linguistic assumptions and data requirements. We evaluate performance against gold standard corpora and also against human subjects.
Mirella Lapata, Alex Lascarides
J. Artif. Intell. Res.1
2005 Cross-Lingual Bootstrapping of Semantic Lexicons: The Case of FrameNet
Sebastian Padó, Mirella Lapata
AAAI2
2005 Modeling Local Coherence: An Entity-Based Approach
abstract
This paper considers the problem of automatic assessment of local coherence.We present a novel entity-based representation of discourse which is inspired by Centering Theory and can be computed automatically from raw text.We view coherence assessment as a ranking learning problem and show that the proposed discourse representation supports the effective learning of a ranking function.Our experiments demonstrate that the induced model achieves significantly higher accuracy than a state-of-the-art coherence model.
Regina Barzilay, Mirella Lapata
ACL2
2005 Automatic Evaluation of Text Coherence: Models and Representations
Mirella Lapata, Regina Barzilay
IJCAI1
2005 A comparison of parsing technologies for the biomedical domain
abstract
This paper reports on a number of experiments which are designed to investigate the extent to which current NLP resources are able to syntactically and semantically analyse biomedical text. We address two tasks: (a) parsing a real corpus with a hand-built wide-coverage grammar, producing both syntactic analyses and logical forms and (b) automatically computing the interpretation of compound nouns where the head is a nominalisation (e.g. hospital arrival means an arrival at hospital, while patient arrival means an arrival of a patient). For the former task we demonstrate that flexible and yet constrained pre-processing techniques are crucial to success: these enable us to use part-of-speech tags to overcome inadequate lexical coverage, and to package up complex technical expressions prior to parsing so that they are blocked from creating misleading amounts of syntactic complexity. We argue that the XML-processing paradigm is ideally suited for automatically preparing the corpus for parsing. For the latter task, we compute interpretations of the compounds by exploiting surface cues and meaning paraphrases, which in turn are extracted from the parsed corpus. This provides an empirical setting in which we can compare the utility of a comparatively deep parser vs. a shallow one, exploring the trade-off between resolving attachment ambiguities on the one hand and generating errors in the parses on the other. We demonstrate that a model of the meaning of compound nominalisations is achievable with the aid of current broad-coverage parsers.
Claire Grover, Alex Lascarides, Mirella Lapata
Nat. Lang. Eng.3
2004 Automatic Paragraph Identification: A Study across Languages and Domains
Caroline Sporleder, Mirella Lapata
EMNLP2
2004 The Web as a Baseline: Evaluating the Performance of Unsupervised Web-based Models for a Range of NLP Tasks
Mirella Lapata, Frank Keller
HLT-NAACL1
2004 Inferring Sentence-internal Temporal Relations
Mirella Lapata, Alex Lascarides
HLT-NAACL1
2004 Verb Class Disambiguation Using Informative Priors
abstract
Levin's (1993) study of verb classes is a widely used resource for lexical semantics. In her framework, some verbs, such as give, exhibit no class ambiguity. But other verbs, such as write, have several alternative classes. We extend Levin's inventory to a simple statistical model of verb class ambiguity. Using this model we are able to generate preferences for ambiguous verbs without the use of a disambiguated corpus. We additionally show that these preferences are useful as priors for a verb sense disambiguator.
Mirella Lapata, Chris Brew
Comput. Linguistics1
2003 Probabilistic Text Structuring: Experiments with Sentence Ordering
abstract
Ordering information is a critical task for natural language generation applications. In this paper we propose an approach to information ordering that is particularly suited for text-to-text generation. We describe a model that learns constraints on sentence order from a corpus of domain-specific texts and an algorithm that yields the most likely order among several alternatives. We evaluate the automatically generated orderings against authored texts from our corpus and against human subjects that are asked to mimic the model's task. We also assess the appropriateness of such a model for multidocument summarization.
Mirella Lapata
ACL1
2003 Constructing Semantic Space Models from Parsed Corpora
abstract
Traditional vector-based models use word co-occurrence counts from large corpora to represent lexical meaning. In this paper we present a novel approach for constructing semantic spaces that takes syntactic relations into account. We introduce a formalisation for this class of models and evaluate their adequacy on two modelling tasks: semantic priming and automatic discrimination of lexical relations.
Sebastian Padó, Mirella Lapata
ACL2
2003 Evaluating and Combining Approaches to Selectional Preference Acquisition
Carsten Brockmann, Mirella Lapata
EACL2
2003 Detecting Novel Compounds: The Role of Distributional Evidence
Mirella Lapata, Alex Lascarides
EACL1
2003 Using the Web to Obtain Frequencies for Unseen Bigrams
abstract
This article shows that the Web can be employed to obtain frequencies for bigrams that are unseen in a given corpus. We describe a method for retrieving counts for adjective-noun, noun-noun, and verb-object bigrams from the Web by querying a search engine. We evaluate this method by demonstrating: (a) a high correlation between Web frequencies and corpus frequencies; (b) a reliable correlation between Web frequencies and plausibility judgments; (c) a reliable correlation between Web frequencies and frequencies recreated using class-based smoothing; (d) a good performance of Web frequencies in a pseudo disambiguation task.
Frank Keller, Mirella Lapata
Comput. Linguistics2
2003 A Probabilistic Account of Logical Metonymy
abstract
In this article we investigate logical metonymy, that is, constructions in which the argument of a word in syntax appears to be different from that argument in logical form (e.g., enjoy the book means enjoy reading the book, and easy problem means a problem that is easy to solve). The systematic variation in the interpretation of such constructions suggests a rich and complex theory of composition on the syntax/semantics interface. Linguistic accounts of logical metonymy typically fail to describe exhaustively all the possible interpretations, or they don't rank those interpretations in terms of their likelihood. In view of this, we acquire the meanings of metonymic verbs and adjectives from a large corpus and propose a probabilistic model that provides a ranking on the set of possible interpretations. We identify the interpretations automatically by exploiting the consistent correspondences between surface syntactic cues and meaning. We evaluate our results against paraphrase judgments elicited experimentally from humans and show that the model's ranking of meanings correlates reliably with human intuitions.
Mirella Lapata, Alex Lascarides
Comput. Linguistics1
1999 Using Subcategorization to Resolve Verb Class Ambiguity
Mirella Lapata, Chris Brew
EMNLP1