Claire Gardent

dblp:71/6819 · DBLP profile ↗
← Back
85ranked-venue papers
24as first author
15since 2021 · last 2026
0000-0002-3805-6662ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 24 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Privacy-Preserving Generation of Synthetic Pathology Reports for Information Extraction
Alejandra Lorenzo, Adrien Coulet, Claire Gardent
AIME (1)3
2025 MuCAL: Contrastive Alignment for Preference-Driven KG-to-Text Generation
abstract
We propose MuCAL (Multilingual Contrastive Alignment Learning) to tackle the challenge of Knowledge Graphs (KG)-to-Text generation using preference learning, where reliable preference data is scarce.MuCAL is a multilingual KG/Text alignment model achieving robust cross-modal retrieval across multiple languages and difficulty levels.Building on Mu-CAL, we automatically create preference data by ranking candidate texts from three LLMs (Qwen2.5 , DeepSeek-v3, Llama-3).We then apply Direct Preference Optimisation (DPO) on these preference data, bypassing typical reward modelling steps to directly align generation outputs with graph semantics.Extensive experiments on KG-to-English Text generation show two main advantages: (1) Our KG/Text alignment model provides a better signal for DPO than similar existing metrics, and (2) significantly better generalisation on out-of-domain datasets compared to standard instruction tuning.Our results highlight MuCAL's effectiveness in supporting preference learning for KGto-English Text generation and lay the foundation for future multilingual extensions.Code and data are available at https://github. com
Claire Gardent
EMNLP2
2025 Fine-Tuning, Prompting and RAG for Knowledge Graph-to-Russian Text Generation. How do these Methods generalise to Out-of-Distribution Data?
abstract
Prior work on Knowledge Graph-to-Text generation has mostly evaluated models on in-domain test sets and/or with English as the target language. In contrast, we focus on Russian and we assess how various generation methods perform on out-of-domain, unseen data. Previous studies have shown that enriching the input with target-language verbalisations of entities and properties substantially improves the performance of fine-tuned models for Russian. We compare multiple variants of two contemporary paradigms — LLM prompting and Retrieval-Augmented Generation (RAG) — and investigate alternative ways to integrate such external knowledge into the generation process. Using automatic metrics and human evaluation, we find that on unseen data the fine-tuned model consistently underperforms, revealing limited generalisation capacity; that while it outperforms RAG by a small margin on most datasets, prompting generates less fluent text; and conversely, that RAG generates text that is less faithful to the input. Overall, both LLM prompting and RAG outperform Fine-Tuning across all unseen testsets. The code for this paper is available at https://github.com/Javanochka/KG-to-text-fine-tuning-prompting-rag
Anna Nikiforovskaya, William Soto Martinez, Evan Parker Kelly Chapple, Claire Gardent
INLG4
2025 Generating Complex Question Decompositions in the Face of Distribution Shifts
abstract
Kelvin Han, Claire Gardent. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kelvin Han, Claire Gardent
NAACL (Long Papers)2
2024 KGConv, a Conversational Corpus Grounded in Wikidata
abstract
We present KGConv, a large corpus of 71k English conversations where each question-answer pair is grounded in a Wikidata fact. Conversations contain on average 8.6 questions and for each Wikidata fact, we provide multiple variants (12 on average) of the corresponding question using templates, human annotations, hand-crafted rules and a question rewriting neural model. We provide baselines for the task of Knowledge-Based, Conversational Question Generation. KGConv can further be used for other generation and analysis tasks such as single-turn question generation from Wikidata triples, question rewriting, question answering from conversation or from knowledge graphs and quiz generation.
Quentin Brabant, Lina Maria Rojas-Barahona, Gwénolé Lecorvé, Claire Gardent
LREC/COLING4
2024 Generating from AMRs into High and Low-Resource Languages using Phylogenetic Knowledge and Hierarchical QLoRA Training (HQL)
abstract
Previous work on multilingual generation from Abstract Meaning Representations has mostly focused on High-and Medium-Resource languages relying on large amounts of training data.In this work, we consider both Highand Low-Resource languages capping training data size at the lower bound set by our Low-Resource languages i.e., 31K training instances.We propose two straightforward techniques to enhance generation results on Low-Resource while preserving performance on High-and Medium-Resource languages.First, we iteratively refine a multilingual model to a set of monolingual models using Low-Rank Adaptation -this enables cross-lingual transfer while reducing over-fitting for High-Resource languages as the monolingual models are trained last.Second, we base our training curriculum on a tree structure which permits investigating how the languages used at each iteration impact generation performance on High and Low-Resource languages.We show an improvement over both mono and multilingual approaches.Comparing different ways of grouping languages at each iteration step we find two beneficial configurations: grouping related languages which promotes transfer, or grouping distant languages which facilitates regularisation.
William Soto Martinez, Yannick Parmentier 0001, Claire Gardent
INLG3
2024 Evaluating RDF-to-text Generation Models for English and Russian on Out Of Domain Data
abstract
While the WebNLG dataset has prompted much research on generation from knowledge graphs, little work has examined how well models trained on the WebNLG data generalise to unseen data and work has mostly been focused on English.In this paper, we introduce novel benchmarks for both English and Russian which contain various ratios of unseen entities and properties.These benchmarks also differ from WebNLG in that some of the graphs stem from Wikidata rather than DBpedia.Evaluating various models for English and Russian on these benchmarks shows a strong decrease in performance while a qualitative analysis highlights the various types of errors induced by non i.i.d data.
Anna Nikiforovskaya, Claire Gardent
INLG2
2023 Document-Level Planning for Text Simplification
abstract
Most existing work on text simplification is limited to sentence-level inputs, with attempts to iteratively apply these approaches to document-level simplification failing to coherently preserve the discourse structure of the document.We hypothesise that by providing a high-level view of the target document, a simplification plan might help to guide generation.Building upon previous work on controlled, sentence-level simplification, we view a plan as a sequence of labels, each describing one of four sentence-level simplification operations (copy, rephrase, split, or delete).We propose a planning model that labels each sentence in the input document while considering both its context (a window of surrounding sentences) and its internal structure (a token-level representation).Experiments on two simplification benchmarks (Newsela-auto and Wikiauto) show that our model outperforms strong baselines both on the planning task and when used to guide document-level simplification models.
Liam Cripwell, Joël Legrand, Claire Gardent
EACL3
2023 Simplicity Level Estimate (SLE): A Learned Reference-Less Metric for Sentence Simplification
abstract
Automatic evaluation for sentence simplification remains a challenging problem.Most popular evaluation metrics require multiple high-quality references -something not readily available for simplification -which makes it difficult to test performance on unseen domains.Furthermore, most existing metrics conflate simplicity with correlated attributes such as fluency or meaning preservation.We propose a new learned evaluation metric (SLE) which focuses on simplicity, outperforming almost all existing metrics in terms of correlation with human judgements.
Liam Cripwell, Joël Legrand, Claire Gardent
EMNLP3
2023 Generating and Answering Simple and Complex Questions from Text and from Knowledge Graphs
abstract
Kelvin Han, Claire Gardent. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kelvin Han, Claire Gardent
IJCNLP (1)2
2023 Phylogeny-Inspired Soft Prompts For Data-to-Text Generation in Low-Resource Languages
abstract
William Soto Martinez, Yannick Parmentier, Claire Gardent. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
William Soto Martinez, Yannick Parmentier 0001, Claire Gardent
IJCNLP (1)3
2022 Generating Biographies on Wikipedia: The Impact of Gender Bias on the Retrieval-Based Generation of Women Biographies
abstract
Generating factual, long-form text such as Wikipedia articles raises three key challenges: how to gather relevant evidence, how to structure information into well-formed text, and how to ensure that the generated text is factually correct.We address these by developing a model for English text that uses a retrieval mechanism to identify relevant supporting information on the web and a cache-based pre-trained encoderdecoder to generate long-form biographies section by section, including citation information.To assess the impact of available web evidence on the output text, we compare the performance of our approach when generating biographies about women (for which less information is available on the web) vs. biographies generally.To this end, we curate a dataset of 1,500 biographies about women.We analyze our generated text to understand how differences in available web evidence data affect generation.We evaluate the factuality, fluency, and quality of the generated texts using automatic metrics and human evaluation.We hope that these techniques can be used as a starting point for human writers, to aid in reducing the complexity inherent in the creation of long-form, factual text.
Angela Fan, Claire Gardent
ACL (1)2
2022 Generating Questions from Wikidata Triples
abstract
Question generation from knowledge bases (or knowledge base question generation, KBQG) is the task of generating questions from structured database information, typically in the form of triples representing facts. To handle rare entities and generalize to unseen properties, previous work on KBQG resorted to extensive, often ad-hoc pre- and post-processing of the input triple. We revisit KBQG – using pre training, a new (triple, question) dataset and taking question type into account – and show that our approach outperforms previous work both in a standard and in a zero-shot setting. We also show that the extended KBQG dataset (also helpful for knowledge base question answering) we provide allows not only for better coverage in terms of knowledge base (KB) properties but also for increased output variability in that it permits the generation of multiple questions from the same KB triple.
Kelvin Han, Thiago Castro Ferreira, Claire Gardent
LREC3
2021 Augmenting Transformers with KNN-Based Composite Memory for Dialog
abstract
Various machine learning tasks can benefit from access to external information of different modalities, such as text and images. Recent work has focused on learning architectures with large memories capable of storing this knowledge. We propose augmenting generative Transformer neural networks with KNN-based Information Fetching (KIF) modules. Each KIF module learns a read operation to access fixed external knowledge. We apply these modules to generative dialog modeling, a challenging task where information must be flexibly retrieved and incorporated to maintain the topic and flow of conversation. We demonstrate the effectiveness of our approach by identifying relevant knowledge required for knowledgeable but engaging dialog from Wikipedia, images, and human-written dialog utterances, and show that leveraging this retrieved information improves model performance, measured by automatic and human evaluation.
Angela Fan, Claire Gardent, Chloé Braud, Antoine Bordes
Trans. Assoc. Comput. Linguistics2
2021 An Error Analysis Framework for Shallow Surface Realisation
abstract
Abstract The metrics standardly used to evaluate Natural Language Generation (NLG) models, such as BLEU or METEOR, fail to provide information on which linguistic factors impact performance. Focusing on Surface Realization (SR), the task of converting an unordered dependency tree into a well-formed sentence, we propose a framework for error analysis which permits identifying which features of the input affect the models’ results. This framework consists of two main components: (i) correlation analyses between a wide range of syntactic metrics and standard performance metrics and (ii) a set of techniques to automatically identify syntactic constructs that often co-occur with low performance scores. We demonstrate the advantages of our framework by performing error analysis on the results of 174 system runs submitted to the Multilingual SR shared tasks; we show that dependency edge accuracy correlate with automatic metrics thereby providing a more interpretable basis for evaluation; and we suggest ways in which our framework could be used to improve models and data. The framework is available in the form of a toolkit which can be used both by campaign organizers to provide detailed, linguistically interpretable feedback on the state of the art in multilingual SR, and by individual researchers to improve models and datasets.1
Anastasia Shimorina, Claire Gardent, Yannick Parmentier 0001
Trans. Assoc. Comput. Linguistics2
2020 Learning Health-Bots from Training Data that was Automatically Created using Paraphrase Detection and Expert Knowledge
abstract
A key bottleneck for developing dialog models is the lack of adequate training data.Due to privacy issues, dialog data is even scarcer in the health domain.We propose a novel method for creating dialog corpora which we apply to create doctor-patient interaction data.We use this data to learn both a generation and a hybrid classification/retrieval model and find that the generation model consistently outperforms the hybrid model.We show that our data creation method has several advantages.Not only does it allow for the semi-automatic creation of large quantities of training data.It also provides a natural way of guiding learning and a novel method for assessing the quality of human-machine interactions.
Anna Liednikova, Philippe Jolivet, Alexandre Durand-Salmon, Claire Gardent
COLING4
2020 Multilingual AMR-to-Text Generation
abstract
Generating text from structured data is challenging because it requires bridging the gap between (i) structure and natural language (NL) and (ii) semantically underspecified input and fully specified NL output.Multilingual generation brings in an additional challenge: that of generating into languages with varied word order and morphological properties.In this work, we focus on Abstract Meaning Representations (AMRs) as structured input, where previous research has overwhelmingly focused on generating only into English.We leverage advances in cross-lingual embeddings, pretraining, and multilingual models to create multilingual AMR-to-text models that generate in twenty one different languages.For eighteen languages, based on automatic metrics, our multilingual models surpass baselines that generate into a single language.We analyse the ability of our multilingual models to accurately capture morphology and word order using human evaluation, and find that native speakers judge our generations to be fluent.
Angela Fan, Claire Gardent
EMNLP (1)2
2020 Modeling Global and Local Node Contexts for Text Generation from Knowledge Graphs
abstract
Recent graph-to-text models generate text from graph-based data using either global or local aggregation to learn node representations. Global node encoding allows explicit communication between two distant nodes, thereby neglecting graph topology as all nodes are directly connected. In contrast, local node encoding considers the relations between neighbor nodes capturing the graph structure, but it can fail to capture long-range relations. In this work, we gather both encoding strategies, proposing novel neural models that encode an input graph combining both global and local node contexts, in order to learn better contextualized node embeddings. In our experiments, we demonstrate that our approaches lead to significant improvements on two graph-to-text datasets achieving BLEU scores of 18.01 on the AGENDA dataset, and 63.69 on the WebNLG dataset for seen categories, outperforming state-of-the-art models by 3.7 and 3.1 points, respectively. 1
Leonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych
Trans. Assoc. Comput. Linguistics3
2019 Using Local Knowledge Graph Construction to Scale Seq2Seq Models to Multi-Document Inputs
abstract
Angela Fan, Claire Gardent, Chloé Braud, Antoine Bordes. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Angela Fan, Claire Gardent, Chloé Braud, Antoine Bordes
EMNLP/IJCNLP (1)2
2019 Enhancing AMR-to-Text Generation with Dual Graph Representations
abstract
Leonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Leonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych
EMNLP/IJCNLP (1)2
2019 Surface Realisation Using Full Delexicalisation
abstract
Anastasia Shimorina, Claire Gardent. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Anastasia Shimorina, Claire Gardent
EMNLP/IJCNLP (1)2
2019 RL Extraction of Syntax-Based Chunks for Sentence Compression
Hoa T. Le, Christophe Cerisara, Claire Gardent
ICANN (4)3
2019 Generating Text from Anonymised Structures
abstract
Surface realisation maps a meaning representation (MR) to a text, usually a single sentence.In this paper, we introduce a new parallel dataset of deep meaning representations and French sentences and we present a novel method for MR-to-text generation which seeks to generalise by abstracting away from lexical content.Most current work on natural language generation focuses on generating text that matches a reference using BLEU as evaluation criteria.In this paper, we additionally consider the model's ability to reintroduce the function words that are absent from the deep input meaning representations.We show that our approach increases both BLEU score and the scores used to assess function words generation.
Émilie Colin, Claire Gardent
INLG2
2019 Revisiting the Binary Linearization Technique for Surface Realization
abstract
End-to-end neural approaches have achieved state-of-the-art performance in many natural language processing (NLP) tasks.Yet, they often lack transparency of the underlying decision-making process, hindering error analysis and certain model improvements.In this work, we revisit the binary linearization approach to surface realization, which exhibits more interpretable behavior, but was falling short in terms of prediction accuracy.We show how enriching the training data to better capture word order constraints almost doubles the performance of the system.We further demonstrate that encoding both local and global prediction contexts yields another considerable performance boost.With the proposed modifications, the system which ranked low in the latest shared task on multilingual surface realization now achieves best results in five out of ten languages, while being on par with the state-of-the-art approaches in others. 1
Yevgeniy Puzikov, Claire Gardent, Ido Dagan, Iryna Gurevych
INLG2
2018 Generating Syntactic Paraphrases
abstract
We study the automatic generation of syntactic paraphrases using four different models for generation: data-to-text generation, textto-text generation, text reduction and text expansion, We derive training data for each of these tasks from the WebNLG dataset and we show (i) that conditioning generation on syntactic constraints effectively permits the generation of syntactically distinct paraphrases for the same input and (ii) that exploiting different types of input (data, text or data+text) further increases the number of distinct paraphrases that can be generated for a given input.
Émilie Colin, Claire Gardent
EMNLP2
2018 Handling Rare Items in Data-to-Text Generation
abstract
Neural approaches to data-to-text generation generally handle rare input items using either delexicalisation or a copy mechanism.We investigate the relative impact of these two methods on two datasets (E2E and WebNLG) and using two evaluation settings.We show (i) that rare items strongly impact performance; (ii) that combining delexicalisation and copying yields the strongest improvement; (iii) that copying underperforms for rare and unseen items and (iv) that the impact of these two mechanisms greatly varies depending on how the dataset is constructed and on how it is split into train, dev and test 1 .
Anastasia Shimorina, Claire Gardent
INLG2
2017 Creating Training Corpora for NLG Micro-Planners
abstract
In this paper, we present a novel framework for semi-automatically creating linguistically challenging microplanning data-to-text corpora from existing Knowledge Bases.Because our method pairs data of varying size and shape with texts ranging from simple clauses to short texts, a dataset created using this framework provides a challenging benchmark for microplanning.Another feature of this framework is that it can be applied to any large scale knowledge base and can therefore be used to train and learn KB verbalisers.We apply our framework to DBpedia data and compare the resulting dataset with Wen et al. (2016)'s.We show that while Wen et al.'s dataset is more than twice larger than ours, it is less diverse both in terms of input and in terms of text.We thus propose our corpus generation framework as a novel method for creating challenging data sets from which NLG models can be learned which are capable of handling the complex interactions occurring during in micro-planning between lexicalisation, aggregation, surface realisation, referring expression generation and sentence segmentation.To encourage researchers to take up this challenge, we recently made available a dataset created using this framework in the context of the WEBNLG shared task.
Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini
ACL (1)1
2017 Split and Rephrase
abstract
We propose a new sentence simplification task (Split-and-Rephrase) where the aim is to split a complex sentence into a meaning preserving sequence of shorter sentences.Like sentence simplification, splitting-and-rephrasing has the potential of benefiting both natural language processing and societal applications.Because shorter sentences are generally better processed by NLP systems, it could be used as a preprocessing step which facilitates and improves the performance of parsers, semantic role labelers and machine translation systems.It should also be of use for people with reading disabilities because it allows the conversion of longer sentences into shorter ones.This paper makes two contributions towards this new task.First, we create and make available a benchmark consisting of 1,066,115 tuples mapping a single complex sentence to a sequence of sentences expressing the same meaning.1 Second, we propose five models (vanilla sequence-to-sequence to semantically-motivated models) to understand the difficulty of the proposed task.
Shashi Narayan, Claire Gardent, Shay B. Cohen, Anastasia Shimorina
EMNLP2
2017 Mapping Natural Language to Description Logic
Bikash Gyawali, Anastasia Shimorina, Claire Gardent, Samuel Cruz-Lara, Mariem Mahfoudh
ESWC (1)3
2017 Symbolic Priors for RNN-based Semantic Parsing
abstract
Seq2seq models based on Recurrent Neural Networks (RNNs) have recently received a lot of attention in the domain of Semantic Parsing. While in principle they can be trained directly on pairs (natural language utterances, logical forms), their performance is limited by the amount of available data. To alleviate this problem, we propose to exploit various sources of prior knowledge: the well-formedness of the logical forms is modeled by a weighted context-free grammar; the likelihood that certain entities present in the input utterance are also present in the logical form is modeled by weighted finite-state automata. The grammar and automata are combined together through an efficient intersection algorithm to form a soft guide (“background”) to the RNN.We test our method on an extension of the Overnight dataset and show that it not only strongly improves over an RNN baseline, but also outperforms non-RNN models based on rich sets of hand-crafted features.
Chunyang Xiao, Marc Dymetman, Claire Gardent
IJCAI3
2017 The WebNLG Challenge: Generating Text from RDF Data
abstract
The WebNLG challenge consists in mapping sets of RDF triples to text.It provides a common benchmark on which to train, evaluate and compare "microplanners", i.e. generation systems that verbalise a given content by making a range of complex interacting choices including referring expression generation, aggregation, lexicalisation, surface realisation and sentence segmentation.In this paper, we introduce the microplanning task, describe data preparation, introduce our evaluation methodology, analyse participant results and provide a brief description of the participating systems.(3) a. LOCATION-COUNTRY-STARTDATE ⇒ Passive-Apposition-Active
Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini
INLG1
2017 Analysing Data-To-Text Generation Benchmarks
abstract
A generation system can only be as good as the data it is trained on.In this short paper, we propose a methodology for analysing data-to-text corpora used for training microplanner i.e., systems which given some input must produce a text verbalising exactly this input.We apply this methodology to three existing benchmarks and we elicite a set of criteria for the creation of a data-to-text benchmark which could help better support the development, evaluation and comparison of linguistically sophisticated data-to-text generators.
Laura Perez-Beltrachini, Claire Gardent
INLG2
2017 ModelWriter: text and model-synchronized document engineering platform
abstract
The ModelWriter platform provides a generic framework for automated traceability analysis. In this paper, we demonstrate how this framework can be used to trace the consistency and completeness of technical documents that consist of a set of System Installation Design Principles used by Airbus to ensure the correctness of aircraft system installation. We show in particular, how the platform allows the integration of two types of reasoning: reasoning about the meaning of text using semantic parsing and description logic theorem proving; and reasoning about document structure using first-order relational logic and finite model finding for traceability analysis.
Ferhat Erata, Claire Gardent, Bikash Gyawali, Anastasia Shimorina, Yvan Lussaud, Bedir Tekinerdogan, Geylani Kardas, Anne Monceaux
ASE2
2017 A Statistical, Grammar-Based Approach to Microplanning
abstract
Although there has been much work in recent years on data-driven natural language generation, little attention has been paid to the fine-grained interactions that arise during microplanning between aggregation, surface realization, and sentence segmentation. In this article, we propose a hybrid symbolic/statistical approach to jointly model the constraints regulating these interactions. Our approach integrates a small handwritten grammar, a statistical hypertagger, and a surface realization algorithm. It is applied to the verbalization of knowledge base queries and tested on 13 knowledge bases to demonstrate domain independence. We evaluate our approach in several ways. A quantitative analysis shows that the hybrid approach outperforms a purely symbolic approach in terms of both speed and coverage. Results from a human study indicate that users find the output of this hybrid statistic/symbolic system more fluent than both a template-based and a purely symbolic grammar-based approach. Finally, we illustrate by means of examples that our approach can account for various factors impacting aggregation, sentence segmentation, and surface realization.
Claire Gardent, Laura Perez-Beltrachini
Comput. Linguistics1
2016 Sequence-based Structured Prediction for Semantic Parsing
abstract
International audience
Chunyang Xiao, Marc Dymetman, Claire Gardent
ACL (1)3
2016 Building RDF Content for Data-to-Text Generation
abstract
In Natural Language Generation (NLG), one important limitation is the lack of common benchmarks on which to train, evaluate and compare data-to-text generators. In this paper, we make one step in that direction and introduce a method for automatically creating an arbitrary large repertoire of data units that could serve as input for generation. Using both automated metrics and a human evaluation, we show that the data units produced by our method are both diverse and coherent.
Laura Perez-Beltrachini, Rania Sayed, Claire Gardent
COLING3
2016 The WebNLG Challenge: Generating Text from DBPedia Data
Émilie Colin, Claire Gardent, Yassine Mrabet, Shashi Narayan, Laura Perez-Beltrachini
INLG2
2016 Category-Driven Content Selection
abstract
In this paper, we introduce a content selection method where the communicative goal is to describe entities of different categories (e.g., astronauts, universities or monuments).We argue that this method provides an interesting basis both for generating descriptions of entities and for semi-automatically constructing a benchmark on which to train, test and compare data-to-text generation systems.
Rania Mohammed, Laura Perez-Beltrachini, Claire Gardent
INLG3
2016 Unsupervised Sentence Simplification Using Deep Semantics
abstract
We present a novel approach to sentence simplification which departs from previous work in two main ways.First, it requires neither hand written rules nor a training corpus of aligned standard and simplified sentences.Second, sentence splitting operates on deep semantic structure.We show (i) that the unsupervised framework we propose is competitive with four state-of-the-art supervised systems and (ii) that our semantic based approach allows for a principled and effective handling of sentence splitting.
Shashi Narayan, Claire Gardent
INLG2
2015 Towards Knowledge-Driven Annotation
abstract
While the Web of data is attracting increasing interest and rapidly growing in size, the major support of information on the surface Web are still multimedia documents. Semantic annotation of texts is one of the main processes that are intended to facilitate meaning-based information exchange between computational agents. However, such annotation faces several challenges such as the heterogeneity of natural language expressions, the heterogeneity of documents structure and context dependencies. While a broad range of annotation approaches rely mainly or partly on the target textual context to disambiguate the extracted entities, in this paper we present an approach that relies mainly on formalized-knowledge expressed in RDF datasets to categorize and disambiguate noun phrases. In the proposed method, we represent the reference knowledge bases as co-occurrence matrices and the disambiguation problem as a 0-1 Integer Linear Programming (ILP) problem. The proposed approach is unsupervised and can be ported to any RDF knowledge base. The system implementing this approach, called KODA, shows very promising results w.r.t. state-of-the-art annotation tools in cross-domain experimentations.
Yassine Mrabet, Claire Gardent, Muriel Foulonneau, Elena Simperl, Eric Ras
AAAI2
2015 Multiple Adjunction in Feature-Based Tree-Adjoining Grammar
abstract
In parsing with Tree Adjoining Grammar (TAG), independent derivations have been shown by Schabes and Shieber (1994) to be essential for correctly supporting syntactic analysis, semantic interpretation, and statistical language modeling. However, the parsing algorithm they propose is not directly applicable to Feature-Based TAGs (FB-TAG). We provide a recognition algorithm for FB-TAG that supports both dependent and independent derivations. The resulting algorithm combines the benefits of independent derivations with those of Feature-Based grammars. In particular, we show that it accounts for a range of interactions between dependent vs. independent derivation on the one hand, and syntactic constraints, linear ordering, and scopal vs. nonscopal semantic dependencies on the other hand.
Claire Gardent, Shashi Narayan
Comput. Linguistics1
2015 Federating clustering and cluster labelling capabilities with a single approach based on feature maximization: French verb classes identification with IGNGF neural clustering
Jean-Charles Lamirel, Ingrid Falk, Claire Gardent
Neurocomputing3
2014 Surface Realisation from Knowledge-Bases
abstract
We present a simple, data-driven approach to generation from knowledge bases (KB).A key feature of this approach is that grammar induction is driven by the extended domain of locality principle of TAG (Tree Adjoining Grammar); and that it takes into account both syntactic and semantic information.The resulting extracted TAG includes a unification based semantics and can be used by an existing surface realiser to generate sentences from KB data.Experimental evaluation on the KBGen data shows that our model outperforms a data-driven generate-and-rank approach based on an automatically induced probabilistic grammar; and is comparable with a handcrafted symbolic approach.
Bikash Gyawali, Claire Gardent
ACL (1)2
2014 Hybrid Simplification using Deep Semantics and Machine Translation
abstract
We present a hybrid approach to sentence simplification which combines deep semantics and monolingual machine translation to derive simple sentences from complex ones. The approach differs from previous work in two main ways. First, it is semantic based in that it takes as input a deep semantic representation rather than e.g., a sentence or a parse tree. Second, it combines a simplification model for splitting and deletion with a monolingual translation model for phrase substitution and reordering. When compared against current state of the art methods, our model yields significantly simpler output that is both grammatical and meaning preserving.
Shashi Narayan, Claire Gardent
ACL (1)2
2014 Incremental Query Generation
abstract
We present a natural language genera-tion system which supports the incremen-tal specification of ontology-based queries in natural language. Our contribution is two fold. First, we introduce a chart based surface realisation algorithm which supports the kind of incremental process-ing required by ontology-based querying. Crucially, this algorithm avoids confusing the end user by preserving a consistent ordering of the query elements through-out the incremental query formulation pro-cess. Second, we show that grammar based surface realisation better supports the generation of fluent, natural sounding queries than previous template-based ap-proaches. 1
Laura Perez-Beltrachini, Claire Gardent, Enrico Franconi
EACL2
2013 Using Paraphrases and Lexical Semantics to Improve the Accuracy and the Robustness of Supervised Models in Situated Dialogue Systems
abstract
This paper explores to what extent lemmatisation, lexical resources, distributional semantics and paraphrases can increase the accuracy of supervised models for dialogue management.The results suggest that each of these factors can help improve performance but that the impact will vary depending on their combination and on the evaluation mode.
Claire Gardent, Lina Maria Rojas-Barahona
EMNLP1
2013 Weakly and Strongly Constrained Dialogues for Language Learning
Claire Gardent, Alejandra Lorenzo, Laura Perez-Beltrachini, Lina Maria Rojas-Barahona
SIGDIAL Conference1
2013 XMG: eXtensible MetaGrammar
abstract
In this article, we introduce eXtensible MetaGrammar (XMG), a framework for specifying tree-based grammars such as Feature-Based Lexicalized Tree-Adjoining Grammars (FB-LTAG) and Interaction Grammars (IG). We argue that XMG displays three features that facilitate both grammar writing and a fast prototyping of tree-based grammars. Firstly, XMG is fully declarative. For instance, it permits a declarative treatment of diathesis that markedly departs from the procedural lexical rules often used to specify tree-based grammars. Secondly, the XMG language has a high notational expressivity in that it supports multiple linguistic dimensions, inheritance, and a sophisticated treatment of identifiers. Thirdly, XMG is extensible in that its computational architecture facilitates the extension to other linguistic formalisms. We explain how this architecture naturally supports the design of three linguistic formalisms, namely, FB-LTAG, IG, and Multi-Component Tree-Adjoining Grammar (MC-TAG). We further show how it permits a straightforward integration of additional mechanisms such as linguistic and formal principles. To further illustrate the declarativity, notational expressivity, and extensibility of XMG, we describe the methodology used to specify an FB-LTAG for French augmented with a unification-based compositional semantics. This illustrates both how XMG facilitates the modeling of the tree fragment hierarchies required to specify tree-based grammars and of a syntax/semantics interface between semantic representations and syntactic trees. Finally, we briefly report on several grammars for French, English, and German that were implemented using XMG and compare XMG with other existing grammar specification frameworks for tree-based grammars.
Benoît Crabbé, Denys Duchier, Claire Gardent, Joseph Le Roux, Yannick Parmentier 0001
Comput. Linguistics3
2012 Classifying French Verbs Using French and English Lexical Resources
Ingrid Falk, Claire Gardent, Jean-Charles Lamirel
ACL (1)2
2012 Error Mining on Dependency Trees
Claire Gardent, Shashi Narayan
ACL (1)1
2012 Error Mining with Suspicion Trees: Seeing the Forest for the Trees
Shashi Narayan, Claire Gardent
COLING2
2012 Structure-Driven Lexicalist Generation
Shashi Narayan, Claire Gardent
COLING2
2012 KBGen - Text Generation from Knowledge Bases as a New Shared Task
Eva Banik, Claire Gardent, Donia Scott, Nikhil Dinesh, Fennie Liang
INLG2
2012 Generation for Grammar Engineering
Claire Gardent, Germán Kruszewski
INLG1
2012 Representation of linguistic and domain knowledge for second language learning in virtual worlds
Alexandre Denis 0002, Ingrid Falk, Claire Gardent, Laura Perez-Beltrachini
LREC3
2012 Building and Exploiting a Corpus of Dialog Interactions between French Speaking Virtual and Human Agents
Lina Maria Rojas-Barahona, Alejandra Lorenzo, Claire Gardent
LREC3
2012 An End-to-End Evaluation of Two Situated Dialog Systems
Lina Maria Rojas-Barahona, Alejandra Lorenzo, Claire Gardent
SIGDIAL Conference3
2011 Deep Semantics for Dependency Structures
Paul Bédaride, Claire Gardent
CICLing (1)2
2011 A Serious Game for Second Language Acquisition
Marilisa Amoia, Claire Gardent, Laura Perez-Beltrachini
CSEDU (1)2
2011 The JSafran Platform for Semi-Automatic Speech Processing
abstract
International audience
Christophe Cerisara, Claire Gardent
INTERSPEECH2
2011 Commas Recovery with Syntactic Features in French and in Czech
abstract
Automatic speech transcripts can be made more readable and useful for further processing by enriching them with punctuation marks and other meta-linguistic information. We study in this work how to improve automatic recovery of one of the most difficult punctuation marks, commas, in French and in Czech. We show that commas detection performances are largely improved in both languages by integrating into our baseline Conditional Random Field model syntactic features derived from dependency structures. We further study the relative impact of language-independent vs. specific features, and show that a combination of both of them gives the largest improvement. Robustness of these features to speech recognition errors is finally discussed.
Christophe Cerisara, Pavel Král, Claire Gardent
INTERSPEECH3
2011 Using regular tree grammars to enhance sentence realisation
abstract
Abstract Feature-based regular tree grammars (FRTG) can be used to generate the derivation trees of a feature-based tree adjoining grammar (FTAG). We make use of this fact to specify and implement both an FTAG-based sentence realiser and a benchmark generator for this realiser. We argue furthermore that the FRTG encoding enables us to improve on other proposals based on a grammar of TAG derivation trees in several ways. It preserves the compositional semantics that can be encoded in feature-based TAGs; it increases efficiency and restricts overgeneration; and it provides a uniform resource for generation, benchmark construction and parsing.
Claire Gardent, Benjamin Gottesman, Laura Perez-Beltrachini
Nat. Lang. Eng.1
2010 RTG based surface realisation for TAG
Claire Gardent, Laura Perez-Beltrachini
COLING1
2010 Memory-based active learning for French broadcast news
abstract
International audience
Frédéric Tantini, Christophe Cerisara, Claire Gardent
INTERSPEECH3
2010 Syntactic Testsuites and Textual Entailment Recognition
Paul Bédaride, Claire Gardent
LREC2
2010 Identifying Sources of Weakness in Syntactic Lexicon Extraction
Claire Gardent, Alejandra Lorenzo
LREC1
2008 Integrating a Unification-Based Semantics in a Large Scale Lexicalised Tree Adjoining Grammar for French
Claire Gardent
COLING1
2008 A Test Suite for Inference Involving Adjectives
Marilisa Amoia, Claire Gardent
LREC2
2007 A Symbolic Approach to Near-Deterministic Surface Realisation using Tree Adjoining Grammar
Claire Gardent, Eric Kow
ACL1
2007 SemTAG: a platform for specifying Tree Adjoining Grammars and performing TAG-based Semantic Construction
Claire Gardent, Yannick Parmentier 0001
ACL1
2006 Coreference Handling in XMG
Claire Gardent, Yannick Parmentier 0001
ACL1
2003 Semantic construction in F-TAG
Claire Gardent, Laura Kallmeyer
EACL1
2002 Generating Minimal Definite Descriptions
abstract
The incremental algorithm introduced in (Dale and Reiter, 1995) for producing distinguishing descriptions does not always generate a minimal description.In this paper, I show that when generalised to sets of individuals and disjunctive properties, this approach might generate unnecessarily long and ambiguous and/or epistemically redundant descriptions.I then present an alternative, constraint-based algorithm and show that it builds on existing related algorithms in that (i) it produces minimal descriptions for sets of individuals using positive, negative and disjunctive properties, (ii) it straightforwardly generalises to n-ary relations and (iii) it is integrated with surface realisation.
Claire Gardent
ACL1
2002 Improving Machine Learning Approaches to Coreference Resolution
abstract
We present a noun phrase coreference system that extends the work of Soon et al. (2001) and, to our knowledge, produces the best results to date on the MUC-6 and MUC-7 coreference resolution data sets -F-measures of 70.4 and 63.4,respectively.Improvements arise from two sources: extra-linguistic changes to the learning framework and a large-scale expansion of the feature set to include more sophisticated linguistic knowledge.
Vincent Ng 0001, Claire Gardent
ACL2
2001 Generating with a Grammar Based on Tree Descriptions: a Constraint-Based Approach
abstract
While the generative view of language processing builds bigger units out of smaller ones by means of rewriting steps, the axiomatic view eliminates invalid linguistic structures out of a set of possible structures by means of well formedness principles. We present a generator based on the axiomatic view and argue that when combined with a TAG-like grammar and a flat semantics, this axiomatic view permits avoiding drawbacks known to hold either of top-down or of bottom-up generators.
Claire Gardent, Stefan Thater
ACL1
1997 Computing Parallelism in Discourse
Claire Gardent, Michael Kohlhase
IJCAI (2)1
1996 Higher-Order Coloured Unification and Natural Language Semantics
abstract
In this paper, we show that Higher-Order Coloured Unification - a form of unification developed for automated theorem proving - provides a general theory for modeling the interface between the interpretation process and other sources of linguistic, non semantic information. In particular, it provides the general theory for the Primary Occurrence Restriction which (Dalrymple et al., 1991)'s analysis called for.
Claire Gardent, Michael Kohlhase
ACL1
1996 Focus and Higher-Order Unification
Claire Gardent, Michael Kohlhase
COLING1
1995 A Specification Language for Lexical Functional Grammars
Patrick Blackburn, Claire Gardent
EACL2
1993 Talking About Trees
Patrick Blackburn, Claire Gardent, Wilfried Meyer-Viol
EACL2
1993 A unification-based approach to multiple VP Ellipsis resolution
Claire Gardent
EACL1
1990 Generating from a Deep Structure
Claire Gardent, Agnès Plainfossé
COLING1
1990 The General Architecture of Generation in ACORD
Dieter Kohl, Agnès Plainfossé, Claire Gardent
COLING3
1989 Efficient Parsing for French
abstract
Parsing with categorial grammars often leads to problems such as proliferating lexical ambiguity, spurious parses and overgeneration. This paper presents a parser for French developed on an unification based categorial grammar (FG) which avoids these problems. This parser is a bottom-up chart parser augmented with a heuristic eliminating spurious parses. The unicity and completeness of parsing are proved.
Claire Gardent, Gabriel G. Bès, Pierre-François Jurie, Karine Baschung
ACL1
1989 French Order Without Order
Gabriel G. Bès, Claire Gardent
EACL2