EDBT 2026 Demo / reviewers in the wild / expert
Claire Gardent
dblp:71/6819
· DBLP profile ↗
85ranked-venue papers
24as first author
15since 2021 · last 2026
0000-0002-3805-6662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 81 · 24 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy-Preserving Generation of Synthetic Pathology Reports for Information Extraction
Alejandra Lorenzo, Adrien Coulet, Claire Gardent |
AIME (1) | 3 |
| 2025 | MuCAL: Contrastive Alignment for Preference-Driven KG-to-Text GenerationabstractWe propose MuCAL (Multilingual Contrastive Alignment Learning) to tackle the challenge of Knowledge Graphs (KG)-to-Text generation using preference learning, where reliable preference data is scarce.MuCAL is a multilingual KG/Text alignment model achieving robust cross-modal retrieval across multiple languages and difficulty levels.Building on Mu-CAL, we automatically create preference data by ranking candidate texts from three LLMs (Qwen2.5 , DeepSeek-v3, Llama-3).We then apply Direct Preference Optimisation (DPO) on these preference data, bypassing typical reward modelling steps to directly align generation outputs with graph semantics.Extensive experiments on KG-to-English Text generation show two main advantages: (1) Our KG/Text alignment model provides a better signal for DPO than similar existing metrics, and (2) significantly better generalisation on out-of-domain datasets compared to standard instruction tuning.Our results highlight MuCAL's effectiveness in supporting preference learning for KGto-English Text generation and lay the foundation for future multilingual extensions.Code and data are available at https://github. com Claire Gardent |
EMNLP | 2 |
| 2025 | Fine-Tuning, Prompting and RAG for Knowledge Graph-to-Russian Text Generation. How do these Methods generalise to Out-of-Distribution Data?abstractPrior work on Knowledge Graph-to-Text generation has mostly evaluated models on in-domain test sets and/or with English as the target language. In contrast, we focus on Russian and we assess how various generation methods perform on out-of-domain, unseen data. Previous studies have shown that enriching the input with target-language verbalisations of entities and properties substantially improves the performance of fine-tuned models for Russian. We compare multiple variants of two contemporary paradigms — LLM prompting and Retrieval-Augmented Generation (RAG) — and investigate alternative ways to integrate such external knowledge into the generation process. Using automatic metrics and human evaluation, we find that on unseen data the fine-tuned model consistently underperforms, revealing limited generalisation capacity; that while it outperforms RAG by a small margin on most datasets, prompting generates less fluent text; and conversely, that RAG generates text that is less faithful to the input. Overall, both LLM prompting and RAG outperform Fine-Tuning across all unseen testsets. The code for this paper is available at https://github.com/Javanochka/KG-to-text-fine-tuning-prompting-rag Anna Nikiforovskaya, William Soto Martinez, Evan Parker Kelly Chapple, Claire Gardent |
INLG | 4 |
| 2025 | Generating Complex Question Decompositions in the Face of Distribution ShiftsabstractKelvin Han, Claire Gardent. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kelvin Han, Claire Gardent |
NAACL (Long Papers) | 2 |
| 2024 | KGConv, a Conversational Corpus Grounded in WikidataabstractWe present KGConv, a large corpus of 71k English conversations where each question-answer pair is grounded in a Wikidata fact. Conversations contain on average 8.6 questions and for each Wikidata fact, we provide multiple variants (12 on average) of the corresponding question using templates, human annotations, hand-crafted rules and a question rewriting neural model. We provide baselines for the task of Knowledge-Based, Conversational Question Generation. KGConv can further be used for other generation and analysis tasks such as single-turn question generation from Wikidata triples, question rewriting, question answering from conversation or from knowledge graphs and quiz generation. Quentin Brabant, Lina Maria Rojas-Barahona, Gwénolé Lecorvé, Claire Gardent |
LREC/COLING | 4 |
| 2024 | Generating from AMRs into High and Low-Resource Languages using Phylogenetic Knowledge and Hierarchical QLoRA Training (HQL)abstractPrevious work on multilingual generation from Abstract Meaning Representations has mostly focused on High-and Medium-Resource languages relying on large amounts of training data.In this work, we consider both Highand Low-Resource languages capping training data size at the lower bound set by our Low-Resource languages i.e., 31K training instances.We propose two straightforward techniques to enhance generation results on Low-Resource while preserving performance on High-and Medium-Resource languages.First, we iteratively refine a multilingual model to a set of monolingual models using Low-Rank Adaptation -this enables cross-lingual transfer while reducing over-fitting for High-Resource languages as the monolingual models are trained last.Second, we base our training curriculum on a tree structure which permits investigating how the languages used at each iteration impact generation performance on High and Low-Resource languages.We show an improvement over both mono and multilingual approaches.Comparing different ways of grouping languages at each iteration step we find two beneficial configurations: grouping related languages which promotes transfer, or grouping distant languages which facilitates regularisation. William Soto Martinez, Yannick Parmentier 0001, Claire Gardent |
INLG | 3 |
| 2024 | Evaluating RDF-to-text Generation Models for English and Russian on Out Of Domain DataabstractWhile the WebNLG dataset has prompted much research on generation from knowledge graphs, little work has examined how well models trained on the WebNLG data generalise to unseen data and work has mostly been focused on English.In this paper, we introduce novel benchmarks for both English and Russian which contain various ratios of unseen entities and properties.These benchmarks also differ from WebNLG in that some of the graphs stem from Wikidata rather than DBpedia.Evaluating various models for English and Russian on these benchmarks shows a strong decrease in performance while a qualitative analysis highlights the various types of errors induced by non i.i.d data. Anna Nikiforovskaya, Claire Gardent |
INLG | 2 |
| 2023 | Document-Level Planning for Text SimplificationabstractMost existing work on text simplification is limited to sentence-level inputs, with attempts to iteratively apply these approaches to document-level simplification failing to coherently preserve the discourse structure of the document.We hypothesise that by providing a high-level view of the target document, a simplification plan might help to guide generation.Building upon previous work on controlled, sentence-level simplification, we view a plan as a sequence of labels, each describing one of four sentence-level simplification operations (copy, rephrase, split, or delete).We propose a planning model that labels each sentence in the input document while considering both its context (a window of surrounding sentences) and its internal structure (a token-level representation).Experiments on two simplification benchmarks (Newsela-auto and Wikiauto) show that our model outperforms strong baselines both on the planning task and when used to guide document-level simplification models. Liam Cripwell, Joël Legrand, Claire Gardent |
EACL | 3 |
| 2023 | Simplicity Level Estimate (SLE): A Learned Reference-Less Metric for Sentence SimplificationabstractAutomatic evaluation for sentence simplification remains a challenging problem.Most popular evaluation metrics require multiple high-quality references -something not readily available for simplification -which makes it difficult to test performance on unseen domains.Furthermore, most existing metrics conflate simplicity with correlated attributes such as fluency or meaning preservation.We propose a new learned evaluation metric (SLE) which focuses on simplicity, outperforming almost all existing metrics in terms of correlation with human judgements. Liam Cripwell, Joël Legrand, Claire Gardent |
EMNLP | 3 |
| 2023 | Generating and Answering Simple and Complex Questions from Text and from Knowledge GraphsabstractKelvin Han, Claire Gardent. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kelvin Han, Claire Gardent |
IJCNLP (1) | 2 |
| 2023 | Phylogeny-Inspired Soft Prompts For Data-to-Text Generation in Low-Resource LanguagesabstractWilliam Soto Martinez, Yannick Parmentier, Claire Gardent. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. William Soto Martinez, Yannick Parmentier 0001, Claire Gardent |
IJCNLP (1) | 3 |
| 2022 | Generating Biographies on Wikipedia: The Impact of Gender Bias on the Retrieval-Based Generation of Women BiographiesabstractGenerating factual, long-form text such as Wikipedia articles raises three key challenges: how to gather relevant evidence, how to structure information into well-formed text, and how to ensure that the generated text is factually correct.We address these by developing a model for English text that uses a retrieval mechanism to identify relevant supporting information on the web and a cache-based pre-trained encoderdecoder to generate long-form biographies section by section, including citation information.To assess the impact of available web evidence on the output text, we compare the performance of our approach when generating biographies about women (for which less information is available on the web) vs. biographies generally.To this end, we curate a dataset of 1,500 biographies about women.We analyze our generated text to understand how differences in available web evidence data affect generation.We evaluate the factuality, fluency, and quality of the generated texts using automatic metrics and human evaluation.We hope that these techniques can be used as a starting point for human writers, to aid in reducing the complexity inherent in the creation of long-form, factual text. Angela Fan, Claire Gardent |
ACL (1) | 2 |
| 2022 | Generating Questions from Wikidata TriplesabstractQuestion generation from knowledge bases (or knowledge base question generation, KBQG) is the task of generating questions from structured database information, typically in the form of triples representing facts. To handle rare entities and generalize to unseen properties, previous work on KBQG resorted to extensive, often ad-hoc pre- and post-processing of the input triple. We revisit KBQG – using pre training, a new (triple, question) dataset and taking question type into account – and show that our approach outperforms previous work both in a standard and in a zero-shot setting. We also show that the extended KBQG dataset (also helpful for knowledge base question answering) we provide allows not only for better coverage in terms of knowledge base (KB) properties but also for increased output variability in that it permits the generation of multiple questions from the same KB triple. Kelvin Han, Thiago Castro Ferreira, Claire Gardent |
LREC | 3 |
| 2021 | Augmenting Transformers with KNN-Based Composite Memory for DialogabstractVarious machine learning tasks can benefit from access to external information of different modalities, such as text and images. Recent work has focused on learning architectures with large memories capable of storing this knowledge. We propose augmenting generative Transformer neural networks with KNN-based Information Fetching (KIF) modules. Each KIF module learns a read operation to access fixed external knowledge. We apply these modules to generative dialog modeling, a challenging task where information must be flexibly retrieved and incorporated to maintain the topic and flow of conversation. We demonstrate the effectiveness of our approach by identifying relevant knowledge required for knowledgeable but engaging dialog from Wikipedia, images, and human-written dialog utterances, and show that leveraging this retrieved information improves model performance, measured by automatic and human evaluation. Angela Fan, Claire Gardent, Chloé Braud, Antoine Bordes |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | An Error Analysis Framework for Shallow Surface RealisationabstractAbstract The metrics standardly used to evaluate Natural Language Generation (NLG) models, such as BLEU or METEOR, fail to provide information on which linguistic factors impact performance. Focusing on Surface Realization (SR), the task of converting an unordered dependency tree into a well-formed sentence, we propose a framework for error analysis which permits identifying which features of the input affect the models’ results. This framework consists of two main components: (i) correlation analyses between a wide range of syntactic metrics and standard performance metrics and (ii) a set of techniques to automatically identify syntactic constructs that often co-occur with low performance scores. We demonstrate the advantages of our framework by performing error analysis on the results of 174 system runs submitted to the Multilingual SR shared tasks; we show that dependency edge accuracy correlate with automatic metrics thereby providing a more interpretable basis for evaluation; and we suggest ways in which our framework could be used to improve models and data. The framework is available in the form of a toolkit which can be used both by campaign organizers to provide detailed, linguistically interpretable feedback on the state of the art in multilingual SR, and by individual researchers to improve models and datasets.1 Anastasia Shimorina, Claire Gardent, Yannick Parmentier 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Learning Health-Bots from Training Data that was Automatically Created using Paraphrase Detection and Expert KnowledgeabstractA key bottleneck for developing dialog models is the lack of adequate training data.Due to privacy issues, dialog data is even scarcer in the health domain.We propose a novel method for creating dialog corpora which we apply to create doctor-patient interaction data.We use this data to learn both a generation and a hybrid classification/retrieval model and find that the generation model consistently outperforms the hybrid model.We show that our data creation method has several advantages.Not only does it allow for the semi-automatic creation of large quantities of training data.It also provides a natural way of guiding learning and a novel method for assessing the quality of human-machine interactions. Anna Liednikova, Philippe Jolivet, Alexandre Durand-Salmon, Claire Gardent |
COLING | 4 |
| 2020 | Multilingual AMR-to-Text GenerationabstractGenerating text from structured data is challenging because it requires bridging the gap between (i) structure and natural language (NL) and (ii) semantically underspecified input and fully specified NL output.Multilingual generation brings in an additional challenge: that of generating into languages with varied word order and morphological properties.In this work, we focus on Abstract Meaning Representations (AMRs) as structured input, where previous research has overwhelmingly focused on generating only into English.We leverage advances in cross-lingual embeddings, pretraining, and multilingual models to create multilingual AMR-to-text models that generate in twenty one different languages.For eighteen languages, based on automatic metrics, our multilingual models surpass baselines that generate into a single language.We analyse the ability of our multilingual models to accurately capture morphology and word order using human evaluation, and find that native speakers judge our generations to be fluent. Angela Fan, Claire Gardent |
EMNLP (1) | 2 |
| 2020 | Modeling Global and Local Node Contexts for Text Generation from Knowledge GraphsabstractRecent graph-to-text models generate text from graph-based data using either global or local aggregation to learn node representations. Global node encoding allows explicit communication between two distant nodes, thereby neglecting graph topology as all nodes are directly connected. In contrast, local node encoding considers the relations between neighbor nodes capturing the graph structure, but it can fail to capture long-range relations. In this work, we gather both encoding strategies, proposing novel neural models that encode an input graph combining both global and local node contexts, in order to learn better contextualized node embeddings. In our experiments, we demonstrate that our approaches lead to significant improvements on two graph-to-text datasets achieving BLEU scores of 18.01 on the AGENDA dataset, and 63.69 on the WebNLG dataset for seen categories, outperforming state-of-the-art models by 3.7 and 3.1 points, respectively. 1 Leonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Using Local Knowledge Graph Construction to Scale Seq2Seq Models to Multi-Document InputsabstractAngela Fan, Claire Gardent, Chloé Braud, Antoine Bordes. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Angela Fan, Claire Gardent, Chloé Braud, Antoine Bordes |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Enhancing AMR-to-Text Generation with Dual Graph RepresentationsabstractLeonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Leonardo F. R. Ribeiro, Claire Gardent, Iryna Gurevych |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Surface Realisation Using Full DelexicalisationabstractAnastasia Shimorina, Claire Gardent. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Anastasia Shimorina, Claire Gardent |
EMNLP/IJCNLP (1) | 2 |
| 2019 | RL Extraction of Syntax-Based Chunks for Sentence Compression
Hoa T. Le, Christophe Cerisara, Claire Gardent |
ICANN (4) | 3 |
| 2019 | Generating Text from Anonymised StructuresabstractSurface realisation maps a meaning representation (MR) to a text, usually a single sentence.In this paper, we introduce a new parallel dataset of deep meaning representations and French sentences and we present a novel method for MR-to-text generation which seeks to generalise by abstracting away from lexical content.Most current work on natural language generation focuses on generating text that matches a reference using BLEU as evaluation criteria.In this paper, we additionally consider the model's ability to reintroduce the function words that are absent from the deep input meaning representations.We show that our approach increases both BLEU score and the scores used to assess function words generation. Émilie Colin, Claire Gardent |
INLG | 2 |
| 2019 | Revisiting the Binary Linearization Technique for Surface RealizationabstractEnd-to-end neural approaches have achieved state-of-the-art performance in many natural language processing (NLP) tasks.Yet, they often lack transparency of the underlying decision-making process, hindering error analysis and certain model improvements.In this work, we revisit the binary linearization approach to surface realization, which exhibits more interpretable behavior, but was falling short in terms of prediction accuracy.We show how enriching the training data to better capture word order constraints almost doubles the performance of the system.We further demonstrate that encoding both local and global prediction contexts yields another considerable performance boost.With the proposed modifications, the system which ranked low in the latest shared task on multilingual surface realization now achieves best results in five out of ten languages, while being on par with the state-of-the-art approaches in others. 1 Yevgeniy Puzikov, Claire Gardent, Ido Dagan, Iryna Gurevych |
INLG | 2 |
| 2018 | Generating Syntactic ParaphrasesabstractWe study the automatic generation of syntactic paraphrases using four different models for generation: data-to-text generation, textto-text generation, text reduction and text expansion, We derive training data for each of these tasks from the WebNLG dataset and we show (i) that conditioning generation on syntactic constraints effectively permits the generation of syntactically distinct paraphrases for the same input and (ii) that exploiting different types of input (data, text or data+text) further increases the number of distinct paraphrases that can be generated for a given input. Émilie Colin, Claire Gardent |
EMNLP | 2 |
| 2018 | Handling Rare Items in Data-to-Text GenerationabstractNeural approaches to data-to-text generation generally handle rare input items using either delexicalisation or a copy mechanism.We investigate the relative impact of these two methods on two datasets (E2E and WebNLG) and using two evaluation settings.We show (i) that rare items strongly impact performance; (ii) that combining delexicalisation and copying yields the strongest improvement; (iii) that copying underperforms for rare and unseen items and (iv) that the impact of these two mechanisms greatly varies depending on how the dataset is constructed and on how it is split into train, dev and test 1 . Anastasia Shimorina, Claire Gardent |
INLG | 2 |
| 2017 | Creating Training Corpora for NLG Micro-PlannersabstractIn this paper, we present a novel framework for semi-automatically creating linguistically challenging microplanning data-to-text corpora from existing Knowledge Bases.Because our method pairs data of varying size and shape with texts ranging from simple clauses to short texts, a dataset created using this framework provides a challenging benchmark for microplanning.Another feature of this framework is that it can be applied to any large scale knowledge base and can therefore be used to train and learn KB verbalisers.We apply our framework to DBpedia data and compare the resulting dataset with Wen et al. (2016)'s.We show that while Wen et al.'s dataset is more than twice larger than ours, it is less diverse both in terms of input and in terms of text.We thus propose our corpus generation framework as a novel method for creating challenging data sets from which NLG models can be learned which are capable of handling the complex interactions occurring during in micro-planning between lexicalisation, aggregation, surface realisation, referring expression generation and sentence segmentation.To encourage researchers to take up this challenge, we recently made available a dataset created using this framework in the context of the WEBNLG shared task. Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini |
ACL (1) | 1 |
| 2017 | Split and RephraseabstractWe propose a new sentence simplification task (Split-and-Rephrase) where the aim is to split a complex sentence into a meaning preserving sequence of shorter sentences.Like sentence simplification, splitting-and-rephrasing has the potential of benefiting both natural language processing and societal applications.Because shorter sentences are generally better processed by NLP systems, it could be used as a preprocessing step which facilitates and improves the performance of parsers, semantic role labelers and machine translation systems.It should also be of use for people with reading disabilities because it allows the conversion of longer sentences into shorter ones.This paper makes two contributions towards this new task.First, we create and make available a benchmark consisting of 1,066,115 tuples mapping a single complex sentence to a sequence of sentences expressing the same meaning.1 Second, we propose five models (vanilla sequence-to-sequence to semantically-motivated models) to understand the difficulty of the proposed task. Shashi Narayan, Claire Gardent, Shay B. Cohen, Anastasia Shimorina |
EMNLP | 2 |
| 2017 | Mapping Natural Language to Description Logic
Bikash Gyawali, Anastasia Shimorina, Claire Gardent, Samuel Cruz-Lara, Mariem Mahfoudh |
ESWC (1) | 3 |
| 2017 | Symbolic Priors for RNN-based Semantic ParsingabstractSeq2seq models based on Recurrent Neural Networks (RNNs) have recently received a lot of attention in the domain of Semantic Parsing. While in principle they can be trained directly on pairs (natural language utterances, logical forms), their performance is limited by the amount of available data. To alleviate this problem, we propose to exploit various sources of prior knowledge: the well-formedness of the logical forms is modeled by a weighted context-free grammar; the likelihood that certain entities present in the input utterance are also present in the logical form is modeled by weighted finite-state automata. The grammar and automata are combined together through an efficient intersection algorithm to form a soft guide (“background”) to the RNN.We test our method on an extension of the Overnight dataset and show that it not only strongly improves over an RNN baseline, but also outperforms non-RNN models based on rich sets of hand-crafted features. Chunyang Xiao, Marc Dymetman, Claire Gardent |
IJCAI | 3 |
| 2017 | The WebNLG Challenge: Generating Text from RDF DataabstractThe WebNLG challenge consists in mapping sets of RDF triples to text.It provides a common benchmark on which to train, evaluate and compare "microplanners", i.e. generation systems that verbalise a given content by making a range of complex interacting choices including referring expression generation, aggregation, lexicalisation, surface realisation and sentence segmentation.In this paper, we introduce the microplanning task, describe data preparation, introduce our evaluation methodology, analyse participant results and provide a brief description of the participating systems.(3) a. LOCATION-COUNTRY-STARTDATE ⇒ Passive-Apposition-Active Claire Gardent, Anastasia Shimorina, Shashi Narayan, Laura Perez-Beltrachini |
INLG | 1 |
| 2017 | Analysing Data-To-Text Generation BenchmarksabstractA generation system can only be as good as the data it is trained on.In this short paper, we propose a methodology for analysing data-to-text corpora used for training microplanner i.e., systems which given some input must produce a text verbalising exactly this input.We apply this methodology to three existing benchmarks and we elicite a set of criteria for the creation of a data-to-text benchmark which could help better support the development, evaluation and comparison of linguistically sophisticated data-to-text generators. Laura Perez-Beltrachini, Claire Gardent |
INLG | 2 |
| 2017 | ModelWriter: text and model-synchronized document engineering platformabstractThe ModelWriter platform provides a generic framework for automated traceability analysis. In this paper, we demonstrate how this framework can be used to trace the consistency and completeness of technical documents that consist of a set of System Installation Design Principles used by Airbus to ensure the correctness of aircraft system installation. We show in particular, how the platform allows the integration of two types of reasoning: reasoning about the meaning of text using semantic parsing and description logic theorem proving; and reasoning about document structure using first-order relational logic and finite model finding for traceability analysis. Ferhat Erata, Claire Gardent, Bikash Gyawali, Anastasia Shimorina, Yvan Lussaud, Bedir Tekinerdogan, Geylani Kardas, Anne Monceaux |
ASE | 2 |
| 2017 | A Statistical, Grammar-Based Approach to MicroplanningabstractAlthough there has been much work in recent years on data-driven natural language generation, little attention has been paid to the fine-grained interactions that arise during microplanning between aggregation, surface realization, and sentence segmentation. In this article, we propose a hybrid symbolic/statistical approach to jointly model the constraints regulating these interactions. Our approach integrates a small handwritten grammar, a statistical hypertagger, and a surface realization algorithm. It is applied to the verbalization of knowledge base queries and tested on 13 knowledge bases to demonstrate domain independence. We evaluate our approach in several ways. A quantitative analysis shows that the hybrid approach outperforms a purely symbolic approach in terms of both speed and coverage. Results from a human study indicate that users find the output of this hybrid statistic/symbolic system more fluent than both a template-based and a purely symbolic grammar-based approach. Finally, we illustrate by means of examples that our approach can account for various factors impacting aggregation, sentence segmentation, and surface realization. Claire Gardent, Laura Perez-Beltrachini |
Comput. Linguistics | 1 |
| 2016 | Sequence-based Structured Prediction for Semantic ParsingabstractInternational audience Chunyang Xiao, Marc Dymetman, Claire Gardent |
ACL (1) | 3 |
| 2016 | Building RDF Content for Data-to-Text GenerationabstractIn Natural Language Generation (NLG), one important limitation is the lack of common benchmarks on which to train, evaluate and compare data-to-text generators. In this paper, we make one step in that direction and introduce a method for automatically creating an arbitrary large repertoire of data units that could serve as input for generation. Using both automated metrics and a human evaluation, we show that the data units produced by our method are both diverse and coherent. Laura Perez-Beltrachini, Rania Sayed, Claire Gardent |
COLING | 3 |
| 2016 | The WebNLG Challenge: Generating Text from DBPedia Data
Émilie Colin, Claire Gardent, Yassine Mrabet, Shashi Narayan, Laura Perez-Beltrachini |
INLG | 2 |
| 2016 | Category-Driven Content SelectionabstractIn this paper, we introduce a content selection method where the communicative goal is to describe entities of different categories (e.g., astronauts, universities or monuments).We argue that this method provides an interesting basis both for generating descriptions of entities and for semi-automatically constructing a benchmark on which to train, test and compare data-to-text generation systems. Rania Mohammed, Laura Perez-Beltrachini, Claire Gardent |
INLG | 3 |
| 2016 | Unsupervised Sentence Simplification Using Deep SemanticsabstractWe present a novel approach to sentence simplification which departs from previous work in two main ways.First, it requires neither hand written rules nor a training corpus of aligned standard and simplified sentences.Second, sentence splitting operates on deep semantic structure.We show (i) that the unsupervised framework we propose is competitive with four state-of-the-art supervised systems and (ii) that our semantic based approach allows for a principled and effective handling of sentence splitting. Shashi Narayan, Claire Gardent |
INLG | 2 |
| 2015 | Towards Knowledge-Driven AnnotationabstractWhile the Web of data is attracting increasing interest and rapidly growing in size, the major support of information on the surface Web are still multimedia documents. Semantic annotation of texts is one of the main processes that are intended to facilitate meaning-based information exchange between computational agents. However, such annotation faces several challenges such as the heterogeneity of natural language expressions, the heterogeneity of documents structure and context dependencies. While a broad range of annotation approaches rely mainly or partly on the target textual context to disambiguate the extracted entities, in this paper we present an approach that relies mainly on formalized-knowledge expressed in RDF datasets to categorize and disambiguate noun phrases. In the proposed method, we represent the reference knowledge bases as co-occurrence matrices and the disambiguation problem as a 0-1 Integer Linear Programming (ILP) problem. The proposed approach is unsupervised and can be ported to any RDF knowledge base. The system implementing this approach, called KODA, shows very promising results w.r.t. state-of-the-art annotation tools in cross-domain experimentations. Yassine Mrabet, Claire Gardent, Muriel Foulonneau, Elena Simperl, Eric Ras |
AAAI | 2 |
| 2015 | Multiple Adjunction in Feature-Based Tree-Adjoining GrammarabstractIn parsing with Tree Adjoining Grammar (TAG), independent derivations have been shown by Schabes and Shieber (1994) to be essential for correctly supporting syntactic analysis, semantic interpretation, and statistical language modeling. However, the parsing algorithm they propose is not directly applicable to Feature-Based TAGs (FB-TAG). We provide a recognition algorithm for FB-TAG that supports both dependent and independent derivations. The resulting algorithm combines the benefits of independent derivations with those of Feature-Based grammars. In particular, we show that it accounts for a range of interactions between dependent vs. independent derivation on the one hand, and syntactic constraints, linear ordering, and scopal vs. nonscopal semantic dependencies on the other hand. Claire Gardent, Shashi Narayan |
Comput. Linguistics | 1 |
| 2015 | Federating clustering and cluster labelling capabilities with a single approach based on feature maximization: French verb classes identification with IGNGF neural clustering
Jean-Charles Lamirel, Ingrid Falk, Claire Gardent |
Neurocomputing | 3 |
| 2014 | Surface Realisation from Knowledge-BasesabstractWe present a simple, data-driven approach to generation from knowledge bases (KB).A key feature of this approach is that grammar induction is driven by the extended domain of locality principle of TAG (Tree Adjoining Grammar); and that it takes into account both syntactic and semantic information.The resulting extracted TAG includes a unification based semantics and can be used by an existing surface realiser to generate sentences from KB data.Experimental evaluation on the KBGen data shows that our model outperforms a data-driven generate-and-rank approach based on an automatically induced probabilistic grammar; and is comparable with a handcrafted symbolic approach. Bikash Gyawali, Claire Gardent |
ACL (1) | 2 |
| 2014 | Hybrid Simplification using Deep Semantics and Machine TranslationabstractWe present a hybrid approach to sentence simplification which combines deep semantics and monolingual machine translation to derive simple sentences from complex ones. The approach differs from previous work in two main ways. First, it is semantic based in that it takes as input a deep semantic representation rather than e.g., a sentence or a parse tree. Second, it combines a simplification model for splitting and deletion with a monolingual translation model for phrase substitution and reordering. When compared against current state of the art methods, our model yields significantly simpler output that is both grammatical and meaning preserving. Shashi Narayan, Claire Gardent |
ACL (1) | 2 |
| 2014 | Incremental Query GenerationabstractWe present a natural language genera-tion system which supports the incremen-tal specification of ontology-based queries in natural language. Our contribution is two fold. First, we introduce a chart based surface realisation algorithm which supports the kind of incremental process-ing required by ontology-based querying. Crucially, this algorithm avoids confusing the end user by preserving a consistent ordering of the query elements through-out the incremental query formulation pro-cess. Second, we show that grammar based surface realisation better supports the generation of fluent, natural sounding queries than previous template-based ap-proaches. 1 Laura Perez-Beltrachini, Claire Gardent, Enrico Franconi |
EACL | 2 |
| 2013 | Using Paraphrases and Lexical Semantics to Improve the Accuracy and the Robustness of Supervised Models in Situated Dialogue SystemsabstractThis paper explores to what extent lemmatisation, lexical resources, distributional semantics and paraphrases can increase the accuracy of supervised models for dialogue management.The results suggest that each of these factors can help improve performance but that the impact will vary depending on their combination and on the evaluation mode. Claire Gardent, Lina Maria Rojas-Barahona |
EMNLP | 1 |
| 2013 | Weakly and Strongly Constrained Dialogues for Language Learning
Claire Gardent, Alejandra Lorenzo, Laura Perez-Beltrachini, Lina Maria Rojas-Barahona |
SIGDIAL Conference | 1 |
| 2013 | XMG: eXtensible MetaGrammarabstractIn this article, we introduce eXtensible MetaGrammar (XMG), a framework for specifying tree-based grammars such as Feature-Based Lexicalized Tree-Adjoining Grammars (FB-LTAG) and Interaction Grammars (IG). We argue that XMG displays three features that facilitate both grammar writing and a fast prototyping of tree-based grammars. Firstly, XMG is fully declarative. For instance, it permits a declarative treatment of diathesis that markedly departs from the procedural lexical rules often used to specify tree-based grammars. Secondly, the XMG language has a high notational expressivity in that it supports multiple linguistic dimensions, inheritance, and a sophisticated treatment of identifiers. Thirdly, XMG is extensible in that its computational architecture facilitates the extension to other linguistic formalisms. We explain how this architecture naturally supports the design of three linguistic formalisms, namely, FB-LTAG, IG, and Multi-Component Tree-Adjoining Grammar (MC-TAG). We further show how it permits a straightforward integration of additional mechanisms such as linguistic and formal principles. To further illustrate the declarativity, notational expressivity, and extensibility of XMG, we describe the methodology used to specify an FB-LTAG for French augmented with a unification-based compositional semantics. This illustrates both how XMG facilitates the modeling of the tree fragment hierarchies required to specify tree-based grammars and of a syntax/semantics interface between semantic representations and syntactic trees. Finally, we briefly report on several grammars for French, English, and German that were implemented using XMG and compare XMG with other existing grammar specification frameworks for tree-based grammars. Benoît Crabbé, Denys Duchier, Claire Gardent, Joseph Le Roux, Yannick Parmentier 0001 |
Comput. Linguistics | 3 |
| 2012 | Classifying French Verbs Using French and English Lexical Resources
Ingrid Falk, Claire Gardent, Jean-Charles Lamirel |
ACL (1) | 2 |
| 2012 | Error Mining on Dependency Trees
Claire Gardent, Shashi Narayan |
ACL (1) | 1 |
| 2012 | Error Mining with Suspicion Trees: Seeing the Forest for the Trees
Shashi Narayan, Claire Gardent |
COLING | 2 |
| 2012 | Structure-Driven Lexicalist Generation
Shashi Narayan, Claire Gardent |
COLING | 2 |
| 2012 | KBGen - Text Generation from Knowledge Bases as a New Shared Task
Eva Banik, Claire Gardent, Donia Scott, Nikhil Dinesh, Fennie Liang |
INLG | 2 |
| 2012 | Generation for Grammar Engineering
Claire Gardent, Germán Kruszewski |
INLG | 1 |
| 2012 | Representation of linguistic and domain knowledge for second language learning in virtual worlds
Alexandre Denis 0002, Ingrid Falk, Claire Gardent, Laura Perez-Beltrachini |
LREC | 3 |
| 2012 | Building and Exploiting a Corpus of Dialog Interactions between French Speaking Virtual and Human Agents
Lina Maria Rojas-Barahona, Alejandra Lorenzo, Claire Gardent |
LREC | 3 |
| 2012 | An End-to-End Evaluation of Two Situated Dialog Systems
Lina Maria Rojas-Barahona, Alejandra Lorenzo, Claire Gardent |
SIGDIAL Conference | 3 |
| 2011 | Deep Semantics for Dependency Structures
Paul Bédaride, Claire Gardent |
CICLing (1) | 2 |
| 2011 | A Serious Game for Second Language Acquisition
Marilisa Amoia, Claire Gardent, Laura Perez-Beltrachini |
CSEDU (1) | 2 |
| 2011 | The JSafran Platform for Semi-Automatic Speech ProcessingabstractInternational audience Christophe Cerisara, Claire Gardent |
INTERSPEECH | 2 |
| 2011 | Commas Recovery with Syntactic Features in French and in CzechabstractAutomatic speech transcripts can be made more readable and useful for further processing by enriching them with punctuation marks and other meta-linguistic information. We study in this work how to improve automatic recovery of one of the most difficult punctuation marks, commas, in French and in Czech. We show that commas detection performances are largely improved in both languages by integrating into our baseline Conditional Random Field model syntactic features derived from dependency structures. We further study the relative impact of language-independent vs. specific features, and show that a combination of both of them gives the largest improvement. Robustness of these features to speech recognition errors is finally discussed. Christophe Cerisara, Pavel Král, Claire Gardent |
INTERSPEECH | 3 |
| 2011 | Using regular tree grammars to enhance sentence realisationabstractAbstract Feature-based regular tree grammars (FRTG) can be used to generate the derivation trees of a feature-based tree adjoining grammar (FTAG). We make use of this fact to specify and implement both an FTAG-based sentence realiser and a benchmark generator for this realiser. We argue furthermore that the FRTG encoding enables us to improve on other proposals based on a grammar of TAG derivation trees in several ways. It preserves the compositional semantics that can be encoded in feature-based TAGs; it increases efficiency and restricts overgeneration; and it provides a uniform resource for generation, benchmark construction and parsing. Claire Gardent, Benjamin Gottesman, Laura Perez-Beltrachini |
Nat. Lang. Eng. | 1 |
| 2010 | RTG based surface realisation for TAG
Claire Gardent, Laura Perez-Beltrachini |
COLING | 1 |
| 2010 | Memory-based active learning for French broadcast newsabstractInternational audience Frédéric Tantini, Christophe Cerisara, Claire Gardent |
INTERSPEECH | 3 |
| 2010 | Syntactic Testsuites and Textual Entailment Recognition
Paul Bédaride, Claire Gardent |
LREC | 2 |
| 2010 | Identifying Sources of Weakness in Syntactic Lexicon Extraction
Claire Gardent, Alejandra Lorenzo |
LREC | 1 |
| 2008 | Integrating a Unification-Based Semantics in a Large Scale Lexicalised Tree Adjoining Grammar for French
Claire Gardent |
COLING | 1 |
| 2008 | A Test Suite for Inference Involving Adjectives
Marilisa Amoia, Claire Gardent |
LREC | 2 |
| 2007 | A Symbolic Approach to Near-Deterministic Surface Realisation using Tree Adjoining Grammar
Claire Gardent, Eric Kow |
ACL | 1 |
| 2007 | SemTAG: a platform for specifying Tree Adjoining Grammars and performing TAG-based Semantic Construction
Claire Gardent, Yannick Parmentier 0001 |
ACL | 1 |
| 2006 | Coreference Handling in XMG
Claire Gardent, Yannick Parmentier 0001 |
ACL | 1 |
| 2003 | Semantic construction in F-TAG
Claire Gardent, Laura Kallmeyer |
EACL | 1 |
| 2002 | Generating Minimal Definite DescriptionsabstractThe incremental algorithm introduced in (Dale and Reiter, 1995) for producing distinguishing descriptions does not always generate a minimal description.In this paper, I show that when generalised to sets of individuals and disjunctive properties, this approach might generate unnecessarily long and ambiguous and/or epistemically redundant descriptions.I then present an alternative, constraint-based algorithm and show that it builds on existing related algorithms in that (i) it produces minimal descriptions for sets of individuals using positive, negative and disjunctive properties, (ii) it straightforwardly generalises to n-ary relations and (iii) it is integrated with surface realisation. Claire Gardent |
ACL | 1 |
| 2002 | Improving Machine Learning Approaches to Coreference ResolutionabstractWe present a noun phrase coreference system that extends the work of Soon et al. (2001) and, to our knowledge, produces the best results to date on the MUC-6 and MUC-7 coreference resolution data sets -F-measures of 70.4 and 63.4,respectively.Improvements arise from two sources: extra-linguistic changes to the learning framework and a large-scale expansion of the feature set to include more sophisticated linguistic knowledge. Vincent Ng 0001, Claire Gardent |
ACL | 2 |
| 2001 | Generating with a Grammar Based on Tree Descriptions: a Constraint-Based ApproachabstractWhile the generative view of language processing builds bigger units out of smaller ones by means of rewriting steps, the axiomatic view eliminates invalid linguistic structures out of a set of possible structures by means of well formedness principles. We present a generator based on the axiomatic view and argue that when combined with a TAG-like grammar and a flat semantics, this axiomatic view permits avoiding drawbacks known to hold either of top-down or of bottom-up generators. Claire Gardent, Stefan Thater |
ACL | 1 |
| 1997 | Computing Parallelism in Discourse
Claire Gardent, Michael Kohlhase |
IJCAI (2) | 1 |
| 1996 | Higher-Order Coloured Unification and Natural Language SemanticsabstractIn this paper, we show that Higher-Order Coloured Unification - a form of unification developed for automated theorem proving - provides a general theory for modeling the interface between the interpretation process and other sources of linguistic, non semantic information. In particular, it provides the general theory for the Primary Occurrence Restriction which (Dalrymple et al., 1991)'s analysis called for. Claire Gardent, Michael Kohlhase |
ACL | 1 |
| 1996 | Focus and Higher-Order Unification
Claire Gardent, Michael Kohlhase |
COLING | 1 |
| 1995 | A Specification Language for Lexical Functional Grammars
Patrick Blackburn, Claire Gardent |
EACL | 2 |
| 1993 | Talking About Trees
Patrick Blackburn, Claire Gardent, Wilfried Meyer-Viol |
EACL | 2 |
| 1993 | A unification-based approach to multiple VP Ellipsis resolution
Claire Gardent |
EACL | 1 |
| 1990 | Generating from a Deep Structure
Claire Gardent, Agnès Plainfossé |
COLING | 1 |
| 1990 | The General Architecture of Generation in ACORD
Dieter Kohl, Agnès Plainfossé, Claire Gardent |
COLING | 3 |
| 1989 | Efficient Parsing for FrenchabstractParsing with categorial grammars often leads to problems such as proliferating lexical ambiguity, spurious parses and overgeneration. This paper presents a parser for French developed on an unification based categorial grammar (FG) which avoids these problems. This parser is a bottom-up chart parser augmented with a heuristic eliminating spurious parses. The unicity and completeness of parsing are proved. Claire Gardent, Gabriel G. Bès, Pierre-François Jurie, Karine Baschung |
ACL | 1 |
| 1989 | French Order Without Order
Gabriel G. Bès, Claire Gardent |
EACL | 2 |