Gilles Sérasset

dblp:68/5446 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-2761-7353ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
abstract
International audience
Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain, Maryem Bouziane, Mohammed Ghennai, Qianwen Guan, Kirill Milintsevich, Salima Mdhaffar, Aidan Mannion, Nils Defauw, Shuyue Gu, Alexandre Audibert, Marco Dinarelli, Yannick Estève, Lorraine Goeuriot, Steffen Lalande, Nicolas Hervé, Maximin Coavoux, François Portet, Étienne Ollion, Marie Candito, Maxime Peyrard, Solange Rossato, Benjamin Lecouteux, Aurélie Nardy, Gilles Sérasset, Vincent Segonne, Solène Evain, Diandra Fabre, Didier Schwab
LREC26
2025 MOOC on Linguistic Linked Data
Jorge Gracia, Slavko Zitnik, Maxim Ionov, Christian Chiarcos, Dagmar Gromann, Francesco Mambrini, Marco Passarotti, Armando Stellato, John P. McCrae, Gilles Sérasset, Andon Tchechmedjiev, Sara Carvalho, Penny Labropoulou, Rute Costa
ESWC (2)10
2025 Towards Sense to Sense Linking across DBnary Languages
abstract
Since 2012, the DBnary project extracts lexical information from different Wiktionary language editions (26 editions in 2025) and makes it available to the community as queryable RDF data (modeled using ontolex-lemon ontology). This dataset contains more than 12M translations linking languages at the level of Lexical Entries. This paper presents an effort to automatically link the DBnary languages at the Lexical Sense level. For this we explore different ways to compute cross-lingual semantic similarity, using multilingual language models.
Gilles Sérasset
LDK1
2024 Bridging Computational Lexicography and Corpus Linguistics: A Query Extension for OntoLex-FrAC
abstract
OntoLex, the dominant community standard for machine-readable lexical resources in the context of RDF, Linked Data and Semantic Web technologies, is currently extended with a designated module for Frequency, Attestations and Corpus-based Information (OntoLex-FrAC). We propose a novel component for OntoLex-FrAC, addressing the incorporation of corpus queries for (a) linking dictionaries with corpus engines, (b) enabling RDF-based web services to exchange corpus queries and responses data dynamically, and (c) using conventional query languages to formalize the internal structure of collocations, word sketches, and colligations. The primary field of application of the query extension is in digital lexicography and corpus linguistics, and we present a proof-of-principle implementation in backend components of a novel platform designed to support digital lexicography for the Serbian language.
Christian Chiarcos, Ranka Stankovic, Maxim Ionov, Gilles Sérasset
LREC/COLING4
2024 MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations
abstract
Understanding the relation between the meanings of words is an important part of comprehending natural language. Prior work has either focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs), with some exceptions. Given the rarity of highly multilingual benchmarks, it is unclear to what extent PLMs capture relational knowledge and are able to transfer it across languages. To start addressing this question, we propose MultiLexBATS, a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages, such as Bambara, Lithuanian, and Albanian. As experiment on cross-lingual transfer of relational knowledge, we test the PLMs’ ability to (1) capture analogies across languages, and (2) predict translation targets. We find considerable differences across relation types and languages with a clear preference for hypernymy and antonymy as well as romance languages.
Dagmar Gromann, Hugo Gonçalo Oliveira, Lucia Pitarch, Elena Apostol, Jordi Bernad, Eliot Bytyci, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabík, Jorge Gracia, Letizia Granata, Anas Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia di Buono, Ana Ostroski Anic, Sigita Rackeviciene, Ricardo Rodrigues 0001, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stankovic, Ciprian-Octavian Truica, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
LREC/COLING21
2024 On Modelling Corpus Citations in Computational Lexical Resources
abstract
In this article we look at how two different standards for lexical resources, TEI and OntoLex, deal with corpus citations in lexicons. We will focus on how corpus citations in retrodigitised dictionaries can be modelled using each of the two standards since this provides us with a suitably challenging use case. After looking at the structure of an example entry from a legacy dictionary, we examine the two approaches offered by the two different standards by outlining an encoding for the example entry using both of them (note that this article features the first extended discussion of how the Frequency Attestation and Corpus (FrAC) module of OntoLex deals with citations). After comparing the two approaches and looking at the advantages and disadvantages of both, we argue for a combination of both. In the last part of the article we discuss different ways of doing this, giving our preference for a strategy which makes use of RDFa.
Anas Fahad Khan, Maxim Ionov, Christian Chiarcos, Laurent Romary, Gilles Sérasset, Besim Kabashi
LREC/COLING5
2024 From Linguistic Linked Data to Big Data
abstract
With advances in the field of Linked (Open) Data (LOD), language data on the LOD cloud has grown in number, size, and variety. With an increased volume and variety of language data, optimizations of methods for distributing, storing, and querying these data become more central. To this end, this position paper investigates use cases at the intersection of LLOD and Big Data, existing approaches to utilizing Big Data techniques within the context of linked data, and discusses the challenges and benefits of this union.
Dimitar Trajanov, Elena Apostol, Radovan Garabík, Katerina Gkirtzou, Dagmar Gromann, Chaya Liebeskind, Cosimo Palma, Mike Rosner, Alexia Sampri, Gilles Sérasset, Blerina Spahiu, Ciprian-Octavian Truica, Giedre Valunaite Oleskeviciene
LREC/COLING10
2024 POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks
abstract
POPCORN is a research project aiming at maturing Information Extraction (IE) solutions for intelligence services. Due to defense security constraints, reports analyzed by intelligence services are not to be accessible to the scientific community. To address this challenge, we propose a dataset made of “fictional” (handcrafted) and “synthetic” (AI generated) French reports. Those synthetic reports are produced by an innovative approach that generates texts closely resembling real-world intelligence reports, facilitating the training and evaluation of IE tasks such as Entity and Relation Extraction. Experiments demonstrate the interest of synthetic reports to enhance the performance of IE models, showcasing their potential to augment real-world intelligence operations.
Bastien Giordano, Maxime Prieur, Nakanyseth Vuth, Sylvain Verdy, Kévin Cousot, Gilles Sérasset, Guillaume Gadek, Didier Schwab, Cédric Lopez
KES6
2023 Leveraging DBnary Data to Enrich Information of Multiword Terms in Wiktionary
Gilles Sérasset, Thierry Declerck, Lenka Bajcetic
LDK1
2023 DBnary2Vec: Preliminary Study on Lexical Embeddings for Downstream NLP Tasks
Nakanyseth Vuth, Gilles Sérasset
LDK2
2022 Cross-Lingual Link Discovery for Under-Resourced Languages
abstract
In this paper, we provide an overview of current technologies for cross-lingual link discovery, and we discuss challenges, experiences and prospects of their application to under-resourced languages. We rst introduce the goals of cross-lingual linking and associated technologies, and in particular, the role that the Linked Data paradigm (Bizer et al., 2011) applied to language data can play in this context. We de ne under-resourced languages with a speci c focus on languages actively used on the internet, i.e., languages with a digitally versatile speaker community, but limited support in terms of language technology. We argue that languages for which considerable amounts of textual data and (at least) a bilingual word list are available, techniques for cross-lingual linking can be readily applied, and that these enable the implementation of downstream applications for under-resourced languages via the localisation and adaptation of existing technologies and resources.
Mike Rosner, Sina Ahmadi, Elena Apostol, Julia Bosque-Gil, Christian Chiarcos, Milan Dojchinovski, Katerina Gkirtzou, Jorge Gracia, Dagmar Gromann, Chaya Liebeskind, Giedre Valunaite Oleskeviciene, Gilles Sérasset, Ciprian-Octavian Truica
LREC12
2015 METEOR for multiple target languages using DBnary
Zied Elloumi, Hervé Blanchon, Gilles Sérasset, Laurent Besacier
MTSummit3
2012 Dbnary: Wiktionary as a LMF based Multilingual RDF network
Gilles Sérasset
LREC1
2006 Multilingual Legal Terminology on the Jibiki Platform: The LexALP Project
abstract
This paper presents the particular use of "Jibiki" (Papillon's web server development platform) for the LexALP1 project. LexALP's goal is to harmonise the terminology on spatial planning and sustainable development used within the Alpine Convention2, so that the member states are able to cooperate and communicate efficiently in the four official languages (French, German, Italian and Slovene). To this purpose, LexALP uses the Jibiki platform to build a term bank for the contrastive analysis of the specialised terminology used in six different national legal systems and four different languages. In this paper we present how a generic platform like Jibiki can cope with a new kind of dictionary.
Gilles Sérasset, Francis Brunet-Manquat, Elena Chiocchetti
ACL1
2000 On UNL as the future "html of the linguistic content" & the reuse of existing NLP components in UNL-related applications with the example of a UNL-French deconverter
Gilles Sérasset, Christian Boitet
COLING1
1999 UNL-French deconversion as transfer & generation from an interlingua with possible quality enhancement through offline human interaction
abstract
We present the architecture of the UNL-French deconverter, which “generates” from the UNL interlingua by first “localizing” the UNL form for French, within UNL, and then applying slightly adapted but classical transfer and generation techniques, implemented in GETA’s Ariane-G5 environment, supplemented by some UNL-specific tools. Online interaction can be used during deconversion to enhance output quality and is now used for development purposes. We show how interaction could be delayed and embedded in the postedition phase, which would then interact not directly with the output text, but indirectly with several components of the deconverter. Interacting online or offline can improve the quality not only of the utterance at hand, but also of the utterances processed later, as various preferences may be automatically changed to let the deconverter “learn”.
Gilles Sérasset, Christian Boitet
MTSummit1
1994 lnterlinguai Lexical Organisation for Multilingual Lexical Databases in NADIA
Gilles Sérasset
COLING1