Marco Dinarelli

dblp:96/2631 · DBLP profile ↗
← Back
39ranked-venue papers
15as first author
13since 2021 · last 2026
0009-0005-4788-4899ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 11 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
abstract
International audience
Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain, Maryem Bouziane, Mohammed Ghennai, Qianwen Guan, Kirill Milintsevich, Salima Mdhaffar, Aidan Mannion, Nils Defauw, Shuyue Gu, Alexandre Audibert, Marco Dinarelli, Yannick Estève, Lorraine Goeuriot, Steffen Lalande, Nicolas Hervé, Maximin Coavoux, François Portet, Étienne Ollion, Marie Candito, Maxime Peyrard, Solange Rossato, Benjamin Lecouteux, Aurélie Nardy, Gilles Sérasset, Vincent Segonne, Solène Evain, Diandra Fabre, Didier Schwab
LREC13
2026 COME-ALPs: Coreference Annotation with MErging Heuristics Using ALignment-based Projection in Parallel Corpora
abstract
International audience
Gabriela González Sáez, Mariam Nakhlé, Illia Kholosha, Rachel Atherly, Marco Dinarelli
LREC5
2025 Towards Early Prediction of Self-Supervised Speech Model Performance
Ryan Whetten, Lucas Maison, Titouan Parcollet, Marco Dinarelli, Yannick Estève
INTERSPEECH4
2024 Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains
abstract
Pretrained Language Models (PLMs) are the de facto backbone of most state-of-the-art NLP systems. In this paper, we introduce a family of domain-specific pretrained PLMs for French, focusing on three important domains: transcribed speech, medicine, and law. We use a transformer architecture based on efficient methods (LinFormer) to maximise their utility, since these domains often involve processing long documents. We evaluate and compare our models to state-of-the-art models on a diverse set of tasks and datasets, some of which are introduced in this paper. We gather the datasets into a new French-language evaluation benchmark for these three domains. We also compare various training configurations: continued pretraining, pretraining from scratch, as well as single- and multi-domain pretraining. Extensive domain-specific experiments show that it is possible to attain competitive downstream performance even when pre-training with the approximative LinFormer attention mechanism. For full reproducibility, we release the models and pretraining data, as well as contributed datasets.
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Audibert, Cécile Macaire, Adrien Pupier, Yongxin Zhou 0004, Mathilde Aguiar, Felix Herron, Magali Norré, Massih-Reza Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
LREC/COLING25
2024 The MAKE-NMTViz Project: Meaningful, Accurate and Knowledge-limited Explanations of NMT Systems for Translators
abstract
This paper describes MAKE-NMTViz, a project designed to help translators visualize neural machine translation outputs using explainable artificial intelligence visualization tools initially developed for computer vision.
Gabriela González Sáez, Fabien Lopez, Mariam Nakhlé, James Turner, Nicolas Ballier, Marco Dinarelli, Emmanuelle Esperança-Rodier, Sui He, Caroline Rossi, Didier Schwab
EAMT (2)6
2024 Exploring NMT Explainability for Translators Using NMT Visualising Tools
abstract
This paper describes work in progress on Visualisation tools to foster collaborations between translators and computational scientists. We aim to describe how visualisation features can be used to explain translation and NMT outputs. We tested several visualisation functionalities with three NMT models based on Chinese-English, Spanish-English and French-English language pairs. We created three demos containing different visualisation tools and analysed them within the framework of performance-explainability, focusing on the translator’s perspective.
Gabriela González Sáez, Mariam Nakhlé, James Turner, Fabien Lopez, Nicolas Ballier, Marco Dinarelli, Emmanuelle Esperança-Rodier, Sui He, Raheel Qader, Caroline Rossi, Didier Schwab
EAMT (1)6
2024 An Analysis of Linear Complexity Attention Substitutes With Best-RQ
abstract
Self-Supervised Learning (SSL) has proven to be effective in various domains, including speech processing. However, SSL is computationally and memory expensive. This is in part due the quadratic complexity of multi-head self-attention (MHSA). Alternatives for MHSA have been proposed and used in the speech domain, but have yet to be investigated properly in an SSL setting. In this work, we study the effects of replacing MHSA with recent state-of-the-art alternatives that have linear complexity, namely, HyperMixing, Fastformer, SummaryMixing, and Mamba. We evaluate these methods by looking at the speed, the amount of VRAM consumed, and the performance on the SSL MP3S benchmark. Results show that these linear alternatives maintain competitive performance compared to MHSA while, on average, decreasing VRAM consumption by around 20% to 60% and increasing speed from 7% to 65% for input sequences ranging from 20 to 80 seconds.
Ryan Whetten, Titouan Parcollet, Adel Moumen, Marco Dinarelli, Yannick Estève
SLT4
2024 LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier
Comput. Speech Lang.10
2023 An Empirical Analysis of Task Relations in the Multi-Task Annotation of an Arabizi Corpus
Elisa Gugliotta, Marco Dinarelli
LDK2
2022 Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation Models
abstract
Multi-encoder models are a broad family of context-aware neural machine translation systems that aim to improve translation quality by encoding document-level contextual information alongside the current sentence.The context encoding is undertaken by contextual parameters, trained on document-level data.In this work, we discuss the difficulty of training these parameters effectively, due to the sparsity of the words in need of context (i.e., the training signal), and their relevant context.We propose to pre-train the contextual parameters over split sentence pairs, which makes an efficient use of the available data for two reasons.Firstly, it increases the contextual training signal by breaking intra-sentential syntactic relations, and thus pushing the model to search the context for disambiguating clues more frequently.Secondly, it eases the retrieval of relevant context, since context segments become shorter.We propose four different splitting methods, and evaluate our approach with BLEU and contrastive test sets.Results show that it consistently improves learning of contextual parameters, both in low and high resource settings.
Lorenzo Lupo, Marco Dinarelli, Laurent Besacier
ACL (1)2
2022 Toward Low-Cost End-to-End Spoken Language Understanding
abstract
International audience
Marco Dinarelli, Marco Naguib, François Portet
INTERSPEECH1
2022 TArC: Tunisian Arabish Corpus, First complete release
abstract
In this paper we present the final result of a project focused on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the realization of two integrated and independent tools: a linguistic corpus and a neural network architecture created to annotate the former with various levels of linguistic information (code-switching classification, transliteration, tokenization, POS-tagging, lemmatization). We discuss the choices made in terms of computational and linguistic methodology and the strategies adopted to improve our results. We report on the experiments performed in order to outline our research path. Finally, we explain the reasons why we believe in the potential of these tools for both computational and linguistic researches.
Elisa Gugliotta, Marco Dinarelli
LREC2
2021 LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech
abstract
Self-Supervised Learning (SSL) using huge unlabeled data has been successfully explored for image and natural language processing. Recent works also investigated SSL from speech. They were notably successful to improve performance on downstream tasks such as automatic speech recognition (ASR). While these works suggest it is possible to reduce dependence on labeled data for building efficient speech systems, their evaluation was mostly made on ASR and using multiple and heterogeneous experimental settings (most of them for English). This questions the objective comparison of SSL approaches and the evaluation of their impact on building speech systems. In this paper, we propose LeBenchmark: a reproducible framework for assessing SSL from speech. It not only includes ASR (high and low resource) tasks but also spoken language understanding, speech translation and emotion recognition. We also focus on speech technologies in a language different than English: French. SSL models of different sizes are trained from carefully sourced and documented datasets. Experiments show that SSL is beneficial for most but not all tasks which confirms the need for exhaustive and reliable benchmarks to evaluate its real impact. LeBenchmark is shared with the scientific community for reproducible research in SSL from speech.
Solène Evain, Hang Le 0001, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia A. Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Estève, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier
Interspeech9
2020 A Data Efficient End-to-End Spoken Language Understanding Architecture
abstract
End-to-end architectures have been recently proposed for spoken language understanding (SLU) and semantic parsing. Based on a large amount of data, those models learn jointly acoustic and linguistic-sequential features. Such architectures give very good results in the context of domain, intent and slot detection, their application in a more complex semantic chunking and tagging task is less easy. For that, in many cases, models are combined with an external a language model to enhance their performance.In this paper we introduce a data efficient system which is trained end-to-end, with no additional, pre-trained external module. One key feature of our approach is an incremental training procedure where acoustic, language and semantic models are trained sequentially one after the other. The proposed model has a reasonable size and achieves competitive results with respect to state-of-the-art while using a small training dataset. In particular, we reach 24.02% Concept Error Rate (CER) on MEDIA/test while training on MEDIA/train without any additional data.
Marco Dinarelli, Nikita Kapoor, Bassam Jabaian, Laurent Besacier
ICASSP1
2020 TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus
abstract
This article describes the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus (TArC). Arabish, also known as Arabizi, is a spontaneous coding of Arabic dialects in Latin characters and “arithmographs” (numbers used as letters). This code-system was developed by Arabic-speaking users of social media in order to facilitate the writing in the Computer-Mediated Communication (CMC) and text messaging informal frameworks. Arabish differs for each Arabic dialect and each Arabish code-system is under-resourced, in the same way as most of the Arabic dialects. In the last few years, the attention of NLP studies on Arabic dialects has considerably increased. Taking this into consideration, TArC will be a useful support for different types of analyses, computational and linguistic, as well as for NLP tools training. In this article we will describe preliminary work on the TArC semi-automatic construction process and some of the first analyses we developed on TArC. In addition, in order to provide a complete overview of the challenges faced during the building process, we will present the main Tunisian dialect characteristics and its encoding in Tunisian Arabish.
Elisa Gugliotta, Marco Dinarelli
LREC2
2018 ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations
Loïc Grobol, Isabelle Tellier, Éric Villemonte de la Clergerie, Marco Dinarelli, Frédéric Landragin
LREC4
2017 Label-Dependencies Aware Recurrent Neural Networks
Yoann Dupont, Marco Dinarelli, Isabelle Tellier
CICLing (1)2
2017 Structured Named Entity Recognition by Cascading CRFs
Yoann Dupont, Marco Dinarelli, Isabelle Tellier, Christian Lautier
CICLing (1)2
2017 Label-Dependency Coding in Simple Recurrent Networks for Spoken Language Understanding
abstract
Modelling target label dependencies is important for sequence labelling tasks. This may become crucial in the case of Spoken Language Understanding (SLU) applications, especially for the slot-filling task where models have to deal often with a high number of target labels. Conditional Random Fields (CRF) were previously considered as the most efficient algorithm in these conditions. More recently, different architectures of Recurrent Neural Networks (RNNs) have been proposed for the SLU slot-filling task. Most of them, however, have been successfully evaluated on the simple ATIS database, on which it is difficult to draw significant conclusions. In this paper we propose new variants of RNNs able to learn efficiently and effectively label dependencies by integrating label embeddings. We show first that modeling label dependencies is useless on the (simple) ATIS database and unstructured models can produce state-of-the-art results on this benchmark. On ATIS our new variants achieve the same results as state-of-the-art models, while being much simpler. On the other hand, on the MEDIA benchmark, we show that the modification introduced in the proposed RNN outperforms traditional RNNs and CRF models.
Marco Dinarelli, Vedran Vukotic, Christian Raymond
INTERSPEECH1
2016 Coreference Resolution for French Oral Data: Machine Learning Experiments with ANCOR
Adèle Désoyer, Frédéric Landragin, Isabelle Tellier, Anaïs Lefeuvre-Halftermeyer, Jean-Yves Antoine, Marco Dinarelli
CICLing (1)6
2016 New Recurrent Neural Network Variants for Sequence Labeling
Marco Dinarelli, Isabelle Tellier
CICLing (1)1
2016 Domain Adaptation for Named Entity Recognition Using CRFs
Tian Tian 0005, Marco Dinarelli, Isabelle Tellier, Pedro Miguel Dias Cardoso
LREC2
2014 Evaluation of different strategies for domain adaptation in opinion mining
Anne Garcia-Fernandez, Olivier Ferret, Marco Dinarelli
LREC3
2012 Tree Representations in Probabilistic Models for Extended Named Entities Detection
Marco Dinarelli, Sophie Rosset
EACL1
2012 Tree-Structured Named Entity Recognition on OCR Data: Analysis, Processing and Results
Marco Dinarelli, Sophie Rosset
LREC1
2012 Discriminative Reranking for Spoken Language Understanding
abstract
Spoken language understanding (SLU) is concerned with the extraction of meaning structures from spoken utterances. Recent computational approaches to SLU, e.g., conditional random fields (CRFs), optimize local models by encoding several features, mainly based on simple n-grams. In contrast, recent works have shown that the accuracy of CRF can be significantly improved by modeling long-distance dependency features. In this paper, we propose novel approaches to encode all possible dependencies between features and most importantly among parts of the meaning structure, e.g., concepts and their combination. We rerank hypotheses generated by local models, e.g., stochastic finite state transducers (SFSTs) or CRF, with a global model. The latter encodes a very large number of dependencies (in the form of trees or sequences) by applying kernel methods to the space of all meaning (sub) structures. We performed comparative experiments between SFST, CRF, support vector machines (SVMs), and our proposed discriminative reranking models (DRMs) on representative conversational speech corpora in three different languages: the ATIS (English), the MEDIA (French), and the LUNA (Italian) corpora. These corpora have been collected within three different domain applications of increasing complexity: informational, transactional, and problem-solving tasks, respectively. The results show that our DRMs consistently outperform the state-of-the-art models based on CRF.
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi
IEEE Trans. Speech Audio Process.1
2011 Hypotheses Selection Criteria in a Reranking Framework for Spoken Language Understanding
Marco Dinarelli, Sophie Rosset
EMNLP1
2011 Models Cascade for Tree-Structured Named Entity Detection
Marco Dinarelli, Sophie Rosset
IJCNLP1
2011 When Was It Written? Automatically Determining Publication Dates
Anne Garcia-Fernandez, Anne-Laure Ligozat, Marco Dinarelli, Delphine Bernhard
SPIRE3
2011 Comparing Stochastic Approaches to Spoken Language Understanding in Multiple Languages
abstract
One of the first steps in building a spoken language understanding (SLU) module for dialogue systems is the extraction of flat concepts out of a given word sequence, usually provided by an automatic speech recognition (ASR) system. In this paper, six different modeling approaches are investigated to tackle the task of concept tagging. These methods include classical, well-known generative and discriminative methods like Finite State Transducers (FSTs), Statistical Machine Translation (SMT), Maximum Entropy Markov Models (MEMMs), or Support Vector Machines (SVMs) as well as techniques recently applied to natural language processing such as Conditional Random Fields (CRFs) or Dynamic Bayesian Networks (DBNs). Following a detailed description of the models, experimental and comparative results are presented on three corpora in different languages and with different complexity. The French MEDIA corpus has already been exploited during an evaluation campaign and so a direct comparison with existing benchmarks is possible. Recently collected Italian and Polish corpora are used to test the robustness and portability of the modeling approaches. For all tasks, manual transcriptions as well as ASR inputs are considered. Additionally to single systems, methods for system combination are investigated. The best performing model on all tasks is based on conditional random fields. On the MEDIA evaluation corpus, a concept error rate of 12.6% could be achieved. Here, additionally to attribute names, attribute values have been extracted using a combination of a rule-based and a statistical approach. Applying system combination using weighted ROVER with all six systems, the concept error rate (CER) drops to 12.0%.
Stefan Hahn, Marco Dinarelli, Christian Raymond, Fabrice Lefèvre, Patrick Lehnen, Renato De Mori, Alessandro Moschitti, Hermann Ney, Giuseppe Riccardi
IEEE Trans. Speech Audio Process.2
2010 The LUNA Spoken Dialogue System: Beyond utterance classification
abstract
We present a call routing application for complex problem solving tasks. Up to date work on call routing has been mainly dealing with call-type classification. In this paper we take call routing further: Initial call classification is done in parallel with a robust statistical Spoken Language Understanding module. This is followed by a dialogue to elicit further task-relevant details from the user before passing on the call. The dialogue capability also allows us to obtain clarifications of the initial classifier guess. Based on an evaluation, we show that conducting a dialogue significantly improves upon call routing based on call classification alone. We present both subjective and objective evaluation results of the system according to standard metrics on real users.
Marco Dinarelli, Evgeny A. Stepanov, Sebastian Varges, Giuseppe Riccardi
ICASSP1
2010 Hypotheses selection for re-ranking semantic annotations
abstract
Discriminative reranking has been successfully used for several tasks of Natural Language Processing (NLP). Recently it has been applied also to Spoken Language Understanding, imrpoving state-of-the-art for some applications. However, such proposed models can be further improved by considering: (i) a better selection of the initial n-best hypotheses to be re-ranked and (ii) the use of a strategy that decides when the reranking model should be used, i.e. in some cases only the basic approach should be applied. In this paper, we apply a semantic inconsistency metric to select the n-best hypotheses from a large set generated by an SLU basic system. Then we apply a state-of-the-art re-ranker based on the Partial Tree Kernel (PTK), which encodes SLU hypotheses in Support Vector Machines (SVM) with complex structured features. Finally, we apply a decision model based on confidence values to select between the first hypothesis provided by the basic SLU model and the first hypothesis provided by the re-ranker. We show the effectiveness of our approach presenting comparative results obtained by reranking hypotheses generated by two very different models: a simple Stochastic Language Model encoded in Finite State Machines (FSM) and a Conditional Random Field (CRF) model. We evaluate our approach on the French MEDIA corpus and on an Italian corpus acquired in the European Project LUNA. The results show a significant improvement with respect to the current state-of-the-art and previous re-ranking models.
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi
SLT1
2009 Ontology-based grounding of Spoken Language Understanding
abstract
Current Spoken Language Understanding models rely on either hand-written semantic grammars or flat attribute-value sequence labeling. In most cases, no relations between concepts are modeled, and both concepts and relations are domain-specific, making it difficult to expand or port the domain model. In contrast, we expand our previous work on a domain model based on an ontology where concepts follow the predicate-argument semantics and domain-independent classical relations are defined on such concepts. We conduct a thorough study on a spoken dialog corpus collected within a customer care problem-solving domain, and we evaluate the coverage and impact of the ontology for the interpretation, grounding and re-ranking of spoken language understanding interpretations.
Silvia Quarteroni, Marco Dinarelli, Giuseppe Riccardi
ASRU2
2009 Re-Ranking Models for Spoken Language Understanding
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi
EACL1
2009 Re-Ranking Models Based-on Small Training Data for Spoken Language Understanding
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi
EMNLP1
2009 Concept segmentation and labeling for conversational speech
abstract
Spoken Language Understanding performs automatic concept labeling and segmentation of speech utterances. For this task, many approaches have been proposed based on both generative and discriminative models. While all these methods have shown remarkable accuracy on manual transcription of spoken utterances, robustness to noisy automatic transcription is still an open issue. In this paper we study algorithms for Spoken Language Understanding combining complementary learning models: Stochastic Finite State Transducers produce a list of hypotheses, which are re-ranked using a discriminative algorithm based on kernel methods. Our experiments on two different spoken dialog corpora, MEDIA and LUNA, show that the combined generative-discriminative model reaches the state-of-the-art such as Conditional Random Fields (CRF) on manual transcriptions, and it is robust to noisy automatic transcriptions, outperforming, in some cases, the state-of-the-art. Copyright © 2009 ISCA.
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi
INTERSPEECH1
2009 What's in an ontology for spoken language understanding
abstract
Current Spoken Language Understanding systems rely either on hand-written semantic grammars or on flat attribute-value se-quence labeling. In both approaches, concepts and their rela-tions (when modeled at all) are domain-specific, thus making it difficult to expand or port the domain model. To address this issue, we introduce: 1) a domain model based on an ontology where concepts are classified into either predicative or argumentative; 2) the modeling of relations be-tween such concept classes in terms of classical relations as defined in lexical semantics. We study and analyze our ap-proach on a corpus of customer care data, where we evaluate the coverage and relevance of the ontology for the interpreta-tion of speech utterances. Index Terms: Spoken Language Understanding, domain mod-eling, ontology design, semantic relations
Silvia Quarteroni, Giuseppe Riccardi, Marco Dinarelli
INTERSPEECH3
2008 Semantic annotations for conversational speech: From speech transcriptions to predicate argument structures
abstract
In this paper, we describe the semantic content, which can be automatically generated, for the design of advanced dialog systems. Since the latter will be based on machine learning approaches, we created training data by annotating a corpus with the needed content. Given a sentence of our transcribed corpus, domain concepts and other linguistic levels ranging from basic ones, i.e. part-of-speech tagging and constituent chunking level, to more advanced ones, i.e. syntactic and predicate argument structure (PAS) levels are annotated. In particular, the proposed PAS and taxonomy of dialog acts appear to be promising for the design of more complex dialog systems. Statistics about our semantic annotation are reported.
Arianna Bisazza, Marco Dinarelli, Silvia Quarteroni, Sara Tonelli, Alessandro Moschitti, Giuseppe Riccardi
SLT2
2008 Joint generative and discriminative models for spoken language understanding
abstract
Spoken Language Understanding aims at mapping a natural language spoken sentence into a semantic representation. In the last decade two main approaches have been pursued: generative and discriminative models. The former is more robust to overfitting whereas the latter is more robust to many irrelevant features. Additionally, the way in which these approaches encode prior knowledge is very different and their relative performance changes based on the task. In this paper we describe a training framework where both models are used: a generative model produces a list of ranked hypotheses whereas a discriminative model, depending on string kernels and Support Vector Machines, re-ranks such list. We tested such approach on a new corpus produced in the European LUNA project. The results show a large improvement on the state-of-the-art in concept segmentation and labeling.
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi
SLT1