Vincent Labatut

dblp:10/6591 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-2619-2835ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
abstract
While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers.Yet, such corpora are necessary to fully represent the diversity of data found in the wild in the context of NLP tasks.We tackle this issue by proposing a method to lawfully and publicly share the annotations of copyrighted literary texts.The corpus creator shares the annotations in clear, along with a non-reversible hashed version of the source material.The corpus user must own the source material, and apply the same hash function to their own tokens, in order to match them to the shared annotations.Crucially, our method is robust to reasonable divergences in the version of the copyrighted data owned by the user.As an illustration, we present alignment experiments on different editions of novels.Our results show that our method is able to correctly align 98.7 to 99.79% of tokens depending on the novel, provided the user version is sufficiently close to the corpus creator's version.We publicly release novelshare, a Python implementation of our method.
Arthur Amalvy, Vincent Labatut, Xavier Bost, Hen-Hsen Huang
ACL (1)2
2025 Persistent Homology of Topic Networks for the Prediction of Reader Curiosity
abstract
Manuel D. S. Hopp, Vincent Labatut, Arthur Amalvy, Richard Dufour, Hannah Stone, Hayley Jach, Kou Murayama. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Manuel D. S. Hopp, Vincent Labatut, Arthur Amalvy, Richard Dufour, Hannah Stone, Hayley K. Jach, Kou Murayama
ACL (1)2
2025 The Role of Natural Language Processing Tasks in Automatic Literary Character Network Construction
abstract
The automatic extraction of character networks from literary texts is generally carried out using natural language processing (NLP) cascading pipelines. While this approach is widespread, no study exists on the impact of low-level NLP tasks on their performance. In this article, we conduct such a study on a literary dataset, focusing on the role of named entity recognition (NER) and coreference resolution when extracting co-occurrence networks. To highlight the impact of these tasks’ performance, we start with gold-standard annotations, progressively add uniformly distributed errors, and observe their impact in terms of character network quality. We demonstrate that NER performance depends on the tested novel and strongly affects character detection. We also show that NER-detected mentions alone miss a lot of character co-occurrences, and that coreference resolution is needed to prevent this. Finally, we present comparison points with 2 methods based on large language models (LLMs), including a fully end-to-end one, and show that these models are outperformed by traditional NLP pipelines in terms of recall.
Arthur Amalvy, Vincent Labatut, Richard Dufour
COLING2
2025 Pattern-Based Graph Classification: Comparison of Quality Measures and Importance of Preprocessing
abstract
Graph classification aims to categorize graphs based on their structural and attribute features, with applications in diverse fields such as social network analysis and bioinformatics. Among the methods proposed to solve this task, those relying on patterns (i.e., subgraphs) provide good explainability, as the patterns used for classification can be directly interpreted. To identify meaningful patterns, a standard approach is to use a quality measure, i.e., a function that evaluates the discriminative power of each pattern. However, the literature provides tens of such measures, making it difficult to select the most appropriate for a given application. Only a handful of surveys try to provide some insight by comparing these measures, and none of them specifically focuses on graphs. This typically results in the systematic use of the most widespread measures, without thorough evaluation. To address this issue, we present a comparative analysis of 38 quality measures from the literature. We characterize them theoretically, based on four mathematical properties. We leverage publicly available datasets to constitute a benchmark, and propose a method to elaborate a gold standard ranking of the patterns. We exploit these resources to perform an empirical comparison of the measures, both in terms of pattern ranking and classification performance. Moreover, we propose a clustering-based preprocessing step, which groups patterns appearing in the same graphs to enhance classification performance. Our experimental results demonstrate the effectiveness of this step, reducing the number of patterns to be processed while achieving comparable performance. Additionally, we show that some popular measures widely used in the literature are not associated with the best results.
Lucas Potin, Rosa Figueiredo 0001, Vincent Labatut, Christine Largeron
ACM Trans. Knowl. Discov. Data3
2024 FedSV: Byzantine-Robust Federated Learning via Shapley Value
abstract
In Federated Learning (FL), several clients jointly learn a machine learning model: each client maintains a local model for its local learning dataset, while a master server maintains a global model by aggregating the local models of the client devices. However, the repetitive communication between server and clients leaves room for attacks aimed at compromising the integrity of the global model, causing errors in its targeted predictions. In response to such threats on FL, various defense measures have been proposed in the literature [1]. In this paper, we present a powerful defense against malicious clients in FL, called FedSV, using the Shapley Value (SV), which has been proposed recently to measure user conribution in FL by computing the marginal increase of average accuracy of the model due to the addition of local data of a user. Our approach makes the identification of malicious clients more robust, since during the learning phase, it estimates the conribution of each client according to the different groups to which the target client belongs. FedSV's effectiveness is demonstrated by extensive experiments on MNIST datasets in a cross-silo context under various attacks.
Khaoula Otmani, Rachid El Azouzi, Vincent Labatut
ICC3
2024 Link Prediction in Bipartite Networks
abstract
Bipartite networks serve as highly suitable models to represent systems involving interactions between two distinct types of entities, such as online dating platforms, job search services, or e-commerce websites. These models can be leveraged to tackle a number of tasks, including link prediction among the most useful ones, especially to design recommendation systems. However, if this task has garnered much interest when conducted on unipartite (i.e. standard) networks, it is far from being the case for bipartite ones. In this study, we address this gap by performing an experimental comparison of 19 link prediction methods able to handle bipartite graphs. Some come directly from the literature, and some are adapted by us from techniques originally designed for unipartite networks. We also propose to repurpose recommendation systems based on graph convolutional networks (GCN) as a novel link prediction solution for bipartite networks. To conduct our experiments, we constitute a benchmark of 3 real-world bipartite network datasets with various topologies. Our results indicate that GCN-based personalized recommendation systems, which have received significant attention in recent years, can produce successful results for link prediction in bipartite networks. Furthermore, purely heuristic metrics that do not rely on any learning process, like the Structural Perturbation Method (SPM), can also achieve success.
Sükrü Demir Inan Özer, Günce Keziban Orman, Vincent Labatut
KES3
2023 Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset
abstract
While recent pre-trained transformer-based models can perform named entity recognition (NER) with great accuracy, their limited range remains an issue when applied to long documents such as whole novels.To alleviate this issue, a solution is to retrieve relevant context at the document level.Unfortunately, the lack of supervision for such a task means one may have to settle for unsupervised approaches.Instead, we propose to generate a synthetic context retrieval training dataset using Alpaca, an instruction-tuned large language model (LLM).Using this dataset, we train a neural context retriever based on a BERT model that is able to find relevant context for NER.We show that our method outperforms several unsupervised retrieval baselines for the NER task on an English literary dataset composed of the first chapter of 40 books, and that it performs on par with re-rankers trained on manually annotated data, or even better.
Arthur Amalvy, Vincent Labatut, Richard Dufour
EMNLP2
2023 Efficient enumeration of the optimal solutions to the correlation clustering problem
Nejat Arinik, Rosa Figueiredo 0001, Vincent Labatut
J. Glob. Optim.3
2020 Serial Speakers: a Dataset of TV Series
abstract
For over a decade, TV series have been drawing increasing interest, both from the audience and from various academic fields. But while most viewers are hooked on the continuous plots of TV serials, the few annotated datasets available to researchers focus on standalone episodes of classical TV series. We aim at filling this gap by providing the multimedia/speech processing communities with “Serial Speakers”, an annotated dataset of 155 episodes from three popular American TV serials: “Breaking Bad”, “Game of Thrones” and “House of Cards”. “Serial Speakers” is suitable both for investigating multimedia retrieval in realistic use case scenarios, and for addressing lower level speech related tasks in especially challenging conditions. We publicly release annotations for every speech turn (boundaries, speaker) and scene boundary, along with annotations for shot boundaries, recurring shots, and interacting speakers in a subset of episodes. Because of copyright restrictions, the textual content of the speech turns is encrypted in the public version of the dataset, but we provide the users with a simple online tool to recover the plain text from their own subtitle files.
Xavier Bost, Vincent Labatut, Georges Linarès
LREC2
2020 WAC: A Corpus of Wikipedia Conversations for Online Abuse Detection
abstract
With the spread of online social networks, it is more and more difficult to monitor all the user-generated content. Automating the moderation process of the inappropriate exchange content on Internet has thus become a priority task. Methods have been proposed for this purpose, but it can be challenging to find a suitable dataset to train and develop them. This issue is especially true for approaches based on information derived from the structure and the dynamic of the conversation. In this work, we propose an original framework, based on the the Wikipedia Comment corpus, with comment-level abuse annotations of different types. The major contribution concerns the reconstruction of conversations, by comparison to existing corpora, which focus only on isolated messages (i.e. taken out of their conversational context). This large corpus of more than 380k annotated messages opens perspectives for online abuse detection and especially for context-based approaches. We also propose, in addition to this corpus, a complete benchmarking platform to stimulate and fairly compare scientific works around the problem of content abuse detection, trying to avoid the recurring problem of result replication. Finally, we apply two classification methods to our dataset to demonstrate its potential.
Noé Cecillon, Vincent Labatut, Richard Dufour, Georges Linarès
LREC2
2019 Similarity Metric Based on Siamese Neural Networks for Voice Casting
abstract
Dubbing contributes to a larger international distribution of multimedia documents. It aims to replace the original voice in a source language by a new one in a target language. For now, the target voice selection procedure, called voice casting, is manually performed by human experts. This selection is not exclusively based on acoustic similarity between the two voices. Actually, it is also supported by more subjective criteria such as the "color" of the voice, sociocultural choices... The objective of this work is to model a voice similarity metric able to embed all the concerned voice characteristics, including the observers' receptive interests. In this paper, we propose a Siamese Neural Networks-based approach, measuring proximity between the original and dubbed voices. We propose an adapted jackknifing cross-validation method to evaluate our similarity model on unseen voices. The results show that we successfully capture information allowing two voices to be associated, with respect to the character's or role's abstract dimension.
Adrien Gresse, Mathias Quillot, Richard Dufour, Vincent Labatut, Jean-François Bonastre
ICASSP4
2019 Remembering winter was coming - Character-oriented video summaries of TV series
Xavier Bost, Serigne Gueye, Vincent Labatut, Martha A. Larson, Georges Linarès, Damien Malinas, Raphaël Roth
Multim. Tools Appl.3
2019 Conversational Networks for Automatic Online Moderation
abstract
Moderation of user-generated content in an online community is a challenge that has great socio-economic ramifications. However, the costs incurred by delegating this paper to human agents are high. For this reason, an automatic system able to detect abuse in user-generated content is of great interest. There are a number of ways to tackle this problem, but the most commonly seen in practice are word filtering or regular expression matching. The main limitations are their vulnerability to intentional obfuscation on the part of the users, and their context-insensitive nature. Moreover, they are language dependent and may require appropriate corpora for training. In this paper, we propose a system for automatic abuse detection that completely disregards message content. We first extract a conversational network from raw chat logs and characterize it through topological measures. We then use these as features to train a classifier on our abuse detection task. We thoroughly assess our system on a dataset of user comments originating from a French massively multiplayer online game. We identify the most appropriate network extraction parameters and discuss the discriminative power of our features, relatively to their topological and temporal nature. Our method reaches an F -measure of 83.89 when using the full feature set, improving on existing approaches. With a selection of the most discriminative features, we dramatically cut computing time while retaining the most of the performance (82.65).
Etienne Papegnies, Vincent Labatut, Richard Dufour, Georges Linarès
IEEE Trans. Comput. Soc. Syst.2
2017 Impact of Content Features for Automatic Online Abuse Detection
Etienne Papegnies, Vincent Labatut, Richard Dufour, Georges Linarès
CICLing (2)2
2017 Acoustic Pairing of Original and Dubbed Voices in the Context of Video Game Localization
abstract
International audience
Adrien Gresse, Mickael Rouvier, Richard Dufour, Vincent Labatut, Jean-François Bonastre
INTERSPEECH4
2016 Narrative smoothing: Dynamic conversational network for the analysis of TV series plots
abstract
Modern popular TV series often develop complex storylines spanning several seasons, but are usually watched in quite a discontinuous way. As a result, the viewer generally needs a comprehensive summary of the previous season plot before the new one starts. The generation of such summaries requires first to identify and characterize the dynamics of the series subplots. One way of doing so is to study the underlying social network of interactions between the characters involved in the narrative. The standard tools used in the Social Networks Analysis field to extract such a network rely on an integration of time, either over the whole considered period, or as a sequence of several time-slices. However, they turn out to be inappropriate in the case of TV series, due to the fact the scenes showed onscreen alternatively focus on parallel storylines, and do not necessarily respect a traditional chronology. In this article, we introduce narrative smoothing, a novel, still exploratory, network extraction method. It smooths the relationship dynamics based on the plot properties, aiming at solving some of the limitations present in the standard approaches. In order to assess our method, we apply it to a new corpus of 3 popular TV series, and compare it to both standard approaches. Our results are promising, showing narrative smoothing leads to more relevant observations when it comes to the characterization of the protagonists and their relationships. It could be used as a basis for further modeling the intertwined storylines constituting TV series plots.
Xavier Bost, Vincent Labatut, Serigne Gueye, Georges Linarès
ASONAM2
2014 Identifying the community roles of social capitalists in the Twitter network
abstract
In the context of Twitter, social capitalists are users trying to increase their number of followers and interactions by any means. They are not healthy for the service, because they introduce a bias in the way user influence and visibility are perceived. Understanding their behavior and position in the network is thus of important interest. In this work, we propose to do so by focusing on the community structure level. We first extend an existing method based on the notion of community role, on three different points: 1) handling of directed networks, 2) more precise modeling of the community-related connectivity and 3) unsupervised role identification. We then take advantage of an existing tool to detect social capitalists, and apply our method to analyze their organization and how their links spread across the network. The specific community roles they hold in the network let us know that they reach to obtain high visibility.
Vincent Labatut, Nicolas Dugué, Anthony Perez 0001
ASONAM1
2014 A method for characterizing communities in dynamic attributed complex networks
abstract
Many methods have been proposed to detect communities in complex networks, but very little work has been done regarding their interpretation. In this work, we propose an efficient method to tackle this problem. We first define a sequence-based representation of networks, combining temporal information, topological measures and nodal attributes. We then describe how to identify the most emerging sequential patterns of this dataset and use them to characterize the communities. We also show how to highlight outliers. Finally, as an illustration, we apply our method to a network of scientific collaborations.
Günce Keziban Orman, Vincent Labatut, Marc Plantevit, Jean-François Boulicaut
ASONAM2
2010 Business-Oriented Analysis of a Social Network of University Students: Informative Value of Individual and Relational Data Compared through Group Detection
abstract
Despites the great interest caused by social networks in Business Science, their analysis is rarely performed both in a global and systematic way in this field: most authors focus on parts of the studied network, or on a few nodes considered individually. This could be explained by the fact that practical extraction of social networks is a difficult and costly task, since the specific relational data it requires are often difficult to access and thereby expensive. One may ask if equivalent information could be extracted from less expensive individual data, i.e. data concerning single individuals instead of several ones. In this work, we try to tackle this problem through group detection. We gather both types of data from a population of students, and estimate groups separately using individual and relational data, leading to sets of clusters and communities, respectively. We found out there is no strong overlapping between them, meaning both types of data do not convey the same information in this specific context, and can therefore be considered as complementary. However, a link, even if weak, exists and appears when we identify the most discriminant attributes relatively to the communities. Implications in Business Science include community prediction using individual data.
Vincent Labatut, Jean-Michel Balasque
ASONAM1
2010 The Effect of Network Realism on Community Detection Algorithms
abstract
Community detection consists in searching cohesive subgroups in complex networks. It has recently become one of the domain pivotal questions for scientists in many different fields where networks are used as modeling tools. Algorithms performing community detection are usually tested on real, but also on artificial networks, the former being costly and difficult to obtain. In this context, being able to generate networks with realistic properties is crucial for the reliability of the tests. Recently, Lancichinetti et al. designed a method to produce realistic networks, with a community structure and power law distributed degrees and community sizes. However, other realistic properties such as degree correlation and transitivity are missing. In this work, we propose a modification of their approach, based on the preferential attachment model, in order to remedy this limitation. We analyze the properties of the generated networks and compare them to the original approach. We then apply different community detection algorithms and observe significant changes in their performances when compared to results on networks generated with the original approach.
Günce Keziban Orman, Vincent Labatut
ASONAM2
2009 A Comparison of Community Detection Algorithms on Artificial Networks
Günce Keziban Orman, Vincent Labatut
Discovery Science2
2004 Cerebral modeling and dynamic Bayesian networks
Vincent Labatut, Josette Pastor, Serge Ruff, Jean-François Démonet, Pierre Celsis
Artif. Intell. Medicine1
2003 Dynamic Bayesian modeling of the cerebral activity
Vincent Labatut, Josette Pastor, Serge Ruff
IJCAI1