Guillaume Gadek

dblp:199/9463 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0002-3238-4275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FactNET: A Granular Fact-Checking Assistant Framework for Complex Claims
abstract
This paper presents FactNET, a framework for human-centered granular fact-checking. Each claim to check is decomposed into sub claims that are fact-checked independently by using an LLM’s internal knowledge. The user interacts with the framework through an innovative interface, displaying the complex claim in the form of a graph, allowing the input of human knowledge on the topic, and evaluating the trust given in the generated evidence. A video demonstrating the system is available at TO_BE_PUBLISHED (in submission documents for review phase).
Géraud Faye, Wassila Ouerdane, Guillaume Gadek, Sylvain Gatepaille, Céline Hudelot
ECAI3
2024 A Multi-Label Dataset of French Fake News: Human and Machine Insights
abstract
We present a corpus of 100 documents, named OBSINFOX, selected from 17 sources of French press considered unreliable by expert agencies, annotated using 11 labels by 8 annotators. By collecting more labels than usual, by more annotators than is typically done, we can identify features that humans consider as characteristic of fake news, and compare them to the predictions of automated classifiers. We present a topic and genre analysis using Gate Cloud, indicative of the prevalence of satire-like text in the corpus. We then use the subjectivity analyzer VAGO, and a neural version of it, to clarify the link between ascriptions of the label Subjective and ascriptions of the label Fake News. The annotated dataset is available online at the following url: https://github.com/obs-info/obsinfox Keywords: Fake News, Multi-Labels, Subjectivity, Vagueness, Detail, Opinion, Exaggeration, French Press
Benjamin Icard, François Maine, Morgane Casanova, Géraud Faye, Julien Chanson, Guillaume Gadek, Ghislain Auguste Atemezing, François Bancilhon, Paul Egré
LREC/COLING6
2024 POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks
abstract
POPCORN is a research project aiming at maturing Information Extraction (IE) solutions for intelligence services. Due to defense security constraints, reports analyzed by intelligence services are not to be accessible to the scientific community. To address this challenge, we propose a dataset made of “fictional” (handcrafted) and “synthetic” (AI generated) French reports. Those synthetic reports are produced by an innovative approach that generates texts closely resembling real-world intelligence reports, facilitating the training and evaluation of IE tasks such as Entity and Relation Extraction. Experiments demonstrate the interest of synthetic reports to enhance the performance of IE models, showcasing their potential to augment real-world intelligence operations.
Bastien Giordano, Maxime Prieur, Nakanyseth Vuth, Sylvain Verdy, Kévin Cousot, Gilles Sérasset, Guillaume Gadek, Didier Schwab, Cédric Lopez
KES7
2024 Shadowfax: Harnessing Textual Knowledge Base Population
abstract
International audience
Maxime Prieur, Cédric du Mouza, Guillaume Gadek, Bruno Grilhères
SIGIR3
2023 Evaluating and Improving End-to-End Systems for Knowledge Base Population
abstract
International audience
Maxime Prieur, Cédric du Mouza, Guillaume Gadek, Bruno Grilhères
ICAART (3)3
2023 A novel hybrid approach for text encoding: Cognitive Attention To Syntax model to detect online misinformation
Géraud Faye, Wassila Ouerdane, Guillaume Gadek, Souhir Gahbiche-Braham, Sylvain Gatepaille
Data Knowl. Eng.3
2022 Duplicate Detection in a Knowledge Base with PIKA
Maxime Prieur, Guillaume Gadek, Bruno Grilhères
ICAART (3)2
2021 Combining vagueness detection with deep learning to identify fake news
Paul Guélorget, Benjamin Icard, Guillaume Gadek, Souhir Gahbiche-Braham, Sylvain Gatepaille, Ghislain Auguste Atemezing, Paul Egré
FUSION3
2021 Active learning to measure opinion and violence in French newspapers
abstract
News articles analysis may be oversimplified when restricted to detecting classes of interest already benefiting from trustworthy labeled datasets, like political affiliation or fakeness. Behind an apparent neutrality, an editorial slant may be embodied by favoring one-sided interviews, avoiding topics or choosing oriented illustrations. These challenges, seen as machine learning problems, would require a tedious annotation task. We introduce ReALMS, an active learning framework capable of quickly elaborating models which detect arbitrary classes in multi-modal text and image documents. Evidence of this capability is given by a case study on French news outlets: the detection of subjectivity, demonstrations and violence.
Paul Guélorget, Guillaume Gadek, Titus Zaharia, Bruno Grilhères
KES2
2020 Arabizi Language Models for Sentiment Analysis
abstract
Arabizi is a written form of spoken Arabic, relying on Latin characters and digits.It is informal and does not follow any conventional rules, raising many NLP challenges.In particular, Arabizi has recently emerged as the Arabic language in online social networks, becoming of great interest for opinion mining and sentiment analysis.Unfortunately, only few Arabizi resources exist and state-of-the-art language models such as BERT do not consider Arabizi.In this work, we construct and release two datasets: (i) LAD, a corpus of 7.7M tweets written in Arabizi and (ii) SALAD, a subset of LAD, manually annotated for sentiment analysis.Then, a BERT architecture is pre-trained on LAD, in order to create and distribute an Arabizi language model called BAERT.We show that a language model (BAERT) pre-trained on a large corpus (LAD) in the same language (Arabizi) as that of the fine-tuning dataset (SALAD), outperforms a state-of-the-art multi-lingual pretrained model (multilingual BERT) on a sentiment analysis task.
Gaétan Baert, Souhir Gahbiche-Braham, Guillaume Gadek, Alexandre Pauchet
COLING3
2020 An interpretable model to measure fakeness and emotion in news
abstract
Fake news and post-truth are everywhere. The huge number of online news outlets and the frequency of content creation underlines the demand for automatic information evaluation tools. Previous work usually focuses either on automatic fact-checking, or on fake-looking identification: the former tries to match a piece of content with trustable information, enough to confirm or infirm the claims. The latter gathers clues to help the reader’s assessment of the piece of content. In this domain, there is no silver bullet: the reader desires verifiable information, thus the fake news detector should be interpretable or explainable. In this article, we propose TC-CNN: an interpretable text classifier. We use it on two tasks: fake news detection and emotion classification. A second contribution relies on these two classifiers, and on a third-party hate detector, to perform a case study on this year real and fake-news press articles, in a comparison between mainstream and alt-right media.
Guillaume Gadek, Paul Guélorget
KES1
2017 Extracting Contextonyms from Twitter for Stance Detection
Guillaume Gadek, Josefin Betsholtz, Alexandre Pauchet, Stephan Brunessaux, Nicolas Malandain, Laurent Vercouter
ICAART (2)1
2017 Topical cohesion of communities on Twitter
abstract
Nowadays, Online Social Networks (OSN) are commonly used by groups of users to communicate. Members of a family, colleagues, fans of a brand, political groups... There is an increasing demand for a precise identification of these groups, coming from brand monitoring, business intelligence and e-reputation management. However, a gap can be observed between the communities detected by many data analytics algorithms on OSN, and effective groups existing in real life: the detected communities often lack of meaning and internal semantic cohesion. Most of existing literature on OSN either focuses on the community detection problem in graphs without considering the topic of the messages exchanged, or concentrates exclusively on the messages without taking into account the social links. In this article, we support the hypothesis that communities extracted on OSN should be topically coherent. We therefore propose a model to represent the groups of interaction on Twitter, the reference on micro-blogging OSN, and two metrics to evaluate the topical cohesion of the detected communities. As an evaluation, we measure the topical cohesion of the groups of users detected by a baseline community detection algorithm.
Guillaume Gadek, Alexandre Pauchet, Nicolas Malandain, Khaled Khelif, Laurent Vercouter, Stephan Brunessaux
KES1
2017 Measures for topical cohesion of user communities on Twitter
abstract
Nowadays, Online Social Networks (OSN) are commonly used by groups of users to communicate. Members of a family, colleagues, fans of a brand, political groups: the demand for a precise identification of these groups is increasing from brand monitoring, business intelligence and e-reputation management.
Guillaume Gadek, Alexandre Pauchet, Nicolas Malandain, Khaled Khelif, Laurent Vercouter, Stephan Brunessaux
WI1