VLDB 2026 Research / reviewers in the wild / expert
Guillaume Gadek
dblp:199/9463
· DBLP profile ↗
14ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0002-3238-4275ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FactNET: A Granular Fact-Checking Assistant Framework for Complex ClaimsabstractThis paper presents FactNET, a framework for human-centered granular fact-checking. Each claim to check is decomposed into sub claims that are fact-checked independently by using an LLM’s internal knowledge. The user interacts with the framework through an innovative interface, displaying the complex claim in the form of a graph, allowing the input of human knowledge on the topic, and evaluating the trust given in the generated evidence. A video demonstrating the system is available at TO_BE_PUBLISHED (in submission documents for review phase). Géraud Faye, Wassila Ouerdane, Guillaume Gadek, Sylvain Gatepaille, Céline Hudelot |
ECAI | 3 |
| 2024 | A Multi-Label Dataset of French Fake News: Human and Machine InsightsabstractWe present a corpus of 100 documents, named OBSINFOX, selected from 17 sources of French press considered unreliable by expert agencies, annotated using 11 labels by 8 annotators. By collecting more labels than usual, by more annotators than is typically done, we can identify features that humans consider as characteristic of fake news, and compare them to the predictions of automated classifiers. We present a topic and genre analysis using Gate Cloud, indicative of the prevalence of satire-like text in the corpus. We then use the subjectivity analyzer VAGO, and a neural version of it, to clarify the link between ascriptions of the label Subjective and ascriptions of the label Fake News. The annotated dataset is available online at the following url: https://github.com/obs-info/obsinfox Keywords: Fake News, Multi-Labels, Subjectivity, Vagueness, Detail, Opinion, Exaggeration, French Press Benjamin Icard, François Maine, Morgane Casanova, Géraud Faye, Julien Chanson, Guillaume Gadek, Ghislain Auguste Atemezing, François Bancilhon, Paul Egré |
LREC/COLING | 6 |
| 2024 | POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction TasksabstractPOPCORN is a research project aiming at maturing Information Extraction (IE) solutions for intelligence services. Due to defense security constraints, reports analyzed by intelligence services are not to be accessible to the scientific community. To address this challenge, we propose a dataset made of “fictional” (handcrafted) and “synthetic” (AI generated) French reports. Those synthetic reports are produced by an innovative approach that generates texts closely resembling real-world intelligence reports, facilitating the training and evaluation of IE tasks such as Entity and Relation Extraction. Experiments demonstrate the interest of synthetic reports to enhance the performance of IE models, showcasing their potential to augment real-world intelligence operations. Bastien Giordano, Maxime Prieur, Nakanyseth Vuth, Sylvain Verdy, Kévin Cousot, Gilles Sérasset, Guillaume Gadek, Didier Schwab, Cédric Lopez |
KES | 7 |
| 2024 | Shadowfax: Harnessing Textual Knowledge Base PopulationabstractInternational audience Maxime Prieur, Cédric du Mouza, Guillaume Gadek, Bruno Grilhères |
SIGIR | 3 |
| 2023 | Evaluating and Improving End-to-End Systems for Knowledge Base PopulationabstractInternational audience Maxime Prieur, Cédric du Mouza, Guillaume Gadek, Bruno Grilhères |
ICAART (3) | 3 |
| 2023 | A novel hybrid approach for text encoding: Cognitive Attention To Syntax model to detect online misinformation
Géraud Faye, Wassila Ouerdane, Guillaume Gadek, Souhir Gahbiche-Braham, Sylvain Gatepaille |
Data Knowl. Eng. | 3 |
| 2022 | Duplicate Detection in a Knowledge Base with PIKA
Maxime Prieur, Guillaume Gadek, Bruno Grilhères |
ICAART (3) | 2 |
| 2021 | Combining vagueness detection with deep learning to identify fake news
Paul Guélorget, Benjamin Icard, Guillaume Gadek, Souhir Gahbiche-Braham, Sylvain Gatepaille, Ghislain Auguste Atemezing, Paul Egré |
FUSION | 3 |
| 2021 | Active learning to measure opinion and violence in French newspapersabstractNews articles analysis may be oversimplified when restricted to detecting classes of interest already benefiting from trustworthy labeled datasets, like political affiliation or fakeness. Behind an apparent neutrality, an editorial slant may be embodied by favoring one-sided interviews, avoiding topics or choosing oriented illustrations. These challenges, seen as machine learning problems, would require a tedious annotation task. We introduce ReALMS, an active learning framework capable of quickly elaborating models which detect arbitrary classes in multi-modal text and image documents. Evidence of this capability is given by a case study on French news outlets: the detection of subjectivity, demonstrations and violence. Paul Guélorget, Guillaume Gadek, Titus Zaharia, Bruno Grilhères |
KES | 2 |
| 2020 | Arabizi Language Models for Sentiment AnalysisabstractArabizi is a written form of spoken Arabic, relying on Latin characters and digits.It is informal and does not follow any conventional rules, raising many NLP challenges.In particular, Arabizi has recently emerged as the Arabic language in online social networks, becoming of great interest for opinion mining and sentiment analysis.Unfortunately, only few Arabizi resources exist and state-of-the-art language models such as BERT do not consider Arabizi.In this work, we construct and release two datasets: (i) LAD, a corpus of 7.7M tweets written in Arabizi and (ii) SALAD, a subset of LAD, manually annotated for sentiment analysis.Then, a BERT architecture is pre-trained on LAD, in order to create and distribute an Arabizi language model called BAERT.We show that a language model (BAERT) pre-trained on a large corpus (LAD) in the same language (Arabizi) as that of the fine-tuning dataset (SALAD), outperforms a state-of-the-art multi-lingual pretrained model (multilingual BERT) on a sentiment analysis task. Gaétan Baert, Souhir Gahbiche-Braham, Guillaume Gadek, Alexandre Pauchet |
COLING | 3 |
| 2020 | An interpretable model to measure fakeness and emotion in newsabstractFake news and post-truth are everywhere. The huge number of online news outlets and the frequency of content creation underlines the demand for automatic information evaluation tools. Previous work usually focuses either on automatic fact-checking, or on fake-looking identification: the former tries to match a piece of content with trustable information, enough to confirm or infirm the claims. The latter gathers clues to help the reader’s assessment of the piece of content. In this domain, there is no silver bullet: the reader desires verifiable information, thus the fake news detector should be interpretable or explainable. In this article, we propose TC-CNN: an interpretable text classifier. We use it on two tasks: fake news detection and emotion classification. A second contribution relies on these two classifiers, and on a third-party hate detector, to perform a case study on this year real and fake-news press articles, in a comparison between mainstream and alt-right media. Guillaume Gadek, Paul Guélorget |
KES | 1 |
| 2017 | Extracting Contextonyms from Twitter for Stance Detection
Guillaume Gadek, Josefin Betsholtz, Alexandre Pauchet, Stephan Brunessaux, Nicolas Malandain, Laurent Vercouter |
ICAART (2) | 1 |
| 2017 | Topical cohesion of communities on TwitterabstractNowadays, Online Social Networks (OSN) are commonly used by groups of users to communicate. Members of a family, colleagues, fans of a brand, political groups... There is an increasing demand for a precise identification of these groups, coming from brand monitoring, business intelligence and e-reputation management. However, a gap can be observed between the communities detected by many data analytics algorithms on OSN, and effective groups existing in real life: the detected communities often lack of meaning and internal semantic cohesion. Most of existing literature on OSN either focuses on the community detection problem in graphs without considering the topic of the messages exchanged, or concentrates exclusively on the messages without taking into account the social links. In this article, we support the hypothesis that communities extracted on OSN should be topically coherent. We therefore propose a model to represent the groups of interaction on Twitter, the reference on micro-blogging OSN, and two metrics to evaluate the topical cohesion of the detected communities. As an evaluation, we measure the topical cohesion of the groups of users detected by a baseline community detection algorithm. Guillaume Gadek, Alexandre Pauchet, Nicolas Malandain, Khaled Khelif, Laurent Vercouter, Stephan Brunessaux |
KES | 1 |
| 2017 | Measures for topical cohesion of user communities on TwitterabstractNowadays, Online Social Networks (OSN) are commonly used by groups of users to communicate. Members of a family, colleagues, fans of a brand, political groups: the demand for a precise identification of these groups is increasing from brand monitoring, business intelligence and e-reputation management. Guillaume Gadek, Alexandre Pauchet, Nicolas Malandain, Khaled Khelif, Laurent Vercouter, Stephan Brunessaux |
WI | 1 |