VLDB 2026 Research / reviewers in the wild / expert
Sandra Bringay
dblp:26/2289
· DBLP profile ↗
14ranked-venue papers in the field
1as first author
7since 2021 · last 2026
0000-0002-2830-3666ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 7Database Systems & Data Management · 4 (1 first)Data Mining & Knowledge Discovery · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Window: Scaling Listwise LLM Reranking via Candidate Filtering
Louis Remy, Sandra Bringay, Pascal Poncelet, Maximilien Servajean |
DEXA (1) | 2 |
| 2025 | HALIFacts: Evaluating Large Language Models for Domain-Specific Fact-Checking and Their Carbon Impact
Théophile Mandon, Sandra Bringay, Pascal Poncelet, Maximilien Servajean |
WISE (2) | 2 |
| 2025 | An In-depth Analysis of the Linguistic Characteristics of Science Claims on the Web and their Impact on Fact-checkingabstractWeb claims, seen as assertions shared on the web and eligible for fact-checking, are at the heart of online discourse. They have been studied extensively on a variety of downstream tasks such as fact-checking, claim retrieval, bias detection, argument mining, or viewpoint discovery. On the other hand, claims originating from scientific publications have also been the subject of several downstream NLP tasks. However, research carried out so far has yet to focus on scientific web claims, which are scientific claims made on the web (e.g., on social media and news articles). The process of detecting and fact-checking a claim from the web can be very different depending on whether the claim is scientific or not, thus making it crucial for the developed datasets, methods, and models to make a distinction between the two. With this work, we aim at understanding what makes this distinction necessary, by understanding the linguistic differences between scientific and non-scientific claims on the web, and the impact those differences have on existing downstream tasks. To do so, we manually annotate 1,524 web claims from established benchmarks for fact-checking-related tasks, and we run statistical tests to analyze and compare the linguistic features of each group. We find that scientific claims on the web use more analytical speech, but also use more sentiment-related speech, more expressions of physical motion, and have distinct parts of speech (PoS) and punctuation styles. We also conduct experiments showing that BERT-based language models perform worse on scientific web claims by up to 17 F1 points for several downstream tasks. To understand why, we develop a novel methodology to map predictive tokens of language models to explainable linguistic features and find that language models fail to detect a specific subset of predictive features of scientific web claims. We conclude by stating that language models aimed at studying scientific web claims ought to be trained on scientific web discourse, as opposed to being trained only on generic web discourse or only on scientific text from scientific publications. Salim Hafid, Sebastian Schellhammer, Yavuz Selim Kartal, Thomas Papastergiou, Stefan Dietze, Sandra Bringay, Konstantin Todorov |
ACM Trans. Web | 6 |
| 2023 | Explaining controversy through community analysis on TwitterabstractControversy refers to content attracting different point-of-views, as well as positive and negative feedback on a specific event, gathering users into different communities. Research on controversy led to two main categories of works: controversy detection/quantification and controversy explainability. When the former aims to quantify controversy on a topic, the latter aims to understand why a topic is controversial or not. This paper mainly contributes to the controversy explainability. We analyze topic discussions on Twitter from the community perspective to investigate the power of text in classifying tweets into the right community. We propose a SHAP-based pipeline to quantify impactful text features on predictions of three tweet classifiers. We also rely on the use of different text features namely BERT, TF − IDF, and LIWC. The results we obtain from both SHAP plots and statistical analysis show clearly significant impacts of some text features in classifying tweets.It also highlights the relevance of the study as well as the potential benefits of combining text and user interactions to quantify controversy. Samy Benslimane, Thomas Papastergiou, Jérôme Azé, Sandra Bringay, Caroline Mollevi, Maximilien Servajean |
IDEAS | 4 |
| 2023 | Negatively Correlated Noisy Learners for At-Risk User Detection on Social Networks: A Study on Depression, Anorexia, Self-Harm, and SuicideabstractMental and physical health are strongly linked in a bidirectional relationship. Due to the stigma, ignorance, prejudice, fear, and many other reasons, there exists a large universal treatment gap for people with mental disorders. This could motivate those at-risk individuals to find their way into social networks, asking for information or emotional support. Language could provide a natural eyepiece for the study and detection of such at-risk individuals through their writings on social media platforms. In this paper, we consider the problem of detecting at-risk users with clear signs of depression, anorexia, self-harm, and suicidal thoughts. We introduce NCNL, a novel deep learning ensemble architecture that makes use of multiple noisy base learners in Negative Correlation Learning (NCL) configuration for text classification. NCNL is designed to be, backbone-independent, and we examine it with modern Transformer-based architectures. We evaluate our models on six different tasks for at-risk user detection and classification. Our models achieve significant improvements over existing state-of-the-art results reported for five out of the six tasks. Extensive experiments show how NCNL improves diversity over the classical conventional ensemble and the effect of using noisy base learners. Waleed Ragheb, Jérôme Azé, Sandra Bringay, Maximilien Servajean |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | SciTweets - A Dataset and Annotation Framework for Detecting Scientific Online DiscourseabstractScientific topics, claims and resources are increasingly debated as part of online discourse, where prominent examples include discourse related to COVID-19 or climate change. This has led to both significant societal impact and increased interest in scientific online discourse from various disciplines. For instance, communication studies aim at a deeper understanding of biases, quality or spreading patterns of scientific information, whereas computational methods have been proposed to extract, classify or verify scientific claims using NLP and IR techniques. However, research across disciplines currently suffers from both a lack of robust definitions of the various forms of science-relatedness as well as appropriate ground truth data for distinguishing them. In this work, we contribute (a) an annotation framework and corresponding definitions for different forms of scientific relatedness of online discourse in tweets, (b) an expert-annotated dataset of 1261 tweets obtained through our labeling framework reaching an average Fleiss Kappa κ of 0.63, (c) a multi-label classifier trained on our data able to detect science- relatedness with 89% F1 and also able to detect distinct forms of scientific knowledge (claims, references). With this work, we aim to lay the foundation for developing and evaluating robust methods for analysing science as part of large-scale online discourse. Salim Hafid, Sebastian Schellhammer, Sandra Bringay, Konstantin Todorov, Stefan Dietze |
CIKM | 3 |
| 2021 | Controversy Detection: A Text and Graph Neural Network Based Approach
Samy Benslimane, Jérôme Azé, Sandra Bringay, Maximilien Servajean, Caroline Mollevi |
WISE (1) | 3 |
| 2017 | DARE to Care: A Context-Aware Framework to Track Suicidal Ideation on Social Media
Bilel Moulahi, Jérôme Azé, Sandra Bringay |
WISE (2) | 3 |
| 2015 | Collaborative Content-Based Method for Estimating User Reputation in Online Forums
Amine Abdaoui, Jérôme Azé, Sandra Bringay, Pascal Poncelet |
WISE (2) | 3 |
| 2014 | Mining Representative Frequent Patterns in a Hierarchy of Contexts
Julien Rabatel, Sandra Bringay, Pascal Poncelet |
IDA | 2 |
| 2014 | Mining Twitter for Suicide Prevention
Amayas Abboute, Yasser Boudjeriou, Gilles Entringer, Jérôme Azé, Sandra Bringay, Pascal Poncelet |
NLDB | 5 |
| 2013 | OrderSpan: Mining Closed Partially Ordered Patterns
Mickaël Fabrègue, Agnès Braud, Sandra Bringay, Florence Le Ber, Maguelonne Teisseire |
IDA | 3 |
| 2012 | The Pattern Next Door: Towards Spatio-sequential Pattern Discovery
Hugo Alatrista Salas, Sandra Bringay, Frédéric Flouvat, Nazha Selmaoui-Folcher, Maguelonne Teisseire |
PAKDD (2) | 2 |
| 2011 | Towards an On-Line Analysis of Tweets Processing
Sandra Bringay, Nicolas Béchet, Flavien Bouillot, Pascal Poncelet, Mathieu Roche, Maguelonne Teisseire |
DEXA (2) | 1 |