VLDB 2026 Research / reviewers in the wild / expert
Sébastien Fournier
dblp:84/6245
· DBLP profile ↗
29ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-1611-0744ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 10 since 2021Databases, data management, data science and information retrieval · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LegitimNarrate: A Dataset for Analyzing Legitimation Mechanisms in Crowdfunding Narratives
Asmaa Lagrid, Sébastien Fournier, Benedicte Aldebert, Ali Ghods, Daisy Bertrand, Gael Leboeuf |
LREC | 2 |
| 2026 | Spatial Role Labeling Based on Dependency Guided Word EmbeddingsabstractTexts have become an essential spatial data resource in recent years. One leading task to manage the texts’ conveyed spatial data is the Spatial role labeling (SpRL). The latter intends the extraction of formal spatial knowledge from text. Instead of treating the text as a straightforward sequence of words, we incorporate syntactic dependencies to identify entities uttering spatial semantics. First, we investigate the impact of dependency-based word embeddings in SpRL. Then, we propose a dependency-guided LSTM-CRF deep learning model to exploit syntactic relationships among words. Then, we enhance these dependencies features with POS tags and CNN-based character-level representations. Experiments are performed on the standard SpRL-2012 and SpRL-2013 datasets. The experimental results show that the proposed model outperforms other machine learning approaches. The findings show the importance of taking dependencies into account in SpRL tasks. Alaeddine Moussa, Sébastien Fournier, Khaoula Mahmoudi, Bernard Espinasse, Sami Faïz |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2025 | Multimodal Learning with Uncertainty Quantification based on Discounted Belief FusionabstractMultimodal AI models are increasingly used in fields like healthcare, finance, and autonomous driving, where information is drawn from multiple sources or modalities such as images, texts, audios, videos. However, effectively managing uncertainty—arising from noise, insufficient evidence, or conflicts between modalities—is crucial for reliable decision-making. Current uncertainty-aware machine learning methods leveraging, for example, evidence averaging, or evidence accumulation underestimate uncertainties in high-conflict scenarios. Moreover, the state-of-the-art evidence averaging strategy is not order invariant and fails to scale to multiple modalities. To address these challenges, we propose a novel multimodal learning method with order-invariant evidence fusion and introduce a conflict-based discounting mechanism that reallocates uncertain mass when unreliable modalities are detected. We provide both theoretical analysis and experimental validation, demonstrating that unlike the previous work, the proposed approach effectively distinguishes between conflicting and non-conflicting samples based on the provided uncertainty estimates, and outperforms the previous models in uncertainty-based conflict detection. Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier |
AISTATS | 4 |
| 2025 | EM-SEC: Efficient Multi-head Set-Valued Evidential Classification
Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier |
ECML/PKDD (2) | 4 |
| 2025 | LUMA: A Benchmark Dataset for Learning from Uncertain and Multimodal DataabstractMultimodal Deep Learning enhances decision-making by integrating diverse information sources, such as texts, images, audio, and videos. To develop trustworthy multimodal approaches, it is essential to understand how uncertainty impacts these models. We propose LUMA, a unique multimodal dataset, featuring audio, image, and textual data from 50 classes, specifically designed for learning from uncertain data. It extends the well-known CIFAR 10/100 dataset with audio samples extracted from three audio corpora, and text data generated using the Gemma-7B Large Language Model (LLM). The LUMA dataset enables the controlled injection of varying types and degrees of uncertainty to achieve and tailor specific experiments and benchmarking initiatives. LUMA is also available as a Python package including the functions for generating multiple variants of the dataset with controlling the diversity of the data, the amount of noise for each modality, and adding out-of-distribution samples. A baseline pre-trained model is also provided alongside three uncertainty quantification methods: Monte-Carlo Dropout, Deep Ensemble, and Reliable Conflictive Multi-View Learning. This comprehensive dataset and its tools are intended to promote and support the development, evaluation, and benchmarking of trustworthy and robust multimodal deep learning approaches. We anticipate that the LUMA dataset will help the research community to design more trustworthy and robust machine learning approaches for safety critical applications. The code and instructions for downloading and processing the dataset can be found at: https://github.com/bezirganyan/LUMA. Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier |
SIGIR | 4 |
| 2024 | MixMAS: A Framework for Sampling-Based Mixer Architecture Search for Multimodal Fusion and LearningabstractChoosing a suitable deep learning architecture for multimodal data fusion is a challenging task, as it requires the effective integration and processing of diverse data types, each with distinct structures and characteristics. In this paper, we introduce MixMAS, a novel framework for sampling-based mixer architecture search tailored to multimodal learning. Our approach automatically selects the optimal MLP-based architecture for a given multimodal machine learning (MML) task. Specifically, MixMAS utilizes a sampling-based micro-benchmarking strategy to explore various combinations of modality-specific encoders, fusion functions, and fusion networks, systematically identifying the architecture that best meets the task’s performance metrics. Abdelmadjid Chergui, Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier |
IEEE Big Data | 5 |
| 2023 | M2-Mixer: A Multimodal Mixer with Multi-head Loss for Classification from Multimodal DataabstractIn this paper, we propose M2-Mixer, an MLP-Mixer based architecture with multi-head loss for multimodal classification. It achieves better performances than the convolutional, recurrent, or neural architecture search based baseline models with the main advantage of conceptual and computational simplicity. The proposed multi-head loss function addresses the problem of modality predominance (i.e., when one of the modalities is favored over the others by the training algorithm). Our experiments demonstrate that our multimodal mixer architecture, combined with the multi-head loss function, outperforms the baseline models on two benchmark multimodal datasets: AVMNIST and MIMIC-III with respectively, on average, + 0.43% in accuracy and 6. 4 times reduction in training time and + 0.33% in accuracy and 13. 3 times reduction in training time, compared with previous best performing models. Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier |
IEEE Big Data | 4 |
| 2022 | Context-aware Relation Classification based on Deep LearningabstractModern information supports carry heterogeneous data, in such large quantities that the traditional means of processing become obsolete and inefficient to meet today's needs. In addition to the quantity of data, the unstructured nature of this data requires new intelligent, efficient and automated processing techniques. In order to produce automatic systems capable of managing this data and extracting relevant knowledge from it, a number of problems must be solved, including the extraction and classification of relations from textual data. While the extraction of relations is mainly based on syntactic aspects of the text, the classification requires a semantic approach. Such existing relation classification systems deal only with few pre-defined types. These systems don't take into account the context, thus reducing the relevance of this classification. In this paper, we propose a simplified definition of what is context and, based on this definition, we propose an approach to classify relations according to their types while taking into account this context. The system, allowing to obtain a degree of “contextualization” of relations, has been tested on the SemEval-2010 Task-8, New York Times corpora and a contextual dataset, named WikiContext, that we have built for this purpose. The results show that our system outperforms the state-of-the-art relation classification systems, thus demonstrating the relevance of taking context into account in this classification process. Maha Mallek, Ramzi Guetari, Sébastien Fournier, Wided Lejouad Chaari, Bernard Espinasse |
ICTAI | 3 |
| 2022 | Mixing Static Word Embeddings and RoBERTa for Spatial Role LabelingabstractLanguage model pretraining has yielded significant results in diverse natural language processing tasks. RoberTa, an efficient method for pretraining self-supervised NLP systems, is a good example. Our hypothesis in this paper is that the performance of Spatial Role Labeling (SpRL) can be improved by combining static word vectors and bags of features with RoberTa vectors. Furthermore, we show that our method is successful in several SpRL datasets. Alaeddine Moussa, Sébastien Fournier, Khaoula Mahmoudi, Bernard Espinasse, Sami Faïz |
KES | 2 |
| 2022 | DeepREF: A Framework for Optimized Deep Learning-based Relation ClassificationabstractThe Relation Extraction (RE) is an important basic Natural Language Processing (NLP) for many applications, such as search engines, recommender systems, question-answering systems and others. There are many studies in this subarea of NLP that continue to be explored, such as SemEval campaigns (2010 to 2018), or DDI Extraction (2013).For more than ten years, different RE systems using mainly statistical models have been proposed as well as the frameworks to develop them. This paper focuses on frameworks allowing to develop such RE systems using deep learning models. Such frameworks should make it possible to reproduce experiments of various deep learning models and pre-processing techniques proposed in various publications. Currently, there are very few frameworks of this type, and we propose a new open and optimizable framework, called DeepREF, which is inspired by the OpenNRE and REflex existing frameworks. DeepREF allows the employment of various deep learning models, to optimize their use, to identify the best inputs and to get better results with each data set for RE and compare with other experiments, making ablation studies possible. The DeepREF Framework is evaluated on several reference corpora from various application domains. Igor Nascimento, Rinaldo Lima, Adrian-Gabriel Chifu, Bernard Espinasse, Sébastien Fournier |
LREC | 5 |
| 2021 | Spatial Role Labeling based on Improved Pre-trained Word Embeddings and Transfer LearningabstractIn several real-world applications, extracting spatial semantics from text is critical. Spatial Role Labeling (SpRL) introduces a language-independent annotation scheme used in these applications, particularly for reasoning purposes. This paper proposes, first of all, a transfer learning method with a word embeddings-based approach for SpRL. Then, we enhance the word vectors with POS tags and CNN-based character-level representations. Finally, we propose a Residual BiLSTM CRF deep learning model to identify the spatial roles. The experimental results on two datasets: SemEval-2012 and SemEval-2013 Task 3, show that the proposed model outperforms other machine learning approaches. Alaeddine Moussa, Sébastien Fournier, Khaoula Mahmoudi, Bernard Espinasse, Sami Faïz |
KES | 2 |
| 2020 | An Unsupervised Approach for Precise Context Identification from Unstructured Text DocumentsabstractThe majority of the documents produced and exchanged through medias and social networks are unstructured. Due to the amount of these unstructured documents on the Web, their exploitation represents a tedious or even impossible task for human beings without assistance by dedicated algorithms and specialized computer systems in document classification or information extraction. To be efficient and relevant, such systems have to understand the content of these unstructured documents. The context (or topic) of a document is one of the basic information essential for the understanding of its content, and the more precise the context of a document, the more relevant its understanding will be. This paper presents a precise context identification approach that is evaluated quantitatively and qualitatively on several reference corpora and compared to other context identification systems. The contexts identified by our model are much more precise than those identified by these others systems. Maha Mallek, Sébastien Fournier, Ramzi Guetari, Bernard Espinasse, Wided Lejouad Chaari |
ICTAI | 2 |
| 2020 | DeepNLPF: A Framework for Integrating Third Party NLP ToolsabstractNatural Language Processing (NLP) of textual data is usually broken down into a sequence of several subtasks, where the output of one the subtasks becomes the input to the following one, which constitutes an NLP pipeline. Many third-party NLP tools are currently available, each performing distinct NLP subtasks. However, it is difficult to integrate several NLP toolkits into a pipeline due to many problems, including different input/output representations or formats, distinct programming languages, and tokenization issues. This paper presents DeepNLPF, a framework that enables easy integration of third-party NLP tools, allowing the user to preprocess natural language texts at lexical, syntactic, and semantic levels. The proposed framework also provides an API for complete pipeline customization including the definition of input/output formats, integration plugin management, transparent ultiprocessing execution strategies, corpus-level statistics, and database persistence. Furthermore, the DeepNLPF user-friendly GUI allows its use even by a non-expert NLP user. We conducted runtime performance analysis showing that DeepNLPF not only easily integrates existent NLP toolkits but also reduces significant runtime processing compared to executing the same NLP pipeline in a sequential manner. Francisco Rodrigues, Rinaldo Lima, William Domingues, Robson do Nascimento Fidalgo, Adrian-Gabriel Chifu, Bernard Espinasse, Sébastien Fournier |
LREC | 7 |
| 2019 | On the Use of Dependencies in Relation Classification of Text with Deep Learning
Bernard Espinasse, Sébastien Fournier, Adrian-Gabriel Chifu, Gaël Guibon, René Azcurra, Valentin Macé |
CICLing (2) | 2 |
| 2019 | Sentiment Analysis and Sentence Classification in Long Book-Search Queries
Amal Htait, Sébastien Fournier, Patrice Bellot |
CICLing (2) | 2 |
| 2018 | Measuring the Centrality of the References in Scientific PapersabstractCitation analysis is considered as major and one of the most popular branches of bibliometrics. Citation analysis is based on the assumption that all citations have similar values and weights each equally. Specific research fields like content-based citation analysis (CCA) seeks to explain the "how" and "why" of citation behavior. In this paper we tackle to explain the "how" from a centrality indicator based on factors which are built automatically according to the authors' citation behavior. This indicator allows to evaluate bibliographical references' importance for reading the paper with which user interacts. From objective quantitative measurements, factors are computed in order to characterize the level of granularity where citations are used. By the setting of the centrality indicator's factors we can highlight citations which tend towards a partial or a global construction of the authors' discourse. We carry out a pilot study in which we test our approach on some papers and discuss the challenges in carrying out the citation analysis in this context. Our results show interesting and consistent correlations between the level of granularity and the significance of citation influences. Anaïs Ollagnier, Sébastien Fournier, Patrice Bellot |
DocEng | 2 |
| 2018 | Predicting Contradiction Intensity: Low, Strong or Very Strong?abstractReviews on web resources (e.g. courses, movies) become increasingly exploited in text analysis tasks (e.g. opinion detection, controversy detection). This paper investigates contradiction intensity in reviews exploiting different features such as variation of ratings and variation of polarities around specific entities (e.g. aspects, topics). Firstly, aspects are identified according to the distributions of the emotional terms in the vicinity of the most frequent nouns in the reviews collection. Secondly, the polarity of each review segment containing an aspect is estimated. Only resources containing these aspects with opposite polarities are considered. Finally, some features are evaluated, using feature selection algorithms, to determine their impact on the effectiveness of contradiction intensity detection. The selected features are used to learn some state-of-the-art learning approaches. The experiments are conducted on the Massive Open Online Courses data set containing 2244 courses and their 73,873 reviews, collected from coursera.org. Results showed that variation of ratings, variation of polarities, and reviews quantity are the best predictors of contradiction intensity. Also, J48 was the most effective learning approach for this type of classification. Ismail Badache, Sébastien Fournier, Adrian-Gabriel Chifu |
SIGIR | 2 |
| 2017 | Harnessing Ratings and Aspect-Sentiment to Estimate Contradiction Intensity in Temporal-Related ReviewsabstractAnalysis of opinions (reviews) generated by users becomes increasingly exploited by a variety of applications. It allows to follow the evolution of the opinions or to carry out investigations on products. The detection of contradictory opinions about a web resource (e.g., courses, movies, products, etc.) is an important task to evaluate the latter. This paper focuses on the problem of detecting contradictions in reviews based on the sentiment analysis around specific aspects of a resource (document). In general, for web resources such as online courses (e.g. on Coursera or edX ), reviews are often generated during course sessions. Between each session users stop reviewing on the course, and this course may have updates. So, in order to avoid the confusion of contradictory reviews coming from two or more different sessions, the reviews related to a given resource should be firstly grouped according to their session. Secondly, certain aspects are extracted according to the distributions of the emotional terms in the vicinity of the most frequent names in the reviews collection. Thirdly, the polarity of each review segment containing an aspect is identified. Then taking only the resources containing these aspects with opposite polarities (positive, negative). Finally, we propose a measure of contradiction intensity based on the joint dispersion of the polarity and the rating of the reviews containing the aspects within each resource. The evaluation of our approach is conducted on the Massive Open Online Courses (MOOC) collection containing 2244 courses and their 73,873 reviews, collected from Coursera . The results of experiments revealed the effectiveness of the proposed approach to capture and quantify contradiction intensity. Ismail Badache, Sébastien Fournier, Adrian-Gabriel Chifu |
KES | 2 |
| 2016 | SegChainW2V: Towards a Generic Automatic Video Segmentation Framework, Based on Lexical Chains of Audio Transcriptions and Word EmbeddingsabstractWith the advances in multimedia broadcasting through a rich variety of channels and with the vulgarization of video production, it becomes essential to be able to provide reliable means of retrieving information within videos, not only the videos themselves. Research in this area has been widely focused on the context of TV news broadcasts, for which the structure itself provides clues for story segmentation. The systematic employment of these clues would lead to thematically driven systems that would not be easily adaptable in the case of videos of other types. The systems are therefore dependent on the type of videos for which they have been designed. In this paper we aim at introducing SegChainW2V, a generic unsupervised framework for story segmentation, based on lexical chains from transcriptions and their vectorization. SegChainW2V takes into account the topic changes by perceiving the fiuctuations of the most frequent terms throughout the video, as well as their semantics through the word embedding vectorization. Adrian-Gabriel Chifu, Sébastien Fournier |
KES | 2 |
| 2016 | Bilbo-Val: Automatic Identification of Bibliographical Zone in Papers
Amal Htait, Sébastien Fournier, Patrice Bellot |
LREC | 2 |
| 2014 | An Effective TF/IDF-Based Text-to-Text Semantic Similarity Measure for Text Classification
Shereen Albitar, Sébastien Fournier, Bernard Espinasse |
WISE (1) | 2 |
| 2013 | A multi-agent system for learner assessment in serious games: Application to learning processes in crisis managementabstractSerious Games (SG) are more and more used for training in various domains, especially in crisis management domain. In order to improve training results, learner assessment can provide insights on what went right or wrong during a training session. Learner assessment requires monitoring (data acquisition in a virtual environment) and supporting learners' feedback (on learners' decisions and actions), which Intelligent Tutoring Systems address as a main issue, as a mean to individualise learning. In this paper, to enhance learning and pedagogy in virtual environment of collaborative serious games for crisis management, we propose a distributed and multicriteria evaluation approach supported both at modelling and software level by a multi-agent system (MAS). In the context of collaborative serious games for crisis management, this supported assessment has to consider individual and collective assessment. The multi-agent based assessment method proposed is integrated in a serious game developed in the SIMFOR project dedicated to crisis management. An illustrative example is presented to describe the evaluation process. M'hammed Ali Oulhaci, Erwan Tranvouez, Sébastien Fournier, Bernard Espinasse |
RCIS | 3 |
| 2012 | Towards a Supervised Rocchio-based Semantic Classification of Web PagesabstractIn this work, we present and compare several methods for web page classification focusing particularly on supervised classification methods. Rocchio, the method we choose for its efficiency, is tested on the reference corpus “20newsGroups”, adopting different similarity measures. Results illustrate some limitations related mainly to ignoring text semantics. In order to overcome these limitations, this work proposes to extend the original Rocchio through vector conceptualization replacing terms with related concepts and by adopting semantic similarity measures, all using semantic resources as ontologies, towards semantic Rocchio-based classification method. Shereen Albitar, Bernard Espinasse, Sébastien Fournier |
KES | 3 |
| 2012 | Conceptualization Effects on MEDLINE Documents Classification Using Rocchio MethodabstractThe aim of this paper is to propose a supervised text classification method for the biomedical domain using semantic resources. We choose the traditional text classification method, Rocchio, for its scalability and extendibility with semantic knowledge. This paper proposes to integrate semantic aspects into Rocchio through a conceptualization task. This conceptualization is realized by mapping terms that are extracted from text to their corresponding concepts in the UMLS®Metathesaurus®in order to take meaning into consideration during text classification. The proposed classifier is tested on the Ohsumed text corpus, which is composed of abstracts of biomedical articles retrieved from the MEDLINE®database. The effects of Conceptualization on Rocchio's performance are discussed according to different standard similarity measures and to a variety of conceptualization strategies. Shereen Albitar, Sébastien Fournier, Bernard Espinasse |
Web Intelligence | 2 |
| 2012 | The Impact of Conceptualization on Text Classification
Shereen Albitar, Sébastien Fournier, Bernard Espinasse |
WISE | 2 |
| 2011 | Toward a methodology for disruption management - Reactive planning and scheduling based on a repair approachabstractIn the current competitive context, the control of disruptions is becoming a major issue to tackle in the management of industrial organizations. In order to minimize disruption impact in present complex organizations, the paper proposes an approach for repairing plans and schedules, based on distributed cooperative solving method. An agent based model of organizations is proposed, where agents operate cooperative repair behaviours in order to determine the solution which minimizes the disruption impact and propose it to the human decision maker. This paper focuses on the description of the methodology for designing the decision support system, the cooperative repair behaviours that agents develop in order to find solutions to disruptions. Alain Ferrarini, Aline Cauvin, Sébastien Fournier |
ETFA | 3 |
| 2010 | Combining Agents and Wrapper Induction for Information Gathering on Restricted Web DomainsabstractWeb is growing constantly and exponentially every day. Thus, gathering relevant information becomes unfeasible. Existent indexing-based search engines ignore information context, which is essential to deciding on its relevance. Restraining to a single web domain, domain ontology can be used to take into consideration the related context, the fact that might enable treating web pages that belong to the considered domain more intelligently. Nevertheless, symbolic rules that exploit domain's ontology to realize this treatment are delicate and fastidious to develop, especially for information extraction task. This paper presents Boosted Wrapper Induction (BWI), a machine learning method for adaptive information extraction, and its exploitation as a replacement of the symbolic approach for information extraction task in AGATHE, a generic multi-agent architecture for information gathering on restrained web domains. Shereen Albitar, Bernard Espinasse, Sébastien Fournier |
RCIS | 3 |
| 2010 | A multiagent Decision Support System for scheduling repair - application to socio-technical organizationsabstractIn order to minimize disturbance impact in complex organizations, we propose an agent-based decision support system using a distributed cooperative scheduling repair method. We propose an isomorphical modelling approach of actual organization, where agents operate cooperative repair behaviours. Agents implement solving strategies composed of atomic repair operations. A case study in a building site organization illustrates this approach. The resulting multiagent system, developed on the Jade platform, will operate the repair behaviours in order to determine the solution which minimizes the disturbance platform and propose it to the human decision maker. Sébastien Fournier, Alain Ferrarini, Erwan Tranvouez |
RCIS | 1 |
| 2007 | AGATHE: an Agent and Ontology based System for Restricted-Domain Information Gathering on the Web
Bernard Espinasse, Sébastien Fournier, Fred Freitas |
RCIS | 2 |