VLDB 2026 Research / reviewers in the wild / expert
Mathieu Roche
dblp:r/MathieuRoche
· DBLP profile ↗
76ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0003-3272-8568ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 28 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Theory of computation · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing language models with selective masking for thematic and misinformation classification in a One Health contextabstractThe objective of this paper is to address the scarcity of labeled textual data and improve the performance of language models in classification tasks within a One Health context by using small domain-specific labeled corpora. To address these challenges, we propose a two-phase training pipeline for language models, in which the first phase involves post-training guided by selective masking (SM) strategies to adapt the model to a specific domain. For this purpose, we propose two novel masking strategies: SM-Lex-TFIDF, which masks domain lexicon terms with high TF-IDF (term frequency-inverse document frequency) values, and SM-NonLex-TFIDF, which masks non-domain lexicon terms with high TF-IDF values. The second phase focuses on fine-tuning the model for the target classification task using small amounts of labeled data. To demonstrate the effectiveness of our approach, we focus on two related application areas within the One Health context, i.e., (i) thematic content in integrated health, covering the biomedical, plant health, and syndromic surveillance domains, and (ii) epidemic misinformation, to achieve improved One Health monitoring. We conduct extensive evaluations to assess the performance of our approach using three language models: BERT Base , SciBERT, and BioBERT. Additionally, we compare our method with low-resource LLM-based approaches, including zero/few-shot classification. Experimental results demonstrate significant improvements in the performance of the language models across classification tasks in both targeted areas, even with limited labeled data. Our approach outperforms zero/few-shot classification using LLaMA-3.1-8B and Mistral-7B in four out of the five datasets evaluated. Furthermore, we provide a summary mapping each strategy to its most effective context. Youssef Mahdoubi, Najlae Idrissi, Mathieu Roche, Sarah Valentin |
Expert Syst. Appl. | 3 |
| 2025 | Coherent Augmentation of Controversial CommentsabstractSelf-sufficiency has grown in popularity over the past decades as a response to ongoing societal crises. Considering the immediacy and gravity of the challenges at hand, it raises considerable debate within the public sphere, notably on widely used media such as YouTube. We propose to study this phenomenon by performing a sociological analysis of the technical knowledge around self-sufficiency in YouTube comment sections. Such study requires large corpora of annotated data, which are difficult to produce and typically scarce, especially in languages other than English. To overcome the problem, we implement in this work two independent methods to augment a dataset of labeled French-language YouTube comments. A first method, AugArg, relies on the use of argumentative structure as a controversy marker in the original corpus, while the other method, AugLLM, proposes to prompt an LLM to generate controversial comments. The fine-tuning of a CamemBERT model is performed using both the original and augmentated data, enabling us to compare the efficiency of both methods in representing controversiality in the context of self-sufficiency. Our results indicate that utilizing argumentative structure in the corpus provides a more controversy-representative set of comments in comparison with LLM-oriented methods. Marina Musse, Mathieu Roche, Natalia Grabar |
IEEE Big Data | 2 |
| 2025 | Enhancing Domain-Specific Named Entity Recognition via Segmentation and Pseudo-Labeled AnnotationabstractNamed Entity Recognition (NER) in specialized domains poses major challenges due to the scarcity of annotated data and the limitations of existing models. One challenge lies in handling long documents, which often requires segmenting the text into smaller chunks that fit within the model's input window. Although several segmentation strategies exist, their impact on NER performance in new domains remains underexplored. Another challenge is domain adaptation: while openschema NER models can perform reasonably well under zero-shot settings, they often struggle in highly specialized contexts without additional supervision. To address these issues, we combine segmentation strategy selection with a semi-supervised finetuning pipeline based on pseudo-labeled annotations. First, we compare four segmentation strategies to identify which offers the best trade-off between precision and recall in zero-shot settings. Then, we fine-tune two open-schema NER models-GLiNER and NuNER-first on a manually annotated dataset, and subsequently on a large pseudo-labeled dataset built from model agreement. Both models are evaluated on a held-out test set and on the full manually annotated corpus. Experiments on French-language documents specifically focused on the underexplored domain of territorial food systems reveal that fine-tuning on pseudo-labeled data-obtained through cross-model agreement-yields better performance than relying solely on human-annotated data. The results also highlight the strengths and weaknesses of different segmentation strategies and confirm the importance of optimizing segmentation choices for NER in domain-specific low-resource settings. Code related to this work is available on GitHub11https://github.com/ibzodiaz/segmentation-strategies. Pape Ibrahima Thiam, Yohann Chasseray, Josiane Mothe, Mathieu Roche, Maguelonne Teisseire |
ICTAI | 4 |
| 2025 | Evaluation of geographical distortions in language modelsabstractGeographic bias in language models (LMs) is an underexplored dimension of model fairness, despite growing attention being given to other social biases. We investigate whether LMs provide equally accurate representations across all global regions and propose a benchmark of four indicators to detect undertrained and underperforming areas: (i) indirect assessment of geographic training data coverage via tokenizer analysis, (ii) evaluation of basic geographic knowledge, (iii) detection of geographic distortions, and (iv) visualization of performance disparities through maps. Applying this framework to ten widely used encoder- and decoder-based models, we find systematic overrepresentation of Western countries and consistent underrepresentation of several African, Eastern European, and Middle Eastern regions, leading to measurable performance gaps. We further analyse the impact of these biases on downstream tasks, particularly in crisis response, and show that regions most vulnerable to natural disasters are often those with poorer LM coverage. Our findings underscore the need for geographically balanced LMs to ensure equitable and effective global applications. Rémy Decoupes, Roberto Interdonato, Mathieu Roche, Maguelonne Teisseire, Sarah Valentin |
Mach. Learn. | 3 |
| 2024 | Evaluation of Geographical Distortions in Language Models
Rémy Decoupes, Roberto Interdonato, Mathieu Roche, Maguelonne Teisseire, Sarah Valentin |
DS (1) | 3 |
| 2024 | Semantically-Informed Domain Adaptation for Named Entity Recognition
Mariya Borovikova, Arnaud Ferré, Robert Bossy, Mathieu Roche, Claire Nedellec |
ISMIS | 4 |
| 2024 | MUST-AI: Multisource Surveillance Tool - Avian InfluenzaabstractThe multisource surveillance tool (MUST) is a platform for collecting, gathering, and visualizing different sources of information related to health events and highly pathogenic avian influenza in mammals (HPAIM). MUST-AI constitutes the first part of the MUST tool, which centralizes health information relating to cases of HPAIM since January 1, 2021, and comes from 3 different notification sources, an official notification source confirmed by public health institutions (i.e., WAHIS) and two other alternative unofficial sources that collect events from online media (PADI-web) and expert networks (ProMED). Owing to the use of natural language processing (NLP) algorithms, HPAIM events are represented on an interactive map associated with a graph that represents their distribution over a given time interval. This paper presents new tools and approaches for data fusion and experiments for selecting data to integrate into MUST that are related to HPAIM events. Carlène Trevennec, Pierre Pompidor, Samira Bououda, Julien Rabatel, Mathieu Roche |
KES | 5 |
| 2024 | EpidGPT: A Combined Strategy to Discriminate Between Redundant and New Information for Epidemiological Surveillance Systems
Edmond Odhiambo Menya, Mathieu Roche, Roberto Interdonato, Dickson Owuor |
NLDB (1) | 2 |
| 2024 | Explainable epidemiological thematic features for event based disease surveillanceabstractEvent based disease surveillance (EBS) systems are biosurveillance systems that have the ability to detect and alert on (re)-emerging infectious diseases by monitoring acute public or animal health event patterns from sources such as blogs, online news reports and curated expert accounts. These information rich sources, however, are largely unstructured text data requiring novel text mining techniques to achieve EBS goals such as epidemiological text classification. The main objective of this research was to improve epidemiological text classification by proposing a novel technique of enriching thematic features using a weak supervision approach. In our approach, we train and test a mixed domain language model named EpidBioELECTRA to first enrich thematic features which are then used to improve epidemiological text classification. We train EpidBioELECTRA on a large dataset which we create consisting of 70,700 annotated documents that includes 70,400 labelled thematic features. We empirically compare EpidBioELECTRA with both general purpose language models and domain specific language models in the task of epidemiological corpus classification. Our findings shows that epidemiological classification systems work best with language models pre-trained using both epidemiological and biomedical corpora with a continual pre-training strategy. EpidBioELECTRA improves epidemiological document classification by 19.2 F1 score points as compared to its vanilla implementation BioELECTRA. We observe this by the comparison of BioELECTRA verses EpidBioELECTRA on our most challenging dataset PADI-WebXL where our approach records 92.33 precision score, 94.62 recall score and 93.46 F1 score. We also experiment the impact of increasing context length of train documents in epidemiological document classification and found out that this improves the classification task by 7.79 F1 score points as recorded by EpidBioELECTRA’s performance. We also compute Almost Stochastic Order (ASO) scores to track EpidBioELECTRA’s statistical dominance. In addition, we carry out ablation studies on our proposed thematic feature enrichment approach using explainable AI techniques. We present explanations for the most critical thematic features and how they influence epidemiological classification task We found out that biomedical features (such as mentions of names of diseases and symptoms) are the most influential while spatio-temporal features (such as the mention of date of a given disease outbreak) are the least influential in epidemiological document classification. Our model can easily be extended to fit other domains. Edmond Odhiambo Menya, Roberto Interdonato, Dickson Owuor, Mathieu Roche |
Expert Syst. Appl. | 4 |
| 2024 | GeoNLPlify: A spatial data augmentation enhancing text classification for crisis monitoringabstractCrises such as natural disasters and public health emergencies generate vast amounts of text data, making it challenging to classify the information into relevant categories. Acquiring expert-labeled data for such scenarios can be difficult, leading to limited training datasets for text classification by fine-tuning BERT-like models. Unfortunately, traditional data augmentation techniques only slightly improve F1-scores. How can data augmentation be used to obtain better results in this applied domain? In this paper, using neural network explicability methods, we aim to highlight that fine-tuned BERT-like models on crisis corpora give too much importance to spatial information to make their predictions. This overfitting of spatial information limits their ability to generalize especially when the event which occurs in a place has evolved and changed since the training dataset has been built. To reduce this bias, we propose GeoNLPlify,1 a novel data augmentation technique that leverages spatial information to generate new labeled data for text classification related to crises. Our approach aims to address overfitting without necessitating modifications to the underlying model architecture, distinguishing it from other prevalent methods employed to combat overfitting. Our results show that GeoNLPlify significantly improves F1-scores, demonstrating the potential of the spatial information for data augmentation for crisis-related text classification tasks. In order to evaluate the contribution of our method, GeoNLPlify is applied to three public datasets (PADI-web, CrisisNLP and SST2) and compared with classical natural language processing data augmentations. Rémy Decoupes, Mathieu Roche, Maguelonne Teisseire |
Intell. Data Anal. | 2 |
| 2024 | How can text mining improve the explainability of Food security situations?
Hugo Deléglise, Agnès Bégué, Roberto Interdonato, Elodie Maître d'Hôtel, Mathieu Roche, Maguelonne Teisseire |
J. Intell. Inf. Syst. | 5 |
| 2024 | Fusion of BERT embeddings and elongation-driven features
Abderrahim Rafae, Mohammed Erritali, Mathieu Roche |
Multim. Tools Appl. | 3 |
| 2023 | Unsupervised Key-Phrase Extraction from Long Texts with Multilingual Sentence Transformers
Hélder Dias, Artur Guimarães, Bruno Martins 0001, Mathieu Roche |
DS | 4 |
| 2023 | Towards a (Semi-)Automatic Urban Planning Rule Identification in the French LanguageabstractOne of the objectives of the Hérelles project is to find new mechanisms to facilitate the labeling (or semantization) of clusters from time series of satellite images. To achieve this, a proposed solution is to associate textual elements of interest with satellite data. The first step in this process consists of an automatic extraction of the information in the form of rules from urban planning documents composed in the French language. To address this challenge, we propose a method which is based on the multi-label classification of textual segments. It includes a special format for representing segments, in which each segment has a title and a subtitle. In addition, we propose a cascade approach aiming to deal with hierarchy of class labels. Finally, we develop several text augmentation techniques for the texts in French, which are able to improve the prediction results. We demonstrate experimentally that the resulting framework correctly classifies each type of segment with more than 90% of accuracy. Maksim Koptelov, Margaux Holveck, Bruno Crémilleux, Justine Reynaud, Mathieu Roche, Maguelonne Teisseire |
DSAA | 5 |
| 2023 | Could KeyWord Masking Strategy Improve Language Model?
Mariya Borovikova, Arnaud Ferré, Robert Bossy, Mathieu Roche, Claire Nedellec |
NLDB | 4 |
| 2022 | Mining News Articles Dealing with Food Security
Hugo Deléglise, Agnès Bégué, Roberto Interdonato, Elodie Maître d'Hôtel, Mathieu Roche, Maguelonne Teisseire |
ISMIS | 5 |
| 2022 | Enriching Epidemiological Thematic Features For Disease Surveillance Corpora ClassificationabstractWe present EpidBioBERT, a biosurveillance epidemiological document tagger for disease surveillance over PADI-Web system. Our model is trained on PADI-Web corpus which contains news articles on Animal Diseases Outbreak extracted from the web. We train a classifier to discriminate between relevant and irrelevant documents based on their epidemiological thematic feature content in preparation for further epidemiology information extraction. Our approach proposes a new way to perform epidemiological document classification by enriching epidemiological thematic features namely disease, host, location and date, which are used as inputs to our epidemiological document classifier. We adopt a pre-trained biomedical language model with a novel fine tuning approach that enriches these epidemiological thematic features. We find these thematic features rich enough to improve epidemiological document classification over a smaller data set than initially used in PADI-Web classifier. This improves the classifiers ability to avoid false positive alerts on disease surveillance systems. To further understand information encoded in EpidBioBERT, we experiment the impact of each epidemiology thematic feature on the classifier under ablation studies. We compare our biomedical pre-trained approach with a general language model based model finding that thematic feature embeddings pre-trained on general English documents are not rich enough for epidemiology classification task. Our model achieves an F1-score of 95.5% over an unseen test set, with an improvement of +5.5 points on F1-Score on the PADI-Web classifier with nearly half the training data set. Edmond Odhiambo Menya, Mathieu Roche, Roberto Interdonato, Dickson Owuor |
LREC | 2 |
| 2022 | How Textual Datasets Enhance the PADI-Web Tool?abstractInternational audience Mathieu Roche, Elena Arsevska, Sarah Valentin, Sylvain Falala, Julien Rabatel, Renaud Lancelot |
WEBIST | 1 |
| 2022 | Food security prediction from heterogeneous data combining machine and deep learning methods
Hugo Deléglise, Roberto Interdonato, Agnès Bégué, Elodie Maître d'Hôtel, Maguelonne Teisseire, Mathieu Roche |
Expert Syst. Appl. | 6 |
| 2022 | A new method to extract n-Ary relation instances from scientific documentsabstractA new method to extract knowledge structured as n-Ary relations from scientific articles is presented. We designed and assessed different approaches to reconstruct instances of n-Ary relations extracted from scientific articles in experimental domains, driven by an Ontological and Terminological Resource (OTR) and based on multi-feature representation of relations and their arguments. The proposed method starts with the identification of partial n-Ary relations in tables of scientific articles and then seeks to reconstruct them with argument instances in the article texts. Based on the so-called Scientific Publication Representation (SciPuRe) of textual arguments and Scientific Table Representation (STaRe) of n-Ary relations representation of an n-Ary relation called STaRe (Scientific Table Representation, originating from partial n-Ary relations extracted from document tables), here we propose and evaluate different approaches for the selection of textual argument instances that could complement partial n-Ary relations: structural, frequentist and word embedding models. The application domain concerns food packaging, especially composition and permeability data. Experiments were conducted on a corpus of 332 relation instances composed of 1547 arguments. Corpora of full and partial relations recognized in document tables and argument instances extracted from texts are available online. Different methods and strategies were measured with an f-score ranging from .34 to .74. These results show that n-Ary relations reconstruction approach depends on the number of selected candidate argument instances. Martin Lentschat, Patrice Buche, Juliette Dibie, Mathieu Roche |
Expert Syst. Appl. | 4 |
| 2021 | Integrating Textual Data into Heterogeneous Data Ingestion ProcessingabstractIn this abstract, two methods for integrating textual data and textual features into ingestion processing are summarized. The first method involves integrating all features, including textual features, into dedicated frameworks, such as by using machine learning techniques. In the second method, text and textual features, such as keywords, are used to explain results returned by heterogeneous data mining. In this context, it is necessary to link data (e.g., databases, images, etc.) and/or obtained results with textual data (e.g., documents and keywords). Mathieu Roche, Maguelonne Teisseire |
IEEE BigData | 1 |
| 2021 | WEIR-P: An Information Extraction Pipeline for the Wastewater Domain
Nanee Chahinian, Thierry Bonnabaud La Bruyère, Francesca Frontini, Carole Delenne, Marin Julien, Rachel Panckhurst, Mathieu Roche, Lucile Sautot, Laurent Deruelle, Maguelonne Teisseire |
RCIS | 7 |
| 2021 | KEOPS: Knowledge ExtractOr Pipeline System
Pierre Martin 0001, Thierry Helmer, Julien Rabatel, Mathieu Roche |
RCIS | 4 |
| 2020 | Could spatial features help the matching of textual data?abstractTextual data is available to an increasing extent through different media (social networks, companies data, data catalogues, etc.). New information extraction methods are needed since these new resources are highly heterogeneous. In this article, we propose a text matching process based on spatial features and assessed through heterogeneous textual data. Besides being compatible with heterogeneous data, it comprises two contributions: first, spatial information is extracted for comparison purposes and subsequently stored in a dedicated spatial textual representation (STR); and then two transformations are applied on STR to improve the spatial similarity estimation. This article outlines the proposed approach with new contributions: (i) a new geocoding methods using general co-occurrences between entities, and (ii) a thorough evaluation followed by (iii) an in-depth discussion. The results obtained on two corpora demonstrate that good spatial matches (≈ 80% precision on major criteria) can be obtained between the most similar STRs with further enhancement achieved via STR transformation. Jacques Fize, Mathieu Roche, Maguelonne Teisseire |
Intell. Data Anal. | 2 |
| 2018 | Environmental and Geo-Spatial Data Analytics (EnGeoData'2018)abstractThe following topics are dealt with: learning (artificial intelligence); data analysis; social networking (online); pattern classification; regression analysis; data mining; Internet; neural nets; graph theory; trees (mathematics). Maguelonne Teisseire, Mathieu Roche, Diana Inkpen |
DSAA | 2 |
| 2018 | Readitopics: Make Your Topic Models Readable via Labeling and BrowsingabstractReaditopics provides a new tool for browsing a textual corpus that showcases several recent work on topic labeling and topic coherence. We demonstrate the potential of these techniques to get a deeper understanding of the topics that structure different datasets. This tool is provided as a Web demo but it can be installed to experiment with your own dataset. It can be further extended to deal with more advanced topic modeling techniques. Julien Velcin, Antoine Gourru, Erwan Giry-Fouquet, Christophe Gravier, Mathieu Roche, Pascal Poncelet |
IJCAI | 5 |
| 2018 | How to combine spatio-temporal and thematic features in online news for enhanced animal disease surveillance?abstractEarly detection of outbreaks of emerging and exotic pathogens is one of the means of preventing the introduction of infectious diseases into unaffected territories. In that context, since 2016, the French Animal Health Epidemic Intelligence team (Veille Sanitaire Internationale, VSI) monitors the online news sources through a designated Platform for Automated extraction of Disease Information from the web (PADI-web). The tool automatically detects, categorizes, and extracts information from online news reports. We focus on the combination of epidemiological features (locations, dates, diseases and hosts) extracted from free text of the news in order to automatically find similarity between different news reports. We describe an original approach based on text mining and data fusion methods and evaluate its performance on a specialized corpus. Sarah Valentin, Renaud Lancelot, Mathieu Roche |
KES | 3 |
| 2018 | Automatic Identification of Research Fields in Scientific Papers
Eric Kergosien, Mohammad Amin Farvardin, Maguelonne Teisseire, Marie-Noëlle Bessagnet, Joachim Schöpfel, Stéphane Chaudiron, Bernard Jacquemin, Annig Lacayrelle, Mathieu Roche, Christian Sallaberry, Jean-Philippe Tonneau |
LREC | 9 |
| 2018 | [Demo] Integration of Text- and Web-Mining Results in EpidVis
Samiha Fadloun, Arnaud Sallaberry, Alizé Mercier, Elena Arsevska, Pascal Poncelet, Mathieu Roche |
NLDB | 6 |
| 2018 | Gemedoc: A Text Similarity Annotation Platform
Jacques Fize, Mathieu Roche, Maguelonne Teisseire |
NLDB | 2 |
| 2018 | United We Stand: Using Multiple Strategies for Topic Labeling
Antoine Gourru, Julien Velcin, Mathieu Roche, Christophe Gravier, Pascal Poncelet |
NLDB | 3 |
| 2018 | EpidNews: An Epidemiological News Explorer for Monitoring Animal DiseasesabstractIn the recent years, there has been a massive increase in the amount of data being produced about human and animal health related events. Epidemiologists have to analyze this epidemiological data on a regular basis. They use this spatio-temporal information, most of which is shared online, to detect, observe, and track geographic locations of disease outbreaks over time. Unfortunately, manually retrieving the data from a website like Google News and then deriving sensible insights from the huge dataset consumes a lot of time and effort. We present EpidNews, a new visual analytics tool that helps to visualize and explore epidemiological news data for animals. The tool uses several views depicting various levels of abstraction, which helps fulfill almost all the data analysis requirements of epidemiologists. We also present the case study of an epidemiology expert, wherein she assesses the usability and productivity of EpidNews by using the tool in her daily work. Rohan Goel, Samiha Fadloun, Sarah Valentin, Arnaud Sallaberry, Mathieu Roche, Pascal Poncelet |
VINCI | 5 |
| 2018 | Spatial Information Extraction from Short Messages
Sarah Zenasni, Eric Kergosien, Mathieu Roche, Maguelonne Teisseire |
Expert Syst. Appl. | 3 |
| 2018 | The role of location and social strength for friendship prediction in location-based social networksabstractInternational audience Jorge Carlos Valverde-Rebaza, Mathieu Roche, Pascal Poncelet, Alneu de Andrade Lopes |
Inf. Process. Manag. | 2 |
| 2018 | A novel framework for biomedical entity sense induction
Juan Antonio Lossio-Ventura, Jiang Bian 0001, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire |
J. Biomed. Informatics | 4 |
| 2017 | Node Overlap Removal for 1D Graph LayoutabstractEnergy based algorithms are powerful techniques for laying out graphs. They tend to generate aesthetically pleasing graph embeddings, exhibiting symmetries and community structures. When dealing with large graphs, an important drawback of these algorithms is to produce embeddings where many nodes overlap, leading to cluttering issues. While several approaches have been proposed for node overlap removal on 2D graph layouts, to the best of our knowledge, there is no work dedicated to 1D graph layouts. In this paper, we first define 4 requirements for 1D graph node overlap removal. Then, we propose a O(|V|log(|V|)) time algorithm meeting these requirements. We illustrate our approach with two case studies based on arc diagrams where nodes are positioned by applying a MDS technique to highlight community structures. Finally, we compare our technique with alternatives from 2D graph techniques, and a discussion highlights some properties of the results. Samiha Fadloun, Pascal Poncelet, Julien Rabatel, Mathieu Roche, Arnaud Sallaberry |
IV | 4 |
| 2017 | Towards a Bio-inspired Approach to Match Heterogeneous DocumentsabstractMatching heterogeneous text documents coming from different sources means matching data extracted from these documents, generally structured in the form of vectors. The accuracy of matching directly depends on the right choice of the content of these vectors. That's why we need to select the best features. In this paper, we present a new approach to select the minimum set of features that represents the semantics of a set of text documents, using a quantum inspired genetic algorithm. Among different Vs characterizing the big data we focus on 'Variety' criterion, therefore, we used three sets of different sources that are semantically similar to retrieve their best features which describe the semantics of the corpus. In the matching phase, our approach shows significant improvement compared with the classic 'Bag-of-words' approach. Nourelhouda Yahi, Hacene Belhadef, Mathieu Roche, Amer Draa |
WEBIST | 3 |
| 2017 | Xart: Discovery of correlated arguments of n-ary relations in text
Soumia Lilia Berrahou, Patrice Buche, Juliette Dibie, Mathieu Roche |
Expert Syst. Appl. | 4 |
| 2016 | MultiLingMine 2016: Modeling, Learning and Mining for Cross/Multilinguality
Dino Ienco, Mathieu Roche, Salvatore Romeo, Paolo Rosso, Andrea Tagarelli |
ECIR | 2 |
| 2016 | A Way to Automatically Enrich Biomedical OntologiesabstractBiomedical ontologies play an important role for information extraction in the biomedical domain. We present a workflow for updating automatically biomedical ontologies, composed of four steps. We detail two contributions concerning the concept extraction and semantic linkage of extracted terminology. Juan Antonio Lossio-Ventura, Mathieu Roche, Clément Jonquet, Maguelonne Teisseire |
EDBT | 2 |
| 2016 | Exploiting social and mobility patterns for friendship prediction in location-based social networksabstractLink prediction is a “hot topic” in network analysis and has been largely used for friendship recommendation in social networks. With the increased use of location-based services, it is possible to improve the accuracy of link prediction methods by using the mobility of users. The majority of the link prediction methods focus on the importance of location for their visitors, disregarding the strength of relationships existing between these visitors. We, therefore, propose three new methods for friendship prediction by combining, efficiently, social and mobility patterns of users in location-based social networks (LBSNs). Experiments conducted on real-world datasets demonstrate that our proposals achieve a competitive performance with methods from the literature and, in most of the cases, outperform them. Moreover, our proposals use less computational resources by reducing considerably the number of irrelevant predictions, making the link prediction task more efficient and applicable for real world applications. Jorge Carlos Valverde-Rebaza, Mathieu Roche, Pascal Poncelet, Alneu de Andrade Lopes |
ICPR | 2 |
| 2016 | Monitoring Disease Outbreak Events on the Web Using Text-mining Approach and Domain Expert Knowledge
Elena Arsevska, Mathieu Roche, Sylvain Falala, Renaud Lancelot, David Chavernac, Pascal Hendrikx, Barbara Dufour |
LREC | 2 |
| 2016 | Integration of Lexical and Semantic Knowledge for Sentiment Analysis in SMS
Wejdene Khiari, Mathieu Roche, Asma Bouhafs Hafsia |
LREC | 2 |
| 2016 | Automatic Biomedical Term Polysemy Detection
Juan Antonio Lossio-Ventura, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire |
LREC | 3 |
| 2016 | Extracting new spatial entities and relations from short messages
Sarah Zenasni, Eric Kergosien, Mathieu Roche, Maguelonne Teisseire |
MEDES | 3 |
| 2016 | Biomedical term extraction: overview and a new methodology
Juan Antonio Lossio-Ventura, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire |
Inf. Retr. J. | 3 |
| 2015 | SentiCompass: Interactive visualization for exploring and comparing the sentiments of time-varying twitter dataabstractIn this work, we introduce SentiCompass for exploring and comparing the sentiments of time-varying Twitter data. Our visualization design combines 2D psychology model of affect (i.e. emotion) with a time tunnel representation. To illustrate our visualization design, two case studies are conducted. They demonstrate the effectiveness of SentiCompass in achieving various tasks related to temporal sentiment and affective analysis of tweets. The interactive demo of our system is available at: http://youtu.be/ZaMF6VNO7tA Florence Ying Wang, Arnaud Sallaberry, Karsten Klein 0001, Masahiro Takatsuka, Mathieu Roche |
PacificVis | 5 |
| 2015 | Discovering Types of Spatial Relations with a Text Mining Approach
Sarah Zenasni, Eric Kergosien, Mathieu Roche, Maguelonne Teisseire |
ISMIS | 3 |
| 2015 | Application of natural language to information systems (NLDB'14)
Elisabeth Métais, Mathieu Roche, Maguelonne Teisseire |
Data Knowl. Eng. | 2 |
| 2015 | Recognition of logical units in log filesabstractWith the development of new technologies more and more information is stored in log files. Analyzing such logs can be very useful for the decision maker. One of the probably best known example is the Web log file analysis where lots of efficient tool Hassan Saneifar, Stéphane Bonniol, Pascal Poncelet, Mathieu Roche |
Intell. Data Anal. | 4 |
| 2014 | Looking for Opinion in Land-Use Planning Corpora
Eric Kergosien, Cédric Lopez, Mathieu Roche, Maguelonne Teisseire |
CICLing (2) | 3 |
| 2014 | Integration of linguistic and web information to improve biomedical terminology extractionabstractComprehensive terminology is essential for a community to describe, exchange, and retrieve data. In multiple domain, the explosion of text data produced has reached a level for which automatic terminology extraction and enrichment is mandatory. Automatic Term Extraction (or Recognition) methods use natural language processing to do so. Methods featuring linguistic and statistical aspects as often proposed in the literature, solve some problems related to term extraction as low frequency, complexity of the multi-word term extraction, human effort to validate candidate terms. In contrast, we present two new measures for extracting and ranking muli-word terms from domain-specific corpora, covering the all mentioned problems. In addition we demonstrate how the use of the Web to evaluate the significance of a multi-word term candidate, helps us to outperform precision results obtain on the biomedical GENIA corpus with previous reported measures such as C-value. Juan Antonio Lossio-Ventura, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire |
IDEAS | 3 |
| 2014 | Classification of Small Datasets: Why Using Class-Based Weighting Measures?
Flavien Bouillot, Pascal Poncelet, Mathieu Roche |
ISMIS | 3 |
| 2014 | Towards Electronic SMS Dictionary Construction: An Alignment-based Approach
Cédric Lopez, Reda Bestandji, Mathieu Roche, Rachel Panckhurst |
LREC | 3 |
| 2014 | How can catchy titles be generated without loss of informativeness?
Cédric Lopez, Violaine Prince, Mathieu Roche |
Expert Syst. Appl. | 3 |
| 2014 | Are opinions expressed in land-use planning documents?abstractA great deal of research on information extraction from textual datasets has been performed in specific data contexts, such as movie reviews, commercial product evaluations, campaign speeches, etc. In this paper, we raise the question on how appropriate these methods are for documents related to land-use planning. The kind of information sought concerns the stakeholders, sentiments, geographic information, and everything else related to the territory. However, it is extremely challenging to link sentiments to the three dimensions that constitute geographic information (location, time, and theme). After highlighting the limitations of existing proposals and discussing issues related to textual data, we present a method called OPILAND (OPinion mIning from LAND-use planning documents) designed to semi-automatically mine opinions related to named-entities in specialized contexts. Experiments are conducted on a Thau lagoon dataset (France), and then applied on three datasets that are related to different areas in order to highlight the relevance and the broader applications of our proposal. Eric Kergosien, Bernard Laval, Mathieu Roche, Maguelonne Teisseire |
Int. J. Geogr. Inf. Sci. | 3 |
| 2013 | Approaches of Anonymisation of an SMS Corpus
Namrata Patel, Pierre Accorsi, Diana Inkpen, Cédric Lopez, Mathieu Roche |
CICLing (1) | 5 |
| 2013 | GenDesc: A Partial Generalization of Linguistic Features for Text Classification
Guillaume Tisserant, Violaine Prince, Mathieu Roche |
NLDB | 3 |
| 2012 | An Unsupervised Framework for Topological Relations Extraction from Geographic Documents
Corrado Loglisci, Dino Ienco, Mathieu Roche, Maguelonne Teisseire, Donato Malerba |
DEXA (2) | 3 |
| 2012 | Just Title It! (by an Online Application)
Cédric Lopez, Violaine Prince, Mathieu Roche |
EACL | 3 |
| 2012 | NOMIT: Automatic Titling by Nominalizing
Cédric Lopez, Violaine Prince, Mathieu Roche |
HLT-NAACL | 3 |
| 2012 | Lexical Knowledge Acquisition Using Spontaneous Descriptions in Texts
Augusta Mela, Mathieu Roche, Mohamed el Amine Bekhtaoui |
NLDB | 2 |
| 2012 | A hybrid approach to managing job offers and candidates
Rémy Kessler, Nicolas Béchet, Mathieu Roche, Juan-Manuel Torres-Moreno, Marc El-Bèze |
Inf. Process. Manag. | 3 |
| 2011 | Towards an On-Line Analysis of Tweets Processing
Sandra Bringay, Nicolas Béchet, Flavien Bouillot, Pascal Poncelet, Mathieu Roche, Maguelonne Teisseire |
DEXA (2) | 5 |
| 2011 | Towards an Automatic Characterization of Criteria
Benjamin Duthil, François Trousset, Mathieu Roche, Gérard Dray, Michel Plantié, Jacky Montmain, Pascal Poncelet |
DEXA (1) | 3 |
| 2011 | How Statistical Information from the Web can Help Identify Named Entities
Mathieu Roche |
WEBIST | 1 |
| 2011 | Sequential patterns mining and gene sequence visualization to discover novelty from microarray data
Arnaud Sallaberry, Nicolas Pecheur, Sandra Bringay, Mathieu Roche, Maguelonne Teisseire |
J. Biomed. Informatics | 4 |
| 2010 | Extraction of unexpected sentences: A sentiment classification assessed approachabstractSentiment classification in text documents is an active data mining research topic in opinion retrieval and analysis. Different from previous studies concentrating on the development of effective classifiers, in this paper, we focus on the extraction Dominique Li, Anne Laurent, Pascal Poncelet, Mathieu Roche |
Intell. Data Anal. | 4 |
| 2009 | Terminology Extraction from Log Files
Hassan Saneifar, Stéphane Bonniol, Anne Laurent, Pascal Poncelet, Mathieu Roche |
DEXA | 5 |
| 2009 | Towards the Selection of Induced Syntactic Relations
Nicolas Béchet, Mathieu Roche, Jacques Chauché |
ECIR | 2 |
| 2009 | How to Rank Terminology Extracted by Exterlog
Hassan Saneifar, Stéphane Bonniol, Anne Laurent, Pascal Poncelet, Mathieu Roche |
IC3K | 5 |
| 2009 | Job Offer Management: How Improve the Ranking of Candidates
Rémy Kessler, Nicolas Béchet, Juan-Manuel Torres-Moreno, Mathieu Roche, Marc El-Bèze |
ISMIS | 4 |
| 2008 | Is a Voting Approach Accurate for Opinion Mining?
Michel Plantié, Mathieu Roche, Gérard Dray, Pascal Poncelet |
DaWaK | 2 |
| 2008 | Extraction of Opposite Sentiments in Classified Free Format Text Reviews
Dominique Li, Anne Laurent, Mathieu Roche, Pascal Poncelet |
DEXA | 3 |
| 2007 | A Context-based Measure for Discovering Approximate Semantic Matching between Schema Elements
Fabien Duchateau, Zohra Bellahsene, Mathieu Roche |
RCIS | 3 |
| 2007 | A Flexible Approach Based on the user Preferences for Schema Matching
Wided Guédria, Zohra Bellahsene, Mathieu Roche |
RCIS | 3 |