Maguelonne Teisseire

dblp:t/MTeisseire · DBLP profile ↗
← Back
92ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-9313-6414ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 49 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorTheory of computation · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Enhancing Domain-Specific Named Entity Recognition via Segmentation and Pseudo-Labeled Annotation
abstract
Named Entity Recognition (NER) in specialized domains poses major challenges due to the scarcity of annotated data and the limitations of existing models. One challenge lies in handling long documents, which often requires segmenting the text into smaller chunks that fit within the model's input window. Although several segmentation strategies exist, their impact on NER performance in new domains remains underexplored. Another challenge is domain adaptation: while openschema NER models can perform reasonably well under zero-shot settings, they often struggle in highly specialized contexts without additional supervision. To address these issues, we combine segmentation strategy selection with a semi-supervised finetuning pipeline based on pseudo-labeled annotations. First, we compare four segmentation strategies to identify which offers the best trade-off between precision and recall in zero-shot settings. Then, we fine-tune two open-schema NER models-GLiNER and NuNER-first on a manually annotated dataset, and subsequently on a large pseudo-labeled dataset built from model agreement. Both models are evaluated on a held-out test set and on the full manually annotated corpus. Experiments on French-language documents specifically focused on the underexplored domain of territorial food systems reveal that fine-tuning on pseudo-labeled data-obtained through cross-model agreement-yields better performance than relying solely on human-annotated data. The results also highlight the strengths and weaknesses of different segmentation strategies and confirm the importance of optimizing segmentation choices for NER in domain-specific low-resource settings. Code related to this work is available on GitHub11https://github.com/ibzodiaz/segmentation-strategies.
Pape Ibrahima Thiam, Yohann Chasseray, Josiane Mothe, Mathieu Roche, Maguelonne Teisseire
ICTAI5
2025 Evaluation of geographical distortions in language models
abstract
Geographic bias in language models (LMs) is an underexplored dimension of model fairness, despite growing attention being given to other social biases. We investigate whether LMs provide equally accurate representations across all global regions and propose a benchmark of four indicators to detect undertrained and underperforming areas: (i) indirect assessment of geographic training data coverage via tokenizer analysis, (ii) evaluation of basic geographic knowledge, (iii) detection of geographic distortions, and (iv) visualization of performance disparities through maps. Applying this framework to ten widely used encoder- and decoder-based models, we find systematic overrepresentation of Western countries and consistent underrepresentation of several African, Eastern European, and Middle Eastern regions, leading to measurable performance gaps. We further analyse the impact of these biases on downstream tasks, particularly in crisis response, and show that regions most vulnerable to natural disasters are often those with poorer LM coverage. Our findings underscore the need for geographically balanced LMs to ensure equitable and effective global applications.
Rémy Decoupes, Roberto Interdonato, Mathieu Roche, Maguelonne Teisseire, Sarah Valentin
Mach. Learn.4
2024 Evaluation of Geographical Distortions in Language Models
Rémy Decoupes, Roberto Interdonato, Mathieu Roche, Maguelonne Teisseire, Sarah Valentin
DS (1)4
2024 GeoNLPlify: A spatial data augmentation enhancing text classification for crisis monitoring
abstract
Crises such as natural disasters and public health emergencies generate vast amounts of text data, making it challenging to classify the information into relevant categories. Acquiring expert-labeled data for such scenarios can be difficult, leading to limited training datasets for text classification by fine-tuning BERT-like models. Unfortunately, traditional data augmentation techniques only slightly improve F1-scores. How can data augmentation be used to obtain better results in this applied domain? In this paper, using neural network explicability methods, we aim to highlight that fine-tuned BERT-like models on crisis corpora give too much importance to spatial information to make their predictions. This overfitting of spatial information limits their ability to generalize especially when the event which occurs in a place has evolved and changed since the training dataset has been built. To reduce this bias, we propose GeoNLPlify,1 a novel data augmentation technique that leverages spatial information to generate new labeled data for text classification related to crises. Our approach aims to address overfitting without necessitating modifications to the underlying model architecture, distinguishing it from other prevalent methods employed to combat overfitting. Our results show that GeoNLPlify significantly improves F1-scores, demonstrating the potential of the spatial information for data augmentation for crisis-related text classification tasks. In order to evaluate the contribution of our method, GeoNLPlify is applied to three public datasets (PADI-web, CrisisNLP and SST2) and compared with classical natural language processing data augmentations.
Rémy Decoupes, Mathieu Roche, Maguelonne Teisseire
Intell. Data Anal.3
2024 How can text mining improve the explainability of Food security situations?
Hugo Deléglise, Agnès Bégué, Roberto Interdonato, Elodie Maître d'Hôtel, Mathieu Roche, Maguelonne Teisseire
J. Intell. Inf. Syst.6
2023 Towards a (Semi-)Automatic Urban Planning Rule Identification in the French Language
abstract
One of the objectives of the Hérelles project is to find new mechanisms to facilitate the labeling (or semantization) of clusters from time series of satellite images. To achieve this, a proposed solution is to associate textual elements of interest with satellite data. The first step in this process consists of an automatic extraction of the information in the form of rules from urban planning documents composed in the French language. To address this challenge, we propose a method which is based on the multi-label classification of textual segments. It includes a special format for representing segments, in which each segment has a title and a subtitle. In addition, we propose a cascade approach aiming to deal with hierarchy of class labels. Finally, we develop several text augmentation techniques for the texts in French, which are able to improve the prediction results. We demonstrate experimentally that the resulting framework correctly classifies each type of segment with more than 90% of accuracy.
Maksim Koptelov, Margaux Holveck, Bruno Crémilleux, Justine Reynaud, Mathieu Roche, Maguelonne Teisseire
DSAA6
2022 Mining News Articles Dealing with Food Security
Hugo Deléglise, Agnès Bégué, Roberto Interdonato, Elodie Maître d'Hôtel, Mathieu Roche, Maguelonne Teisseire
ISMIS6
2022 Food security prediction from heterogeneous data combining machine and deep learning methods
Hugo Deléglise, Roberto Interdonato, Agnès Bégué, Elodie Maître d'Hôtel, Maguelonne Teisseire, Mathieu Roche
Expert Syst. Appl.5
2021 Integrating Textual Data into Heterogeneous Data Ingestion Processing
abstract
In this abstract, two methods for integrating textual data and textual features into ingestion processing are summarized. The first method involves integrating all features, including textual features, into dedicated frameworks, such as by using machine learning techniques. In the second method, text and textual features, such as keywords, are used to explain results returned by heterogeneous data mining. In this context, it is necessary to link data (e.g., databases, images, etc.) and/or obtained results with textual data (e.g., documents and keywords).
Mathieu Roche, Maguelonne Teisseire
IEEE BigData2
2021 WEIR-P: An Information Extraction Pipeline for the Wastewater Domain
Nanee Chahinian, Thierry Bonnabaud La Bruyère, Francesca Frontini, Carole Delenne, Marin Julien, Rachel Panckhurst, Mathieu Roche, Lucile Sautot, Laurent Deruelle, Maguelonne Teisseire
RCIS10
2020 Could spatial features help the matching of textual data?
abstract
Textual data is available to an increasing extent through different media (social networks, companies data, data catalogues, etc.). New information extraction methods are needed since these new resources are highly heterogeneous. In this article, we propose a text matching process based on spatial features and assessed through heterogeneous textual data. Besides being compatible with heterogeneous data, it comprises two contributions: first, spatial information is extracted for comparison purposes and subsequently stored in a dedicated spatial textual representation (STR); and then two transformations are applied on STR to improve the spatial similarity estimation. This article outlines the proposed approach with new contributions: (i) a new geocoding methods using general co-occurrences between entities, and (ii) a thorough evaluation followed by (iii) an in-depth discussion. The results obtained on two corpora demonstrate that good spatial matches (≈ 80% precision on major criteria) can be obtained between the most similar STRs with further enhancement achieved via STR transformation.
Jacques Fize, Mathieu Roche, Maguelonne Teisseire
Intell. Data Anal.3
2018 Environmental and Geo-Spatial Data Analytics (EnGeoData'2018)
abstract
The following topics are dealt with: learning (artificial intelligence); data analysis; social networking (online); pattern classification; regression analysis; data mining; Internet; neural nets; graph theory; trees (mathematics).
Maguelonne Teisseire, Mathieu Roche, Diana Inkpen
DSAA1
2018 Automatic Identification of Research Fields in Scientific Papers
Eric Kergosien, Mohammad Amin Farvardin, Maguelonne Teisseire, Marie-Noëlle Bessagnet, Joachim Schöpfel, Stéphane Chaudiron, Bernard Jacquemin, Annig Lacayrelle, Mathieu Roche, Christian Sallaberry, Jean-Philippe Tonneau
LREC3
2018 Gemedoc: A Text Similarity Annotation Platform
Jacques Fize, Mathieu Roche, Maguelonne Teisseire
NLDB3
2018 Spatial Information Extraction from Short Messages
Sarah Zenasni, Eric Kergosien, Mathieu Roche, Maguelonne Teisseire
Expert Syst. Appl.4
2018 A novel framework for biomedical entity sense induction
Juan Antonio Lossio-Ventura, Jiang Bian 0001, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire
J. Biomed. Informatics5
2016 A Way to Automatically Enrich Biomedical Ontologies
abstract
Biomedical ontologies play an important role for information extraction in the biomedical domain. We present a workflow for updating automatically biomedical ontologies, composed of four steps. We detail two contributions concerning the concept extraction and semantic linkage of extracted terminology.
Juan Antonio Lossio-Ventura, Mathieu Roche, Clément Jonquet, Maguelonne Teisseire
EDBT4
2016 Automatic Biomedical Term Polysemy Detection
Juan Antonio Lossio-Ventura, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire
LREC4
2016 Extracting new spatial entities and relations from short messages
Sarah Zenasni, Eric Kergosien, Mathieu Roche, Maguelonne Teisseire
MEDES4
2016 Spatio-sequential patterns mining: Beyond the boundaries
abstract
Data mining methods extract knowledge from huge amounts of data. Recently with the explosion of mobile technologies, a new type of data appeared. The resulting databases can be described as spatiotemporal data in which spatial information (e.g., the location of an event) and temporal information (e .g., the date of the event) are included. In this article, we focus on spatiotemporal patterns extraction from this kind of databases. These patterns can be considered as sequences representing changes of events localized in areas and its near surrounding over time. Two algorithms are proposed to tackle this problem: the first one uses \emph{a priori} strategy and the second one is based on pattern-growth approach. We have applied our generic method on two different real datasets related to: 1) pollution of rivers in France; and 2) monitoring of dengue epidemics in New Caledonia. Additionally, experiments on synthetic data have been conducted to measure the performance of the proposed algorithms.
Hugo Alatrista Salas, Sandra Bringay, Frédéric Flouvat, Nazha Selmaoui-Folcher, Maguelonne Teisseire
Intell. Data Anal.5
2016 Biomedical term extraction: overview and a new methodology
Juan Antonio Lossio-Ventura, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire
Inf. Retr. J.4
2015 Discovering Types of Spatial Relations with a Text Mining Approach
Sarah Zenasni, Eric Kergosien, Mathieu Roche, Maguelonne Teisseire
ISMIS4
2015 Mining Multi-Relational Gradual Patterns
abstract
Gradual patterns highlight covariations of attributes of the form “The more/less X, the more/less Y”. Their usefulness in several applications has recently stimulated the synthesis of several algorithms for their automated discovery from large datasets. However, existing techniques require all the interesting data to be in a single database relation or table. This paper extends the notion of gradual pattern to the case in which the co-variations are possibly expressed between attributes of different database relations. The interestingness measure for this class of “relational gradual patterns” is defined on the basis of both Kendall's τ and gradual supports. Moreover, this paper proposes two algorithms, named τRGP Miner and gRGP Miner, for the discovery of relational gradual rules. Three pruning strategies to reduce the search space are proposed. The efficiency of the algorithms is empirically validated, and the usefulness of relational gradual patterns is proved on some real-world databases.
NhatHai Phan, Dino Ienco, Donato Malerba, Pascal Poncelet, Maguelonne Teisseire
SDM5
2015 Application of natural language to information systems (NLDB'14)
Elisabeth Métais, Mathieu Roche, Maguelonne Teisseire
Data Knowl. Eng.3
2015 Spatio-temporal data classification through multidimensional sequential patterns: Application to crop mapping in complex landscape
Yoann Pitarch, Dino Ienco, Elodie Vintrou, Agnès Bégué, Anne Laurent, Pascal Poncelet, Michel Sala, Maguelonne Teisseire
Eng. Appl. Artif. Intell.8
2015 Mining closed partially ordered patterns, a new optimized algorithm
Mickaël Fabrègue, Agnès Braud, Sandra Bringay, Florence Le Ber, Maguelonne Teisseire
Knowl. Based Syst.5
2014 Looking for Opinion in Land-Use Planning Corpora
Eric Kergosien, Cédric Lopez, Mathieu Roche, Maguelonne Teisseire
CICLing (2)4
2014 Integration of linguistic and web information to improve biomedical terminology extraction
abstract
Comprehensive terminology is essential for a community to describe, exchange, and retrieve data. In multiple domain, the explosion of text data produced has reached a level for which automatic terminology extraction and enrichment is mandatory. Automatic Term Extraction (or Recognition) methods use natural language processing to do so. Methods featuring linguistic and statistical aspects as often proposed in the literature, solve some problems related to term extraction as low frequency, complexity of the multi-word term extraction, human effort to validate candidate terms. In contrast, we present two new measures for extracting and ranking muli-word terms from domain-specific corpora, covering the all mentioned problems. In addition we demonstrate how the use of the Web to evaluate the significance of a multi-word term candidate, helps us to outperform precision results obtain on the biomedical GENIA corpus with previous reported measures such as C-value.
Juan Antonio Lossio-Ventura, Clément Jonquet, Mathieu Roche, Maguelonne Teisseire
IDEAS4
2014 Evaluation of fusion methods for crop monitoring purposes
abstract
In this contribution we present a local evaluation procedure of Landsat-MODIS fusion methods for crop monitoring purposes. Two fusion methods are applied to obtain a two-year time series of Landsat-resolution images. The validation is applied at pixel level in order to analyze if the simulated images are capable of unmixing coarse-resolution pixels and obtaining an accurate temporal profile of the high-resolution pixels near the boundaries between two fields. The experiment has been conducted in two neighbor fields and results have shown that the temporal profile of these high-resolution pixels agrees with the temporal profile of a pure coarse-resolution pixel in the corresponding field. They highlight that the simulated images allow identifying the membership of the high-resolution simulated pixels to one field or the neighbor one.
Mar Bisquert, Agnès Bégué, Pascal Poncelet, Maguelonne Teisseire
IGARSS4
2014 Monitoring the phenology of mediterranean natural habitats with multispectral sensors - An analysis based on multiseasonal field spectra
abstract
Due to their high degree of vegetation heterogeneity, fragmentation and biodiversity, Mediterranean natural habitats are difficult to assess and monitor with in-situ observations solely. Together with standardized ground plots and regular in-situ measurements, remote sensing is a powerful device that can contribute to a better understanding of the diversity of natural and semi-natural habitats and to monitor their phenology. In this paper, we implemented a systematic test of the suitability of multiseasonal remote sensing data for monitoring the phenological variations of natural habitats in a Mediterranean landscape. Six multispectral sensor signals were simulated for comparison based on their spectral response curves and in-situ averaged spectra collected at monthly intervals between February and October 2013 (IKONOS, Landsat 5 TM, Landsat 8, Pléiades, Sentinel-2, and Worldview-2). The simulations and comparisons performed in this study showed that Sentinel-2 sensor has the higher sensitivity to the variations in the coverage of photosynthetic vegetation, thus offering interesting perspectives for operational monitoring of natural habitats.
Christina Corbane, Fabio Guttler, Samuel Alleaume, Dino Ienco, Maguelonne Teisseire
IGARSS5
2014 Exploring high repetitivity remote sensing time series for mapping and monitoring natural habitats - A new approach combining OBIA and k-partite graphs
abstract
High repetitivity remote sensing could substantially improve natural habitats monitoring and mapping in the next years. However, dense time series of satellite images require new processing methodologies. In this paper we proposed an approach which combines Object Based Image Analysis (OBIA) and k-partite graphs for detecting spatiotemporal evolutions in a Mediterranean protected site composed of several types of natural and semi-natural habitats. The method was applied over a recent dataset (SPOT4 Take-5) specially conceived to simulate the acquisition frequency of the future Sentinel-2 satellites. The results indicate our method is capable to synthesize complex spatiotemporal evolutions in a semi-automatic way, therefore offering a new tool to analyze high repetitivity satellite time series.
Fabio Guttler, Samuel Alleaume, Christina Corbane, Dino Ienco, Jordi Nin, Pascal Poncelet, Maguelonne Teisseire
IGARSS7
2014 Soft Fusion of Heterogeneous Image Time Series
Mar Bisquert, Gloria Bordogna, Mirco Boschetti, Pascal Poncelet, Maguelonne Teisseire
IPMU (1)5
2014 Towards the Use of Sequential Patterns for Detection and Characterization of Natural and Agricultural Areas
Fabio Guttler, Dino Ienco, Maguelonne Teisseire, Jordi Nin, Pascal Poncelet
IPMU (1)3
2014 Are opinions expressed in land-use planning documents?
abstract
A great deal of research on information extraction from textual datasets has been performed in specific data contexts, such as movie reviews, commercial product evaluations, campaign speeches, etc. In this paper, we raise the question on how appropriate these methods are for documents related to land-use planning. The kind of information sought concerns the stakeholders, sentiments, geographic information, and everything else related to the territory. However, it is extremely challenging to link sentiments to the three dimensions that constitute geographic information (location, time, and theme). After highlighting the limitations of existing proposals and discussing issues related to textual data, we present a method called OPILAND (OPinion mIning from LAND-use planning documents) designed to semi-automatically mine opinions related to named-entities in specialized contexts. Experiments are conducted on a Thau lagoon dataset (France), and then applied on three datasets that are related to different areas in order to highlight the relevance and the broader applications of our proposal.
Eric Kergosien, Bernard Laval, Mathieu Roche, Maguelonne Teisseire
Int. J. Geogr. Inf. Sci.4
2014 A contribution to the discovery of multidimensional patterns in healthcare trajectories
Elias Egho, Nicolas Jay, Chedy Raïssi, Dino Ienco, Pascal Poncelet, Maguelonne Teisseire, Amedeo Napoli
J. Intell. Inf. Syst.6
2013 Knowledge-Free Table Summarization
Dino Ienco, Yoann Pitarch, Pascal Poncelet, Maguelonne Teisseire
DaWaK4
2013 A Density-Based Backward Approach to Isolate Rare Events in Large-Scale Applications
Enikö Székely, Pascal Poncelet, Florent Masseglia, Maguelonne Teisseire, Renaud Cezar
Discovery Science4
2013 OrderSpan: Mining Closed Partially Ordered Patterns
Mickaël Fabrègue, Agnès Braud, Sandra Bringay, Florence Le Ber, Maguelonne Teisseire
IDA5
2013 Mining Representative Movement Patterns through Compression
NhatHai Phan, Dino Ienco, Pascal Poncelet, Maguelonne Teisseire
PAKDD (1)4
2012 Mining Fuzzy Moving Object Clusters
NhatHai Phan, Dino Ienco, Pascal Poncelet, Maguelonne Teisseire
ADMA4
2012 An Unsupervised Framework for Topological Relations Extraction from Geographic Documents
Corrado Loglisci, Dino Ienco, Mathieu Roche, Maguelonne Teisseire, Donato Malerba
DEXA (2)4
2012 Including Spatial Relations and Scales within Sequential Pattern Extraction
Mickaël Fabrègue, Agnès Braud, Sandra Bringay, Florence Le Ber, Maguelonne Teisseire
Discovery Science5
2012 Mining time relaxed gradual moving object clusters
abstract
One of the objectives of spatio-temporal data mining is to analyze moving object datasets to exploit interesting patterns. Traditionally, existing methods only focus on an unchanged group of moving objects during a time period. Thus, they cannot capture object moving trends which can be very useful for better understanding the natural moving behavior in various real world applications. In this paper, we present a novel concept of "time relaxed gradual trajectory pattern", denoted real-Gpattern, which captures the object movement tendency. Additionally, we also propose an efficient algorithm, called ClusterGrowth, designed to extract the complete set of all interesting maximal real-Gpatterns. Conducted experiments on real and large synthetic datasets demonstrate the effectiveness, parameter sensitiveness and efficiency of our methods.
NhatHai Phan, Dino Ienco, Pascal Poncelet, Maguelonne Teisseire
SIGSPATIAL/GIS4
2012 GeT_Move: An Efficient and Unifying Spatio-temporal Pattern Mining Algorithm for Moving Objects
NhatHai Phan, Pascal Poncelet, Maguelonne Teisseire
IDA3
2012 The Pattern Next Door: Towards Spatio-sequential Pattern Discovery
Hugo Alatrista Salas, Sandra Bringay, Frédéric Flouvat, Nazha Selmaoui-Folcher, Maguelonne Teisseire
PAKDD (2)5
2012 Extracting Trajectories through an Efficient and Unifying Spatio-temporal Pattern Mining System
NhatHai Phan, Dino Ienco, Pascal Poncelet, Maguelonne Teisseire
ECML/PKDD (2)4
2012 A Fuzzy Associative Classification Approach for Recommender Systems
abstract
Despite the existence of different methods, including data mining techniques, available to be used in recommender systems, such systems still contain numerous limitations. They are in a constant need for personalization in order to make effective suggestions and to provide valuable information of items available. A way to reach such personalization is by means of an alternative data mining technique called classification based on association, which uses association rules in a prediction perspective. In this work we propose a hybrid methodology for recommender systems, which uses collaborative filtering and content-based approaches in a joint method taking advantage from the strengths of both approaches. Moreover, we also employ fuzzy logic to enhance recommendations' quality and effectiveness. In order to analyze the behavior of the techniques used in our methodology, we accomplished a case study using real data gathered from two recommender systems. Results revealed that such techniques can be applied effectively in recommender systems, minimizing the effects typical drawbacks they present.
Joel Pinho Lucas, Anne Laurent, María N. Moreno García, Maguelonne Teisseire
Int. J. Uncertain. Fuzziness Knowl. Based Syst.4
2011 Towards an On-Line Analysis of Tweets Processing
Sandra Bringay, Nicolas Béchet, Flavien Bouillot, Pascal Poncelet, Mathieu Roche, Maguelonne Teisseire
DEXA (2)6
2011 Mining microarray data to predict the histological grade of a breast cancer
abstract
BACKGROUND: The aim of this study was to develop an original method to extract sets of relevant molecular biomarkers (gene sequences) that can be used for class prediction and can be included as prognostic and predictive tools. MATERIALS AND METHODS: The method is based on sequential patterns used as features for class prediction. We applied it to classify breast cancer tumors according to their histological grade. RESULTS: We obtained very good recall and precision for grades 1 and 3 tumors, but, like other authors, our results were less satisfactory for grade 2 tumors. CONCLUSIONS: We demonstrated the interest of sequential patterns for class prediction of microarrays and we now have the material to use them for prognostic and predictive applications.
Mickaël Fabrègue, Sandra Bringay, Pascal Poncelet, Maguelonne Teisseire, Beatrice Orsetti
J. Biomed. Informatics4
2011 Sequential patterns mining and gene sequence visualization to discover novelty from microarray data
Arnaud Sallaberry, Nicolas Pecheur, Sandra Bringay, Mathieu Roche, Maguelonne Teisseire
J. Biomed. Informatics5
2010 Intelligent Energy Data Warehouse: What Challenges?
abstract
The RIDER -Réseau et Inter connectivité Des Energies classiques et Renouvelables (Network and Inter-Connectivity of Classical and Renewable Energies) project gathers a pool of university laboratories, national and international companies to design intelligent energy management platforms. We present here our contribution on designing data warehouse architectural models for massive and heterogeneous data, to be integrated in a multi-building intelligent energy management platform. Although a lot of work has been done on the subject, present tools and techniques still do not cover all the challenges we face when confronted to real time energy-related data management. This paper presents an overview of the related research work on data streams, data warehousing and ETL (Extract Transform Load) processes. ETL data exceptions will be our main point of focus. This critical subject has been rather left aside in research works so far.
Lucie Copin, Herve Rey, Xavier Vasques, Anne Laurent, Maguelonne Teisseire
ICTAI (2)5
2010 Mining multidimensional and multilevel sequential patterns
abstract
Multidimensional databases have been designed to provide decision makers with the necessary tools to help them understand their data. This framework is different from transactional data as the datasets contain huge volumes of historicized and aggregated data defined over a set of dimensions that can be arranged through multiple levels of granularities. Many tools have been proposed to query the data and navigate through the levels of granularity. However, automatic tools are still missing to mine this type of data in order to discover regular specific patterns. In this article, we present a method for mining sequential patterns from multidimensional databases, at the same time taking advantage of the different dimensions and levels of granularity, which is original compared to existing work. The necessary definitions and algorithms are extended from regular sequential patterns to this particular case. Experiments are reported, showing the significance of this approach.
Marc Plantevit, Anne Laurent, Dominique Laurent 0001, Maguelonne Teisseire, Yeow Wei Choong
ACM Trans. Knowl. Discov. Data4
2009 Mining Discriminant Sequential Patterns for Aging Brain
Paola Salle, Sandra Bringay, Maguelonne Teisseire
AIME3
2009 Using OWA operators for gene sequential pattern clustering
abstract
Nowadays, the management of sequential patterns data becomes an increasing need in biological knowledge discovery processes. An important task in these processes is the restitution of the results obtained by using data mining methods. In a complex domain as biomedical, an efficient interpretation of the patterns without any assistance is difficult. One of the most common knowledge discovery process is clustering. But the application of clustering to gene sequential patterns is far from easy on biomedical data. In this paper, we introduce a new gene sequential patterns similarity function and summarization algorithm.
Jordi Nin, Paola Salle, Sandra Bringay, Maguelonne Teisseire
CBMS4
2009 Handling fuzzy gaps in sequential patterns: Application to health
abstract
Dealing with digital data for mining novel knowledge is a non trivial task that has received much attention in the last years. However, it is still not easy to handle such data, especially when large volumes of values must be analyzed. In our work, we focus on biological data from DNA chips that biologists study in order to try and discover new gene correlations that could help understanding diseases like breast cancer. In this framework, we consider the values from the DNA microarrays, which convey the behavior of some genes, and we want to discover how these behaviors are correlated. This data are digital values that can be ordered and sorted. In previous work, sequential patterns like {(1 5)(2)} have been discovered, meaning that genes 1 and 5 have the same expression level followed by gene 2 that has a higher expression value. However, such data are very noisy and considering close values as ordered is often false. We thus consider here fuzzy rankings based on a fuzzy partition provided by the experts. Rules can then better characterize how genes are correlated.
Sandra Bringay, Anne Laurent, Beatrice Orsetti, Paola Salle, Maguelonne Teisseire
FUZZ-IEEE5
2009 Mining Frequent Gradual Itemsets from Large Databases
Lisa Di-Jorio, Anne Laurent, Maguelonne Teisseire
IDA3
2009 Efficient mining of sequential patterns with time constraints: Reducing the combinations
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
Expert Syst. Appl.3
2009 FTMnodes: Fuzzy tree mining based on partial inclusion
Federico Del Razo López, Anne Laurent, Pascal Poncelet, Maguelonne Teisseire
Fuzzy Sets Syst.4
2009 Tree mining: Equivalence classes for candidate generation
abstract
With the rise of active research fields such as bioinformatics, taxonomies and the growing use of XML documents, tree data are playing a more and more important role. Mining for frequent subtrees from these data is thus an active research problem and
Federico Del Razo López, Anne Laurent, Maguelonne Teisseire, Pascal Poncelet
Intell. Data Anal.3
2009 Evolution patterns and gradual trends
abstract
Nowadays, many databases record ordered or temporally annotated data, such as Web access logs or genomic sequences. Therefore, sequence mining has become an important research area. Among these data mining approaches, sequential patterns aim at describing frequent behaviors. In the access data of a commercial Web site, one may, for instance, discover that “35% of customers successively buy a PSP then a memory stick and PSP games”. To provide more complete information, fuzzy sequential patterns were designed, including quantitative values within the mining task. Such patterns, considering the previous example, would be “35% of customers buy a PSP, then they buy few games many times, and then they buy a high-capacity memory stick once.” However, symbolic or fuzzy sequential patterns, in their current form, do not allow to extract temporal tendencies that are typical of sequential data. By means of temporal tendency mining, one may discover in the same access data that “An increasing number of purchases of PSP games during a very short period is frequently followed by a purchase of a high-capacity memory stick a few days later.” It would be easy to conclude that the users either quickly succeed in registering or make several attempts before they look at the help page within a few seconds. To the best of our knowledge, no method has been designed for discovering this kind of patterns. Therefore, we propose, in this paper, two approaches that extract pattern-expressing trends or evolution. First, we define evolution patterns that summarize the evolution of the quantities in the data. We explain how they can be obtained from a quantitative sequence database. Second, we define gradual trends in fuzzy sequential data. These trends describe variations in the fulfillment of fuzzy properties according to time. For both kinds of patterns, we developed algorithms that were implemented and tested on real data. © 2009 Wiley Periodicals, Inc.
Céline Fiot, Florent Masseglia, Anne Laurent, Maguelonne Teisseire
Int. J. Intell. Syst.4
2008 Learning Bayesian Network Structure from Incomplete Data without Any Assumption
Céline Fiot, G. A. Putri Saptawati, Anne Laurent, Maguelonne Teisseire
DASFAA4
2008 Up and Down: Mining Multidimensional Sequential Patterns Using Hierarchies
Marc Plantevit, Anne Laurent, Maguelonne Teisseire
DaWaK3
2008 TED and EVA: Expressing temporal tendencies among quantitative variables using fuzzy sequential patterns
abstract
Temporal data can be handled in many ways for discovering specific knowledge. Sequential pattern mining is one of these relevant approaches when dealing with temporally annotated data. It allows discovering frequent sequences embedded in the records. In the access data of a commercial Web site, one may, for instance, discover that ldquo5% of the users request the page register.php 3 times and then request the page help.htmlrdquo. However, symbolic or fuzzy sequential patterns, in their current form, do not allow extracting temporal tendencies that are typical of sequential data. By means of temporal tendency mining, one may discover in the same access data that ldquoan increasing number of accesses to the register form preceeds an increasing number of accesses to the help page a few seconds laterrdquo. It would be easy to conclude that the users either quickly succeed in registering or make several attempts before they look at the help page within a few seconds. In this paper, we propose the definition of evolution patterns that allow discovering such knowledge. We show how to extract evolution patterns thanks to fuzzy sequential pattern mining techniques. We introduce our algorithms TED and EVA, designed for evolution pattern mining. Our proposal is validated by experiments and a sample of extracted knowledge is discussed.
Céline Fiot, Florent Masseglia, Anne Laurent, Maguelonne Teisseire
FUZZ-IEEE4
2008 Detection of Sequential Outliers Using a Variable Length Markov Model
abstract
The problem of mining for outliers in sequential datasets is crucial to forward appropriate analysis of data. Therefore, many approaches for the discovery of such anomalies have been proposed. However, most of them use a sample of known typical sequences to build the model. Besides, they remain greedy in terms of memory usage. In this paper we propose an extension of one such approach, based on a Probabilistic Suffix Tree and on a measure of similarity. We add a pruning criterion which reduces the size of the tree while improving the model, and a sharp inequality for the concentration of the measure of similarity, to better sort the outliers. We prove the feasability of our approach through a set of experiments over a protein database.
Cécile Low-Kam, Anne Laurent, Maguelonne Teisseire
ICMLA3
2008 Web usage mining: extracting unexpected periods from web logs
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire, Alice Marascu
Data Min. Knowl. Discov.3
2007 Mining unexpected multidimensional rules
abstract
Discovering unexpected rules is essential, particularly for industrial applications with marketing stakes. In this context, many works have been done for association rules. However, non of them address sequences. In this paper, we thus propose to discover unexpected multidimensional sequential rules in data cubes. We define the concept of multidimensional sequential rule, and then unexpectedness. We formalize these concepts and define an algorithm for mining this kind of rules. Experiments on a real data cube are reported and highlight the interest of our approach. Categories and Subject Descriptors H.2.8 [Database Management]: Database applications, data mining
Marc Plantevit, Sabine Goutier, Françoise Guisnel, Anne Laurent, Maguelonne Teisseire
DOLAP5
2007 Approximate Sequential Patterns for Incomplete Sequence Database Mining
abstract
Databases available from many industrial or research fields are often imperfect. In particular, they are most of the time incomplete in the sense that some of the values are missing. When facing this kind of imperfect data, two techniques can be investigated: either using only the available information or estimating the missing values. In this paper we propose an estimation-based approach for sequence mining. This approach considers partial inclusion of an item within a record using fuzzy sets. Experiments run on various synthetic datasets show the feasibility and validity of our proposal as well in terms of quality as in terms of the robustness to the rate of missing values.
Céline Fiot, Anne Laurent, Maguelonne Teisseire
FUZZ-IEEE3
2007 On Transversal Hypergraph Enumeration in Mining Sequential Patterns
abstract
The transversal hypergraph enumeration based algorithms can be efficient in mining frequent itemsets, however it is difficult to apply them to sequence mining problems. In this paper we first analyze the constraints of using transversal hypergraph enumeration in itemset mining, then propose the ordered pattern model for representing and mining sequences with respect to these constraints. We show that the problem of mining sequential patterns can be transformed to the problem of mining frequent ordered patterns, and therefore we propose an application of the Dualize and Advance algorithm, which is transversal hypergraph enumeration based, in mining sequential patterns.
Dominique Li, Anne Laurent, Maguelonne Teisseire
IDEAS3
2007 Fuzzy Tree Mining: Go Soft on Your Nodes
Federico Del Razo López, Anne Laurent, Pascal Poncelet, Maguelonne Teisseire
IFSA (1)4
2007 Extended Time Constraints for Sequence Mining
abstract
Many applications require techniques for temporal knowledge discovery. Some of those approaches can handle time constraints between events. In particular some work has been done to mine generalized sequential patterns. However, such constraints are often too crisp or need a very precise assessment to avoid erroneous information. Therefore, in this paper we propose to soften temporal constraints used for generalized sequential pattern mining. To handle these constraints while data mining, we design an algorithm based on sequence graphs. Moreover, as these relaxed constraints may extract more generalized patterns, we propose temporal accuracy measure for helping the analysis of the numerous discovered patterns.
Céline Fiot, Anne Laurent, Maguelonne Teisseire
TIME3
2007 Towards a new approach for mining frequent itemsets on data stream
Chedy Raïssi, Pascal Poncelet, Maguelonne Teisseire
J. Intell. Inf. Syst.3
2007 From Crispness to Fuzziness: Three Algorithms for Soft Sequential Pattern Mining
abstract
Most real world databases consist of historical and numerical data such as sensor, scientific or even demographic data. In this context, classical algorithms extracting sequential patterns, which are well adapted to the temporal aspect of data, do not allow numerical information processing. Therefore, the data are pre-processed to be transformed into a binary representation, which leads to a loss of information. Fuzzy algorithms have been proposed to process numerical data using intervals, particularly fuzzy intervals, but none of these methods is satisfactory. Therefore this paper completely defines the concepts linked to fuzzy sequential pattern mining. Using different fuzzification levels, we propose three methods to mine fuzzy sequential patterns and detail the resulting algorithms (SpeedyFuzzy, MiniFuzzy, and TotallyFuzzy). Finally, we assess them through different experiments, thus revealing the robustness and the relevancy of this work.
Céline Fiot, Anne Laurent, Maguelonne Teisseire
IEEE Trans. Fuzzy Syst.3
2006 Peer-to-Peer Usage Analysis: a Distributed Mining Approach
abstract
With the huge number of information sources available on the Internet, peer-to-peer (P2P) systems offer a novel kind of system architecture providing the large-scale community with applications for file sharing, distributed file systems, distributed computing, messaging and real-time communication. P2P applications also provide a good infrastructure for data and compute intensive operations such as data mining. In this paper we propose a new approach for improving resource searching in a dynamic and distributed database such as an unstructured P2P system. This approach takes advantage of data mining techniques. By using a genetic-inspired algorithm, we propose to extract patterns or relationships occurring in a large number of nodes. Such a knowledge is very useful for proposing the user with often downloaded or requested files according to a majority of behaviors. It may also be useful in order to avoid extra bandwidth consumption
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
AINA (1)3
2006 Privacy preserving sequential pattern mining in distributed databases
abstract
Research in the areas of privacy preserving techniques in databases and subsequently in privacy enhancement technologies have witnessed an explosive growth-spurt in recent years. This escalation has been fueled by the growing mistrust of individuals towards organizations collecting and disbursing their Personally Identifiable Information (PII). Digital repositories have become increasingly susceptible to intentional or unintentional abuse, resulting in organizations to be liable under the privacy legislations that are being adopted by governments the world over. These privacy concerns have necessitated new advancements in the field of distributed data mining wherein, collaborating parties may be legally bound not to reveal the private information of their customers. In this paper, we present a new algorithm PriPSeP (Privacy Preserving SEquential Patterns) for the mining of sequential patterns from distributed databases while preserving privacy. A salient feature of PriPSeP is that due to its flexibility it is more pertinent to mining operations for real world applications in terms of efficiency and functionality. Under some reasonable assumptions, we prove that our architecture and protocol employed by our algorithm for multi-party computation is secure.
Vishal Kapoor, Pascal Poncelet, François Trousset, Maguelonne Teisseire
CIKM4
2006 HYPE: mining hierarchical sequential patterns
abstract
Mining data warehouses is still an open problem as few approaches really take the specificities of this framework into account (e.g. multidimensionality, hierarchies, historized data). Multidimensional sequential patterns have been studied but they do not provide any way to handle hierarchies. In this paper, we propose an original sequential pattern extraction method that takes the hierarchies into account. This method extracts more accurate knowledge and extends our preceding M2SP approach. We define the concepts related to our problems as well as the associated algorithms. The results of our experiments confirm the relevance of our proposal.
Marc Plantevit, Anne Laurent, Maguelonne Teisseire
DOLAP3
2006 Why Fuzzy Sequential Patterns can Help Data Summarization: An Application to the INPI Trade Mark Database
abstract
Mining fuzzy rules is one of the best ways to summarize large databases while keeping information as clear and understandable as possible for the end-user. Several approaches have been proposed to mine such fuzzy rules, in particular to mine fuzzy association rules. However, we argue that it is important to mine rules that convey information about the order. For instance, it is very interesting to convey the idea of time running in rules, which is done in fuzzy sequential patterns. In this paper, we thus focus on fuzzy sequential patterns. We show that mining such rules requires to manage a lot of information and we propose algorithms to remain efficient in both memory use and computation time. Our proposition is assessed by experiments. Particularly, we apply our algorithms on the INPI database which stores almost 2 million trademarks.
Céline Fiot, Anne Laurent, Maguelonne Teisseire, Bénédicte Laurent
FUZZ-IEEE3
2006 Sequential patterns for text categorization
Simon Jaillet, Anne Laurent, Maguelonne Teisseire
Intell. Data Anal.3
2006 Mining spatio-temporal data
Gennady L. Andrienko, Donato Malerba, Michael May 0001, Maguelonne Teisseire
J. Intell. Inf. Syst.4
2005 M2SP: Mining Sequential Patterns Among Several Dimensions
Marc Plantevit, Yeow Wei Choong, Anne Laurent, Dominique Laurent 0001, Maguelonne Teisseire
PKDD5
2004 Where's Charlie: Family-Based Heuristics for Peer-to-Peer Schema Integration
John Tranier, Renaud Baraër, Zohra Bellahsene, Maguelonne Teisseire
IDEAS4
2004 Pre-Processing Time Constraints for Efficiently Mining Generalized Sequential Patterns
abstract
In this paper we consider the problem of discovering sequential patterns by handling time constraints. While sequential patterns could be seen as temporal relationships between facts embedded in the database, generalized sequential patterns aim at providing the end user with a more flexible handling of the transactions embedded in the database. We propose a new efficient algorithm, called GTC (graph for time constraints) for mining such patterns in very large databases. It is based on the idea that handling time constraints in the earlier stage of the algorithm can be highly beneficial since it minimizes computational costs by preprocessing data sequences. Our test shows that the proposed algorithm performs significantly faster than a state-of-the-art sequence mining algorithm.
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
TIME3
2003 AUSMS: An Environment for Frequent Sub-structures Extraction in a Semi-structured Object Collection
Pierre-Alain Laur, Maguelonne Teisseire, Pascal Poncelet
DEXA2
2003 Incremental mining of sequential patterns in large databases
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
Data Knowl. Eng.3
2003 HDM: A Client/Server/Engine Architecture for Real-Time Web Usage Mining
Florent Masseglia, Maguelonne Teisseire, Pascal Poncelet
Knowl. Inf. Syst.2
2001 Real-Time Web Usage Mining: A Heuristic Based Distributed Miner
abstract
The behaviour of a Web site's users may change so quickly that attempting to make predictions, according to the frequent patterns coming from the analysis of an access log file, becomes challenging. In order for the obsolescence of the behavioural patterns to become as null as possible, the ideal method would provide frequent patterns in real time, allowing the result to be available immediately. We propose, in this paper a method allowing to find frequent behavioural patterns in real time, whatever the number of connected users is. Considering how fast the frequent behaviour patterns can change since the last analysis of the access log file, this result thus provide completely adapted navigation schemas for user behaviour predictions. Based on a distributed heuristic, our method also answers several tackled problems within the data mining framework: Discovering "interesting zones" (a great number of frequent patterns concentrated over a period of time, or the discovering of "super-frequent" patterns), discovering very long sequential patterns and interactive data mining ("on the fly" modification of the minimum support).
Florent Masseglia, Maguelonne Teisseire, Pascal Poncelet
WISE (1)2
2000 Web Usage Mining: How to Efficiently Manage New Transactions and New Clients
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
PKDD3
1997 Activity Threads: A Unified Framework for Aiding Behavioural Modelling
Bernard Faure, Maguelonne Teisseire, Rosine Cicchetti
DEXA2
1996 Views for Information System Design without Reorganization
Zohra Bellahsene, Pascal Poncelet, Maguelonne Teisseire
CAiSE3
1994 Dynamic Modelling with Events
Maguelonne Teisseire, Pascal Poncelet, Rosine Cicchetti
CAiSE1
1994 An Algebraic Language for Event-Driven Modelling
Maguelonne Teisseire, Rosine Cicchetti
DEXA1
1994 Towards Event-Driven Modelling for Database Design
Maguelonne Teisseire, Pascal Poncelet, Rosine Cicchetti
VLDB1
1993 Towards a Formal Approach for Object Database Design
Pascal Poncelet, Maguelonne Teisseire, Rosine Cicchetti, Lotfi Lakhal
VLDB2