Florent Masseglia

dblp:m/FlorentMasseglia · DBLP profile ↗
← Back
59ranked-venue papers
11as first author
7since 2021 · last 2026
0000-0002-1149-585XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 43 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 34 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Efficient Detection of Seasonal Bursts: Applications to Climate Time Series
Guillaume Coulaud, Audrey Brouillet, Reza Akbarinia, Florent Masseglia, Dennis E. Shasha
DEXA (2)4
2025 ClimBurst: A Dynamic Visualization Tool to Display Climatological Anomalies over Time and Space
abstract
Detecting abnormal climate events across temporal and spatial scales is crucial to the understanding of local and regional climate trends. This demonstration introduces ClimBurst, a dynamic tool to detect climate bursts, which are unusually high or low values of one or more climate variables over some time interval. ClimBurst detects bursts without prior assumptions about their temporal duration. The demonstration will allow users to interact directly with our system to see both a summary showing the presence/absence of bursts over a user-specified year and spatial range. The demonstration will also allow users to perform time-travel queries to see how bursts propagate over space and time.
Guillaume Coulaud, Benoit Lange, Dennis E. Shasha, Audrey Brouillet, Reza Akbarinia, Florent Masseglia
CIKM6
2025 Scalable and accurate online multivariate anomaly detection
Rebecca Salles, Benoit Lange, Reza Akbarinia, Florent Masseglia, Eduardo S. Ogasawara, Esther Pacitti
Inf. Syst.4
2024 A One-Health Platform for Antimicrobial Resistance Data Analytics
abstract
Antimicrobial resistance (AMR) poses potentially critical health issues for human and animal populations in the near future. To meet this challenge, we need to adopt a "One Health" strategy, which involves studying and linking information from human and animal populations, as well as from the environment.
Benoit Lange, Reza Akbarinia, Florent Masseglia
CIKM3
2023 kNN matrix profile for knowledge discovery from time series
Tanmoy Mondal, Reza Akbarinia, Florent Masseglia
Data Min. Knowl. Discov.3
2022 Parallel Techniques for Variable Size Segmentation of Time Series Datasets
Lamia Djebour, Reza Akbarinia, Florent Masseglia
ADBIS3
2021 BestNeighbor: efficient evaluation of kNN queries on large time series databases
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas, Dennis E. Shasha, Patrick Valduriez
Knowl. Inf. Syst.5
2020 Space-time series clustering: Algorithms, taxonomy, and case study on urban smart cities
abstract
This paper provides a short overview of space–time series clustering, which can be generally grouped into three main categories such as: hierarchical, partitioning-based, and overlapping clustering. The first hierarchical category is to identify hierarchies in space–time series data. The second partitioning-based category focuses on determining disjoint partitions among the space–time series data, whereas the third overlapping category explores fuzzy logic to determine the different correlations between the space–time series clusters. We also further describe solutions for each category in this paper. Furthermore, we show the applications of these solutions in an urban traffic data captured on two urban smart cities (e.g., Odense in Denmark and Beijing in China). The perspectives on open questions and research challenges are also mentioned and discussed that allow to obtain a better understanding of the intuition, limitations, and benefits for the various space–time series clustering methods. This work can thus provide the guidances to practitioners for selecting the most suitable methods for their used cases, domains, and applications.
Asma Belhadi, Youcef Djenouri, Kjetil Nørvåg, Heri Ramampiaro, Florent Masseglia, Jerry Chun-Wei Lin
Eng. Appl. Artif. Intell.5
2020 Spatial-time motifs discovery
abstract
Discovering motifs in time series data has been widely explored. Various techniques have been developed to tackle this problem. However, when it comes to spatial-time series, a clear gap can be observed according to the literature review. This paper tackles such a gap by presenting an approach to discover and rank motifs in spatial-time series, denominated Combined Series Approach (CSA). CSA is based on partitioning the spatial-time series into blocks. Inside each block, subsequences of spatial-time series are combined in a way that hash-based motif discovery algorithm is applied. Motifs are validated according to both temporal and spatial constraints. Later, motifs are ranked according to their entropy, the number of occurrences, and the proximity of their occurrences. The approach was evaluated using both synthetic and seismic datasets. CSA outperforms traditional methods designed only for time series. CSA was also able to prioritize motifs that were meaningful both in the context of synthetic data and also according to seismic specialists.
Heraldo Borges, Murillo Dutra, Amin Bazaz, Rafaelli de C. Coutinho, Fabio Perosi, Fábio Porto 0001, Florent Masseglia, Esther Pacitti, Eduardo S. Ogasawara
Intell. Data Anal.7
2020 Massively Distributed Time Series Indexing and Querying
abstract
Indexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on four billion time series in less than five hours, while the state of the art centralized algorithms do not scale and have their limit on 1 billion time series, where they need more than five days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas
IEEE Trans. Knowl. Data Eng.3
2019 High Dimensional Data Clustering by means of Distributed Dirichlet Process Mixture Models
abstract
Clustering is a data mining technique intensively used for data analytics, with applications to marketing, security, text/document analysis, or sciences like biology, astronomy, and many more. Dirichlet Process Mixture (DPM) is a model used for multivariate clustering with the advantage of discovering the number of clusters automatically and offering favorable characteristics. However, in the case of high dimensional data, it becomes an important challenge with numerical and theoretical pitfalls. The advantages of DPM come at the price of prohibitive running times, which impair its adoption and makes centralized DPM approaches inefficient, especially with high dimensional data. We propose HD4C (High Dimensional Data Distributed Dirichlet Clustering), a parallel clustering solution that addresses the curse of dimensionality by two means. First it gracefully scales to massive datasets by distributed computing, while remaining DPM-compliant. Second, it performs clustering of high dimensional data such as time series (as a function of time), hyperspectral data (as a function of wavelength) etc. Our experiments, on both synthetic and real world data, illustrate the high performance of our approach.
Khadidja Meguelati, Benedicte Fontez, Nadine Hilgert, Florent Masseglia
IEEE BigData4
2019 Parallel Streaming Implementation of Online Time Series Correlation Discovery on Sliding Windows with Regression Capabilities
abstract
International audience
Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez
CLOSER5
2019 Pipelined Implementation of a Parallel Streaming Method for Time Series Correlation Discovery on Sliding Windows
Boyan Kolev, Reza Akbarinia, Ricardo Jiménez-Peris, Oleksandra Levchenko, Florent Masseglia, Marta Patiño-Martínez, Patrick Valduriez
DATA5
2019 Distributed Algorithms to Find Similar Time Series
abstract
International audience
Oleksandra Levchenko, Boyan Kolev, Djamel Edine Yagoubi, Dennis E. Shasha, Themis Palpanas, Patrick Valduriez, Reza Akbarinia, Florent Masseglia
ECML/PKDD (3)8
2018 Spark-parSketch: A Massively Distributed Indexing of Time Series Datasets
abstract
A growing number of domains (finance, seismology, internet-of-things, etc.) collect massive time series. When the number of series grow to the hundreds of millions or even billions, similarity queries become intractable on a single machine. Further, naive (quadratic) parallelization won't work well. So, we need both efficient indexing and parallelization. We propose a demonstration of Spark-parSketch, a complete solution based on sketches / random projections to efficiently perform both the parallel indexing of large sets of time series and a similarity search on them. Because our method is approximate, we explore the tradeoff between time and precision. A video showing the dynamics of the demonstration can be found by the link http://parsketch.gforge.inria.fr/video/parSketchdemo_720p.mov.
Oleksandra Levchenko, Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Boyan Kolev, Dennis E. Shasha
CIKM4
2018 Discovering Tight Space-Time Sequences
Riccardo Campisano, Heraldo Borges, Fábio Porto 0001, Fabio Perosi, Esther Pacitti, Florent Masseglia, Eduardo S. Ogasawara
DaWaK6
2018 ParCorr: efficient parallel methods to identify similar time series pairs across sliding windows
Djamel Edine Yagoubi, Reza Akbarinia, Boyan Kolev, Oleksandra Levchenko, Florent Masseglia, Patrick Valduriez, Dennis E. Shasha
Data Min. Knowl. Discov.5
2017 Massively Distributed Environments and Closed Itemset Mining: The DCIM Approach
Mehdi Zitouni, Reza Akbarinia, Sadok Ben Yahia, Florent Masseglia
CAiSE4
2017 RadiusSketch: Massively Distributed Indexing of Time Series
abstract
Performing similarity queries on hundreds of millions of time series is a challenge requiring both efficient indexing techniques and parallelization. We propose a sketch/random projection-based approach that scales nearly linearly in parallel environments, and provides high quality answers. We illustrate the performance of our approach, called RadiusSketch, on real and synthetic datasets of up to 1 Terabytes and 500 million time series. The sketch method, as we have implemented, is superior in both quality and response time compared with the state of the art approach, iSAX2+. Already, in the sequential case it improves recall and precision by a factor of two, while giving shorter response times. In a parallel environment with 32 processors, on both real and synthetic data, our parallel approach improves by a factor of up to 100 in index time construction and up to 15 in query answering time. Finally, our data structure makes use of idle computing time to improve the recall and precision yet further.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Dennis E. Shasha
DSAA3
2017 DPiSAX: Massively Distributed Partitioned iSAX
abstract
Indexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on 1 billion time series in less than 2 hours, while the state of the art centralized algorithms need more than 5 days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism.
Djamel Edine Yagoubi, Reza Akbarinia, Florent Masseglia, Themis Palpanas
ICDM3
2017 A highly scalable parallel algorithm for maximally informative k-itemset mining
Saber Salah, Reza Akbarinia, Florent Masseglia
Knowl. Inf. Syst.3
2017 Data placement in massively distributed environments for fast parallel mining of frequent itemsets
Saber Salah, Reza Akbarinia, Florent Masseglia
Knowl. Inf. Syst.3
2016 A new privacy-preserving solution for clustering massively distributed personal times-series
abstract
New personal data fields are currently emerging due to the proliferation of on-body/at-home sensors connected to personal devices. However, strong privacy concerns prevent individuals to benefit from large-scale analytics that could be performed on this fine-grain highly sensitive wealth of data. We propose a demonstration of Chiaroscuro, a complete solution for clustering massively-distributed sensitive personal data while guaranteeing their privacy. The demonstration scenario highlights the affordability of the privacy vs. quality and privacy vs. performance tradeoffs by dissecting the inner working of Chiaroscuro - launched over energy consumption times-series -, by exposing the results obtained by the individuals participating in the clustering process, and by illustrating possible uses.
Tristan Allard, Georges Hébrail, Florent Masseglia, Esther Pacitti
ICDE3
2015 Data Partitioning for Fast Mining of Frequent Itemsets in Massively Distributed Environments
Saber Salah, Reza Akbarinia, Florent Masseglia
DEXA (1)3
2015 A Prime Number Based Approach for Closed Frequent Itemset Mining in Big Data
Mehdi Zitouni, Reza Akbarinia, Sadok Ben Yahia, Florent Masseglia
DEXA (1)4
2015 Fast Parallel Mining of Maximally Informative k-Itemsets in Big Data
abstract
The discovery of informative itemsets is a fundamental building block in data analytics and information retrieval. While the problem has been widely studied, only few solutions scale. This is particularly the case when i) the data set is massive, calling for large-scale distribution, and/or ii) the length k of the informative itemset to be discovered is high. In this paper, we address the problem of parallel mining of maximally informative k-itemsets (miki) based on joint entropy. We propose PHIKS (Parallel Highly Informative K-ItemSet) a highly scalable, parallel miki mining algorithm. PHIKS renders the mining process of large scale databases (up to terabytes of data) succinct and effective. Its mining process is made up of only two efficient parallel jobs. With PHIKS, we provide a set of significant optimizations for calculating the joint entropies of miki having different sizes, which drastically reduces the execution time of the mining process. PHIKS has been extensively evaluated using massive real-world data sets. Our experimental results confirm the effectiveness of our proposal by the significant scale-up obtained with high itemsets length and over very large databases.
Saber Salah, Reza Akbarinia, Florent Masseglia
ICDM3
2015 Chiaroscuro: Transparency and Privacy for Massive Personal Time-Series Clustering
abstract
The advent of on-body/at-home sensors connected to personal devices leads to the generation of fine grain highly sensitive personal data at an unprecendent rate. However, despite the promises of large scale analytics there are obvious privacy concerns that prevent individuals to share their personnal data. In this paper, we propose Chiaroscuro, a complete solution for clustering personal data with strong privacy guarantees. The execution sequence produced by Chiaroscuro is massively distributed on personal devices, coping with arbitrary connections and disconnections. Chiaroscuro builds on our novel data structure, called Diptych, which allows the participating devices to collaborate privately by combining encryption with differential privacy. Our solution yields a high clustering quality while minimizing the impact of the differentially private perturbation. Chiaroscuro is both correct and secure. Finally, we provide an experimental validation of our approach on both real and synthetic sets of time-series.
Tristan Allard, Georges Hébrail, Florent Masseglia, Esther Pacitti
SIGMOD Conference3
2014 The anti-bouncing data stream model for web usage streams with intralinkings
Chongsheng Zhang, Florent Masseglia, Yves Lechevallier
Inf. Sci.2
2014 Autonomic intrusion detection: Adaptively detecting anomalies over unlabeled audit data streams in computer networks
Wei Wang 0012, Thomas Guyet, Rene Quiniou, Marie-Odile Cordier, Florent Masseglia, Xiangliang Zhang 0001
Knowl. Based Syst.5
2013 A Density-Based Backward Approach to Isolate Rare Events in Large-Scale Applications
Enikö Székely, Pascal Poncelet, Florent Masseglia, Maguelonne Teisseire, Renaud Cezar
Discovery Science3
2013 Fast and Exact Mining of Probabilistic Data Streams
Reza Akbarinia, Florent Masseglia
ECML/PKDD (1)2
2012 Discovering Highly Informative Feature Set over High Dimensions
abstract
For many textual collections, the number of features is often overly large. These features can be very redundant, it is therefore desirable to have a small, succinct, yet highly informative collection of features that describes the key characteristics of a dataset. Information theory is one such tool for us to obtain this feature collection. With this paper, we mainly contribute to the improvement of efficiency for the process of selecting the most informative feature set over high-dimensional unlabeled data. We propose a heuristic theory for informative feature set selection from high dimensional data. Moreover, we design data structures that enable us to compute the entropies of the candidate feature sets efficiently. We also develop a simple pruning strategy that eliminates the hopeless candidates at each forward selection step. We test our method through experiments on real-world data sets, showing that our proposal is very efficient.
Chongsheng Zhang, Florent Masseglia, Xiangliang Zhang 0001
ICTAI2
2012 Modeling and Clustering Users with Evolving Profiles in Usage Streams
abstract
Today, there is an increasing need of data stream mining technology to discover important patterns on the fly. Existing data stream models and algorithms commonly assume that users' records or profiles in data streams will not be updated or revised once they arrive. Nevertheless, in various applications such as Web usage, the records/profiles of the users can evolve along time. This kind of streaming data evolves in two forms, the streaming of tuples or transactions as in the case of traditional data streams, and more importantly, the evolving of user records/profiles inside the streams. Such data streams bring difficulties on modeling and clustering for exploringusers' behaviors. In this paper, we propose three models to summarize this kind of data streams, which are the batch model, the Evolving Objects (EO) model and the Dynamic Data Stream (DDS) model. Through creating, updating and deleting user profiles, these models summarize the behaviors of each user as a profile object. Based upon these models, clustering algorithms are employed to discover interesting user groups from the profile objects. We have evaluated all the proposed models on a large real-world data set, showing that the DDS model summarizes the data streams with evolving tuples more efficiently and effectively, and provides better basis for clustering users than the other two models.
Chongsheng Zhang, Florent Masseglia, Xiangliang Zhang 0001
TIME2
2011 Atypicity detection in data streams: A self-adjusting approach
abstract
Outlyingness is a subjective concept relying on the isolation level of a (set of) record(s). Clustering-based outlier detection is a field that aims to cluster data and to detect outliers depending on their characteristics (i.e. small, tight and/or dense clusters might be considered as outliers). E xisting methods require a parameter standing for the “level of outlyingness”, such as the maximum size or a percentage of small clusters, in order to build the set of outliers. Unfortunately, manually setting this parameter in a streaming environment should not be possible, given the fast time response usually needed. In this paper we propose Wod, a method that separates outliers from clusters thanks to a natural and effective principle. The main advantages of Wod are its ability to automatically adjust to any clustering result and to be parameterless.
Alice Marascu, Florent Masseglia
Intell. Data Anal.2
2011 Discovering Significant Evolution Patterns from Satellite Image Time Series
abstract
Satellite Image Time Series (SITS) provide us with precious information on land cover evolution. By studying these series of images we can both understand the changes of specific areas and discover global phenomena that spread over larger areas. Changes that can occur throughout the sensing time can spread over very long periods and may have different start time and end time depending on the location, which complicates the mining and the analysis of series of images. This work focuses on frequent sequential pattern mining (FSPM) methods, since this family of methods fits the above-mentioned issues. This family of methods consists of finding the most frequent evolution behaviors, and is actually able to extract long-term changes as well as short term ones, whenever the change may start and end. However, applying FSPM methods to SITS implies confronting two main challenges, related to the characteristics of SITS and the domain's constraints. First, satellite images associate multiple measures with a single pixel (the radiometric levels of different wavelengths corresponding to infra-red, red, etc.), which makes the search space multi-dimensional and thus requires specific mining algorithms. Furthermore, the non evolving regions, which are the vast majority and overwhelm the evolving ones, challenge the discovery of these patterns. We propose a SITS mining framework that enables discovery of these patterns despite these constraints and characteristics. Our proposal is inspired from FSPM and provides a relevant visualization principle. Experiments carried out on 35 images sensed over 20 years show the proposed approach makes it possible to extract relevant evolution behaviors.
François Petitjean, Florent Masseglia, Pierre Gançarski, Germain Forestier
Int. J. Neural Syst.2
2011 Discovering frequent behaviors: time is an essential element of the context
Bashar Saleh, Florent Masseglia
Knowl. Inf. Syst.2
2010 Discovering Highly Informative Feature Sets from Data Streams
Chongsheng Zhang, Florent Masseglia
DEXA (1)2
2010 ABS: The Anti Bouncing Model for Usage Data Streams
abstract
Usage data mining is an important research area with applications in various fields. However, usage data is usually considered streaming, due to its high volumes and rates. Because of these characteristics, we only have access, at any point in time, to a small fraction of the stream. When the data is observed through such a limited window, it is challenging to give a reliable description of the recent usage data. We study the important consequences of these constraints, through the “bounce rate” problem and the clustering of usage data streams. Then, we propose the ABS (Anti-Bouncing Stream) model which combines the advantages of previous models but discards their drawbacks. First, under the same resource constraints as existing models in the literature, ABS can better model the recent data. Second, owing to its simple but effective management approach, the data in ABS is available at any time for analysis. We demonstrate its superiority through a theoretical study and experiments on two real-world data sets.
Chongsheng Zhang, Florent Masseglia, Yves Lechevallier
ICDM2
2010 Analysing Satellite Image Time Series by Means of Pattern Mining
François Petitjean, Pierre Gançarski, Florent Masseglia, Germain Forestier
IDEAL3
2009 A Multi-resolution Approach for Atypical Behaviour Mining
Alice Marascu, Florent Masseglia
PAKDD2
2009 Data Mining for Intrusion Detection: From Outliers to True Intrusions
Goverdhan Singh, Florent Masseglia, Céline Fiot, Alice Marascu, Pascal Poncelet
PAKDD2
2009 A general framework for adaptive and online detection of web attacks
abstract
Detection of web attacks is an important issue in current defense-in-depth security framework. In this paper, we propose a novel general framework for adaptive and online detection of web attacks. The general framework can be based on any online clustering methods. A detection model based on the framework is able to learn online and deal with "concept drift" in web audit data streams. Str-DBSCAN that we extended DBSCAN to streaming data as well as StrAP are both used to validate the framework. The detection model based on the framework automatically labels the web audit data and adapts to normal behavior changes while identifies attacks through dynamical clustering of the streaming data. A very large size of real HTTP Log data collected in our institute is used to validate the framework and the model. The preliminary testing results demonstrated its effectiveness.
Wei Wang 0012, Florent Masseglia, Thomas Guyet, Rene Quiniou, Marie-Odile Cordier
WWW2
2009 Efficient mining of sequential patterns with time constraints: Reducing the combinations
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
Expert Syst. Appl.1
2009 Evolution patterns and gradual trends
abstract
Nowadays, many databases record ordered or temporally annotated data, such as Web access logs or genomic sequences. Therefore, sequence mining has become an important research area. Among these data mining approaches, sequential patterns aim at describing frequent behaviors. In the access data of a commercial Web site, one may, for instance, discover that “35% of customers successively buy a PSP then a memory stick and PSP games”. To provide more complete information, fuzzy sequential patterns were designed, including quantitative values within the mining task. Such patterns, considering the previous example, would be “35% of customers buy a PSP, then they buy few games many times, and then they buy a high-capacity memory stick once.” However, symbolic or fuzzy sequential patterns, in their current form, do not allow to extract temporal tendencies that are typical of sequential data. By means of temporal tendency mining, one may discover in the same access data that “An increasing number of purchases of PSP games during a very short period is frequently followed by a purchase of a high-capacity memory stick a few days later.” It would be easy to conclude that the users either quickly succeed in registering or make several attempts before they look at the help page within a few seconds. To the best of our knowledge, no method has been designed for discovering this kind of patterns. Therefore, we propose, in this paper, two approaches that extract pattern-expressing trends or evolution. First, we define evolution patterns that summarize the evolution of the quantities in the data. We explain how they can be obtained from a quantitative sequence database. Second, we define gradual trends in fuzzy sequential data. These trends describe variations in the fulfillment of fuzzy properties according to time. For both kinds of patterns, we developed algorithms that were implemented and tested on real data. © 2009 Wiley Periodicals, Inc.
Céline Fiot, Florent Masseglia, Anne Laurent, Maguelonne Teisseire
Int. J. Intell. Syst.2
2008 TED and EVA: Expressing temporal tendencies among quantitative variables using fuzzy sequential patterns
abstract
Temporal data can be handled in many ways for discovering specific knowledge. Sequential pattern mining is one of these relevant approaches when dealing with temporally annotated data. It allows discovering frequent sequences embedded in the records. In the access data of a commercial Web site, one may, for instance, discover that ldquo5% of the users request the page register.php 3 times and then request the page help.htmlrdquo. However, symbolic or fuzzy sequential patterns, in their current form, do not allow extracting temporal tendencies that are typical of sequential data. By means of temporal tendency mining, one may discover in the same access data that ldquoan increasing number of accesses to the register form preceeds an increasing number of accesses to the help page a few seconds laterrdquo. It would be easy to conclude that the users either quickly succeed in registering or make several attempts before they look at the help page within a few seconds. In this paper, we propose the definition of evolution patterns that allow discovering such knowledge. We show how to extract evolution patterns thanks to fuzzy sequential pattern mining techniques. We introduce our algorithms TED and EVA, designed for evolution pattern mining. Our proposal is validated by experiments and a sample of extracted knowledge is discussed.
Céline Fiot, Florent Masseglia, Anne Laurent, Maguelonne Teisseire
FUZZ-IEEE2
2008 Time Aware Mining of Itemsets
abstract
Frequent behavioural pattern mining is a very important topic of knowledge discovery, intended to extract correlations between items recorded in large databases or Web access logs. However, those databases are usually considered as a whole and hence, itemsets are extracted over the entire set of records. Our claim is that possible periods, hidden within the structure of the data and containing compact itemsets, may exist. These periods, as well as the itemsets they contain, might not be found by traditional data mining methods due to their very weak support. Furthermore, these periods might be lost depending on an arbitrary division of the data. The goal of our work is to find itemsets that are frequent over a specific period but would not be extracted by traditional methods since their support is very low over the whole dataset. In this paper, we introduce the definition of solid itemsets, which represent a coherent and compact behavior over a specific period, and we propose SIM, an algorithm for their extraction. This work may find many applications in sensitive domains such as fraud or intrusion detection.
Bashar Saleh, Florent Masseglia
TIME2
2008 Web usage mining: extracting unexpected periods from web logs
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire, Alice Marascu
Data Min. Knowl. Discov.1
2008 Editorial: Introduction to the Special Issue on Multimedia Data Mining
abstract
The twelve papers in this special issue focus on multimedia data mining. The special issue evolved from a successful workshop organized in conjunction with the 2006 ACM KDD conference, but the special issue was open to the whole community.
Zhongfei Zhang, Florent Masseglia, Ramesh Jain 0001, Alberto Del Bimbo
IEEE Trans. Multim.2
2006 Peer-to-Peer Usage Analysis: a Distributed Mining Approach
abstract
With the huge number of information sources available on the Internet, peer-to-peer (P2P) systems offer a novel kind of system architecture providing the large-scale community with applications for file sharing, distributed file systems, distributed computing, messaging and real-time communication. P2P applications also provide a good infrastructure for data and compute intensive operations such as data mining. In this paper we propose a new approach for improving resource searching in a dynamic and distributed database such as an unstructured P2P system. This approach takes advantage of data mining techniques. By using a genetic-inspired algorithm, we propose to extract patterns or relationships occurring in a large number of nodes. Such a knowledge is very useful for proposing the user with often downloaded or requested files according to a majority of behaviors. It may also be useful in order to avoid extra bandwidth consumption
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
AINA (1)1
2006 Mining sequential patterns from data streams: a centroid approach
Alice Marascu, Florent Masseglia
J. Intell. Inf. Syst.2
2004 Web Usage Mining: Sequential Pattern Extraction with a Very Low Support
Florent Masseglia, Doru Tanasa, Brigitte Trousse
APWeb1
2004 Pre-Processing Time Constraints for Efficiently Mining Generalized Sequential Patterns
abstract
In this paper we consider the problem of discovering sequential patterns by handling time constraints. While sequential patterns could be seen as temporal relationships between facts embedded in the database, generalized sequential patterns aim at providing the end user with a more flexible handling of the transactions embedded in the database. We propose a new efficient algorithm, called GTC (graph for time constraints) for mining such patterns in very large databases. It is based on the idea that handling time constraints in the earlier stage of the algorithm can be highly beneficial since it minimizes computational costs by preprocessing data sequences. Our test shows that the proposed algorithm performs significantly faster than a state-of-the-art sequence mining algorithm.
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
TIME1
2003 Incremental mining of sequential patterns in large databases
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
Data Knowl. Eng.1
2003 HDM: A Client/Server/Engine Architecture for Real-Time Web Usage Mining
Florent Masseglia, Maguelonne Teisseire, Pascal Poncelet
Knowl. Inf. Syst.1
2001 Real-Time Web Usage Mining: A Heuristic Based Distributed Miner
abstract
The behaviour of a Web site's users may change so quickly that attempting to make predictions, according to the frequent patterns coming from the analysis of an access log file, becomes challenging. In order for the obsolescence of the behavioural patterns to become as null as possible, the ideal method would provide frequent patterns in real time, allowing the result to be available immediately. We propose, in this paper a method allowing to find frequent behavioural patterns in real time, whatever the number of connected users is. Considering how fast the frequent behaviour patterns can change since the last analysis of the access log file, this result thus provide completely adapted navigation schemas for user behaviour predictions. Based on a distributed heuristic, our method also answers several tackled problems within the data mining framework: Discovering "interesting zones" (a great number of frequent patterns concentrated over a period of time, or the discovering of "super-frequent" patterns), discovering very long sequential patterns and interactive data mining ("on the fly" modification of the minimum support).
Florent Masseglia, Maguelonne Teisseire, Pascal Poncelet
WISE (1)1
2000 Schema Mining: Finding Structural Regularity among Semistructured Data
Pierre-Alain Laur, Florent Masseglia, Pascal Poncelet
PKDD2
2000 Web Usage Mining: How to Efficiently Manage New Transactions and New Clients
Florent Masseglia, Pascal Poncelet, Maguelonne Teisseire
PKDD1
1999 WebTool: An Integrated Framework for Data Mining
Florent Masseglia, Pascal Poncelet, Rosine Cicchetti
DEXA1
1998 The PSP Approach for Mining Sequential Patterns
Florent Masseglia, Fabienne Cathala, Pascal Poncelet
PKDD1