Luiz F. Mendes

dblp:76/5633 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 50% Data stream processing · 44% Query processing and optimization · 6%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.112008
Stream Sequential Pattern Mining with Precise Error Bounds · ICDM 2008
Data mining › pattern mining
sequential pattern mining
0.112008
Stream Sequential Pattern Mining with Precise Error Bounds · ICDM 2008
Data stream processing
stream mining
0.112008
Stream Sequential Pattern Mining with Precise Error Bounds · ICDM 2008
Data stream processing › stream mining
stream pattern mining
0.112008
Stream Sequential Pattern Mining with Precise Error Bounds · ICDM 2008
Data mining
approximation algorithm
0.012008
Stream Sequential Pattern Mining with Precise Error Bounds · ICDM 2008
Query processing and optimization › approximate query processing
error bounds
0.012008
Stream Sequential Pattern Mining with Precise Error Bounds · ICDM 2008

Methods — techniques the papers use, named apart from their topics

pruning strategies · 0.1batch processing · 0.1
YearPublicationVenuePosition
2008 Stream Sequential Pattern Mining with Precise Error Bounds
abstract
Sequential pattern mining is an interesting data mining problem with many real-world applications. This problem has been studied extensively in static databases. However, in recent years, emerging applications have introduced a new form of data called data stream. In a data stream, new elements are generated continuously. This poses additional constraints on the methods used for mining such data: memory usage is restricted, the infinitely flowing original dataset cannot be scanned multiple times, and current results should be available on demand.This paper introduces two effective methods for mining sequential patterns from data streams: the SS-BE method and the SS-MB method. The proposed methods break the stream into batches and only process each batch once. The two methods use different pruning strategies that restrict the memory usage but can still guarantee that all true sequential patterns are output at the end of any batch. Both algorithms scale linearly in execution time as the number of sequences grows, making them effective methods for sequential pattern mining in data streams. The experimental results also show that our methods are very accurate in that only a small fraction of the patterns that are output are false positives. Even for these false positives, SS-BE guarantees that their true support is above a pre-defined threshold.
Luiz F. Mendes, Bolin Ding, Jiawei Han 0001
ICDM1