Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Helen Pinto

dblp:16/2974 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.122004
Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach · IEEE Trans. Knowl. Data Eng. 2004
PrefixSpan: Mining Sequential Patterns by Prefix-Projected Growth · ICDE 2001
Data mining › pattern mining
sequential pattern mining
0.122004
Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach · IEEE Trans. Knowl. Data Eng. 2004
PrefixSpan: Mining Sequential Patterns by Prefix-Projected Growth · ICDE 2001
Data mining › pattern mining
pattern-growth
0.012004
Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach · IEEE Trans. Knowl. Data Eng. 2004

Methods — techniques the papers use, named apart from their topics

pseudo projection · 0.0prefix-projection · 0.0
YearPublicationVenuePosition
2016 A Symbolic Tree Model for Oil and Gas Production Prediction Using Time-Series Production Data
abstract
Oil and gas well production prediction takes place in early stages of production to estimate future recovery. A data driven workflow is proposed in this paper to construct a symbolic tree model to predict new well production using historic time-series production data of analogous wells. Production data are firstly aggregated and symbolized for dimensionality reduction and data discretization of time-series data. A symbolic tree is constructed on time-series symbol sequences, and pre-pruning mechanisms - minimum node size and spatial information gain - are integrated to achieve a compact and informative tree. A coverage index is used to assess the tree size. A case study was conducted applying the proposed workflow to shale gas wells in Montney-A pool in Canada. It has proved the feasibility and accuracy of the proposed method.
Bingjie Wei, Helen Pinto, Xin Wang 0002
DSAA2
2004 Mining Sequential Patterns by Pattern-Growth: The PrefixSpan Approach
abstract
Sequential pattern mining is an important data mining problem with broad applications. However, it is also a difficult problem since the mining may have to generate or examine a combinatorially explosive number of intermediate subsequences. Most of the previously developed sequential pattern mining methods, such as GSP, explore a candidate generation-and-test approach [R. Agrawal et al. (1994)] to reduce the number of candidates to be examined. However, this approach may not be efficient in mining large sequence databases having numerous patterns and/or long patterns. In this paper, we propose a projection-based, sequential pattern-growth approach for efficient mining of sequential patterns. In this approach, a sequence database is recursively projected into a set of smaller projected databases, and sequential patterns are grown in each projected database by exploring only locally frequent fragments. Based on an initial study of the pattern growth-based sequential pattern mining, FreeSpan [J. Han et al. (2000)], we propose a more efficient method, called PSP, which offers ordered growth and reduced projected databases. To further improve the performance, a pseudoprojection technique is developed in PrefixSpan. A comprehensive performance study shows that PrefixSpan, in most cases, outperforms the a priori-based algorithm GSP, FreeSpan, and SPADE [M. Zaki, (2001)] (a sequential pattern mining algorithm that adopts vertical data format), and PrefixSpan integrated with pseudoprojection is the fastest among all the tested algorithms. Furthermore, this mining methodology can be extended to mining sequential patterns with user-specified constraints. The high promise of the pattern-growth approach may lead to its further extension toward efficient mining of other kinds of frequent patterns, such as frequent substructures.
Jian Pei 0001, Jiawei Han 0001, Behzad Mortazavi-Asl, Jianyong Wang 0001, Helen Pinto, Umeshwar Dayal, Meichun Hsu
IEEE Trans. Knowl. Data Eng.5
2001 Multi-Dimensional Sequential Pattern Mining
abstract
Sequential pattern mining, which finds the set of frequent subsequences in sequence databases, is an important data-mining task and has broad applications. Usually, sequence patterns are associated with different circumstances, and such circumstances form a multiple dimensional space. For example, customer purchase sequences are associated with region, time, customer group, and others. It is interesting and useful to mine sequential patterns associated with multi-dimensional information.In this paper, we propose the theme of multi-dimensional sequential pattern mining, which integrates the multidimensional analysis and sequential data mining. We also thoroughly explore efficient methods for multi-dimensional sequential pattern mining. We examine feasible combinations of efficient sequential pattern mining and multi-dimensional analysis methods, as well as develop uniform methods for high-performance mining. Extensive experiments show the advantages as well as limitations of these methods. Some recommendations on selecting proper method with respect to data set properties are drawn.
Helen Pinto, Jiawei Han 0001, Jian Pei 0001, Ke Wang 0001, Umeshwar Dayal
CIKM1
2001 PrefixSpan: Mining Sequential Patterns by Prefix-Projected Growth
abstract
Sequential pattern mining is an important data mining problem with broad applications. It is challenging since one may need to examine a combinatorially explosive number of possible subsequence patterns. Most of the previously developed sequential pattern mining methods follow the methodology of \t which may substantially reduce the number of combinations to be examined. However, \t still encounters problems when a sequence database is large and/or when sequential patterns to be mined are numerous and/or long. In this paper, we propose a novel sequential pattern mining method, called PrefixSpan (i.e., Prefix-projected Sequential pattern mining), which explores prefixprojection in sequential pattern mining. PrefixSpan mines the complete set of patterns but greatly reduces the efforts of candidate subsequence generation. Moreover, prefix-projection substantially reduces the size of projected databases and leads to efficient processing. Our performance study shows that PrefixSpan outperforms both the -based GSP algorithm and another recently proposed method, FreeSpan, in mining large sequence databases. 1
Jian Pei 0001, Jiawei Han 0001, Behzad Mortazavi-Asl, Helen Pinto, Umeshwar Dayal, Meichun Hsu
ICDE4