Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shengnan Cong

dblp:74/3510 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
0since 2021 · last 2005
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 84% High-performance computing · 16%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.122005
A sampling-based framework for parallel data mining · PPoPP 2005
Parallel mining of closed sequential patterns · KDD 2005
Data mining › pattern mining
sequential pattern mining
0.122005
A sampling-based framework for parallel data mining · PPoPP 2005
Parallel mining of closed sequential patterns · KDD 2005
Data mining › pattern mining › sequential pattern mining
closed sequential pattern mining
0.112005
Parallel mining of closed sequential patterns · KDD 2005
Data mining › pattern mining › itemset mining
frequent itemset mining
0.112005
A sampling-based framework for parallel data mining · PPoPP 2005
Data mining › big data analytics › large-scale data mining
parallel data mining
0.112005
A sampling-based framework for parallel data mining · PPoPP 2005
Parallel and multicore computing
parallel data mining
0.112005
Parallel mining of closed sequential patterns · KDD 2005
High-performance computing
distributed memory systems
0.012005
Parallel mining of closed sequential patterns · KDD 2005
Parallel and multicore computing › parallel algorithms › parallel algorithm design
divide-and-conquer parallelization
0.012005
A sampling-based framework for parallel data mining · PPoPP 2005
Parallel and multicore computing
parallel programming models
0.012005
A sampling-based framework for parallel data mining · PPoPP 2005

Methods — techniques the papers use, named apart from their topics

selective sampling · 0.2load balancing · 0.1dynamic scheduling · 0.1divide-and-conquer · 0.1
YearPublicationVenuePosition
2005 Parallel mining of closed sequential patterns
abstract
Discovery of sequential patterns is an essential data mining task with broad applications. Among several variations of sequential patterns, closed sequential pattern is the most useful one since it retains all the information of the complete pattern set but is often much more compact than it. Unfortunately, there is no parallel closed sequential pattern mining method proposed yet. In this paper we develop an algorithm, called Par-CSP (Parallel Closed Sequential Pattern mining), to conduct parallel mining of closed sequential patterns on a distributed memory system. Par-CSP partitions the work among the processors by exploiting the divide-and-conquer property so that the overhead of interprocessor communication is minimized. Par-CSP applies dynamic scheduling to avoid processor idling. Moreover, it employs a technique, called selective sampling to address the load imbalance problem. We implement Par-CSP using MPI on a 64-node Linux cluster. Our experimental results show that Par-CSP attains good parallelization efficiencies on various input datasets.
Shengnan Cong, Jiawei Han 0001, David A. Padua
KDD1
2005 A sampling-based framework for parallel data mining
abstract
The goal of data mining algorithm is to discover useful information embedded in large databases. Frequent itemset mining and sequential pattern mining are two important data mining problems with broad applications. Perhaps the most efficient way to solve these problems sequentially is to apply a pattern-growth algorithm, which is a divide-and-conquer algorithm [9, 10]. In this paper, we present a framework for parallel mining frequent itemsets and sequential patterns based on the divide-and-conquer strategy of pattern growth. Then, we discuss the load balancing problem and introduce a sampling technique, called selective sampling, to address this problem. We implemented parallel versions of both frequent itemsets and sequential pattern mining algorithms following our framework. The experimental results show that our parallel algorithms usually achieve excellent speedups.
Shengnan Cong, Jiawei Han 0001, Jay P. Hoeflinger, David A. Padua
PPoPP1