VLDB 2026 Research / reviewers in the wild / expert
Shengnan Cong
dblp:74/3510
· DBLP profile ↗
2ranked-venue papers
2as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 84% High-performance computing · 16% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.1 | 2 | 2005 | A sampling-based framework for parallel data mining · PPoPP 2005 Parallel mining of closed sequential patterns · KDD 2005 |
Data mining › pattern mining
sequential pattern mining |
0.1 | 2 | 2005 | A sampling-based framework for parallel data mining · PPoPP 2005 Parallel mining of closed sequential patterns · KDD 2005 |
Data mining › pattern mining › sequential pattern mining
closed sequential pattern mining |
0.1 | 1 | 2005 | Parallel mining of closed sequential patterns · KDD 2005 |
Data mining › pattern mining › itemset mining
frequent itemset mining |
0.1 | 1 | 2005 | A sampling-based framework for parallel data mining · PPoPP 2005 |
Data mining › big data analytics › large-scale data mining
parallel data mining |
0.1 | 1 | 2005 | A sampling-based framework for parallel data mining · PPoPP 2005 |
Parallel and multicore computing
parallel data mining |
0.1 | 1 | 2005 | Parallel mining of closed sequential patterns · KDD 2005 |
High-performance computing
distributed memory systems |
0.0 | 1 | 2005 | Parallel mining of closed sequential patterns · KDD 2005 |
Parallel and multicore computing › parallel algorithms › parallel algorithm design
divide-and-conquer parallelization |
0.0 | 1 | 2005 | A sampling-based framework for parallel data mining · PPoPP 2005 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2005 | A sampling-based framework for parallel data mining · PPoPP 2005 |
Methods — techniques the papers use, named apart from their topics
selective sampling · 0.2load balancing · 0.1dynamic scheduling · 0.1divide-and-conquer · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Parallel mining of closed sequential patternsabstractDiscovery of sequential patterns is an essential data mining task with broad applications. Among several variations of sequential patterns, closed sequential pattern is the most useful one since it retains all the information of the complete pattern set but is often much more compact than it. Unfortunately, there is no parallel closed sequential pattern mining method proposed yet. In this paper we develop an algorithm, called Par-CSP (Parallel Closed Sequential Pattern mining), to conduct parallel mining of closed sequential patterns on a distributed memory system. Par-CSP partitions the work among the processors by exploiting the divide-and-conquer property so that the overhead of interprocessor communication is minimized. Par-CSP applies dynamic scheduling to avoid processor idling. Moreover, it employs a technique, called selective sampling to address the load imbalance problem. We implement Par-CSP using MPI on a 64-node Linux cluster. Our experimental results show that Par-CSP attains good parallelization efficiencies on various input datasets. Shengnan Cong, Jiawei Han 0001, David A. Padua |
KDD | 1 |
| 2005 | A sampling-based framework for parallel data miningabstractThe goal of data mining algorithm is to discover useful information embedded in large databases. Frequent itemset mining and sequential pattern mining are two important data mining problems with broad applications. Perhaps the most efficient way to solve these problems sequentially is to apply a pattern-growth algorithm, which is a divide-and-conquer algorithm [9, 10]. In this paper, we present a framework for parallel mining frequent itemsets and sequential patterns based on the divide-and-conquer strategy of pattern growth. Then, we discuss the load balancing problem and introduce a sampling technique, called selective sampling, to address this problem. We implemented parallel versions of both frequent itemsets and sequential pattern mining algorithms following our framework. The experimental results show that our parallel algorithms usually achieve excellent speedups. Shengnan Cong, Jiawei Han 0001, Jay P. Hoeflinger, David A. Padua |
PPoPP | 1 |