Wen K. Lee

dblp:20/2603 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 1994
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.011994
A Decomposition-Based Simulated Annealing Technique for Data Clustering · PODS 1994
Storage systems › i/o optimization
disk i/o reduction
0.011994
A Decomposition-Based Simulated Annealing Technique for Data Clustering · PODS 1994

Methods — techniques the papers use, named apart from their topics

statistical sampling · 0.0simulated annealing · 0.0decomposition-based approach · 0.0
YearPublicationVenuePosition
1994 A Decomposition-Based Simulated Annealing Technique for Data Clustering
abstract
It has been demonstrated that simulated annealing provides high-quality results for the data clustering problem. However, existing simulated annealing schemes are memory-based algorithms; they are not suited for solving large problems such as data clustering which typically are too big to fit in the memory space in its entirety. Various buffer replacement policies, assuming either temporal or spatial locality, are not useful in this case since simulated annealing is based on a randomized search process. Poor locality of references will cause the memory to thrash because too many replacements are required. This phenomenon will incur excessive disk accesses and force the machine to run at the speed of the I/O subsystem. In this paper, we formulate the data clustering problem as a graph partition problem (GPP), and propose a decomposition-based approach to address the issue of excessive disk accesses during annealing. We apply the statistical sampling technique to randomly select subgraphs of the GPP into memory for annealing. Both the analytical and experimental studies indicate that the decomposition-based approach can dramatically reduce the costly disk I/O activities while obtaining excellent optimized results.
Kien A. Hua, Sheau-Dong Lang, Wen K. Lee
PODS3
1992 Parallel Simulated Annealing for Efficient Data Clustering
Kien A. Hua, Wen K. Lee, Sheau-Dong Lang
DEXA2