Nhat-Phuong Tran

dblp:64/10422 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0002-4374-9068ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.612022
Incremental Density-Based Clustering on Multicore Processors · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Data mining › clustering
density-based clustering
0.612022
Incremental Density-Based Clustering on Multicore Processors · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Parallel and multicore computing › parallel data mining
parallel clustering
0.612022
Incremental Density-Based Clustering on Multicore Processors · IEEE Trans. Pattern Anal. Mach. Intell. 2022

Methods — techniques the papers use, named apart from their topics

parallelization · 1.1object node graph · 1.1incremental clustering · 1.1
YearPublicationVenuePosition
2022 Incremental Density-Based Clustering on Multicore Processors
abstract
The density-based clustering algorithm is a fundamental data clustering technique with many real-world applications. However, when the database is frequently changed, how to effectively update clustering results rather than reclustering from scratch remains a challenging task. In this work, we introduce IncAnyDBC, a unique parallel incremental data clustering approach to deal with this problem. First, IncAnyDBC can process changes in bulks rather than batches like state-of-the-art methods for reducing update overheads. Second, it keeps an underlying cluster structure called the object node graph during the clustering process and uses it as a basis for incrementally updating clusters wrt. inserted or deleted objects in the database by propagating changes around affected nodes only. In additional, IncAnyDBC actively and iteratively examines the graph and chooses only a small set of most meaningful objects to produce exact clustering results of DBSCAN or to approximate results under arbitrary time constraints. This makes it more efficient than other existing methods. Third, by processing objects in blocks, IncAnyDBC can be efficiently parallelized on multicore CPUs, thus creating a work-efficient method. It runs much faster than existing techniques using one thread while still scaling well with multiple threads. Experiments are conducted on various large real datasets for demonstrating the performance of IncAnyDBC.
Son T. Mai, Jon Jacobsen, Sihem Amer-Yahia, Ivor T. A. Spence, Nhat-Phuong Tran, Ira Assent, Nguyen Quoc Viet Hung
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 Memory-Efficient Parallelization of 3D Lattice Boltzmann Flow Solver on a GPU
abstract
Lattice Boltzmann Method (LBM) is a powerful numerical simulation method of the fluid flow. With its data parallel nature and the simple kernel structure, it is a promising candidate for a parallel implementation on a GPU. The LBM, however, is heavily data-intensive and memory bound. In particular, moving the data to the adjacent cells in the streaming computation phase of the LBM incurs a lot of uncoalesced accesses on the GPU which affects the overall performance. In this paper, we parallelize the LBM on a GPU by incorporating memory-efficient techniques such as the tiling optimization with the data layout changes and the data update scheme so called a pull scheme. Furthermore, we developed optimization techniques such as removing branch divergences, reducing the register uses, and reducing the number of double precision floating-point instructions. Experimental results on Nvidia Tesla K20 GPU show that our approach delivers up to 1105 MLUPS (Million Lattice Updates Per Second) and 156-times speedup compared with a serial implementation.
Nhat-Phuong Tran, Myungho Lee, Dong Hoon Choi
HiPC1