Sunhong Min

dblp:298/8683 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 61% Hardware accelerators and domain-specific architectures · 34% Storage systems · 5%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 74% Machine learning and data management · 26%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.612022
ANNA: Specialized Architecture for Approximate Nearest Neighbor Search · HPCA 2022
Machine learning and data management › deep learning
graph neural network training
0.612022
Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory Caching · Proc. VLDB Endow. 2022
Information retrieval › similarity search
nearest neighbor search
0.612022
ANNA: Specialized Architecture for Approximate Nearest Neighbor Search · HPCA 2022
Hardware accelerators and domain-specific architectures › approximate computing accelerator
approximate nearest neighbor search accelerator
0.612022
ANNA: Specialized Architecture for Approximate Nearest Neighbor Search · HPCA 2022
Memory systems › cache
in-memory caching
0.612022
Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory Caching · Proc. VLDB Endow. 2022
Hardware accelerators and domain-specific architectures › domain-specific accelerator
vector search accelerator
0.612022
ANNA: Specialized Architecture for Approximate Nearest Neighbor Search · HPCA 2022
Information retrieval
search engines
0.512021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021
Memory systems › in-memory computing
in-memory search accelerator
0.512021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021
Memory systems › processing-in-memory
near-data processing
0.512021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021
Memory systems › data locality
data reuse
0.212022
ANNA: Specialized Architecture for Approximate Nearest Neighbor Search · HPCA 2022
Memory systems
non-volatile memory
0.212022
Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory Caching · Proc. VLDB Endow. 2022
Storage systems › flash and SSD › solid-state drive
NVMe SSD
0.212022
Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory Caching · Proc. VLDB Endow. 2022
Memory systems › non-volatile memory
storage class memory
0.112021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021

Methods — techniques the papers use, named apart from their topics

specialized dataflow pipeline · 1.1inspector-executor model · 1.1belady's algorithm · 1.1early-termination search · 1.0decompression · 1.0
YearPublicationVenuePosition
2022 ANNA: Specialized Architecture for Approximate Nearest Neighbor Search
abstract
Similarity search or nearest neighbor search is a task of retrieving a set of vectors in the (vector) database that are most similar to the provided query vector. It has been a key kernel for many applications for a long time. However, it is becoming especially more important in recent days as modern neural networks and machine learning models represent the semantics of images, videos, and documents as high-dimensional vectors called embeddings. Finding a set of similar embeddings for the provided query embedding is now the critical operation for modern recommender systems and semantic search engines. Since exhaustively searching for the most similar vectors out of billion vectors is such a prohibitive task, approximate nearest neighbor search (ANNS) is often utilized in many real-world use cases. Unfortunately, we find that utilizing the server-class CPUs and GPUs for the ANNS task leads to suboptimal performance and energy efficiency. To address such limitations, we propose a specialized architecture named ANNA (Approximate Nearest Neighbor search Accelerator), which is compatible with state-of-the-art ANNS algorithms such as Google ScaNN and Facebook Faiss. By combining the benefits of a specialized dataflow pipeline and efficient data reuse, ANNA achieves multiple orders of magnitude higher energy efficiency, 2.3-61.6× higher throughput, and 4.3-82.1× lower latency than the conventional CPU or GPU for both million- and billion-scale datasets.
Yejin Lee 0001, Hyunji Choi, Sunhong Min, Hyunseung Lee 0001, Sangwon Beak, Jae W. Lee, Tae Jun Ham
HPCA3
2022 Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory Caching
abstract
Graph Neural Networks (GNNs) are receiving a spotlight as a powerful tool that can effectively serve various inference tasks on graph structured data. As the size of real-world graphs continues to scale, the GNN training system faces a scalability challenge. Distributed training is a popular approach to address this challenge by scaling out CPU nodes. However, not much attention has been paid to disk-based GNN training, which can scale up the single-node system in a more cost-effective manner by leveraging high-performance storage devices like NVMe SSDs. We observe that the data movement between the main memory and the disk is the primary bottleneck in the SSD-based training system, and that the conventional GNN training pipeline is sub-optimal without taking this overhead into account. Thus, we propose Ginex, the first SSD-based GNN training system that can process billion-scale graph datasets on a single machine. Inspired by the inspector-executor execution model in compiler optimization, Ginex restructures the GNN training pipeline by separating sample and gather stages. This separation enables Ginex to realize a provably optimal replacement algorithm, known as Belady's algorithm , for caching feature vectors in memory, which account for the dominant portion of I/O accesses. According to our evaluation with four billion-scale graph datasets and two GNN models, Ginex achieves 2.11X higher training throughput on average (2.67X at maximum) than the SSD-extended PyTorch Geometric.
Yeonhong Park, Sunhong Min, Jae W. Lee
Proc. VLDB Endow.2
2021 BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory
abstract
Search is one of the most popular and important web services. The inverted index is the standard data structure adopted by most full-text search engines. Recently, custom hardware accelerators for inverted index search have emerged to demonstrate much higher throughput than the conventional CPU or GPU. However, less attention has been paid to addressing the memory capacity pressure with inverted index. The conventional DDRx DRAM memory system significantly increases the system cost to make a terabyte-scale main memory. Instead, a shared memory pool composed of storage-class memory (SCM) devices is a promising alternative for scaling memory capacity at a much lower cost. However, this SCM-based pooled memory poses new challenges caused by the limited bandwidth of both SCM devices and the shared interconnect to the host CPU. Thus, we propose BOSS, the first near-data processing (NDP) architecture for inverted index search on SCM-based pooled memory, which maintains high throughput of query processing in this bandwidth- constrained environment. BOSS mitigates the impact of low bandwidth of SCM devices by employing early-termination search algorithms, reducing the footprint of intermediate data, and introducing a programmable decompression module that can select the best compression scheme for a given inverted index. Furthermore, BOSS includes a top-k selection module in hardware to substantially reduce the host-accelerator bandwidth consumption. Compared to Apache Lucene, a production-grade search engine library, running on 8 CPU cores, BOSS achieves a geomean speedup of 8.1× on various complex query types, while reducing the average energy consumption by 189×.
Jun Heo 0001, Seung Yul Lee, Sunhong Min, Yeonhong Park, Sungjun Jung, Tae Jun Ham, Jae W. Lee
ISCA3