Sungjun Jung

dblp:227/6030 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0005-3232-4709ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 49% Hardware accelerators and domain-specific architectures · 32% Performance modeling and evaluation · 8%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 72% Indexing and storage engines · 28%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › similarity search › nearest neighbor search › approximate nearest neighbor search
graph-based approximate nearest neighbor search
0.912025
Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025
Information retrieval › similarity search
nearest neighbor search
0.912025
Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025
Indexing and storage engines
vector index
0.912025
Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025
Information retrieval
search engines
0.512021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021
Memory systems › in-memory computing
in-memory search accelerator
0.512021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021
Memory systems › processing-in-memory
near-data processing
0.512021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021
Hardware accelerators and domain-specific architectures
approximate computing accelerator
0.412020
A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020
Performance modeling and evaluation
approximation algorithms
0.412020
A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention accelerator
0.412020
A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020
Electronic design automation
hardware/software co-design
0.412020
A Specialized Architecture for Object Serialization with Applications to Big Data Analytics · ISCA 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.412020
A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020
Memory systems › DRAM › DRAM architecture
3D-stacked DRAM
0.412019
Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019
Memory systems › processing-in-memory
near-memory processing
0.412019
Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019
Memory systems
non-volatile memory
0.412019
Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019
Memory systems
processing-in-memory
0.412019
Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019
Memory systems › non-volatile memory
storage class memory
0.112021
BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021
Machine learning › Deep learning architectures and training
attention mechanism
0.112020
A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020
Storage systems › computational storage
in-storage computing
0.112020
A Specialized Architecture for Object Serialization with Applications to Big Data Analytics · ISCA 2020

Methods — techniques the papers use, named apart from their topics

early-termination search · 1.0decompression · 1.0neighbor selection · 0.9greedy search · 0.9angular distance estimation · 0.9hardware specialization · 0.9algorithmic approximation · 0.9cycle-level simulation · 0.4RTL design · 0.4specialized processing unit · 0.4near-memory processing · 0.4
YearPublicationVenuePosition
2025 Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search
abstract
Graph-based approximate nearest neighbor search (ANNS) algorithms are widely used to identify the most similar vectors to a given query vector. Graph-based ANNS consists of two stages: constructing a graph and searching on the graph for a given query vector. While reducing the query response time is of great practical importance, less attention has been paid to improving the online search method than the offline graph construction method. This paper provides an extensive experimental analysis on the popular greedy search and other search optimization strategies. We also propose a novel angular distance-guided search method for graph-based ANNS (ADA-NNS) to improve search efficiency. The key innovation of ADA-NNS is introducing a low-cost neighbor selection mechanism based on approximate similarity score derived from angular distance estimation, which effectively filters out less relevant neighbors. We compare state-of-the-art search techniques, including FINGER, on six datasets using different similarity metrics. It provides a comprehensive perspective on their tradeoffs in terms of throughput, latency, and recall. Our evaluation shows that ADA-NNS achieves 34%-107% higher queries per second (QPS) than the greedy search at 95% recall@10 on HNSW, one of the most popular graph structures for ANNS.
Sungjun Jung, Yongsang Park, Young H. Oh, Jae W. Lee
WWW1
2021 BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory
abstract
Search is one of the most popular and important web services. The inverted index is the standard data structure adopted by most full-text search engines. Recently, custom hardware accelerators for inverted index search have emerged to demonstrate much higher throughput than the conventional CPU or GPU. However, less attention has been paid to addressing the memory capacity pressure with inverted index. The conventional DDRx DRAM memory system significantly increases the system cost to make a terabyte-scale main memory. Instead, a shared memory pool composed of storage-class memory (SCM) devices is a promising alternative for scaling memory capacity at a much lower cost. However, this SCM-based pooled memory poses new challenges caused by the limited bandwidth of both SCM devices and the shared interconnect to the host CPU. Thus, we propose BOSS, the first near-data processing (NDP) architecture for inverted index search on SCM-based pooled memory, which maintains high throughput of query processing in this bandwidth- constrained environment. BOSS mitigates the impact of low bandwidth of SCM devices by employing early-termination search algorithms, reducing the footprint of intermediate data, and introducing a programmable decompression module that can select the best compression scheme for a given inverted index. Furthermore, BOSS includes a top-k selection module in hardware to substantially reduce the host-accelerator bandwidth consumption. Compared to Apache Lucene, a production-grade search engine library, running on 8 CPU cores, BOSS achieves a geomean speedup of 8.1× on various complex query types, while reducing the average energy consumption by 189×.
Jun Heo 0001, Seung Yul Lee, Sunhong Min, Yeonhong Park, Sungjun Jung, Tae Jun Ham, Jae W. Lee
ISCA5
2020 A3: Accelerating Attention Mechanisms in Neural Networks with Approximation
abstract
With the increasing computational demands of the neural networks, many hardware accelerators for the neural networks have been proposed. Such existing neural network accelerators often focus on popular neural network types such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs); however, not much attention has been paid to attention mechanisms, an emerging neural network primitive that enables neural networks to retrieve most relevant information from a knowledge-base, external memory, or past states. The attention mechanism is widely adopted by many state-of-the-art neural networks for computer vision, natural language processing, and machine translation, and accounts for a large portion of total execution time. We observe today's practice of implementing this mechanism using matrix-vector multiplication is suboptimal as the attention mechanism is semantically a content-based search where a large portion of computations ends up not being used. Based on this observation, we design and architect A3, which accelerates attention mechanisms in neural networks with algorithmic approximation and hardware specialization. Our proposed accelerator achieves multiple orders of magnitude improvement in energy efficiency (performance/watt) as well as substantial speedup over the state-of-the-art conventional hardware.
Tae Jun Ham, Sungjun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee 0002, Kyoung Park, Jae W. Lee, Deog-Kyoon Jeong
HPCA2
2020 A Specialized Architecture for Object Serialization with Applications to Big Data Analytics
abstract
Object serialization and deserialization (S/D) is an essential feature for efficient communication between distributed computing nodes with potentially non-uniform execution environments. S/D operations are widely used in big data analytics frameworks for remote procedure calls and massive data transfers like shuffles. However, frequent S/D operations incur significant performance and energy overheads as they must traverse and process a large object graph. Prior approaches improve S/D throughput by effectively hiding disk or network I/O latency with computation, increasing compression ratio, and/or application-specific customization. However, inherent dependencies in the existing (de)serialization formats and algorithms eventually become the major performance bottleneck. Thus, we propose Cereal, a specialized hardware accelerator for memory object serialization. By co-designing the serialization format with hardware architecture, Cereal effectively utilizes abundant parallelism in the S/D process to deliver high throughput. Cereal also employs an efficient object packing scheme to compress metadata such as object reference offsets and a space-efficient bitmap representation for the object layout. Our evaluation of Cereal using both a cycle-level simulator and synthesizable Chisel RTL demonstrates that Cereal delivers 43.4× higher average S/D throughput than 88 other S/D libraries on Java Serialization Benchmark Suite. For six Spark applications Cereal achieves 7.97× and 4.81× speedups on average for S/D operations over Java built-in serializer and Kryo, respectively, while saving S/D energy by 227.75× and 136.28×.
Jaeyoung Jang, Sungjun Jung, Sunmin Jeong, Jun Heo 0001, Hoon Shin, Tae Jun Ham, Jae W. Lee
ISCA2
2019 Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory
abstract
Garbage collection (GC) is a standard feature for high productivity programming, saving a programmer from many nasty memory-related bugs. However, these productivity benefits come with a cost in terms of application throughput, worst-case latency, and energy consumption. Since the first introduction of GC by the Lisp programming language in the 1950s, a myriad of hardware and software techniques have been proposed to reduce this cost. While the idea of accelerating GC in hardware is appealing, its impact has been very limited due to narrow coverage, lack of flexibility, intrusive system changes, and significant hardware cost. Even with specialized hardware GC performance is eventually limited by memory bandwidth bottleneck. Fortunately, emerging 3D stacked DRAM technologies shed new light on this decades-old problem by enabling efficient near-memory processing with ample memory bandwidth. Thus, we propose Charon1, the first 3D stacked memory-based GC accelerator. Through a detailed performance analysis of HotSpot JVM, we derive a set of key algorithmic primitives based on their GC time coverage and implementation complexity in hardware. Then we devise a specialized processing unit to substantially improve their memory-level parallelism and throughput with a low hardware cost. Our evaluation of Charon with the full-production HotSpot JVM running two big data analytics frameworks, Spark and GraphChi, demonstrates a 3.29× geomean speedup and 60.7% energy savings for GC over the baseline 8-core out-of-order processor.
Jaeyoung Jang, Jun Heo 0001, Yejin Lee 0001, Jaeyeon Won, Seonghak Kim, Sungjun Jung, Hakbeom Jang, Tae Jun Ham, Jae W. Lee
MICRO6
2018 A portable, automatic data qantizer for deep neural networks
abstract
With the proliferation of AI-based applications and services, there are strong demands for efficient processing of deep neural networks (DNNs). DNNs are known to be both compute-and memory-intensive as they require a tremendous amount of computation and large memory space. Quantization is a popular technique to boost efficiency of DNNs by representing a number with fewer bits, hence reducing both computational strength and memory footprint. However, it is a difficult task to find an optimal number representation for a DNN due to a combinatorial explosion in feasible number representations with varying bit widths, which is only exacerbated by layer-wise optimization. Besides, existing quantization techniques often target a specific DNN framework and/or hardware platform, lacking portability across various execution environments. To address this, we propose libnumber, a portable, automatic quantization framework for DNNs. By introducing Number abstract data type (ADT), libnumber encapsulates the internal representation of a number from the user. Then the auto-tuner of libnumber finds a compact representation (type, bit width, and bias) for the number that minimizes the user-supplied objective function, while satisfying the accuracy constraint. Thus, libnumber effectively separates the concern of developing an effective DNN model from low-level optimization of number representation. Our evaluation using eleven DNN models on two DNN frameworks targeting an FPGA platform demonstrates over 8× (7×) reduction in the parameter size on average when up to 7% (1%) loss of relative accuracy is tolerable, with a maximum reduction of 16×, compared to the baseline using 32-bit floating-point numbers. This leads to an geomean speedup of 3.79× with a maximum speedup of 12.77× over the baseline, while requiring only minimal programmer effort.
Young H. Oh, Quan Quan, Seonghak Kim, Jun Heo 0001, Sungjun Jung, Jaeyoung Jang, Jae W. Lee
PACT6