EDBT 2026 Demo / reviewers in the wild / expert
Sungjun Jung
dblp:227/6030
· DBLP profile ↗
6ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0005-3232-4709ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 49% Hardware accelerators and domain-specific architectures · 32% Performance modeling and evaluation · 8% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 72% Indexing and storage engines · 28% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › similarity search › nearest neighbor search › approximate nearest neighbor search
graph-based approximate nearest neighbor search |
0.9 | 1 | 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025 |
Information retrieval › similarity search
nearest neighbor search |
0.9 | 1 | 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025 |
Indexing and storage engines
vector index |
0.9 | 1 | 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025 |
Information retrieval
search engines |
0.5 | 1 | 2021 | BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021 |
Memory systems › in-memory computing
in-memory search accelerator |
0.5 | 1 | 2021 | BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021 |
Memory systems › processing-in-memory
near-data processing |
0.5 | 1 | 2021 | BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021 |
Hardware accelerators and domain-specific architectures
approximate computing accelerator |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Performance modeling and evaluation
approximation algorithms |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention accelerator |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Electronic design automation
hardware/software co-design |
0.4 | 1 | 2020 | A Specialized Architecture for Object Serialization with Applications to Big Data Analytics · ISCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Memory systems › DRAM › DRAM architecture
3D-stacked DRAM |
0.4 | 1 | 2019 | Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019 |
Memory systems › processing-in-memory
near-memory processing |
0.4 | 1 | 2019 | Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019 |
Memory systems
non-volatile memory |
0.4 | 1 | 2019 | Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019 |
Memory systems
processing-in-memory |
0.4 | 1 | 2019 | Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory · MICRO 2019 |
Memory systems › non-volatile memory
storage class memory |
0.1 | 1 | 2021 | BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory · ISCA 2021 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Storage systems › computational storage
in-storage computing |
0.1 | 1 | 2020 | A Specialized Architecture for Object Serialization with Applications to Big Data Analytics · ISCA 2020 |
Methods — techniques the papers use, named apart from their topics
early-termination search · 1.0decompression · 1.0neighbor selection · 0.9greedy search · 0.9angular distance estimation · 0.9hardware specialization · 0.9algorithmic approximation · 0.9cycle-level simulation · 0.4RTL design · 0.4specialized processing unit · 0.4near-memory processing · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor SearchabstractGraph-based approximate nearest neighbor search (ANNS) algorithms are widely used to identify the most similar vectors to a given query vector. Graph-based ANNS consists of two stages: constructing a graph and searching on the graph for a given query vector. While reducing the query response time is of great practical importance, less attention has been paid to improving the online search method than the offline graph construction method. This paper provides an extensive experimental analysis on the popular greedy search and other search optimization strategies. We also propose a novel angular distance-guided search method for graph-based ANNS (ADA-NNS) to improve search efficiency. The key innovation of ADA-NNS is introducing a low-cost neighbor selection mechanism based on approximate similarity score derived from angular distance estimation, which effectively filters out less relevant neighbors. We compare state-of-the-art search techniques, including FINGER, on six datasets using different similarity metrics. It provides a comprehensive perspective on their tradeoffs in terms of throughput, latency, and recall. Our evaluation shows that ADA-NNS achieves 34%-107% higher queries per second (QPS) than the greedy search at 95% recall@10 on HNSW, one of the most popular graph structures for ANNS. Sungjun Jung, Yongsang Park, Young H. Oh, Jae W. Lee |
WWW | 1 |
| 2021 | BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class MemoryabstractSearch is one of the most popular and important web services. The inverted index is the standard data structure adopted by most full-text search engines. Recently, custom hardware accelerators for inverted index search have emerged to demonstrate much higher throughput than the conventional CPU or GPU. However, less attention has been paid to addressing the memory capacity pressure with inverted index. The conventional DDRx DRAM memory system significantly increases the system cost to make a terabyte-scale main memory. Instead, a shared memory pool composed of storage-class memory (SCM) devices is a promising alternative for scaling memory capacity at a much lower cost. However, this SCM-based pooled memory poses new challenges caused by the limited bandwidth of both SCM devices and the shared interconnect to the host CPU. Thus, we propose BOSS, the first near-data processing (NDP) architecture for inverted index search on SCM-based pooled memory, which maintains high throughput of query processing in this bandwidth- constrained environment. BOSS mitigates the impact of low bandwidth of SCM devices by employing early-termination search algorithms, reducing the footprint of intermediate data, and introducing a programmable decompression module that can select the best compression scheme for a given inverted index. Furthermore, BOSS includes a top-k selection module in hardware to substantially reduce the host-accelerator bandwidth consumption. Compared to Apache Lucene, a production-grade search engine library, running on 8 CPU cores, BOSS achieves a geomean speedup of 8.1× on various complex query types, while reducing the average energy consumption by 189×. Jun Heo 0001, Seung Yul Lee, Sunhong Min, Yeonhong Park, Sungjun Jung, Tae Jun Ham, Jae W. Lee |
ISCA | 5 |
| 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with ApproximationabstractWith the increasing computational demands of the neural networks, many hardware accelerators for the neural networks have been proposed. Such existing neural network accelerators often focus on popular neural network types such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs); however, not much attention has been paid to attention mechanisms, an emerging neural network primitive that enables neural networks to retrieve most relevant information from a knowledge-base, external memory, or past states. The attention mechanism is widely adopted by many state-of-the-art neural networks for computer vision, natural language processing, and machine translation, and accounts for a large portion of total execution time. We observe today's practice of implementing this mechanism using matrix-vector multiplication is suboptimal as the attention mechanism is semantically a content-based search where a large portion of computations ends up not being used. Based on this observation, we design and architect A3, which accelerates attention mechanisms in neural networks with algorithmic approximation and hardware specialization. Our proposed accelerator achieves multiple orders of magnitude improvement in energy efficiency (performance/watt) as well as substantial speedup over the state-of-the-art conventional hardware. Tae Jun Ham, Sungjun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee 0002, Kyoung Park, Jae W. Lee, Deog-Kyoon Jeong |
HPCA | 2 |
| 2020 | A Specialized Architecture for Object Serialization with Applications to Big Data AnalyticsabstractObject serialization and deserialization (S/D) is an essential feature for efficient communication between distributed computing nodes with potentially non-uniform execution environments. S/D operations are widely used in big data analytics frameworks for remote procedure calls and massive data transfers like shuffles. However, frequent S/D operations incur significant performance and energy overheads as they must traverse and process a large object graph. Prior approaches improve S/D throughput by effectively hiding disk or network I/O latency with computation, increasing compression ratio, and/or application-specific customization. However, inherent dependencies in the existing (de)serialization formats and algorithms eventually become the major performance bottleneck. Thus, we propose Cereal, a specialized hardware accelerator for memory object serialization. By co-designing the serialization format with hardware architecture, Cereal effectively utilizes abundant parallelism in the S/D process to deliver high throughput. Cereal also employs an efficient object packing scheme to compress metadata such as object reference offsets and a space-efficient bitmap representation for the object layout. Our evaluation of Cereal using both a cycle-level simulator and synthesizable Chisel RTL demonstrates that Cereal delivers 43.4× higher average S/D throughput than 88 other S/D libraries on Java Serialization Benchmark Suite. For six Spark applications Cereal achieves 7.97× and 4.81× speedups on average for S/D operations over Java built-in serializer and Kryo, respectively, while saving S/D energy by 227.75× and 136.28×. Jaeyoung Jang, Sungjun Jung, Sunmin Jeong, Jun Heo 0001, Hoon Shin, Tae Jun Ham, Jae W. Lee |
ISCA | 2 |
| 2019 | Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in MemoryabstractGarbage collection (GC) is a standard feature for high productivity programming, saving a programmer from many nasty memory-related bugs. However, these productivity benefits come with a cost in terms of application throughput, worst-case latency, and energy consumption. Since the first introduction of GC by the Lisp programming language in the 1950s, a myriad of hardware and software techniques have been proposed to reduce this cost. While the idea of accelerating GC in hardware is appealing, its impact has been very limited due to narrow coverage, lack of flexibility, intrusive system changes, and significant hardware cost. Even with specialized hardware GC performance is eventually limited by memory bandwidth bottleneck. Fortunately, emerging 3D stacked DRAM technologies shed new light on this decades-old problem by enabling efficient near-memory processing with ample memory bandwidth. Thus, we propose Charon1, the first 3D stacked memory-based GC accelerator. Through a detailed performance analysis of HotSpot JVM, we derive a set of key algorithmic primitives based on their GC time coverage and implementation complexity in hardware. Then we devise a specialized processing unit to substantially improve their memory-level parallelism and throughput with a low hardware cost. Our evaluation of Charon with the full-production HotSpot JVM running two big data analytics frameworks, Spark and GraphChi, demonstrates a 3.29× geomean speedup and 60.7% energy savings for GC over the baseline 8-core out-of-order processor. Jaeyoung Jang, Jun Heo 0001, Yejin Lee 0001, Jaeyeon Won, Seonghak Kim, Sungjun Jung, Hakbeom Jang, Tae Jun Ham, Jae W. Lee |
MICRO | 6 |
| 2018 | A portable, automatic data qantizer for deep neural networksabstractWith the proliferation of AI-based applications and services, there are strong demands for efficient processing of deep neural networks (DNNs). DNNs are known to be both compute-and memory-intensive as they require a tremendous amount of computation and large memory space. Quantization is a popular technique to boost efficiency of DNNs by representing a number with fewer bits, hence reducing both computational strength and memory footprint. However, it is a difficult task to find an optimal number representation for a DNN due to a combinatorial explosion in feasible number representations with varying bit widths, which is only exacerbated by layer-wise optimization. Besides, existing quantization techniques often target a specific DNN framework and/or hardware platform, lacking portability across various execution environments. To address this, we propose libnumber, a portable, automatic quantization framework for DNNs. By introducing Number abstract data type (ADT), libnumber encapsulates the internal representation of a number from the user. Then the auto-tuner of libnumber finds a compact representation (type, bit width, and bias) for the number that minimizes the user-supplied objective function, while satisfying the accuracy constraint. Thus, libnumber effectively separates the concern of developing an effective DNN model from low-level optimization of number representation. Our evaluation using eleven DNN models on two DNN frameworks targeting an FPGA platform demonstrates over 8× (7×) reduction in the parameter size on average when up to 7% (1%) loss of relative accuracy is tolerable, with a maximum reduction of 16×, compared to the baseline using 32-bit floating-point numbers. This leads to an geomean speedup of 3.79× with a maximum speedup of 12.77× over the baseline, while requiring only minimal programmer effort. Young H. Oh, Quan Quan, Seonghak Kim, Jun Heo 0001, Sungjun Jung, Jaeyoung Jang, Jae W. Lee |
PACT | 6 |