EDBT 2026 Demo / reviewers in the wild / expert
Young H. Oh
dblp:149/4002
· DBLP profile ↗
8ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0001-5971-9093ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 40% Cloud and datacenter computing · 20% Processor architecture and microarchitecture · 19% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 67% Indexing and storage engines · 33% | |
| Software engineering, system software, and programming languages
2 papers |
Runtime systems and virtual machines · 74% Programming languages and type systems · 26% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 54% Deep learning architectures and training · 46% |
Topics — the 21 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › similarity search › nearest neighbor search › approximate nearest neighbor search
graph-based approximate nearest neighbor search |
0.9 | 1 | 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025 |
Information retrieval › similarity search
nearest neighbor search |
0.9 | 1 | 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025 |
Indexing and storage engines
vector index |
0.9 | 1 | 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor Search · WWW 2025 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.5 | 1 | 2021 | Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise Scheduling · HPCA 2021 |
Cloud and datacenter computing
inference serving |
0.5 | 1 | 2021 | Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise Scheduling · HPCA 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit |
0.5 | 1 | 2021 | Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise Scheduling · HPCA 2021 |
Hardware accelerators and domain-specific architectures
approximate computing accelerator |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Performance modeling and evaluation
approximation algorithms |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention accelerator |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Hardware accelerators and domain-specific architectures › spatial architecture
dataflow accelerator |
0.4 | 1 | 2020 | Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.4 | 1 | 2020 | Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.4 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Processor architecture and microarchitecture
instruction set architecture |
0.3 | 1 | 2017 | Typed Architectures: Architectural Support for Lightweight Scripting · ASPLOS 2017 |
Runtime systems and virtual machines
interpreter |
0.2 | 1 | 2016 | Short-Circuit Dispatch: Accelerating Virtual Machine Interpreters on Embedded Processors · ISCA 2016 |
Processor architecture and microarchitecture
branch prediction |
0.2 | 1 | 2016 | Short-Circuit Dispatch: Accelerating Virtual Machine Interpreters on Embedded Processors · ISCA 2016 |
Processor architecture and microarchitecture › branch prediction
branch target buffer |
0.2 | 1 | 2016 | Short-Circuit Dispatch: Accelerating Virtual Machine Interpreters on Embedded Processors · ISCA 2016 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with Approximation · HPCA 2020 |
Bioinformatics and computational biology › genomics
genomic data analysis |
0.1 | 1 | 2020 | Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020 |
Cloud and datacenter computing › cloud deployment
cloud FPGA deployment |
0.1 | 1 | 2020 | Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020 |
Reconfigurable computing and FPGAs › cloud FPGA
FPGA-as-a-service |
0.1 | 1 | 2020 | Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020 |
Embedded and real-time systems
embedded processor |
0.1 | 1 | 2016 | Short-Circuit Dispatch: Accelerating Virtual Machine Interpreters on Embedded Processors · ISCA 2016 |
Methods — techniques the papers use, named apart from their topics
layer-wise scheduling · 1.0neighbor selection · 0.9greedy search · 0.9angular distance estimation · 0.9on-chip scratchpad · 0.9non-blocking APIs · 0.9hardware specialization · 0.9extended SQL · 0.9dataflow architecture · 0.9algorithmic approximation · 0.9polymorphic instructions · 0.6hardware type checking · 0.6time-multiplexing · 0.5time multiplexing · 0.5type tag extraction · 0.3short-circuit dispatch · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Angular Distance-Guided Neighbor Selection for Graph-Based Approximate Nearest Neighbor SearchabstractGraph-based approximate nearest neighbor search (ANNS) algorithms are widely used to identify the most similar vectors to a given query vector. Graph-based ANNS consists of two stages: constructing a graph and searching on the graph for a given query vector. While reducing the query response time is of great practical importance, less attention has been paid to improving the online search method than the offline graph construction method. This paper provides an extensive experimental analysis on the popular greedy search and other search optimization strategies. We also propose a novel angular distance-guided search method for graph-based ANNS (ADA-NNS) to improve search efficiency. The key innovation of ADA-NNS is introducing a low-cost neighbor selection mechanism based on approximate similarity score derived from angular distance estimation, which effectively filters out less relevant neighbors. We compare state-of-the-art search techniques, including FINGER, on six datasets using different similarity metrics. It provides a comprehensive perspective on their tradeoffs in terms of throughput, latency, and recall. Our evaluation shows that ADA-NNS achieves 34%-107% higher queries per second (QPS) than the greedy search at 95% recall@10 on HNSW, one of the most popular graph structures for ANNS. Sungjun Jung, Yongsang Park, Young H. Oh, Jae W. Lee |
WWW | 4 |
| 2021 | Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingabstractTo meet surging demands for deep learning inference services, many cloud computing vendors employ high-performance specialized accelerators, called neural processing units (NPUs). One important challenge for effective use of NPUs is to achieve high resource utilization over a wide spectrum of deep neural network (DNN) models with diverse arithmetic intensities. There is often an intrinsic mismatch between the compute-to-memory bandwidth ratio of an NPU and the arithmetic intensity of the model it executes, leading to under-utilization of either compute resources or memory bandwidth. Ideally, we want to saturate both compute TOP/s and DRAM bandwidth to achieve high system throughput. Thus, we propose Layerweaver, an inference serving system with a novel multi-model time-multiplexing scheduler for NPUs. Layerweaver reduces the temporal waste of computation resources by interweaving layer execution of multiple different models with opposing characteristics: compute-intensive and memory-intensive. Layerweaver hides the memory time of a memory-intensive model by overlapping it with the relatively long computation time of a compute-intensive model, thereby minimizing the idle time of the computation units waiting for off-chip data transfers. For a two-model serving scenario of batch 1 with 16 different pairs of compute- and memory-intensive models, Layerweaver improves the temporal utilization of computation units and memory channels by 44.0% and 28.7%, respectively, to increase the system throughput by 60.1% on average, over the baseline executing one model at a time. Young H. Oh, Seonghak Kim, Yunho Jin, Sam Son, Jonghyun Bae, Jongsung Lee 0001, Yeonhong Park, Dong Uk Kim, Tae Jun Ham, Jae W. Lee |
HPCA | 1 |
| 2020 | A3: Accelerating Attention Mechanisms in Neural Networks with ApproximationabstractWith the increasing computational demands of the neural networks, many hardware accelerators for the neural networks have been proposed. Such existing neural network accelerators often focus on popular neural network types such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs); however, not much attention has been paid to attention mechanisms, an emerging neural network primitive that enables neural networks to retrieve most relevant information from a knowledge-base, external memory, or past states. The attention mechanism is widely adopted by many state-of-the-art neural networks for computer vision, natural language processing, and machine translation, and accounts for a large portion of total execution time. We observe today's practice of implementing this mechanism using matrix-vector multiplication is suboptimal as the attention mechanism is semantically a content-based search where a large portion of computations ends up not being used. Based on this observation, we design and architect A3, which accelerates attention mechanisms in neural networks with algorithmic approximation and hardware specialization. Our proposed accelerator achieves multiple orders of magnitude improvement in energy efficiency (performance/watt) as well as substantial speedup over the state-of-the-art conventional hardware. Tae Jun Ham, Sungjun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee 0002, Kyoung Park, Jae W. Lee, Deog-Kyoon Jeong |
HPCA | 4 |
| 2020 | Genesis: A Hardware Acceleration Framework for Genomic Data AnalysisabstractIn this paper, we describe our vision to accelerate algorithms in the domain of genomic data analysis by proposing a framework called Genesis (genome analysis) that contains an interface and an implementation of a system that processes genomic data efficiently. This framework can be deployed in the cloud and exploit the FPGAs-as-a-service paradigm to provide cost-efficient secondary DNA analysis. We propose conceptualizing genomic reads and associated read attributes as a very large relational database and using extended SQL as a domain-specific language to construct queries that form various data manipulation operations. To accelerate such queries, we design a Genesis hardware library which consists of primitive hardware modules that can be composed to construct a dataflow architecture specialized for those queries. As a proof of concept for the Genesis framework, we present the architecture and the hardware implementation of several genomic analysis stages in the secondary analysis pipeline corresponding to the best known software analysis toolkit, GATK4 workflow proposed by the Broad Institute. We walk through the construction of genomic data analysis operations using a sequence of SQL-style queries and show how Genesis hardware library modules can be utilized to construct the hardware pipelines designed to accelerate such queries. We exploit parallelism and data reuse by utilizing a dataflow architecture along with the use of on-chip scratchpads as well as non-blocking APIs to manage the accelerators, allowing concurrent execution of the accelerator and the host. Our accelerated system deployed on the cloud FPGA performs up to 19.3× better than GATK4 running on a commodity multi-core Xeon server and obtains up to 15× better cost savings. We believe that if a software algorithm can be mapped onto a hardware library to utilize the underlying accelerator(s) using an already-standardized software interface such as SQL, while allowing the efficient mapping of such interface to primitive hardware modules as we have demonstrated here, it will expedite the acceleration of domainspecific algorithms and allow the easy adaptation of algorithm changes. Tae Jun Ham, David Bruns-Smith, Brendan Sweeney, Yejin Lee 0001, Seong Hoon Seo, U. Gyeong Song, Young H. Oh, Krste Asanovic, Jae W. Lee, Lisa Wu Wills |
ISCA | 7 |
| 2018 | A portable, automatic data qantizer for deep neural networksabstractWith the proliferation of AI-based applications and services, there are strong demands for efficient processing of deep neural networks (DNNs). DNNs are known to be both compute-and memory-intensive as they require a tremendous amount of computation and large memory space. Quantization is a popular technique to boost efficiency of DNNs by representing a number with fewer bits, hence reducing both computational strength and memory footprint. However, it is a difficult task to find an optimal number representation for a DNN due to a combinatorial explosion in feasible number representations with varying bit widths, which is only exacerbated by layer-wise optimization. Besides, existing quantization techniques often target a specific DNN framework and/or hardware platform, lacking portability across various execution environments. To address this, we propose libnumber, a portable, automatic quantization framework for DNNs. By introducing Number abstract data type (ADT), libnumber encapsulates the internal representation of a number from the user. Then the auto-tuner of libnumber finds a compact representation (type, bit width, and bias) for the number that minimizes the user-supplied objective function, while satisfying the accuracy constraint. Thus, libnumber effectively separates the concern of developing an effective DNN model from low-level optimization of number representation. Our evaluation using eleven DNN models on two DNN frameworks targeting an FPGA platform demonstrates over 8× (7×) reduction in the parameter size on average when up to 7% (1%) loss of relative accuracy is tolerable, with a maximum reduction of 16×, compared to the baseline using 32-bit floating-point numbers. This leads to an geomean speedup of 3.79× with a maximum speedup of 12.77× over the baseline, while requiring only minimal programmer effort. Young H. Oh, Quan Quan, Seonghak Kim, Jun Heo 0001, Sungjun Jung, Jaeyoung Jang, Jae W. Lee |
PACT | 1 |
| 2017 | Typed Architectures: Architectural Support for Lightweight ScriptingabstractDynamic scripting languages are becoming more and more widely adopted not only for fast prototyping but also for developing production-grade applications. They provide high-productivity programming environments featuring high levels of abstraction with powerful built-in functions, automatic memory management, object-oriented programming paradigm and dynamic typing. However, their flexible, dynamic type systems easily become the source of inefficiency in terms of instruction count, memory footprint, and energy consumption. This overhead makes it challenging to deploy these high-productivity programming technologies on emerging single-board computers for IoT applications. Addressing this challenge, this paper introduces Typed Architectures, a high-efficiency, low-cost execution substrate for dynamic scripting languages, where each data variable retains high-level type information at an ISA level. Typed Architectures calculate and check the dynamic type of each variable implicitly in hardware, rather than explicitly in software, hence significantly reducing instruction count for dynamic type checking. Besides, Typed Architectures introduce polymorphic instructions (e.g., xadd), which are bound to the correct native instruction at runtime within the pipeline (e.g., add or fadd) to efficiently implement polymorphic operators. Finally, Typed Architectures provide hardware support for flexible yet efficient type tag extraction and insertion, capturing common data layout patterns of tag-value pairs. Our evaluation using a fully synthesizable RISC-V RTL design on FPGA shows that Typed Architectures achieve geomean speedups of 11.2% and 9.9% with maximum speedups of 32.6% and 43.5% for two production-grade scripting engines for JavaScript and Lua, respectively. Moreover, Typed Architectures improve the energy-delay product (EDP) by 19.3% for JavaScript and 16.5% for Lua with an area overhead of 1.6% at a 40nm technology node. Channoh Kim, Jaehyeok Kim, Sungmin Kim, Namho Kim, Gitae Na, Young H. Oh, Hyeon-Gyu Cho, Jae W. Lee |
ASPLOS | 7 |
| 2016 | Short-Circuit Dispatch: Accelerating Virtual Machine Interpreters on Embedded ProcessorsabstractInterpreters are widely used to implement high-level language virtual machines (VMs), especially on resource-constrained embedded platforms. Many scripting languages employ interpreter-based VMs for their advantages over native code compilers, such as portability, smaller resource footprint, and compact codes. For efficient interpretation a script (program) is first compiled into an intermediate representation, or bytecodes. The canonical interpreter then runs an infinite loop that fetches, decodes, and executes one bytecode at a time. This bytecode dispatch loop is a well-known source of inefficiency, typically featuring a large jump table with a hard-to-predict indirect jump. Most existing techniques to optimize this loop focus on reducing the misprediction rate of this indirect jump in both hardware and software. However, these techniques are much less effective on embedded processors with shallow pipelines and low IPCs. Instead, we tackle another source of inefficiency more prominent on embedded platforms - redundant computation in the dispatch loop. To this end, we propose Short-Circuit Dispatch (SCD), a low cost architectural extension that enables fast, hardware-based bytecode dispatch with fewer instructions. The key idea of SCD is to overlay the software-created bytecode jump table on a branch target buffer (BTB). Once a bytecode is fetched, the BTB is looked up using the bytecode, instead of PC, as key. If it hits, the interpreter directly jumps to the target address retrieved from the BTB, otherwise, it goes through the original dispatch path. This effectively eliminates redundant computation in the dispatcher code for decode, bound check, and target address calculation, thus significantly reducing total instruction count. Our simulation results demonstrate that SCD achieves geomean speedups of 19.9% and 14.1% for two production-grade script interpreters for Lua and JavaScript, respectively. Moreover, our fully synthesizable RTL design based on a RISC-V embedded processor shows that SCD improves the EDP of the Lua interpreter by 24.2%, while increasing the chip area by only 0.72% at a 40nm technology node. Channoh Kim, Sungmin Kim, Hyeon-Gyu Cho, Jaehyeok Kim, Young H. Oh, Hakbeom Jang, Jae W. Lee |
ISCA | 6 |
| 2014 | eDRAM-based tiered-reliability memory with applications to low-power frame buffersabstractEmbedded DRAM (eDRAM) is becoming more and more popular as a low-cost alternative to on-chip SRAM. eDRAM is particularly attractive for frame buffers in video applications with ever increasing screen resolutions. However, eDRAM suffers short retention time and high refresh power, which prevents its widespread adoption. To save the refresh power of eDRAM-based frame buffers, we propose Tiered-Reliability Memory (TRM), where the frame buffer is divided into multiple segments with different refresh periods and hence different error rates. By allocating most-significant bits to the most reliable segment, our four-tier TRM reduces refresh power by 48% without degrading user experience. Kyungsang Cho, Yongjun Lee, Young H. Oh, Gyoo-Cheol Hwang, Jae W. Lee |
ISLPED | 3 |