EDBT 2026 Demo / reviewers in the wild / expert
Eunjin Baek
dblp:250/8937
· DBLP profile ↗
6ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0003-4089-2392ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 27% Reconfigurable computing and FPGAs · 19% Electronic design automation · 13% | |
| Computer networks
1 paper |
Transport protocols and congestion control · 87% Datacenter networks · 13% |
Topics — the 11 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
1.1 | 2 | 2023 | STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023 A Multi-Neural Network Acceleration Architecture · ISCA 2020 |
Transport protocols and congestion control › transport protocol implementation
TCP offload |
0.7 | 1 | 2023 | F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework · ISCA 2023 |
Transport protocols and congestion control
transport protocols |
0.7 | 1 | 2023 | F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework · ISCA 2023 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.7 | 1 | 2023 | F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework · ISCA 2023 |
Memory systems › memory management
on-chip memory management |
0.7 | 1 | 2023 | STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023 |
GPUs and heterogeneous computing › GPU sharing
spatio-temporal sharing |
0.7 | 1 | 2023 | STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023 |
Electronic design automation › hardware verification and test
hardware verification |
0.5 | 1 | 2021 | DifuzzRTL: Differential Fuzz Testing to Find CPU Bugs · SP 2021 |
Hardware accelerators and domain-specific architectures › scientific computing accelerator
brain simulation accelerator |
0.4 | 1 | 2019 | FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning · MICRO 2019 |
Emerging computing paradigms
neuromorphic computing |
0.4 | 1 | 2019 | FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning · MICRO 2019 |
Reconfigurable computing and FPGAs
reconfigurable computing |
0.4 | 1 | 2019 | FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning · MICRO 2019 |
Electronic design automation › high-level synthesis
scheduling |
0.2 | 1 | 2023 | STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023 |
Methods — techniques the papers use, named apart from their topics
hardware offload · 1.3FPGA implementation · 1.3register-coverage guided fuzzing · 1.0differential fuzzing · 1.0on-chip learning · 0.8design space exploration · 0.7runtime scheduling · 0.4compile-time task partitioning · 0.4learning rules · 0.4learning rule · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration FrameworkabstractAs complex workloads that run on many servers are pursuing higher networking throughput, more CPU cycles are consumed to support the TCP stack. To mitigate the high CPU burden from executing the compute-intensive TCP, prior works have proposed to offload TCP processing to the embedded processors, ASICs, or FPGAs in network devices. However, none of the approaches satisfy all of the critical requirements of TCP simultaneously, which are high performance, many connections, and high flexibility. Embedded processors do not provide enough performance to fully offload the TCP stack, while ASICs fail to provide enough flexibility. Meanwhile, existing FPGA-based TCP accelerators either fail to provide high performance or give up some of the critical features and requirements to achieve high performance due to their inefficient processing architecture. Junehyuk Boo, Yujin Chung, Eunjin Baek, Seongmin Na, Changsu Kim 0004, Jangwoo Kim |
ISCA | 3 |
| 2023 | STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory ManagementabstractTo maximize the cost-effectiveness of neural network (NN) accelerators, architects are actively developing single-chip accelerators which can execute many NNs simultaneously. However, previous approaches fail to achieve full performance potential by exploiting only spatial or temporal resource sharing (SS or TS). They also do not consider memory management that can significantly affect performance. This limitation leads to the dire need for a new multi-NN accelerator taking both opportunities with careful memory management. But, it is extremely challenging to design an ideal spatio-temporal sharing accelerator because it requires (1) an algorithm that determines the degree of SS/TS in large exploration spaces, (2) a new STS-enabled accelerator devised with diverse design points, and (3) carefully-designed memory management that minimizes resource contention during numerous data transfers upon reconfiguration. To this end, we propose STfusion, a fast and flexible multi-NN execution architecture. First, STfusion partitions an accelerator into multiple smaller TS-enabled accelerators. Second, STfusion dynamically fuses small accelerators to adjust the accelerator sizes. Third, STfusion manages on-chip buffer in a page-granularity for stall-free data transfers. Lastly, STfusion provides an algorithm that determines the degree of SS/TS to achieve high throughput while satisfying QoS goals. Our evaluation shows that STfusion significantly outperforms state-of-the-art multi-NN accelerators. Eunjin Baek, Eunbok Lee, Taehun Kang, Jangwoo Kim |
IEEE Trans. Computers | 1 |
| 2021 | DifuzzRTL: Differential Fuzz Testing to Find CPU BugsabstractSecurity bugs in CPUs have critical security impacts to all the computation related hardware and software components as it is the core of the computation. In spite of the fact that architecture and security communities have explored a vast number of static or dynamic analysis techniques to automatically identify such bugs, the problem remains unsolved and challenging largely due to the complex nature of CPU RTL designs.This paper proposes DIFUZZRTL, an RTL fuzzer to automatically discover unknown bugs in CPU RTLs. DIFUZZRTL develops a register-coverage guided fuzzing technique, which efficiently yet correctly identifies a state transition in the finite state machine of RTL designs. DIFUZZRTL also develops several new techniques in consideration of unique RTL design characteristics, including cycle-sensitive register coverage guiding, asynchronous interrupt events handling, a unified CPU input format with Tilelink protocols, and drop-in-replacement designs to support various CPU RTLs. We implemented DIFUZZRTL, and performed the evaluation with three real-world open source CPU RTLs: OpenRISC Mor1kx Cappuccino, RISC-V Rocket Core, and RISC-V Boom Core. During the evaluation, DIFUZZRTL identified 16 new bugs from these CPU RTLs, all of which were confirmed by the respective development communities and vendors. Six of those are assigned with CVE numbers, and to the best of our knowledge, we reported the first and the only CVE of RISC-V cores, demonstrating its strong practical impacts to the security community. Jaewon Hur, Suhwan Song, Dongup Kwon, Eunjin Baek, Jangwoo Kim, Byoungyoung Lee |
SP | 4 |
| 2021 | An accurate and fair evaluation methodology for SNN-based inferencing with full-stack hardware design space explorations
Hunjun Lee, Chanmyeong Kim, Eunjin Baek, Jangwoo Kim |
Neurocomputing | 4 |
| 2020 | A Multi-Neural Network Acceleration ArchitectureabstractA cost-effective multi-tenant neural network execution is becoming one of the most important design goals for modern neural network accelerators. For example, as emerging AI services consist of many heterogeneous neural network executions, a cloud provider wants to serve a large number of clients using a single AI accelerator for improving its cost effectiveness. Therefore, an ideal next-generation neural network accelerator should support a simultaneous multi-neural network execution, while fully utilizing its hardware resources. However, existing accelerators which are optimized for a single neural network execution can suffer from severe resource underutilization when running multiple neural networks, mainly due to the load imbalance between computation and memory-access tasks from different neural networks.In this paper, we propose AI-MultiTasking (AI-MT), a novel accelerator architecture which enables a cost-effective, high-performance multi-neural network execution. The key idea of AI-MT is to fully utilize the accelerator’s computation resources and memory bandwidth by matching compute- and memory-intensive tasks from different networks and executing them in parallel. However, it is highly challenging to find and schedule the best load-matching tasks from different neural networks during runtime, without significantly increasing the size of on-chip memory. To overcome the challenges, AI-MT first creates fine-grain tasks at compile time by dividing each layer into multiple identical sub-layers. During runtime, AI-MT dynamically applies three sub-layer scheduling methods: memory block prefetching and compute block merging for the best resource load matching, and memory block eviction for the minimum on-chip memory footprint. Our evaluations using MLPerf benchmarks show that AI-MT achieves up to 1.57x speedup over the baseline scheduling method. Eunjin Baek, Dongup Kwon, Jangwoo Kim |
ISCA | 1 |
| 2019 | FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip LearningabstractTo understand how the human brain works, neuroscientists heavily rely on brain simulations which incorporate the concept of time to their operating model. In the simulations, neurons transmit their signals through synapses whose weights change over time and by the activity of the associated neurons. Such changes in synaptic weights, known as learning, are thought to contribute to memory, and various learning rules exist to model different behaviors of the human brain. Due to the diverse neurons and learning rules, neuroscientists perform the simulations using highly programmable general-purpose processors. Unfortunately, the processors greatly suffer from the high computational overheads of the learning rules. As an alternative, brain simulation accelerators achieve orders of magnitude higher performance; however, they have limited flexibility and cannot support the diverse neurons and learning rules. Eunjin Baek, Hunjun Lee, Youngsok Kim, Jangwoo Kim |
MICRO | 1 |