Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Eunjin Baek

dblp:250/8937 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0003-4089-2392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware accelerators and domain-specific architectures · 27% Reconfigurable computing and FPGAs · 19% Electronic design automation · 13%
Computer networks
1 paper
Transport protocols and congestion control · 87% Datacenter networks · 13%

Topics — the 11 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
1.122023
STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023
A Multi-Neural Network Acceleration Architecture · ISCA 2020
Transport protocols and congestion control › transport protocol implementation
TCP offload
0.712023
F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework · ISCA 2023
Transport protocols and congestion control
transport protocols
0.712023
F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework · ISCA 2023
Reconfigurable computing and FPGAs
FPGA accelerator
0.712023
F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework · ISCA 2023
Memory systems › memory management
on-chip memory management
0.712023
STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023
GPUs and heterogeneous computing › GPU sharing
spatio-temporal sharing
0.712023
STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023
Electronic design automation › hardware verification and test
hardware verification
0.512021
DifuzzRTL: Differential Fuzz Testing to Find CPU Bugs · SP 2021
Hardware accelerators and domain-specific architectures › scientific computing accelerator
brain simulation accelerator
0.412019
FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning · MICRO 2019
Emerging computing paradigms
neuromorphic computing
0.412019
FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning · MICRO 2019
Reconfigurable computing and FPGAs
reconfigurable computing
0.412019
FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning · MICRO 2019
Electronic design automation › high-level synthesis
scheduling
0.212023
STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management · IEEE Trans. Computers 2023

Methods — techniques the papers use, named apart from their topics

hardware offload · 1.3FPGA implementation · 1.3register-coverage guided fuzzing · 1.0differential fuzzing · 1.0on-chip learning · 0.8design space exploration · 0.7runtime scheduling · 0.4compile-time task partitioning · 0.4learning rules · 0.4learning rule · 0.4
YearPublicationVenuePosition
2023 F4T: A Fast and Flexible FPGA-based Full-stack TCP Acceleration Framework
abstract
As complex workloads that run on many servers are pursuing higher networking throughput, more CPU cycles are consumed to support the TCP stack. To mitigate the high CPU burden from executing the compute-intensive TCP, prior works have proposed to offload TCP processing to the embedded processors, ASICs, or FPGAs in network devices. However, none of the approaches satisfy all of the critical requirements of TCP simultaneously, which are high performance, many connections, and high flexibility. Embedded processors do not provide enough performance to fully offload the TCP stack, while ASICs fail to provide enough flexibility. Meanwhile, existing FPGA-based TCP accelerators either fail to provide high performance or give up some of the critical features and requirements to achieve high performance due to their inefficient processing architecture.
Junehyuk Boo, Yujin Chung, Eunjin Baek, Seongmin Na, Changsu Kim 0004, Jangwoo Kim
ISCA3
2023 STfusion: Fast and Flexible Multi-NN Execution Using Spatio-Temporal Block Fusion and Memory Management
abstract
To maximize the cost-effectiveness of neural network (NN) accelerators, architects are actively developing single-chip accelerators which can execute many NNs simultaneously. However, previous approaches fail to achieve full performance potential by exploiting only spatial or temporal resource sharing (SS or TS). They also do not consider memory management that can significantly affect performance. This limitation leads to the dire need for a new multi-NN accelerator taking both opportunities with careful memory management. But, it is extremely challenging to design an ideal spatio-temporal sharing accelerator because it requires (1) an algorithm that determines the degree of SS/TS in large exploration spaces, (2) a new STS-enabled accelerator devised with diverse design points, and (3) carefully-designed memory management that minimizes resource contention during numerous data transfers upon reconfiguration. To this end, we propose STfusion, a fast and flexible multi-NN execution architecture. First, STfusion partitions an accelerator into multiple smaller TS-enabled accelerators. Second, STfusion dynamically fuses small accelerators to adjust the accelerator sizes. Third, STfusion manages on-chip buffer in a page-granularity for stall-free data transfers. Lastly, STfusion provides an algorithm that determines the degree of SS/TS to achieve high throughput while satisfying QoS goals. Our evaluation shows that STfusion significantly outperforms state-of-the-art multi-NN accelerators.
Eunjin Baek, Eunbok Lee, Taehun Kang, Jangwoo Kim
IEEE Trans. Computers1
2021 DifuzzRTL: Differential Fuzz Testing to Find CPU Bugs
abstract
Security bugs in CPUs have critical security impacts to all the computation related hardware and software components as it is the core of the computation. In spite of the fact that architecture and security communities have explored a vast number of static or dynamic analysis techniques to automatically identify such bugs, the problem remains unsolved and challenging largely due to the complex nature of CPU RTL designs.This paper proposes DIFUZZRTL, an RTL fuzzer to automatically discover unknown bugs in CPU RTLs. DIFUZZRTL develops a register-coverage guided fuzzing technique, which efficiently yet correctly identifies a state transition in the finite state machine of RTL designs. DIFUZZRTL also develops several new techniques in consideration of unique RTL design characteristics, including cycle-sensitive register coverage guiding, asynchronous interrupt events handling, a unified CPU input format with Tilelink protocols, and drop-in-replacement designs to support various CPU RTLs. We implemented DIFUZZRTL, and performed the evaluation with three real-world open source CPU RTLs: OpenRISC Mor1kx Cappuccino, RISC-V Rocket Core, and RISC-V Boom Core. During the evaluation, DIFUZZRTL identified 16 new bugs from these CPU RTLs, all of which were confirmed by the respective development communities and vendors. Six of those are assigned with CVE numbers, and to the best of our knowledge, we reported the first and the only CVE of RISC-V cores, demonstrating its strong practical impacts to the security community.
Jaewon Hur, Suhwan Song, Dongup Kwon, Eunjin Baek, Jangwoo Kim, Byoungyoung Lee
SP4
2021 An accurate and fair evaluation methodology for SNN-based inferencing with full-stack hardware design space explorations
Hunjun Lee, Chanmyeong Kim, Eunjin Baek, Jangwoo Kim
Neurocomputing4
2020 A Multi-Neural Network Acceleration Architecture
abstract
A cost-effective multi-tenant neural network execution is becoming one of the most important design goals for modern neural network accelerators. For example, as emerging AI services consist of many heterogeneous neural network executions, a cloud provider wants to serve a large number of clients using a single AI accelerator for improving its cost effectiveness. Therefore, an ideal next-generation neural network accelerator should support a simultaneous multi-neural network execution, while fully utilizing its hardware resources. However, existing accelerators which are optimized for a single neural network execution can suffer from severe resource underutilization when running multiple neural networks, mainly due to the load imbalance between computation and memory-access tasks from different neural networks.In this paper, we propose AI-MultiTasking (AI-MT), a novel accelerator architecture which enables a cost-effective, high-performance multi-neural network execution. The key idea of AI-MT is to fully utilize the accelerator’s computation resources and memory bandwidth by matching compute- and memory-intensive tasks from different networks and executing them in parallel. However, it is highly challenging to find and schedule the best load-matching tasks from different neural networks during runtime, without significantly increasing the size of on-chip memory. To overcome the challenges, AI-MT first creates fine-grain tasks at compile time by dividing each layer into multiple identical sub-layers. During runtime, AI-MT dynamically applies three sub-layer scheduling methods: memory block prefetching and compute block merging for the best resource load matching, and memory block eviction for the minimum on-chip memory footprint. Our evaluations using MLPerf benchmarks show that AI-MT achieves up to 1.57x speedup over the baseline scheduling method.
Eunjin Baek, Dongup Kwon, Jangwoo Kim
ISCA1
2019 FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning
abstract
To understand how the human brain works, neuroscientists heavily rely on brain simulations which incorporate the concept of time to their operating model. In the simulations, neurons transmit their signals through synapses whose weights change over time and by the activity of the associated neurons. Such changes in synaptic weights, known as learning, are thought to contribute to memory, and various learning rules exist to model different behaviors of the human brain. Due to the diverse neurons and learning rules, neuroscientists perform the simulations using highly programmable general-purpose processors. Unfortunately, the processors greatly suffer from the high computational overheads of the learning rules. As an alternative, brain simulation accelerators achieve orders of magnitude higher performance; however, they have limited flexibility and cannot support the diverse neurons and learning rules.
Eunjin Baek, Hunjun Lee, Youngsok Kim, Jangwoo Kim
MICRO1