Sungyeob Yoo

dblp:322/1062 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-7783-9176ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 An End-to-End Diffusion Accelerator With Reconfigurable Hyper-Precision and Unified Non-Matrix Processing Engine
abstract
Diffusion models have emerged as state-of-the-art generative AI models but face significant computational challenges due to their iterative denoising process. While quantization techniques help reduce computation, conventional methods often degrade accuracy, and non-matrix operations remain a latency bottleneck. We propose Picasso, an end-to-end diffusion accelerator featuring the novel Hyper-Precision 8 (HYP8) data type that balances numerical precision and hardware efficiency by extending dynamic range while preserving resolution for near-zero values. Picasso integrates Hyper-Efficient Reconfigurable Arrays (HERA) for matrix operations and Unified Non-Matrix Processing Engines (UNPE) for normalization, softmax, and element-wise computations. Fabricated in a 28-nm CMOS process, Picasso achieves 9.83 TOPS with 4.96 TOPS/W energy efficiency, outperforming prior works by up to$26.8\times $in speed,$2.8\times $in energy efficiency, and$30.5\times $in area efficiency, while maintaining FP16-equivalent generation quality with reduced memory footprint.
Sungyeob Yoo, Seeyeon Kim, Geonwoo Ko, Seri Ham, Yi Chen 0035, Joo-Young Kim 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 Picasso: An Area/Energy-Efficient End-to-End Diffusion Accelerator with Hyper-Precision Data Type
abstract
This work presents Picasso, an end-to-end diffusion accelerator designed for enhancing the efficiency of diffusion-based machine learning models used in applications such as image and video generation, and inpainting. Picasso introduces a novel hyper-precision 8 (HYP8) data type and a reconfigurable architecture designed to significantly enhance hardware efficiency, providing an extended dynamic range without sacrificing accuracy. It also features a unified engine that streamlines the processing of all non-matrix operations and employs sub-block pipeline scheduling to reduce overall latency. Fabricated in 28nm CMOS technology, this accelerator achieves an energy efficiency of 4.96 TOPS/W and a peak performance of 9.83 TOPS. Compared to previous works, Picasso demonstrates speedups ranging from 8.4× to 26.8× while also improving energy and area efficiency by 1.1× to 2.8× and 3.6× to 30.5×, respectively.
Sungyeob Yoo, Geonwoo Ko, Seri Ham, Seeyeon Kim, Yi Chen 0035, Joo-Young Kim 0001
HCS1
2023 LightTrader: A Standalone High-Frequency Trading System with Deep Learning Inference Accelerators and Proactive Scheduler
abstract
Recent research shows that artificial intelligence (AI) algorithms can dramatically improve the profitability of high-frequency trading (HFT) with accurate market prediction, overcoming the limitation of conventional latency-oriented approaches. However, it is challenging to integrate the computationally intensive AI algorithm into the existing trading pipeline due to its excessively long latency and insufficient throughput, necessitating a breakthrough in hardware. Furthermore, harsh HFT environments such as bursty data traffic and stringent power constraint make it even more difficult to achieve system-level performance without missing crucial market signals.In this paper, we present LightTrader, the world’s first AI-enabled HFT system that incorporates an FPGA and custom AI accelerators for short-latency-high-throughput trading systems. Leveraging the computing power of brand-new AI accelerators fabricated in TSMC’s 7nm FinFET technology, LightTrader optimizes the tick-to-trade latency and response rate for stock market data. The AI accelerators, adopting Coarse-Grained Reconfigurable Array (CGRA) architecture, which maximizes the hardware utilization from the flexible dataflow architecture, achieve a throughput of 16 TFLOPS and 64 TOPS. In addition, we propose both workload scheduling and dynamic voltage and frequency scaling (DVFS) scheduling algorithms to find an optimal offloading strategy under bursty market data traffic and limited power condition. Finally, we build a reliable and rerunnable simulation framework that can back-test the historical market data, such as Chicago Mercantile Exchange (CME), to evaluate the LightTrader system. We thoroughly explore the performance of LightTrader when the number of AI accelerators, power conditions, and complexity of deep neural network models change. As a result, LightTrader achieves 13.92× and 7.28× speed-up of AI algorithm processing compared to existing GPU-based, FPGA-based systems, respectively. LightTrader with multiple AI accelerators achieves up to 99.5% response rates, while LightTrader with the proposed workload scheduling and DVFS scheduling algorithm relieves the miss rate from 17.1% to 23.1%.
Sungyeob Yoo, Hyunsung Kim 0003, Jinseok Kim 0006, Sunghyun Park 0006, Joo-Young Kim 0001, Jinwook Oh
HPCA1
2022 LightTrader : World's first AI-enabled High-Frequency Trading Solution with 16 TFLOPS / 64 TOPS Deep Learning Inference Accelerators
abstract
We present the world’s first AI-enabled high-frequency trading (HFT) system, LightTrader , which integrates the custom AI accelerators and the FPGA-based conventional HFT pipeline for the low-latency-high-throughput trading solutions with a reduced query miss rate. For better utilization, adaptive job scheduling methods are also proposed to further improve the performance, where layer-wise workload scaling and dynamic voltage-frequency scaling (DVFS) techniques progressively adjust the workloads of AI accelerators, in conjunction with the architecture support. LightTrader integrating TSMC 7nm tape-out accelerators solely achieves 6x speed-up of DNN processing and 30-50x reduction of query miss rate without the scheduling method while the scheduling scheme further improves the energy efficiency by 25% and reduces the query miss rate by 2.4x .
Hyunsung Kim 0003, Sungyeob Yoo, Jaewan Bae, Kyeongryeol Bong, Yoonho Boo, Karim Charfi, Hyo-Eun Kim, Hyun Suk Kim, Jinseok Kim 0006, Byungjae Lee, Myeongbo Shim, Sungho Shin, Jeong Seok Woo, Joo-Young Kim 0001, Sunghyun Park 0006, Jinwook Oh
HCS2