EDBT 2026 Demo / reviewers in the wild / expert
Joo-Young Kim 0001
dblp:43/2609-1
· DBLP profile ↗
54ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0003-1099-1496ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 3 first-author · 35 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache RetrievalabstractStreaming video large language models (LLMs) are increasingly used for real-time multimodal tasks such as video captioning, question answering, conversational agents, and augmented reality. However, these models face fundamental memory and computational challenges because their key-value (KV) caches grow substantially with continuous streaming video input. This process requires an iterative prefill stage, which is a unique feature of streaming video LLMs. Prior works reduce excessive cache overhead by utilizing the KV cache retrieval algorithm, which offloads the full KV cache to CPU memory or storage, then selectively fetches the most relevant entries. Nevertheless, due to its iterative prefill stage, they suffer from significant limitations, including extensive computation, substantial data transfer, and degradation in accuracy. Crucially, this issue is exacerbated for edge deployment, which is the primary target for these models. The memory footprint exceeds the memory capacity within minutes of video streams, making low-latency, energy-efficient inference infeasible. In this work, we propose V-Rex, the first software-hardware co-designed accelerator that comprehensively addresses both algorithmic and hardware bottlenecks in streaming video LLM inference. At its core, V-Rex introduces ReSV, a training-free dynamic KV cache retrieval algorithm. ReSV exploits temporal and spatial similarity-based token clustering to reduce excessive KV cache memory across video frames, and dynamically adjusts token selection per transformer layer and attention head to minimize the number of selected tokens. To fully realize these algorithmic benefits, V-Rex offers a compact, low-latency hardware accelerator with a dynamic KV cache retrieval engine (DRE), featuring bit-level and early-exit based computing units, as well as hierarchical KV cache memory management. Evaluated on COIN benchmarks, V-Rex achieves unprecedented real-time of 3.9-8.3 FPS and energy-efficient streaming video LLM inference on edge deployment with negligible accuracy loss. While DRE only accounts for 2.2% power and 2.0% area, the system delivers 1.9-19.7× speedup and 3.1-18.5× energy efficiency improvements over AGX Orin GPU. This work is the first to comprehensively tackle KV cache retrieval across algorithm and hardware, enabling real-time streaming video LLM inference on resource-constrained edge devices, with clear potential for scalable deployment in large-scale server environments. Sejeong Yang, Wonjin Shin, Joo-Young Kim 0001 |
HPCA | 4 |
| 2026 | SCRec: A Scalable Computational Storage System With Statistical Sharding and Tensor-Train Decomposition for Recommendation ModelsabstractDeep Learning Recommendation Models (DLRMs) are essential for personalized content delivery in web applications, but their parameter sizes have grown to terabytes, with memory bandwidth demands exceeding TB/s. Furthermore, the workload intensity within the model varies based on the target mechanism, making it difficult to build an optimized recommendation system. In this paper, we propose SCRec, a scalable computational storage recommendation system that can handle TB-scale industrial DLRMs while guaranteeing high bandwidth requirements. SCRec utilizes a software framework that features a mixed-integer programming (MIP)-based cost model, efficiently fetching data based on data access patterns and adaptively configuring memory-centric and compute-centric cores. Additionally, SCRec integrates hardware acceleration cores to enhance DLRM computations, particularly allowing for the high-performance reconstruction of approximated embedding vectors from extremely compressed tensor-train (TT) format. By combining its software framework and hardware accelerators, while eliminating data communication overhead by being implemented on a single server, SCRec achieves substantial improvements in DLRM inference performance. It delivers up to 55.77× speedup compared to a CPU-DRAM system with no loss in accuracy and up to 13.35× energy efficiency gains over a multi-GPU system. Ji-Hoon Kim 0004, Joo-Young Kim 0001 |
IEEE Trans. Computers | 3 |
| 2026 | An End-to-End Diffusion Accelerator With Reconfigurable Hyper-Precision and Unified Non-Matrix Processing EngineabstractDiffusion models have emerged as state-of-the-art generative AI models but face significant computational challenges due to their iterative denoising process. While quantization techniques help reduce computation, conventional methods often degrade accuracy, and non-matrix operations remain a latency bottleneck. We propose Picasso, an end-to-end diffusion accelerator featuring the novel Hyper-Precision 8 (HYP8) data type that balances numerical precision and hardware efficiency by extending dynamic range while preserving resolution for near-zero values. Picasso integrates Hyper-Efficient Reconfigurable Arrays (HERA) for matrix operations and Unified Non-Matrix Processing Engines (UNPE) for normalization, softmax, and element-wise computations. Fabricated in a 28-nm CMOS process, Picasso achieves 9.83 TOPS with 4.96 TOPS/W energy efficiency, outperforming prior works by up to$26.8\times $in speed,$2.8\times $in energy efficiency, and$30.5\times $in area efficiency, while maintaining FP16-equivalent generation quality with reduced memory footprint. Sungyeob Yoo, Seeyeon Kim, Geonwoo Ko, Seri Ham, Yi Chen 0035, Joo-Young Kim 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | SeeSSD: Computational Storage for Energy-Efficient Real-Time Object DetectionabstractIn this work, we present our intelligent SSD, SeeSSD , an energy-efficient computational SSD for a real-time object detection system. SeeSSD embeds an FPGA-based CNN processing engine and the firmware that performs the convolutional operation on the target image. SeeSSD processes the image data at the storage before sending it to the host. This reduces the amount of data transferred to the host and lowers the data movement overhead, thus reducing transfer time and saving power. By using our SeeSSD system and YOLO_Embed, an object detection neural network model, we are able to outperform the fastest YOLO model for an embedded controller, YOLO-Lite, in terms of performance, accuracy, and energy efficiency. YOLO (You Only Look Once) models are a series of one-stage object detection neural models that have become very popular due to their fast speed and high accuracy. The contribution of this work includes designing and implementing our SeeSSD system with a lightweight object detection model, YOLO_Embed, for reducing the data movement overhead, performing real-time inference, and lowering the overall power consumption. We implemented the entire software stack associated with the SeeSSD system; on-device CNN acceleration engine implemented on FPGA, object identification interface for SeeSSD using YOLO_Embed, and embedded software layer in SeeSSD for on-device convolutional processing. We calculated our YOLO_Embed model’s accuracy on object detection dataset benchmarks such as PASCAL VOC 2012, which came out to be 38.1% mAP (mean Accuracy Precision). Our system was able to perform inference in 0.21 seconds while reducing the power consumption by approximately 1.2× and 1.4× for CPU-Only and CPU+GPU systems, respectively. We were also able to reduce the data movement overhead by 24× for a single target image. Muhammad Danish Tehseen, Gyeongcheol Shin, Joo-Young Kim 0001, Youjip Won |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | AoP-SAM: Automation of Prompts for Efficient SegmentationabstractThe Segment Anything Model (SAM) is a powerful foundation model for image segmentation, showing robust zero-shot generalization through prompt engineering. However, relying on manual prompts is impractical for real-world applications, particularly in scenarios where rapid prompt provision and resource efficiency are crucial. In this paper, we propose the Automation of Prompts for SAM (AoP-SAM), a novel approach that learns to generate essential prompts in optimal locations automatically. AoP-SAM enhances SAM’s efficiency and usability by eliminating manual input, making it better suited for real-world tasks. Our approach employs a lightweight yet efficient Prompt Predictor model that detects key entities across images and identifies the optimal regions for placing prompt candidates. This method leverages SAM’s image embeddings, preserving its zero-shot generalization capabilities without requiring fine-tuning. Additionally, we introduce a test-time instance-level Adaptive Sampling and Filtering mechanism that generates prompts in a coarse-to-fine manner. This notably enhances both prompt and mask generation efficiency by reducing computational overhead and minimizing redundant mask refinements. Evaluations of three datasets demonstrate that AoP-SAM substantially improves both prompt generation efficiency and mask generation accuracy, making SAM more effective for automated segmentation tasks. Yi Chen 0035, Muyoung Son, Chuanbo Hua, Joo-Young Kim 0001 |
AAAI | 4 |
| 2025 | ABC-FHE: A Resource-Efficient Accelerator Enabling Bootstrappable Parameters for Client-Side Fully Homomorphic EncryptionabstractAs the demand for privacy-preserving computation continues to grow, fully homomorphic encryption (FHE)—which enables continuous computation on encrypted data—has become a critical solution. However, its adoption is hindered by significant computational overhead, requiring 10000-fold more computation compared to plaintext processing. Recent advancements in FHE accelerators have successfully improved server-side performance, but client-side computations remain a bottleneck, particularly under bootstrappable parameter configurations, which involve combinations of encoding, encrypt, decoding, and decrypt for large-sized parameters. To address this challenge, we propose ABC-FHE, an area- and power-efficient FHE accelerator that supports bootstrappable parameters on the client side. ABCFHE employs a streaming architecture to maximize performance density, minimize area usage, and reduce off-chip memory access. Key innovations include a reconfigurable Fourier engine capable of switching between NTT and FFT modes. Additionally, an onchip pseudo-random number generator and a unified on-the-fly twiddle factor generator significantly reduce memory demands, while optimized task scheduling enhances the CKKS clientside processing, achieving reduced latency. Overall, ABC-FHE occupies a die area of 28.638 mm2 and consumes 5.654 W of power in 28 nm technology. It delivers significant performance improvements, achieving a 1112× speed-up in encoding and encryption execution time compared to a CPU, and 214× over the state-of-the-art client-side accelerator. For decoding and decryption, it achieves a 963× speed-up over the CPU and 82× over the state-of-the-art accelerator. Sungwoong Yune, Adiwena Putra, Hyunjun Cho, Cuong Duong Manh, Joo-Young Kim 0001 |
DAC | 7 |
| 2025 | Adelia: A 4nm LLM Processor for Efficient Generative Al Inference
Seungjae Moon, Juntaek Oh, Jay Kim, Joo-Young Kim 0001 |
HCS | 5 |
| 2025 | EXION: Exploiting Inter-and Intra-Iteration Output Sparsity for Diffusion Models
Jaehoon Heo, Adiwena Putra, Jieon Yoon, Sungwoong Yune, Hangyeol Lee, Ji-Hoon Kim 0004, Joo-Young Kim 0001 |
HPCA | 7 |
| 2025 | Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache QuantizationabstractModern Large Language Model (LLM) serving system batches multiple requests to achieve high throughput, while batching attention operations is challenging, rendering memory bandwidth a critical bottleneck.Today, to mitigate this issue, the community relies on high-end GPUs with multiple high-bandwidth memory (HBM) channels.Unfortunately, HBM's high bandwidth often comes at the expense of limited memory capacity, necessitating systems to scale, which reduces core utilization and increases costs.Moreover, recent advancements enabling longer contexts for LLMs have substantially increased the key-value (KV) cache size, further intensifying the pressures on memory capacity.To lower the pressure, the literature has explored KV cache quantization techniques, which commonly use low bitwidth (e.g., INT4) for most values, selectively using higher bitwidth (e.g., FP16) for outlier values.While this approach helps achieve high accuracy and low bitwidth simultaneously, it comes with the limitation that the cost for online outlier detection is excessively high, negating the advantages of quantization.Inspired by these insights, we propose Oaken, an acceleration solution that achieves high accuracy and high performance simultaneously through co-designing algorithm and hardware.To effectively find a sweet spot in the accuracy-performance trade-off space of KV cache quantization, Oaken employs an online-offline hybrid approach, setting outlier thresholds offline, which are then used to determine the quantization scale online.To translate the proposed algorithmic technique into tangible performance gains, Oaken also comes with custom quantization/dequantization engines and memory management units that can be integrated with any LLM accelerators.We built an Oaken accelerator on top of Minsu Kim 0004, Seongmin Hong, Ryeowook Ko, Soongyu Choi, Hunjong Lee, Junsoo Kim 0002, Joo-Young Kim 0001, Jongse Park |
ISCA | 7 |
| 2025 | LightNobel: Improving Sequence Length Limitation in Protein Structure Prediction Model via Adaptive Activation QuantizationabstractRecent advances in Protein Structure Prediction Models (PPMs), such as AlphaFold2 and ESMFold, have revolutionized computational biology by achieving unprecedented accuracy in predicting three-dimensional protein folding structures.However, these models face significant scalability challenges, particularly when processing proteins with long amino acid sequences (e.g., sequence length > 1,000).The primary bottleneck that arises from the exponential growth in activation sizes is driven by the unique data structure in PPM, which introduces an additional dimension that leads to substantial memory and computational demands.These limitations have hindered the effective scaling of PPM for real-world applications, such as analyzing large proteins or complex multimers with critical biological and pharmaceutical relevance.In this paper, we present LightNobel, the first hardware-software co-designed accelerator developed to overcome scalability limitations on the sequence length in PPM.At the software level, we propose Token-wise Adaptive Activation Quantization (AAQ), which leverages unique token-wise characteristics, such as distogram patterns in PPM activations, to enable fine-grained quantization techniques without compromising accuracy.At the hardware level, LightNobel integrates the multi-precision reconfigurable matrix processing unit (RMPU) and versatile vector processing unit (VVPU) to enable the efficient execution of AAQ.Through these innovations, LightNobel achieves up to 8.44×, 8.41× speedup and 37.29×, 43.35× higher power efficiency over the latest NVIDIA A100 and H100 GPUs, respectively, while maintaining negligible accuracy loss.It also reduces the peak memory requirement up to 120.05× in PPM, enabling scalable processing for proteins with long sequences. Soongyu Choi, Joo-Young Kim 0001 |
ISCA | 3 |
| 2025 | Hybe: GPU-NPU Hybrid System for Efficient LLM Inference with Million-Token Context WindowabstractThe growth of context window size in large language model (LLM) inference poses a very distinct computational challenge of hardware inefficiency.The inefficiency arises from the computational imbalance during LLM inference between the compute-intensive prefill stage, and memory-intensive decode stage.The predominant inference hardware, GPU, boasts large number of cores to excel in the prefill stage, which processes the entire input context at once, but suffers from hardware underutilization in the decode stage, which iteratively generates one output token at a time.In conventional LLM, batching has been able to alleviate the underutilization by generating multiple tokens of different requests.However, batching becomes infeasible in models with large context windows over 100K tokens because the Key-Value (KV) activations dominate the physical memory capacity, surpassing the entire model size.In this paper, we propose Hybe, a GPU-NPU hybrid system for efficient LLM inference with a million-token context window.Hybe utilizes the preexisting GPU for the prefill stage and employs lightweight NPUs during the decode stage.Each NPU includes only the necessary computing resources to fully utilize the given memory bandwidth, thereby achieving maximum hardware efficiency.Furthermore, Hybe introduces fine-grained KV transmission, a kernel scheduling method that immediately offloads partial KV produced from the GPU to the NPU, which significantly reduces the KV memory required in the GPU.Lastly, Hybe scheduler applies stage-wise pipelining that dynamically assigns queued requests to idle hardware to minimize stalls.Hybe utilizes NVIDIA H100 GPU with inference-optimized vLLM library and implement Hybe NPU in 4nm process with equal HBM specification.Hybe achieves 2.1× speedup for Phi-3 with 100K-token context window and 3.9× energy efficiency for Llama-3 with 1M-token context window, over H100 GPUs with equal total device count. Seungjae Moon, Junseo Cha, Joo-Young Kim 0001 |
ISCA | 4 |
| 2025 | ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and ThroughputabstractThe growing adoption of Large Language Models (LLMs) across various domains has driven the demand for efficient and scalable AI-serving solutions. Deploying LLMs requires optimizations to manage their significant computational and data demands. The prefill stage processes large numbers of input tokens in parallel, increasing computational load, while the decoding stage relies heavily on memory bandwidth due to the auto-regressive nature of LLMs. Current hardware, such as GPUs, often fails to balance these demands, leading to inefficient utilization. While batching improves hardware efficiency, it delays response times, degrading Quality-of-Service (QoS). This disconnect between vendors, who aim to maximize resource efficiency, and users, who prioritize low latency, highlights the need for a better solution. To address this, we propose ADOR, a framework that automatically identifies and recommends hardware architectures tailored to LLM serving. By lever-aging predefined architecture templates specialized for heterogeneous dataflows, ADOR optimally balances throughput and latency. It efficiently explores design spaces to suggest architectures that meet the requirements of both vendors and users. ADOR demonstrates substantial performance improvements, achieving$2.51 \times$higher QoS and$4.01 \times$better area efficiency compared to the A100 at high batch sizes, making it a robust solution for scalable and cost-effective LLM serving. Junsoo Kim 0002, Hunjong Lee, Geonwoo Ko, Gyubin Choi, Seri Ham, Seongmin Hong, Joo-Young Kim 0001 |
ISPASS | 7 |
| 2025 | HLX: A Unified Pipelined Architecture for Optimized Performance of Hybrid Transformer-Mamba Language ModelsabstractThe rapid increase in demand for long-context language models has revealed fundamental performance limitations in conventional Transformer architectures, particularly their quadratic computational complexity.Hybrid Transformer-Mamba models, which interleave attention layers with efficient state-space model layers such as Mamba-2, have emerged as promising solutions combining the strengths of both Transformer and Mamba.However, maintaining a high compute utilization and performance across workloads (e.g., varying sequence length and batch size) in the Hybrid models is challenging due to their heterogeneous compute patterns and shifting performance bottlenecks between the two key computational kernels: FlashAttention-2 (FA-2) and State-Space Duality (SSD).In this paper, we introduce HLX, a unified pipelined architecture designed to ensure optimized performance across workloads for Hybrid models.Through detailed kernel-level analysis, we identify two key blockers that limit compute utilization: inter-operation dependencies in FA-2 and excessive memory traffic in SSD.To overcome these hurdles, we propose two novel fine-grained pipelined dataflows named PipeFlash and PipeSSD.PipeFlash effectively hides operational dependencies in attention computations, while PipeSSD firstly introduces the fused pipelined execution for SSD computations, substantially enhancing data reuse and reducing memory traffic.In addition, we propose a unified hardware architecture that can process both PipeFlash and PipeSSD in an efficient pipelining scheme to maximize the compute utilization.Finally, across sequence lengths from 1K to 128K, the proposed HLX architecture achieves up to 97.5% and 78.4% compute utilization for FA-2 and SSD, respectively, resulting in an average speedup of 1.75× and 2.91× over A100, and an average 2.78× (FA-2), 1.84× (FA-3), and 4.95× speedups over H100.For end-to-end latency and batching, HLX achieves a 1.56× and 1.38× speedup over A100 and a 2.08× and 1.76× (1.84× and 1.72×) speedup when running FA-2 (FA-3) on H100.It also significantly reduces area and power consumption by up to 89.8% and 63.8% compared to GPU baselines. In-Jun Jung, Gyeongrok Yang, Jaeha Min, Joo-Young Kim 0001 |
MICRO | 4 |
| 2025 | SAL-PIM: A Subarray-Level Processing-in-Memory Architecture With LUT-Based Linear Interpolation for Transformer-Based Text GenerationabstractText generation is a compelling sub-field of natural language processing, aiming to generate human-readable text from input words. Although many deep learning models have been proposed, the recent emergence of transformer-based large language models advances its academic research and industry development, showing remarkable qualitative results in text generation. In particular, the decoder-only generative models, such as generative pre-trained transformer (GPT), are widely used for text generation, with two major computational stages: summarization and generation. Unlike the summarization stage, which can process the input tokens in parallel, the generation stage is difficult to accelerate due to its sequential generation of output tokens through iteration. Moreover, each iteration requires reading a whole model with little data reuse opportunity. Therefore, the workload of transformer-based text generation is severely memory-bound, making the external memory bandwidth system bottleneck. In this paper, we propose a subarray-level processing-in-memory (PIM) architecture named SAL-PIM, the first HBM-based PIM architecture for the end-to-end acceleration of transformer-based text generation. With optimized data mapping schemes for different operations, SAL-PIM utilizes higher internal bandwidth by integrating multiple subarray-level arithmetic logic units (S-ALUs) next to memory subarrays. To minimize the area overhead for S-ALU, it uses shared MACs leveraging slow clock frequency of commands for the same bank. In addition, a few subarrays in the bank are used as look-up tables (LUTs) to handle non-linear functions in PIM, supporting multiple addressing to select sections for linear interpolation. Lastly, the channel-level arithmetic logic unit (C-ALU) is added in the buffer die of HBM to perform the accumulation and reduce-sum operations of data across multiple banks, completing end-to-end inference on PIM. To validate the SAL-PIM architecture, we built a cycle-accurate simulator based on Ramulator. We also implemented the SAL-PIM’s logic units in 28-nm CMOS technology and scaled the results to DRAM technology to verify its feasibility. We measured the end-to-end latency of SAL-PIM when it runs various text generation workloads on the GPT-2 medium model (with 345 million parameters), in which the input and output token numbers vary from 32 to 128 and from 1 to 256, respectively. As a result, with 4.81% area overhead, SAL-PIM achieves up to 4.72× speedup (1.83× on average) over the Nvidia Titan RTX GPU running FasterTransformer Framework. Wontak Han, Hyunjun Cho, Joo-Young Kim 0001 |
IEEE Trans. Computers | 4 |
| 2024 | ACane: An Efficient FPGA-based Embedded Vision Platform with Accumulation-as-Convolution Packing for Autonomous Mobile RobotsabstractConvolutional Neural Networks (CNNs) have been extensively deployed on autonomous mobile robots in recent years, and embedded platforms based on field-programmable gate arrays (FPGAs) that involve digital signal processors (DSPs) effectively utilize low-precision quantization with DSP-packing methods to implement large CNN models. However, DSP-packing has a limitation in improving computation performance due to zero bits that prevent bit contamination of output operands. In this paper, we propose ACane, a compact FPGA-based vision platform for autonomous mobile robots, based on a novel DSP-packing technique called accumulation-as-convolution packing, which effectively packs low-bit values to a single DSP, with boosting convolution operations. It also applies optimized data mapping and dataflow to improve computation parallelism of the DSP-packing. ACane successfully achieves the highest DSP efficiency (1.465 GOPS/DSP) and energy efficiency (361.8 GOPS/W), which are $1.98-8.32 \times $ and $4.03-25.5 \times $ higher compared to the state-of-the-art FPGA-based vision works, respectively. Sungwoong Yune, Sukbin Lim, Joo-Young Kim 0001 |
ASPDAC | 5 |
| 2024 | Picasso: An Area/Energy-Efficient End-to-End Diffusion Accelerator with Hyper-Precision Data TypeabstractThis work presents Picasso, an end-to-end diffusion accelerator designed for enhancing the efficiency of diffusion-based machine learning models used in applications such as image and video generation, and inpainting. Picasso introduces a novel hyper-precision 8 (HYP8) data type and a reconfigurable architecture designed to significantly enhance hardware efficiency, providing an extended dynamic range without sacrificing accuracy. It also features a unified engine that streamlines the processing of all non-matrix operations and employs sub-block pipeline scheduling to reduce overall latency. Fabricated in 28nm CMOS technology, this accelerator achieves an energy efficiency of 4.96 TOPS/W and a peak performance of 9.83 TOPS. Compared to previous works, Picasso demonstrates speedups ranging from 8.4× to 26.8× while also improving energy and area efficiency by 1.1× to 2.8× and 3.6× to 30.5×, respectively. Sungyeob Yoo, Geonwoo Ko, Seri Ham, Seeyeon Kim, Yi Chen 0035, Joo-Young Kim 0001 |
HCS | 6 |
| 2024 | Morphling: A Throughput-Maximized TFHE-based Accelerator using Transform-domain ReuseabstractFully Homomorphic Encryption (FHE) has become an increasingly important aspect in modern computing, particularly in preserving privacy in cloud computing by enabling computation directly on encrypted data. Despite its potential, FHE generally poses major computational challenges, including huge computational and memory requirements. The bootstrapping operation, which is essential particularly in Torus-Fhe(tfhe) scheme, involves intensive computations characterized by an enormous number of polynomial multiplications. For instance, performing a single bootstrapping at the 128-bit security level requires more than 10,000 polynomial multiplications. Our in-depth analysis reveals that domain-transform operations, i.e., Fast Fourier Transform (FFT), contribute up to 88% of these operations, which is the bottleneck of the TFHE system. To address these challenges, we propose Morphling, an accelerator architecture that combines the 2D systolic array and strategic use of transform-domain reuse in order to reduce the overhead of domain-transform in TFHE. This novel approach effectively reduces the number of required domain-transform operations by up to 83.3 %, allowing more computational cores in a given die area. In addition, we optimize its micro architecture design for end-to-end TFHE operation, such as merge-split pipelined-FFT for efficient domain-transform operation, double-pointer method for high-throughput polynomial rotation, and specialized buffer design. Furthermore, we introduce custom instructions for tiling, batching, and scheduling of multiple ciphertext operations. This facilitates software-hardware co-optimization, effectively mapping high-level applications such as XG-Boost classifier, Neural-Network, and VGG-9. As a result, Morphling, with four 2D systolic arrays and four vector units with domain-transform reuse, takes 74.79 mm2die area and 53.00 W power consumption in 28nm process. It achieves a throughput of up to 147,615 bootstrappings per second, demonstrating improvements of 3440x over the CPU, 143x over the GPU, and 14.7x over the state-of-the-art TFHE accelerator. It can run various deep learning models with sub-second latency. Prasetiyo, Adiwena Putra, Joo-Young Kim 0001 |
HPCA | 3 |
| 2024 | APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled CircuitsabstractAs the importance of Privacy-Preserving Inference of Transformers (PiT) increases, a hybrid protocol that integrates Garbled Circuits (GC) and Homomorphic Encryption (HE) is emerging for its implementation. While this protocol is preferred for its ability to maintain accuracy, it has a severe drawback of excessive latency. To address this, existing protocols primarily focused on reducing HE latency, thus making GC the new latency bottleneck. Furthermore, previous studies only focused on individual computing layers, such as protocol or hardware accelerator, lacking a comprehensive solution at the system level. Hyunjun Cho, Jaehoon Heo, Joo-Young Kim 0001 |
ICCAD | 4 |
| 2024 | BLESS: Bandwidth and Locality Enhanced SMEM Seeding Acceleration for DNA SequencingabstractIn an era marked by the pervasive spread of harmful viruses like COVID-19, the importance of DNA sequencing has grown significantly, given its crucial role in devising effective countermeasures. The seeding process, which aims to find locations of super-maximal exact matches (SMEM) between the DNA samples and reference genome for comparative analysis, has emerged as a major bottleneck due to its memory-intensive characteristics. The learned index approach has been developed that uses machine learning model to partially predict the location of the exact matches, which has effectively reduced the memory access. However, the lack of locality in the current in dexing structure and randomness at runtime of the seeding workload have constrained the memory bandwidth usage and have limited further performance advantage. In this paper, we propose BLESS, a bandwidth and locality enhanced SMEM seeding accelerator for learned-index-based DNA sequence alignment. BLESS is the first domain-specific seeding accelerator to maximize the potential hardware advantage of the learned index approach. We introduce coarse-fine (CF) block data structure, a novel memory mapping of seeding parameters to exploit spatial locality and increase effective bandwidth usage for any memory type, including high bandwidth memory (HBM). We also develop guaranteed search range update (GSRU) algorithm, a method that exploits caching in the search procedure to enable temporal locality and data reuse. Utilizing the CF block and GSRU algorithm, we develop a multi-core seeding accelerator using HBM with context switching and runtime scheduling for maximum core and memory bandwidth utilization. With these improvements, BLESS achieves $35.65 \times$ and $15.49 \times$ speedup over the state-of-the-art seeding system BWA-MEME and ERT-ASIC, respectively, in raw system performance. Seungjae Moon, Teokkyu Suh, Jaehoon Heo, Joo-Young Kim 0001 |
ISCA | 5 |
| 2024 | AdapTiV: Sign-Similarity Based Image-Adaptive Token Merging for Vision Transformer AccelerationabstractThe advent of Vision Transformers (ViT) has set a new performance leap in computer vision by leveraging self-attention mechanisms. However, the computational efficiency of ViTs is limited by the quadratic complexity of self-attention and redundancy among image tokens. To address these issues, token merging strategies have been explored to reduce input size by merging similar tokens. Nonetheless, implementing token merging presents a degradation of latency performance due to its two factors: inefficient computations and fixed merge rate nature. This paper introduces AdapTiV, a novel hardware-software co-designed accelerator that accelerates ViTs through image-adaptive token merging, effectively addressing the afore-mentioned challenges. Under the design philosophy of reducing the overhead of token merging and concealing its latency within the Layer Normalization (LN) process, AdapTiV incorporates algorithmic innovations such as Local Matching, which restricts the search space for token merging, thereby reducing the computational complexity; Sign Similarity, which simplifies the calculation of similarity between tokens; and Dynamic Merge Rate, which enables image-adaptive token merging. Additionally, the hardware component that supports AdapTiV's algorithms, named the Adaptive Token Merging Engine, employs Sign-Driven Scheduling to conceal the overhead of token merging effectively. This engine integrates submodules such as a Sign Similarity Computing Unit, which calculates the similarity between tokens using a newly introduced similarity metric; a Sign Scratchpad, which is a lightweight, image-width-sized memory that stores previous tokens; a Sign Scratchpad Managing Unit, which controls the Sign Scratchpad; and a Token Integration Map to facilitate efficient, image-adaptive token merging. Our evaluations demonstrate that AdapTiV achieves, on average, 309.4 ×, 18.4×, 89.8×, 6.3× speedups and 262.1×, 21.5×, 496.6×, 11.2× improvements in energy efficiency over edge CPUs, edge GPUs, server CPUs, and server GPUs, while maintaining an accuracy loss below 1 % without additional training. Seungjae Yoo, Joo-Young Kim 0001 |
MICRO | 3 |
| 2024 | A DVS-Enabled Distributed Digital LDO Providing Rapid Uniform Power Grid and Ripple Reduction Achieving 20.1-ps FOM in 28 nm CMOSabstractA dynamic voltage scaling (DVS) enabled distributed digital low-dropout voltage regulator (LDO) is described. The proposed distributed LDO utilizes a multi-point average sensing to enable rapid and uniform output voltage regulation across a large-scale power grid, even during unbalanced load transients. A 16-bit thermometer-code flash analog-to-digital converter (FADC) combined with unary passgate configurations and an adaptive on-resistance (R$_{\mathrm {ON}}$) modulation is employed to ensure a small output voltage ripple during DVS operation, using only a small 13.4nF output capacitor. The proposed distributed LDO has been implemented in 28nm CMOS, achieves a 10.4A/mm2 current density, 99.96% current efficiency, and a 20.1ps FOM. It has also been tested under various unbalanced load transient conditions and can rapidly regulate the output voltage back to the target level. Yuli Han, Gunmo Koo, Jaejin Kim, Jusung Kim, Joo-Young Kim 0001, Kunhee Cho |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | LightTrader: A Standalone High-Frequency Trading System with Deep Learning Inference Accelerators and Proactive SchedulerabstractRecent research shows that artificial intelligence (AI) algorithms can dramatically improve the profitability of high-frequency trading (HFT) with accurate market prediction, overcoming the limitation of conventional latency-oriented approaches. However, it is challenging to integrate the computationally intensive AI algorithm into the existing trading pipeline due to its excessively long latency and insufficient throughput, necessitating a breakthrough in hardware. Furthermore, harsh HFT environments such as bursty data traffic and stringent power constraint make it even more difficult to achieve system-level performance without missing crucial market signals.In this paper, we present LightTrader, the world’s first AI-enabled HFT system that incorporates an FPGA and custom AI accelerators for short-latency-high-throughput trading systems. Leveraging the computing power of brand-new AI accelerators fabricated in TSMC’s 7nm FinFET technology, LightTrader optimizes the tick-to-trade latency and response rate for stock market data. The AI accelerators, adopting Coarse-Grained Reconfigurable Array (CGRA) architecture, which maximizes the hardware utilization from the flexible dataflow architecture, achieve a throughput of 16 TFLOPS and 64 TOPS. In addition, we propose both workload scheduling and dynamic voltage and frequency scaling (DVFS) scheduling algorithms to find an optimal offloading strategy under bursty market data traffic and limited power condition. Finally, we build a reliable and rerunnable simulation framework that can back-test the historical market data, such as Chicago Mercantile Exchange (CME), to evaluate the LightTrader system. We thoroughly explore the performance of LightTrader when the number of AI accelerators, power conditions, and complexity of deep neural network models change. As a result, LightTrader achieves 13.92× and 7.28× speed-up of AI algorithm processing compared to existing GPU-based, FPGA-based systems, respectively. LightTrader with multiple AI accelerators achieves up to 99.5% response rates, while LightTrader with the proposed workload scheduling and DVFS scheduling algorithm relieves the miss rate from 17.1% to 23.1%. Sungyeob Yoo, Hyunsung Kim 0003, Jinseok Kim 0006, Sunghyun Park 0006, Joo-Young Kim 0001, Jinwook Oh |
HPCA | 5 |
| 2023 | PRIMO: A Full-Stack Processing-in-DRAM Emulation Framework for Machine Learning WorkloadsabstractRecently, the size of deep learning models has significantly increased, making the excessive memory access between the AI processor and DRAM a major bottleneck of the system. The processing-in-DRAM (DRAM-PIM) concept has emerged as a promising solution, which integrates computing logic within memory, thus saving abundant access to external memory. Although many simulators have been proposed to model and analyze the benefits of DRAM-PIM, they are often too slow to run an entire application. FPGA-based emulators have been introduced to overcome this limitation. However, none of the prior works include the full software stack from the model to DRAM-PIM hardware. This paper presents a full-stack processing-in-DRAM emulation framework named PRIMO, the first emulation framework that can model and analyze DRAM-PIM for end-to-end ML inference. PRIMO enables software developers to develop and test their customized software stacks on various ML workloads without requiring a real DRAM-PIM chip. Moreover, it allows designers to explore design space and monitor memory access patterns, facilitating software and hardware co-design for efficient DRAM-PIM architectures. To achieve these goals, we develop a real-time FPGA emulator that emulates DRAM-PIM architecture and generates experimental results such as predicted cycle information and computed output at incomparably high speeds compared to the CPU-based simulation. In addition, we propose a software stack comprising a PIM compiler that enables the execution of various ML workloads, including end-to-end inference, and a PIM driver that runs the workloads with high bandwidth utilization by leveraging virtual memory scatter-gather DMA. Finally, we demonstrate that PRIMO can successfully emulate DRAM-PIM 106.64-6093.56× faster than the CPU-based simulation framework for ML workloads ranging from small microbenchmarks to end-to-end inference of ResNets. Jaehoon Heo, Yongwon Shin, Sangjin Choi, Sungwoong Yune, Hyojin Sung, Youngjin Kwon, Joo-Young Kim 0001 |
ICCAD | 8 |
| 2023 | Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable BootstrappingabstractHomomorphic encryption (HE) is a type of cryptography that allows computations to be performed on encrypted data. The technique relies on learning with errors problem, where data is hidden under noise for security. To avoid excessive noise, bootstrapping is used to reset the noise level in the ciphertext, but it requires a large key and is computationally expensive. The fully homomorphic encryption over the torus (TFHE) scheme offers a faster and programmable bootstrapping (PBS) algorithm, which is crucial for many privacy-focused applications. Nonetheless, the current TFHE scheme does not support ciphertext packing, resulting in low-throughput performance. To the best of our knowledge, this is the first work that thoroughly analyzes TFHE bootstrapping, identifies the TFHE acceleration bottleneck in GPUs, and proposes a hardware TFHE accelerator to solve the bottleneck. Adiwena Putra, Prasetiyo, Yi Chen 0035, John Kim 0001, Joo-Young Kim 0001 |
MICRO | 5 |
| 2023 | Accelerating Large-Scale Graph-Based Nearest Neighbor Search on a Computational Storage Platformabstract$K$-nearest neighbor search is one of the fundamental tasks in various applications and the hierarchical navigable small world (HNSW) has recently drawn attention in large-scale cloud services, as it easily scales up the database while offering fast search. On the other hand, a computational storage device (CSD) that combines programmable logic and storage modules on a single board becomes popular to address the data bandwidth bottleneck of modern computing systems. In this paper, we propose a computational storage platform that can accelerate a large-scale graph-based nearest neighbor search algorithm based on SmartSSD CSD. To this end, we modify the algorithm more amenable on the hardware and implement two types of accelerators using HLS- and RTL-based methodology with various optimization methods. In addition, we scale up the proposed platform to have 4 SmartSSDs and apply graph parallelism to boost the system performance further. As a result, the proposed computational storage platform achieves 75.59 query per second throughput for the SIFT1B dataset at 258.66W power dissipation, which is 12.83x and 17.91x faster and 10.43x and 24.33x more energy efficient than the conventional CPU-based and GPU-based server platform, respectively. With multi-terabyte storage and custom acceleration capability, we believe that the proposed computational storage platform is a promising solution for cost-sensitive cloud datacenters. Ji-Hoon Kim 0004, Yeo-Reum Park, Jaeyoung Do, Soo-Young Ji, Joo-Young Kim 0001 |
IEEE Trans. Computers | 5 |
| 2023 | Agamotto: A Performance Optimization Framework for CNN Accelerator With Row Stationary DataflowabstractWe propose a software/hardware co-design framework called Agamotto for the complete design automation and performance optimization of the row stationary-based CNN accelerator. We design a scalable accelerator template whose critical design parameters can be configured. Based on the hardware template, Agamotto estimates the performance of the numerous possible hardware implementations for the target FPGA device and CNN model using the latency modeling tool. It chooses the best hardware design and generates the instructions and optimal runtime variables for each target CNN layer. As a result, Agamotto can generate the best hardware design within 61.67 seconds, achieving up to 2.8x higher hardware utilization than the original accelerator. In addition, experimental results show that the performance estimation is accurate, showing only 4.8% difference against the FPGA runtime for the end-to-end CNN model execution. The accelerator implemented on the Xilinx VCU118 evaluation board achieves 402 giga operations per second (GOPS) at 200 MHz, resulting in 13 frames per second (FPS) for the end-to-end execution of VGG-16. It is flexible enough to run more complex CNN models such as ResNet-50 and DarkNet-53, achieving 29.3 FPS and 16.9 FPS, respectively. Sanghyun Jeong, Joo-Young Kim 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Accelerating Deep Convolutional Neural Networks Using Number Theoretic TransformabstractModern deep convolutional neural networks (CNNs) suffer from high computational complexity due to excessive convolution operations. Recently, fast convolution algorithms such as fast Fourier transform (FFT) and Winograd transform have gained attention to address this problem. They reduce the number of multiplications required in the convolution operation by replacing it with element-wise multiplication in the transform domain. However, fast convolution-based CNN accelerators have three major concerns: expensive domain transform, large memory overhead, and limited flexibility in kernel size. In this paper, we present a novel CNN accelerator based on number theoretic transform (NTT), which overcomes the existing limitations. We propose the low-cost NTT and inverse-NTT converter that only use adders and shifters for on-chip domain transform, which solves the inflated bandwidth problem and enables more parallel computations in the accelerator. We also propose the accelerator architecture that includes multiple tile engines with the optimized data flow and mapping. Finally, we implement the proposed NTT-based CNN accelerator on the Xilinx Alveo U50 FPGA and evaluate it for popular deep CNN models. As a result, the proposed accelerator achieves 2859.5, 990.3, and 805.6 GOPS throughput for VGG-16, GoogLeNet, and Darknet-19, respectively. It outperforms the existing fast convolution-based CNN accelerators up to$9.6\times $. Prasetiyo, Seongmin Hong, Yashael Faith Arthanto, Joo-Young Kim 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | A Dual-Mode Similarity Search Accelerator based on Embedding Compression for Online Cross-Modal Image-Text RetrievalabstractImage-text retrieval (ITR) that identifies the relevant images for a given text query, or vice versa, is the fundamental task in emerging vision-and-language machine learning applications. Recently, the cross-modal approach that extracts image and text features in separate reasoning pipelines but performs the similarity search on the same embedding representation is proposed for the real-time ITR system. However, the similarity search that finds the most relevant data in huge data embeddings for a given query becomes the bottleneck of the ITR system.In this paper, we propose a dual-mode similarity search accelerator that can solve the computational hurdle for online image-to-text and text-to-image retrieval service. We propose an embedding compression scheme that removes the sparsity in the text embeddings, further eliminating the time-consuming masking operations in the later processing pipeline. Combining with the data quantization from 32-bit floating-point to 8- bit integer, we reduce the target dataset size by 95.1% with less than 0.1% accuracy loss for 1024-dimensional embedding features. In addition, we propose a streamlined similarity search data flow for both query types, which minimizes the required memory bandwidth with maximal data reuse. The query and data embeddings are guaranteed to be fetched only once from the external memory with the optimized data flow. Based on the proposed data representation and flow, we design a scalable similarity search accelerator that includes multiple ITR kernels. Each ITR kernel has modular design, composed of a separate memory access module and a computing module. The computing module supports pipelined operations of the four similarity search tasks: dot product calculation, data reordering, partial score aggregation, and ranking. We double the number of processing operations in the computing module with the DSP packing technique. Finally, we implement the proposed accelerator with six ITR kernels on the Xilinx Alveo U280 FPGA card. It shows 2.98 tera operations per second (TOPS) performance at 186 MHz, achieving 526/144 and 1163/306 queries per second (QPS) performance for image-to-text and text-to-image retrieval on MS-COCO 1K/5K benchmark. It is up to 359.0 × and 13.9 × faster and 503.6 × and 68.7 × more energy-efficient than the baseline and optimized GPU implementation on Nvidia Titan RTX, respectively. Yeo-Reum Park, Ji-Hoon Kim 0004, Jaeyoung Do, Joo-Young Kim 0001 |
FCCM | 4 |
| 2022 | OpenMDS: An Open-Source Shell Generation Framework for High-Performance Design on Multi-Die FPGAsabstractFPGA is a promising platform in designing a hardware accelerator due to its design flexibility and fast development cycle, despite the device's limited hardware resources. To address this, latest FPGAs have adopted a multi-die architecture providing abundant hardware resources with high yield and cost-benefit. However, the multi-die architecture causes critical timing issues when signal paths cross the die-to-die boundaries, adding another design challenge in using FPGA. We propose OpenMDS, an open-source shell generation framework for high-performance design on multi-die FPGAs. Based on the user's design requirements, it generates an optimized shell for the target FPGA via automated bus pipelining, customized floorplanning, and scalable clocking scheme. Gyeongcheol Shin, Junsoo Kim 0002, Joo-Young Kim 0001 |
FCCM | 3 |
| 2022 | FSHMEM: Supporting Partitioned Global Address Space on FPGAs for Large-Scale Hardware Acceleration InfrastructureabstractBy providing highly efficient one-sided communication with globally shared memory space, Partitioned Global Address Space (PGAS) has become one of the most promising parallel computing models in high-performance computing (HPC). Meanwhile, FPGA is getting attention as an alternative compute platform for HPC systems with the benefit of custom computing and design flexibility. However, the exploration of PGAS has not been conducted on FPGAs, unlike the traditional message passing interface. This paper proposes FSHMEM, a software/hardware framework that enables the PGAS programming model on FPGAs. We implement the core functions of GASNet specification on FPGA for native PGAS integration in hardware, while its programming interface is designed to be highly compatible with legacy software. Our experiments show that FSHMEM achieves the peak bandwidth of 3813 MB/s, which is more than 95% of the theoretical maximum, outperforming the prior works by 9.5×. It records 0.35us and 0.59us latency for remote write and read operations, respectively. Finally, we conduct a case study on the two Intel D5005 FPGA nodes integrating Intel's deep learning accelerator. The two-node system programmed by FSHMEM achieves 1.94× and 1.98× speedup for matrix multiplication and convolution operation, respectively, showing its scalability notential for HPC infrastructure. Yashael Faith Arthanto, David Ojika, Joo-Young Kim 0001 |
FPL | 3 |
| 2022 | LearningGroup: A Real-Time Sparse Training on FPGA via Learnable Weight Grouping for Multi-Agent Reinforcement LearningabstractMulti-agent reinforcement learning (MARL) is a powerful technology to construct interactive artificial intelligent systems in various applications such as multi-robot control and self-driving cars. Unlike supervised model or single-agent rein-forcement learning, which actively exploits network pruning, it is obscure that how pruning will work in multi-agent reinforcement learning with its cooperative and interactive characteristics. In this paper, we present a real-time sparse training accel-eration system named LearningGroup, which adopts network pruning on the training of MARL for the first time with an algorithm/architecture co-design approach. We create spar-sity using a weight grouping algorithm and propose on-chip sparse data encoding loop (OSEL) that enables fast encoding with efficient implementation. Based on the OSEL's encoding format, LearningGroup performs efficient weight compression and computation workload allocation to multiple cores, where each core handles multiple sparse rows of the weight matrix simultaneously with vector processing units. As a result, LearningGroup system minimizes the cycle time and memory footprint for sparse data generation up to 5.72x and 6.81x. Its FPGA accelerator shows 257.40-3629.48 GFLOPS throughput and 7.10-100.12 GFLOPS/W energy efficiency for various conditions in MARL, which are 7.13x higher and 12.43x more energy efficient than Nvidia Titan RTX GPU, thanks to the fully on-chip training and highly optimized dataflow/data format provided by FPGA. Most importantly, the accelerator shows speedup up to 12.52 x for processing sparse data over the dense case, which is the highest among state-of-the-art sparse training accelerators. Je Yang, Jaeuk Kim, Joo-Young Kim 0001 |
FPT | 3 |
| 2022 | DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generationabstract•DFX: a low-latency multi-FPGA appliance for accelerating transformer-based text generation–DFX is a multi-FPGA appliance that accelerates transformer-based text generation–DFX adopts model parallelism to efficiently process the large-scale language model–Xilinx Alveo U280 data center accelerator card provides high performance with low-cost–FPGA-to-FPGA communication is enabled by QSFP cable at 100 Gb/s Seongmin Hong, Seungjae Moon, Junsoo Kim 0002, Sungjae Lee 0002, Minsub Kim, Dongsoo Lee, Joo-Young Kim 0001 |
HCS | 7 |
| 2022 | Trinity: End-to-End In-Database Near-Data Machine Learning Acceleration Platform for Advanced Data AnalyticsabstractThree Important yet Independent Technology Trends Ji-Hoon Kim 0004, Kwanghyun Park 0001, Soo-Young Ji, Joo-Young Kim 0001 |
HCS | 5 |
| 2022 | LightTrader : World's first AI-enabled High-Frequency Trading Solution with 16 TFLOPS / 64 TOPS Deep Learning Inference AcceleratorsabstractWe present the world’s first AI-enabled high-frequency trading (HFT) system, LightTrader , which integrates the custom AI accelerators and the FPGA-based conventional HFT pipeline for the low-latency-high-throughput trading solutions with a reduced query miss rate. For better utilization, adaptive job scheduling methods are also proposed to further improve the performance, where layer-wise workload scaling and dynamic voltage-frequency scaling (DVFS) techniques progressively adjust the workloads of AI accelerators, in conjunction with the architecture support. LightTrader integrating TSMC 7nm tape-out accelerators solely achieves 6x speed-up of DNN processing and 30-50x reduction of query miss rate without the scheduling method while the scheduling scheme further improves the energy efficiency by 25% and reduces the query miss rate by 2.4x . Hyunsung Kim 0003, Sungyeob Yoo, Jaewan Bae, Kyeongryeol Bong, Yoonho Boo, Karim Charfi, Hyo-Eun Kim, Hyun Suk Kim, Jinseok Kim 0006, Byungjae Lee, Myeongbo Shim, Sungho Shin, Jeong Seok Woo, Joo-Young Kim 0001, Sunghyun Park 0006, Jinwook Oh |
HCS | 15 |
| 2022 | DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text GenerationabstractTransformer is a deep learning language model widely used for natural language processing (NLP) services in datacenters. Among transformer models, Generative Pretrained Transformer (GPT) has achieved remarkable performance in text generation, or natural language generation (NLG), which needs the processing of a large input context in the summarization stage, followed by the generation stage that produces a single word at a time. The conventional platforms such as GPU are specialized for the parallel processing of large inputs in the summarization stage, but their performance significantly degrades in the generation stage due to its sequential characteristic. Therefore, an efficient hardware platform is required to address the high latency caused by the sequential characteristic of text generation. In this paper, we present DFX, a multi-FPGA acceleration appliance that executes GPT-2 model inference end-to-end with low latency and high throughput in both summarization and generation stages. DFX uses model parallelism and optimized dataflow that is model-and-hardware-aware for fast simultaneous workload execution among devices. Its compute cores operate on custom instructions and provide GPT-2 operations end-to-end. We implement the proposed hardware architecture on four Xilinx Alveo U280 FPGAs and utilize all of the channels of the high bandwidth memory (HBM) and the maximum number of compute resources for high hardware efficiency. DFX achieves 5.58$\times$ speedup and 3.99$\times$ energy efficiency over four NVIDIA V100 GPUs on the modern GPT-2 model. DFX is also 8.21$\times$ more cost-effective than the GPU appliance, suggesting that it is a promising solution for text generation workloads in cloud datacenters. Seongmin Hong, Seungjae Moon, Junsoo Kim 0002, Sungjae Lee 0002, Minsub Kim, Dongsoo Lee, Joo-Young Kim 0001 |
MICRO | 7 |
| 2021 | FIXAR: A Fixed-Point Deep Reinforcement Learning Platform with Quantization-Aware Training and Adaptive ParallelismabstractDeep reinforcement learning (DRL) is a powerful technology to deal with decision-making problem in various application domains such as robotics and gaming, by allowing an agent to learn its action policy in an environment to maximize a cumulative reward. Unlike supervised models which actively use data quantization, DRL still uses the single-precision floating-point for training accuracy while it suffers from computationally intensive deep neural network (DNN) computations. In this paper, we present a deep reinforcement learning acceleration platform named FIXAR, which employs fixed-point data types and arithmetic units for the first time using a SW/HW co-design approach. We propose a quantization-aware training algorithm in fixed-point, which enables to reduce the data precision by half after a certain amount of training time without losing accuracy. We also design a FPGA accelerator that employs adaptive dataflow and parallelism to handle both inference and training operations. Its processing element has configurable datapath to efficiently support the proposed quantized-aware training. We validate our FIXAR platform, where the host CPU emulates the DRL environment and the FPGA accelerates the agent’s DNN operations, by running multiple benchmarks in continuous action spaces based on a latest DRL algorithm called DDPG. Finally, the FIXAR platform achieves 25293.3 inferences per second (IPS) training throughput, which is 2.7 times higher than the CPU-GPU platform. In addition, its FPGA accelerator shows 53826.8 IPS and 2638.0 IPS/W energy efficiency, which are 5.5 times higher and 15.4 times more energy efficient than those of GPU, respectively. FIXAR also shows the best IPS throughput and energy efficiency among other state-of-the-art acceleration platforms using FPGA, even it targets one of the most complex DNN models. Je Yang, Seongmin Hong, Joo-Young Kim 0001 |
DAC | 3 |
| 2021 | Accelerating Large-Scale Nearest Neighbor Search with Computational Storage DeviceabstractK-nearest neighbor algorithm that searches the K closest samples in a high dimensional feature space is one of the most fundamental tasks in machine learning and image retrieval applications. Computational storage device that combines computing unit and storage module on a single board becomes popular to address the data bandwidth bottleneck of the conventional computing system. In this paper, we propose a nearest neighbor search acceleration platform based on computational storage device, which can process a large-scale image dataset efficiently in terms of speed, energy, and cost. We believe that the proposed acceleration platform is promising to be deployed in cloud datacenters for data-intensive applications. Ji-Hoon Kim 0004, Yeo-Reum Park, Jaeyoung Do, Soo Young Ji, Joo-Young Kim 0001 |
FCCM | 5 |
| 2016 | A cloud-scale acceleration architectureabstractHyperscale datacenter providers have struggled to balance the growing need for specialized hardware (efficiency) with the economic benefits of homogeneity (manageability). In this paper we propose a new cloud architecture that uses reconfigurable logic to accelerate both network plane functions and applications. This Configurable Cloud architecture places a layer of reconfigurable logic (FPGAs) between the network switches and the servers, enabling network flows to be programmably transformed at line rate, enabling acceleration of local applications running on the server, and enabling the FPGAs to communicate directly, at datacenter scale, to harvest remote FPGAs unused by their local servers. We deployed this design over a production server bed, and show how it can be used for both service acceleration (Web search ranking) and network acceleration (encryption of data in transit at high-speeds). This architecture is much more scalable than prior work which used secondary rack-scale networks for inter-FPGA communication. By coupling to the network plane, direct FPGA-to-FPGA messages can be achieved at comparable latency to previous work, without the secondary network. Additionally, the scale of direct inter-FPGA messaging is much larger. The average round-trip latencies observed in our measurements among 24, 1000, and 250,000 machines are under 3, 9, and 20 microseconds, respectively. The Configurable Cloud architecture has been deployed at hyperscale in Microsoft's production datacenters worldwide. Adrian M. Caulfield, Eric S. Chung, Andrew Putnam, Hari Angepat, Jeremy Fowers, Michael Haselman, Stephen Heil, Matt Humphrey, Puneet Kaur, Joo-Young Kim 0001, Daniel Lo, Todd Massengill, Kalin Ovtcharov, Michael Papamichael, Lisa Woods, Sitaram Lanka, Derek Chiou, Doug Burger |
MICRO | 10 |
| 2015 | A Scalable High-Bandwidth Architecture for Lossless Compression on FPGAsabstractData compression techniques have been the subject of intense study over the past several decades due to exponential increases in the quantity of data stored and transmitted by computer systems. Compression algorithms are traditionally forced to make tradeoffs between throughput and compression quality (the ratio of original file size to compressed file size). FPGAs represent a compelling substrate for streaming applications such as data compression thanks to their capacity for deep pipelines and custom caching solutions. Unfortunately, data hazards in compression algorithms such as LZ77 inhibit the creation of deep pipelines without sacrificing some amount of compression quality. In this work we detail a scalable fully pipelined FPGA accelerator that performs LZ77 compression and static Huffman encoding at rates up to 5.6 GB/s. Furthermore, we explore tradeoffs between compression quality and FPGA area that allow the same throughput at a fraction of the logic utilization in exchange for moderate reductions in compression quality. Compared to recent FPGA compression studies, our emphasis on scalability gives our accelerator a 3.0x advantage in resource utilization at equivalent throughput and compression ratio. Jeremy Fowers, Joo-Young Kim 0001, Doug Burger, Scott Hauck |
FCCM | 2 |
| 2015 | Toward accelerating deep learning at scale using specialized hardware in the datacenter
Kalin Ovtcharov, Olatunji Ruwase, Joo-Young Kim 0001, Jeremy Fowers, Karin Strauss, Eric S. Chung |
Hot Chips Symposium | 3 |
| 2014 | Energy efficient canonical huffman encodingabstractAs data centers are increasingly focused on energy efficiency, it becomes important to develop low power implementations of the various applications that run on them. Data compression plays a critical role in data centers to mitigate storage and communication costs. This work focuses on building a low power, high performance implementation for canonical Huffman encoding. We develop a number of different hardware and software implementations targeting Xilinx Zynq FPGA, ARM Cortex-A9, and Intel Core i7. Despite its sequential nature, we show that our hardware accelerated implementation is substantially more energy efficient than both the ARM and Intel Core i7 implementations. When compared to highly optimized software running on the ARM processor, our hardware accelerated implementation has approximately 15 times more throughput with 10% higher power usage, resulting in an 8X benefit in energy efficiency (measured in encodings/Watt). Additionally, our hardware accelerated implementation is up to 80% faster and over 230 times more energy efficient than a highly optimized Core i7 implementation. Janarbek Matai, Joo-Young Kim 0001, Ryan Kastner |
ASAP | 2 |
| 2014 | A Scalable Multi-engine Xpress9 Compressor with Asynchronous Data TransferabstractData compression is crucial in large-scale storage servers to save both storage and network bandwidth, but it suffers from high computational cost. In this work, we present a high throughput FPGA based compressor as a PCIe accelerator to achieve CPU resource saving and high power efficiency. The proposed compressor is differentiated from previous hardware compressors by the following features:1) targeting Xpress9 algorithm, whose compression quality is comparable to the best Gzip implementation (level 9), 2) a scalable multi-engine architecture with various IP blocks to handle algorithmic complexity as well as to achieve high throughput, 3) supporting a heavily multi-threaded server environment with an asynchronous data transfer interface between the host and the accelerator. The implemented Xpress9 compressor on Altera Stratix V GS performs 1.6-2.4Gbps throughput with 7 engines on various compression benchmarks, supporting up to 128 thread contexts. Joo-Young Kim 0001, Scott Hauck, Doug Burger |
FCCM | 1 |
| 2014 | A reconfigurable fabric for accelerating large-scale datacenter servicesabstractDatacenter workloads demand high computational capabilities, flexibility, power efficiency, and low cost. It is challenging to improve all of these factors simultaneously. To advance datacenter capabilities beyond what commodity server designs can provide, we have designed and built a composable, reconfigurable fabric to accelerate portions of large-scale software services. Each instantiation of the fabric consists of a 6×8 2-D torus of high-end Stratix V FPGAs embedded into a half-rack of 48 machines. One FPGA is placed into each server, accessible through PCIe, and wired directly to other FPGAs with pairs of 10 Gb SAS cables. In this paper, we describe a medium-scale deployment of this fabric on a bed of 1,632 servers, and measure its efficacy in accelerating the Bing web search engine. We describe the requirements and architecture of the system, detail the critical engineering challenges and solutions needed to make the system robust in the presence of failures, and measure the performance, power, and resilience of the system when ranking candidate documents. Under high load, the largescale reconfigurable fabric improves the ranking throughput of each server by a factor of 95% for a fixed latency distribution—or, while maintaining equivalent throughput, reduces the tail latency by 29%. Andrew Putnam, Adrian M. Caulfield, Eric S. Chung, Derek Chiou, Kypros Constantinides, John Demme, Hadi Esmaeilzadeh, Jeremy Fowers, Gopi Prashanth Gopal, Jan Gray, Michael Haselman, Scott Hauck, Stephen Heil, Amir Hormati, Joo-Young Kim 0001, Sitaram Lanka, James R. Larus, Eric Peterson, Simon Pope, Aaron Smith, Jason Thong, Phillip Yi Xiao, Doug Burger |
ISCA | 15 |
| 2011 | 24-GOPS 4.5-mm2 Digital Cellular Neural Network for Rapid Visual Attention in an Object-Recognition SoCabstractThis paper presents the Visual Attention Engine (VAE), which is a digital cellular neural network (CNN) that executes the VA algorithm to speed up object-recognition. The proposed time-multiplexed processing element (TMPE) CNN topology achieves high performance and small area by integrating 4800 (80 × 60) cells and 120 PEs. Pipelined operation of the PEs and single-cycle global shift capability of the cells result in a high PE utilization ratio of 93%. The cells are implemented by 6T static random access memory-based register files and dynamic shift registers to enable a small area of 4.5 mm(2). The bus connections between PEs and cells are optimized to minimize power consumption. The VAE is integrated within an object-recognition system-on-chip (SoC) fabricated in the 0.13- μm complementary metal-oxide-semiconductor process. It achieves 24 GOPS peak performance and 22 GOPS sustained performance at 200 MHz enabling one CNN iteration on an 80 × 60 pixel image to be completed in just 4.3 μs. With VA enabled using the VAE, the workload of the object-recognition SoC is significantly reduced, resulting in 83% higher frame rate while consuming 45% less energy per frame without degradation of recognition accuracy. Seungjin Lee 0001, Minsu Kim 0004, Kwanho Kim, Joo-Young Kim 0001, Hoi-Jun Yoo |
IEEE Trans. Neural Networks | 4 |
| 2010 | Familiarity based unified visual attention model for fast and robust object recognition
Seungjin Lee 0001, Kwanho Kim, Joo-Young Kim 0001, Minsu Kim 0004, Hoi-Jun Yoo |
Pattern Recognit. | 3 |
| 2010 | An attention controlled multi-core architecture for energy efficient object recognition
Joo-Young Kim 0001, Sejong Oh, Seungjin Lee 0001, Minsu Kim 0004, Jinwook Oh, Hoi-Jun Yoo |
Signal Process. Image Commun. | 1 |
| 2010 | Visual Image Processing RAM: Memory Architecture With 2-D Data Location Search and Data Consistency Management for a Multicore Object Recognition ProcessorabstractAbstract-Visual image processing random access memory (VIP-RAM) is proposed for a real-time multicore object recognition processor. It has two key features for the overall processor: 1) single cycle local maximum location search (LMLS) for fast key-point localization in object recognition, and 2) data consistency management (DCM) for producer-consumer data transactions among the processors. To achieve single cycle LMLS operation for a 3 x 3 window, the VIP-RAM adopts a hierarchical three-bank architecture that finds the maximum of each row in each bank first, then finds the final maximum of the window and its address in the top level. To this end, each memory bank embeds specialized logic blocks, such as three successive data read logic and bitwise competition logic comparator. With the single cycle LMLS operation, the key-point localization task is accelerated by 2.6 ? with a 27% reduction of power. For the DCM function, the VIP-RAM includes a valid check unit (VCU) that automatically manages the validity of each 32-bit data. It dynamically updates/checks the validity of the shared data when the producer processor writes the data or the consumer processor reads data. With a customized single-ended memory cell and multibit-line selection logic, the VCU can provide a validity check not only for single data access, but also for multiple data accesses such as burst and LMLS operation. Eliminating data synchronization overhead with the DCM, the VIP-RAM reduces the amount of on-chip data transactions and execution time in producer-consumer data transactions by 22.6% and 15.4%, respectively. The overall object recognition processor that includes eight VIP-RAMs and ten processors is fabricated in 0.18/im complementary metal-oxide-semiconductor technology with the chip size of 7.7 mm ? 5 mm. The VIP-RAM occupies a 1.09 mm ? 0.83 mm die area and dissipates 113.2 mW when it performs the LMLS operation in every cycle at 200 MHz frequency and 1.8-V supply. Joo-Young Kim 0001, Donghyun Kim 0014, Seungjin Lee 0001, Kwanho Kim, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | A 60fps 496mW multi-object recognition processor with workload-aware dynamic power managementabstractAn energy efficient object recognition processor is proposed for real-time visual applications. Its energy efficiency is improved by lowering average power consumption while sustaining high frame rate. To this end, the proposed processor features from all levels of chip design. In architecture level, it performs 3-stage task pipelining for high frame rate operation and workload-aware dynamic power management for low power consumption. In block level, energy efficient special purposed engines are employed while software controlled clock gating is exploited for fine-grained clock control. In circuit level, analog-digital mixed design is used to reduce power with the same performance. As a result, the 49mm2 chip in a 0.13mm technology achieves 60fps object recognition for VGA (640x480) input with 496mW power at the supply of 1.2V. It means only 8.2mJ is dissipated per frame, which is 3.2X more energy efficient than the state of the art. Joo-Young Kim 0001, Seungjin Lee 0001, Jinwook Oh, Minsu Kim 0004, Hoi-Jun Yoo |
ISLPED | 1 |
| 2009 | A Configurable Heterogeneous Multicore Architecture With Cellular Neural Network for Real-Time Object RecognitionabstractAs object recognition requires huge computation power to deal with complex image processing tasks, it is very challenging to meet real-time processing demands under low-power constraints for embedded systems. In this paper, a configurable heterogeneous multicore architecture with a dual-mode linear processor array and a cellular neural network on the network-on-chip platform is presented for real-time object recognition. The bio-inspired attention-based object recognition algorithm is devised to reduce computational complexity of the object recognition. The cellular neural network is utilized to accelerate the visual attention algorithm for selecting salient image regions rapidly. The dual-mode parallel processor is configured into single instruction, multiple data (SIMD) or multiple-instruction-multiple-data modes to perform data-intensive image processing operations while exploiting pixel-level and feature-level parallelisms required for the attention-based object recognition. The algorithm's hybrid parallelization strategy on the proposed architecture is adopted to obtain maximum performance improvement. The performance analysis results, using a cycle-accurate architecture simulator, show that the proposed architecture achieves a speedup of 2.8 times for the target algorithm over conventional massively parallel SIMD architecture at low hardware cost overhead. A prototype chip of the proposed architecture, fabricated in 0.13 mum complementary metal-oxide-semiconductor technology, achieves 22 frames/s real-time object recognition with less than 600 mW power consumption. Kwanho Kim, Seungjin Lee 0001, Joo-Young Kim 0001, Minsu Kim 0004, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | 81.6 GOPS Object Recognition Processor Based on a Memory-Centric NoCabstractFor mobile intelligent robot applications, an 81.6 GOPS object recognition processor is implemented. Based on an analysis of the target application, the chip architecture and hardware features are decided. The proposed processor aims to support both task-level and data-level parallelism. Ten processing elements are integrated for the task-level parallelism and single instruction multiple data (SIMD) instruction is added to exploit the data-level parallelism. The memory-centric network-on-chip (NoC) is proposed to support efficient pipelined task execution using the ten processing elements. It also provides coherence and consistency schemes tailored for 1-to-N and M-to-1 data transactions in a task-level pipeline. For further performance gain, the visual image processing memory is also implemented. The chip is fabricated in a 0.18-mum CMOS technology and computes the key-point localization stage of the SIFT object recognition twice faster than the 2.3 GHz Core 2 Duo processor. Donghyun Kim 0014, Kwanho Kim, Joo-Young Kim 0001, Seungjin Lee 0001, Se-Joong Lee, Hoi-Jun Yoo |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Vision platform for mobile intelligent robot based on 81.6 GOPS object recognition processorabstractTo enable power-efficient object recognition of mobile intelligent robots, 81.6GOPS object recognition processor is proposed. Based on analysis of Scale Invariant Feature Transform (SIFT) algorithm, architecture of the proposed processor is designed to support both task and data level parallelism. 10 Processing Elements (PEs) are integrated for task parallelism, and each PE is equipped with SIMD instruction for data parallelism as well. In addition, Visual Image Processing memory replaces complex local maximum pixel search operation with a single read operation for further performance gain. With the proposed processor, we also realized vision platform for real-time SIFT computation of mobile robots. The chip operation is tested up to 200MHz and consumes 540mW in the vision platform at 1.8V supply voltage and 100 MHz operation frequency. Donghyun Kim 0014, Kwanho Kim, Joo-Young Kim 0001, Seungjin Lee 0001, Hoi-Jun Yoo |
DAC | 3 |
| 2008 | A 0.6pJ/b 3Gb/s/ch transceiver in 0.18 µm CMOS for 10mm on-chip interconnectsabstractThis paper presents a high speed and low energy transceiver for 10mm long minimum width on-chip global interconnects. To improve the link bandwidth, the transmitter employs a capacitive-resistive pre-emphasis technique and the receiver employs the AC-coupled Resistive Feedback Inverter (RFI) de-emphasis technique. Exploiting two emphasis techniques, the proposed interconnect achieves 1.26GHz bandwidth which is 20 times improved compared to conventional link. As a result, it achieves error-free 3Gb/s data rate and consumes less than 0.6pJ/b during transmission by using low-swing and pulse signaling. The test chip is designed using 1.8V 0.18 μm 6M CMOS technology. Joonsung Bae, Joo-Young Kim 0001, Hoi-Jun Yoo |
ISCAS | 2 |
| 2007 | Solutions for Real Chip Implementation Issues of NoC and Their Application to Memory-Centric NoCabstractThis paper describes real chip implementation issues of network-on-chip (NoC) and their solutions along with series of chip design examples. The solutions described in this paper cover both architectural aspects and circuit level techniques for practical chip implementation of NoC. As for architecture level solutions, topology selection, chip-aware protocol design, and on-chip serialization (OCS) for link area reduction are explained. For circuit level techniques, SERDES and synchronizer design, crossbar switch partial activation, and low-voltage link are presented as the foundations for power and area efficient NoC implementation. Regarding presented solutions for NoC implementation, this paper proposes memory centric NoC (MC-NoC) for homogeneous multi processor SoC (MPSoC). Flexibility and feasibility of task mapping on homogeneous SoC is the key feature of the MC-NoC. 8 dual port SRAMs connected to crossbar switches in hierarchical star topology network facilitate data communication between processors, regardless of task mapping into the MC-NoC. Experimental result obtained by mapping edge detection tasks on the MC-NoC in various configurations shows almost constant performance. This result proves the effectiveness of the proposed architecture. The MC-NoC based SoC is also implemented on TSMC 0.18 um process technology Donghyun Kim 0014, Kwanho Kim, Joo-Young Kim 0001, Seungjin Lee 0001, Hoi-Jun Yoo |
NOCS | 3 |
| 2006 | A 372 ps 64-bit adder using fast pull-up logic in 0.18µm CMOSabstractThis paper presents a 372 ps 64-bit adder using fast pull-up logic (FPL) in 0.18 mum CMOS technology. Fast pull-up logic is devised and applied to decrease pull-up time which is critical in domino-static adder. The implemented adder measures the worst case delay of 372 ps. The adder has a modified tree architecture using load distribution method and has 6 logic stages Joo-Young Kim 0001, Kangmin Lee, Hoi-Jun Yoo |
ISCAS | 1 |