Jeik Choi

dblp:332/3410 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-6674-450XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 82% Processor architecture and microarchitecture · 11% Energy-efficient computing · 7%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › accelerator architecture
accelerator microarchitecture
0.912025
FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering · ISCA 2025
Hardware accelerators and domain-specific architectures › neural rendering accelerator
neural radiance field accelerator
0.912025
FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering · ISCA 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › low-precision arithmetic
block floating point
0.712023
DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA Extensions · DAC 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator
0.712023
DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA Extensions · DAC 2023
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.312025
FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering · ISCA 2025
Processor architecture and microarchitecture › instruction set architecture
instruction set extension
0.212023
DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA Extensions · DAC 2023
Processor architecture and microarchitecture › instruction set architecture
RISC-V
0.212023
DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA Extensions · DAC 2023

Methods — techniques the papers use, named apart from their topics

sparsity-aware data format · 0.9precision scaling · 0.928nm CMOS layout · 0.9
YearPublicationVenuePosition
2025 FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering
abstract
Neural Radiance Fields (NeRF), an AI-driven approach for 3D view reconstruction, has demonstrated impressive performance, sparking active research across fields.As a result, a range of advanced NeRF models has emerged, leading on-device applications to increasingly adopt NeRF for highly realistic scene reconstructions.With the advent of diverse NeRF models, NeRF-based applications leverage a variety of NeRF frameworks, creating the need for hardware capable of efficiently supporting these models.However, GPUs fail to meet the performance, power, and area (PPA) cost demanded by these on-device applications, or are specialized for specific NeRF algorithms, resulting in lower efficiency when applied to other NeRF models.To address this limitation, in this work, we introduce FlexNeRFer, an energy-efficient versatile NeRF accelerator.The key components enabling the enhancement of FlexNeRFer include: i) a flexible network-on-chip (NoC) supporting multi-dataflow and sparsity on precision-scalable MAC array, and ii) efficient data storage using an optimal sparsity format based on the sparsity ratio and precision modes.To evaluate the effectiveness of FlexNeRFer, we performed a layout implementation using 28nm CMOS technology.Our evaluation shows that FlexNeRFer achieves 8.2∼243.3×speedup and 24.1∼520.3×improvement in energy efficiency over a GPU (i.e., NVIDIA RTX 2080 Ti), while demonstrating 4.2∼86.9×speedup and 2.3∼47.5×improvement in energy efficiency compared to a state-of-the-art NeRF accelerator (i.e., NeuRex).
Seock-Hwan Noh, Banseok Shin, Jeik Choi, Seungpyo Lee, Yeseong Kim
ISCA3
2023 DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA Extensions
abstract
Over the past decade, it has been found that deep neural networks (DNNs) perform better on visual perception and language understanding tasks as their size increases. However, this comes at the cost of high energy consumption and large memory requirement to train such large models. As the training DNNs necessitates a wide dynamic range in representing tensors, floating point formats are normally used. In this work, we utilize a block floating point (BFP) format that significantly reduces the size of tensors and the power consumption of arithmetic units. Unfortunately, prior work on BFP-based DNN training empirically selects the block size and the precision that maintain the training accuracy. To make the BFP-based training more feasible, we propose dynamic block size and precision scaling (DBPS) for highly efficient DNN training. We also present a hardware accelerator, called DBPS core, which supports the DBPS control by configuring arithmetic units with custom instructions extended in a RISC-V processor. As a result, the training time and energy consumption reduce by 67.1% and 72.0%, respectively, without hurting the training accuracy.
Jeik Choi, Seock-Hwan Noh, Jahyun Koo 0002, Jaeha Kung 0001
DAC2
2022 LightNorm: Area and Energy-Efficient Batch Normalization Hardware for On-Device DNN Training
abstract
When training early-stage deep neural networks (DNNs), generating intermediate features via convolution or linear layers occupied most of the execution time. Accordingly, extensive research has been done to reduce the computational burden of the convolution or linear layers. In recent mobile-friendly DNNs, however, the relative number of operations involved in processing these layers has significantly reduced. As a result, the proportion of the execution time of other layers, such as batch normalization layers, has increased. Thus, in this work, we conduct a detailed analysis of the batch normalization layer to efficiently reduce the runtime overhead in the batch normalization process. Backed up by the thorough analysis, we present an extremely efficient batch normalization, named LightNorm, and its associated hardware module. In more detail, we fuse three approximation techniques that are i) low bit-precision, ii) range batch normalization, and iii) block floating point. All these approximate techniques are carefully utilized not only to maintain the statistics of intermediate feature maps, but also to minimize the off-chip memory accesses. By using the proposed LightNorm hardware, we can achieve significant area and energy savings during the DNN training without hurting the training accuracy. This makes the proposed hardware a great candidate for the on-device training.
Seock-Hwan Noh, Junsang Park, Dahoon Park, Jahyun Koo 0002, Jeik Choi, Jaeha Kung 0001
ICCD5