Jeongwoo Park 0001

dblp:88/11394-1 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-9603-9588ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BOA-3DGS: Backward-Striding Optimized Accelerator for Reduced Memory Contention in 3D Gaussian Splatting Training
abstract
D Gaussian Splatting (3DGS) algorithms have gained increasing attention due to their ability to enable realistic 3D scene reconstruction with faster runtime. However, training these models on graphical processing units (GPU) faces unique challenges during the backpropagation stage, primarily due to memory barriers in the execution model, where every pixel within a tile computes the gradient for a single Gaussian per thread. In this paper, we propose the Backward-Striding Optimized Accelerator for 3D Gaussian Splatting (BOA-3DGS), a hardware-software co-optimized accelerator designed to optimize backward rasterization of 3D Gaussian Splatting by enabling pixels to stride through and select only relevant Gaussians for computing. However, such processing styles lead to a large number of memory stalls due to conflicting gradient storage and Gaussian parameter fetches. By introducing a pixel-independent, funnel-like multi-Gaussian alpha computation and a majoritybased gradient accumulation method, we avoid such memory stalls to efficiently accelerate backpropagation of 3DGS. Together, these enhancements improve gradient calculation efficiency and accelerate the backpropagation stage of 3DGS rasterization. Experimental results demonstrate that BOA-3DGS achieves up to $1.58 \times$ speedup during backpropagation compared to prior work, while utilizing only $0.86 \times$ the area and consuming $0.99 \times$ power.
Hyuk Jun Kweon, Jongyeop Kim, Jeongwoo Park 0001
ASP-DAC3
2026 H2NoC: A Hybrid NoC Architecture for FPGAs with Hardened Interconnects
abstract
Modern FPGA platforms with high-bandwidth memory (HBM) subsystems often incorporate hardened interconnects in their static regions to alleviate routing congestion in programmable logic. However, these interconnects typically suffer from up to 67% bandwidth degradation in platforms such as the Alveo U280 under specific access patterns due to their inherent structural limitations. Previous works introduce custom Networks-on-Chip (NoC) in FPGA programmable logic but fail to fully utilize all 32 HBM channels due to excessive resource overhead and severe routing congestion. In this paper, we propose $\mathbf{H}^{2} \mathbf{N o C}$, a hybrid NoC architecture that integrates the hardened interconnect embedded in the static HBM region with a custom Butterfly-Fat Tree (BFT) NoC in programmable logic, enabling full 32-channel utilization with minimum routing congestion. The custom 3-stage NoC is partitioned into four clusters, enabling flexible floorplanning strategies that facilitate optimal implementation across SLRs. To compensate for the reduced number of on-chip stages, an out-of-order generating buffer (OGB) is incorporated, which utilizes lightweight AXI reordering to mitigate read-bandwidth degradation. Compared to prior state-of-the-art NoCs on Alveo series, $\mathbf{H}^{\mathbf{2}} \mathbf{N o C}$ demonstrates an improvement of 134.8% in bandwidth-to-resource efficiency. $\mathrm{H}^{2} \mathrm{NoC}$ is evaluated using a set of synthetic benchmarks designed to emulate various traffic patterns. It achieves a peak bandwidth of $391 \mathrm{~GB} / \mathrm{s}$ under bit permutation traffic and sustains over $200 \mathrm{~GB} / \mathrm{s}$ across most traffic patterns. The proposed hybrid architecture delivers performance comparable to a full custom 5-stage BFT network while requiring significantly fewer FPGA resources.
Jinhyeong Park, Younggil Jeong, Jeongwoo Park 0001
ASP-DAC3
2025 RGHT-Q: Reconfigurable GEMM Unit for Heterogeneous-Homogeneous Tensor Quantization
abstract
The high computational demands of large language models (LLMs) are limited by the lack of GPU hardware support for heterogeneous quantization, which mixes integers and floating points. To address this limitation, we propose an LLM processing element (PE), RGHT-Q, which features reconfigurable general-matrix multiplication (GEMM) operations for both heterogeneous and homogeneous tensor quantization. The RGHT-Q introduces a novel design that leverages butterfly routing and multi-precision multipliers. As a result, we achieve significant performance improvements, offering 3.14× higher energy efficiency, and 1.56× better area efficiency compared to prior designs.
Donghyun Nam, Jeongwoo Park 0001
DATE3
2025 CLAT: A Clustering-Based Attention Transformer Accelerator for Low-Latency Text Generation in LLMs
abstract
Transformer-based large language models (LLMs) excel in text generation but face challenges like memory bandwidth bottlenecks and large key-value (KV) cache sizes as context lengths grow, impacting low-latency performance. Existing accelerators adopt parallelism, model compression, and sparsity exploitation but often fail to fully utilize token-specific sparsity, limiting their effectiveness for long context lengths. CLAT addresses these issues with a low-overhead clustering algorithm that identifies relevant key vector clusters for each token’s query, omitting less relevant vectors with minimal impact. It optimizes memory bandwidth using routing for single-batch inference and introduces scheduling techniques to reduce attention layer latency. Additionally, CLAT compresses model parameters to 4-bit precision and KV caches to 8-bit precision, supported by a multi-precision MAC structure that avoids extra overhead. Validated on Llama2-7B, OPT-6.7B, and Llama3-8B models, CLAT reduces attention layer latency by up to 88.6% and overall text generation latency by up to 34.9%. It improves single-batch text generation throughput by$1.66\times $to$2.42\times $over an A100 GPU, demonstrating significant performance gains.
Sunwoo Lee 0005, Beomseok Kim, Jeongwoo Park 0001, Dongsuk Jeon
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Toward Efficient Low-Precision Training: Data Format Optimization and Hysteresis Quantization
Sunwoo Lee 0005, Jeongwoo Park 0001, Dongsuk Jeon
ICLR2
2021 Activation Sharing with Asymmetric Paths Solves Weight Transport Problem without Bidirectional Connection
abstract
One of the reasons why it is difficult for the brain to perform backpropagation (BP) is the weight transport problem, which argues forward and feedback neurons cannot share the same synaptic weights during learning in biological neural networks. Recently proposed algorithms address the weight transport problem while providing good performance similar to BP in large-scale networks. However, they require bidirectional connections between the forward and feedback neurons to train their weights, which is observed to be rare in the biological brain. In this work, we propose an Activation Sharing algorithm that removes the need for bidirectional connections between the two types of neurons. In this algorithm, hidden layer outputs (activations) are shared across multiple layers during weight updates. By applying this learning rule to both forward and feedback networks, we solve the weight transport problem without the constraint of bidirectional connections, also achieving good performance even on deep convolutional neural networks for various datasets. In addition, our algorithm could significantly reduce memory access overhead when implemented in hardware.
Sunghyeon Woo, Jeongwoo Park 0001, Jiwoo Hong, Dongsuk Jeon
NeurIPS2