Joongho Jo

dblp:276/1950 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0002-0421-7031ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FUSE-GS: Frequency-Aware Pruning and Scale-Driven Rendering for Energy Efficient 3D Gaussian Splatting
Hanjun Choi, Hyerin Lim, Joongho Jo, Byeonghun Kwon
ISLPED3
2025 GS-TG: 3D Gaussian Splatting Accelerator with Tile Grouping for Reducing Redundant Sorting while Preserving Rasterization Efficiency
abstract
3D Gaussian Splatting (3D-GS) has emerged as a promising alternative to neural radiance fields (NeRF) as it offers high speed as well as high image quality in novel view synthesis. Despite these advancements, 3D-GS still struggles to meet the frames per second (FPS) demands of real-time applications. In this paper, we introduce GS-TG, a tile-grouping-based accelerator that enhances 3D-GS rendering speed by reducing redundant sorting operations and preserving rasterization efficiency. GS-TG addresses a critical trade-off issue in 3D-GS rendering: increasing the tile size effectively reduces redundant sorting operations, but it concurrently increases unnecessary rasterization computations. So, during sorting of the proposed approach, GS-TG groups small tiles (for making large tiles) to share sorting operations across tiles within each group, significantly reducing redundant computations. During rasterization, a bitmask assigned to each Gaussian identifies relevant small tiles, to enable efficient sharing of sorting results. Consequently, GS-TG enables sorting to be performed as if a large tile size is used by grouping tiles during the sorting stage, while allowing rasterization to proceed with the original small tiles by using bitmasks in the rasterization stage. GS-TG is a lossless method requiring no retraining or fine-tuning and it can be seamlessly integrated with previous 3D-GS optimization techniques. Experimental results show that GS-TG achieves an average speed-up of 1.54 times over state-of-the-art 3D-GS accelerators.
Joongho Jo, Jongsun Park 0001
DAC1
2025 PS-GS: Group-Wise Parallel Rendering with Stage-Wise Complexity Reductions for Real-Time 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3D-GS) is an emerging rendering technique that surpasses the neural radiance field (NeRF) in both rendering speed and image quality. Despite its advantages, running 3D-GS on mobile or edge devices in real-time remains challenging due to large computational complexity. In this paper, we introduce PS-GS, a specialized low-complexity hardware designed to enhance the pipeline parallelism of 3D-GS rendering pipeline process. In this work, we first observe that 3D-GS rendering can be parallelized when the approximate order of Gaussians, from those closest to the camera to those farthest, is known ahead. But, to enhance 3D-GS rendering speed via parallel processing, an efficient viewpoint-adaptive grouping method with low computational costs is essential. Two key computational bottlenecks of viewpoint-adaptive grouping are the grouping of invisible Gaussians and depth-based sorting. For efficient group-wise parallel rendering with low complexity viewpoint-adaptive grouping, we propose three key techniques—cluster-based preprocessing, sorting, and grouping—all seamlessly incorporated into the PS-GS architecture. Our experimental results demonstrate that PS-GS delivers an average speedup of 1.20x with negligible peak signal-to-noise ratio (PSNR) degradation.
Joongho Jo, Jongsun Park 0001
DATE1
2023 LoCoExNet: Low-Cost Early Exit Network for Energy Efficient CNN Accelerator Design
abstract
Early exit techniques, where the inference process of convolutional neural networks (CNNs) is early terminated by using auxiliary classifiers (called branches) to reduce data processing energy for easy inputs, have been actively researched. However, the conventional early exit works have suffered from the large size of branches, whose memory access energy is significantly large compared to that of the main network. In this article, we propose a low-cost early exit network (LoCoExNet), which significantly improves energy efficiencies by reducing the parameters used in inference with efficient branch structures. To further reduce the energy consumption of branches, based on the observation that the classification difficulties of the images can be predicted using the magnitudes of input activations, we also propose hardware-friendly dynamic branch pruning (DBP). Given the network and target dataset, we can construct an energy-efficient early exit network through the branch structure determination algorithm and DBP algorithm. Finally, we develop several architectural supporting modules to improve the energy efficiency of the proposed LoCoExNet with DBP. The experimental results show that the LoCoExNet uses 28.7% of the total network parameters on average for Tiny-ImageNet dataset using VGG-16 without accuracy loss. The CNN accelerator that implements LoCoExNet with DBP has been implemented using the 65nm CMOS process. The implementation results show that the LoCoExNet accelerator achieves up to 91% of energy savings and up to$4.46\mathbf {\times }$of speedups without accuracy loss for CIFAR-10 dataset using VGG-16 compared to the state-of-the-art CNN accelerators.
Joongho Jo, Geonho Kim, Seungtae Kim, Jongsun Park 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Low Complexity Gradient Computation Techniques to Accelerate Deep Neural Network Training
abstract
Deep neural network (DNN) training is an iterative process of updating network weights, called gradient computation, where (mini-batch) stochastic gradient descent (SGD) algorithm is generally used. Since SGD inherently allows gradient computations with noise, the proper approximation of computing weight gradients within SGD noise can be a promising technique to save energy/time consumptions during DNN training. This article proposes two novel techniques to reduce the computational complexity of the gradient computations for the acceleration of SGD-based DNN training. First, considering that the output predictions of a network (confidence) change with training inputs, the relation between the confidence and the magnitude of the weight gradient can be exploited to skip the gradient computations without seriously sacrificing the accuracy, especially for high confidence inputs. Second, the angle diversity-based approximations of intermediate activations for weight gradient calculation are also presented. Based on the fact that the angle diversity of gradients is small (highly uncorrelated) in the early training epoch, the bit precision of activations can be reduced to 2-/4-/8-bit depending on the resulting angle error between the original gradient and quantized gradient. The simulations show that the proposed approach can skip up to 75.83% of gradient computations with negligible accuracy degradation for CIFAR-10 dataset using ResNet-20. Hardware implementation results using 65-nm CMOS technology also show that the proposed training accelerator achieves up to 1.69× energy efficiency compared with other training accelerators.
Dongyeob Shin, Geonho Kim, Joongho Jo, Jongsun Park 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 A Charge-Sharing based 8T SRAM In-Memory Computing for Edge DNN Acceleration
abstract
This paper presents a charge-sharing based customized 8T SRAM in-memory computing (IMC) architecture. In the proposed IMC approach, the multiply-accumulate (MAC) operation of multi-bit activations and weights is supported using the charge sharing between bit-line (BL) parasitic capacitances. The area-efficient customized 8T SRAM macro can achieve robust and voltage-scalable MAC operations due to the charge-domain computation. We also propose a split capacitor structure-based 5/6-bit reconfigurable successive approximation register analog-to-digital converter (SAR-ADC) to reduce the hardware cost of an analog readout circuit while supporting higher precision MAC operations. The proposed reconfigurable SAR-ADC has been exploited to implement layer-by-layer mixed bit-precisions in convolution layer for increasing energy efficiency with negligible accuracy loss. The 256×64 8T SRAM IMC macro has been implemented using 28nm CMOS process technology. The proposed SRAM macro achieves 11. 20-TOPS/W with a maximum clock frequency of 125MHz at 1. 0V. It also supports supply voltage scaling from 0.5V to 1.1V with the energy efficiency ranging from 8.3-TOPS/W to 35.4-TOPS/W within 1 % accuracy loss.
Kyeongho Lee, Sungsoo Cheon, Joongho Jo, Woong Choi, Jongsun Park 0001
DAC3
2020 Prediction Confidence based Low Complexity Gradient Computation for Accelerating DNN Training
abstract
In deep neural network (DNN) training, network weights are iteratively updated with the weight gradients that are obtained from stochastic gradient descent (SGD). Since SGD inherently allows gradient calculations with noise, approximating weight gradient computations have a large potential of training energy/time savings without degrading accuracy. In this paper, we propose an input-dependent approximation of the weight gradient for improving energy efficiency of training process. Considering that the output predictions of network (confidence) changes with training inputs, the relation between the confidence and the magnitude of weight gradient can be efficiently exploited to skip the gradient computations without accuracy drop, especially for high confidence inputs. With a given squared error constraint, the computation skip rates can be also controlled by changing the confidence threshold. The simulation results show that our approach can skip 72.6% of gradient computations for CIFAR-100 dataset using ResNet-18 without accuracy degradation. Hardware implementation with 65nm CMOS process shows that our design achieves 88.84% and 98.16% of maximum per epoch training energy and time savings, respectively, for CIFAR-100 dataset using ResNet-18 compared to state-of-the-art training accelerator.
Dongyeob Shin, Geonho Kim, Joongho Jo, Jongsun Park 0001
DAC3