EDBT 2026 Demo / reviewers in the wild / expert
Hongil Yoon
dblp:34/5178
· DBLP profile ↗
17ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GS-Scale: Unlocking Large-Scale 3D Gaussian Splatting Training via Host OffloadingabstractThe advent of 3D Gaussian Splatting has revolutionized graphics rendering by delivering high visual quality and fast rendering speeds. However, training large-scale scenes at high quality remains challenging due to the substantial memory demands required to store parameters, gradients, and optimizer states, which can quickly overwhelm GPU memory. To address these limitations, we propose GS-Scale, a fast and memory-efficient training system for 3D Gaussian Splatting. GS-Scale stores all Gaussians in host memory, transferring only a subset to the GPU on demand for each forward and backward pass. While this dramatically reduces GPU memory usage, it requires frustum culling and optimizer updates to be executed on the CPU, introducing slowdowns due to CPU's limited compute and memory bandwidth. To mitigate this, GS-Scale employs three system-level optimizations: (1) selective offloading of geometric parameters for fast frustum culling, (2) parameter forwarding to pipeline CPU optimizer updates with GPU computation, and (3) deferred optimizer update to minimize unnecessary memory accesses for Gaussians with zero gradients. Our extensive evaluations on large-scale datasets demonstrate that GS-Scale significantly lowers GPU memory demands by 3.3-5.6x, while achieving training speeds comparable to GPU without host offloading. This enables large-scale 3D Gaussian Splatting training on consumer-grade GPUs; for instance, GS-Scale can scale the number of Gaussians from 4 million to 18 million on an RTX 4070 Mobile GPU, leading to 23-35% LPIPS (learned perceptual image patch similarity) improvement. Donghyun Lee 0005, Jae W. Lee, Hongil Yoon |
ASPLOS (2) | 4 |
| 2026 | A Compact Low-Voltage and Low-Power Flip-Flop with Static, Contention-Free, and Redundant-Transition-Free Operation
Suna Lee, Taegun Yim, Choongkeun Lee, Hongil Yoon |
ISCAS | 4 |
| 2026 | A Separated Pre-Charge Sense Amplifier With Fast Sensing, Low Power, Small Area, and High Reliability for Hybrid MTJ/CMOS Logic CircuitsabstractThe use of logic circuits combined with emerging devices has been studied to overcome the limitations of complementary metal-oxide-semiconductor (CMOS) transistors. Among a variety of emerging devices, magnetic tunnel junction (MTJ) is a promising candidate owing to its non-volatility, high endurance, and CMOS compatibility. However, process variations in MTJs and CMOS transistors hinder reliable and precise resistance-to-voltage conversion in hybrid MTJ/CMOS logic circuits. To address this issue, this paper proposes a novel separated pre-charge sense amplifier that achieves fast-sensing, low-power, small-area, and high-reliability. The proposed circuit eliminates intermediate inverters between the discharge and evaluation stages. It incorporates P-channel MOS (PMOS) transistors within the inverter latch, whose gates are directly biased by voltages that reflect the resistance difference between a pair of MTJs. It minimizes its nodes to be charged or discharged during operation. Furthermore, it reduces the total transistor count, including clock-driven transistors. Simulations are performed using Cadence and HSPICE tools with the NCSU CMOS 45nm design kit and a physics-based MTJ SPICE model. Monte Carlo simulations are conducted to check the circuit’s reliability under process, voltage, and temperature (PVT) variations. Post-layout simulation results show that the proposed circuit achieves the fastest sensing delay among the compared circuits except for Separated Pre-Charge Sense Amplifier (SPCSA), lowest power consumption, lowest power-delay product, smallest area overhead, and lowest sensing error rate for a viable usage in hybrid MTJ/CMOS logic memory circuits. Taegun Yim, Hongil Yoon |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | FastPoint: Accelerating 3D Point Cloud Model Inference via Sample Point Distance PredictionabstractDeep neural networks have revolutionized 3D point cloud processing, yet efficiently handling large and irregular point clouds remains challenging. To tackle this problem, we introduce FastPoint, a novel software-based acceleration technique that leverages the predictable distance trend between sampled points during farthest point sampling. By predicting the distance curve, we can efficiently identify subsequent sample points without exhaustively computing all pairwise distances. Our proposal substantially accelerates farthest point sampling and neighbor search operations while preserving sampling quality and model performance. By integrating FastPoint into state-of-the-art 3D point cloud models, we achieve 2.55x end-to-end speedup on NVIDIA RTX 3090 GPU without sacrificing accuracy. Donghyun Lee 0005, Jae W. Lee, Hongil Yoon |
ICCV | 4 |
| 2024 | Frugal 3D Point Cloud Model Training via Progressive Near Point Filtering and Fused Aggregation
Donghyun Lee 0005, Yejin Lee 0001, Jae W. Lee, Hongil Yoon |
ECCV (66) | 4 |
| 2023 | Not All Neighbors Matter: Point Distribution-Aware Pruning for 3D Point CloudabstractApplying deep neural networks to 3D point cloud processing has demonstrated a rapid pace of advancement in those domains where 3D geometry information can greatly boost task performance, such as AR/VR, robotics, and autonomous driving. However, as the size of both the neural network model and 3D point cloud continues to scale, reducing the entailed computation and memory access overhead is a primary challenge to meet strict latency and energy constraints of practical applications. This paper proposes a new weight pruning technique for 3D point cloud based on spatial point distribution. We identify that particular groups of neighborhood voxels in 3D point cloud contribute more frequently to actual output features than others. Based on this observation, we propose to selectively prune less contributing groups of neighborhood voxels first to reduce the computation overhead while minimizing the impact on model accuracy. We apply our proposal to three representative sparse 3D convolution libraries. Our proposal reduces the inference latency by 1.60× on average and energy consumption by 1.74× on NVIDIA GV100 GPU with no loss in accuracy metric Yejin Lee 0001, Donghyun Lee 0005, JungUk Hong, Jae W. Lee, Hongil Yoon |
AAAI | 5 |
| 2022 | STT-MRAM-Based Multicontext FPGA for Multithreading Computing EnvironmentabstractThe demand for high-performance computing and rapidly increasing power consumption has increased the necessity for application-specific accelerators. In the datacenter and mobile system, more applications are increasingly relying on accelerators. Field-programmable gate arrays (FPGAs) emerge as a good candidate because they have high programmability and power efficiency. As the number of applications requiring acceleration increases, there is huge demand for FPGAs that support multiple contexts. Previous FPGA designs that support multicontext have various shortcomings such as volatility, poor power efficiency, large performance, area, and reconfiguration overhead. In this article, we propose a spin-transfer torque magnetic RAM (STT-MRAM)-based nonvolatile multicontext FPGA (NVMC-FPGA) that overcomes these shortcomings. We introduce the NVMC-FPGA architecture and operation modes that take advantage of nonvolatility and support multicontext. We also develop the multicontext-aware FPGA computer aided design flow to make the most of the NVMC-FPGA. Compared to the conventional SRAM-based FPGA, when eight identical circuits are mapped, the NVMC-FPGA improves the performance by 15.3% on average and reduces the power consumption by 11.2%–80.7%, depending on the number of simultaneously activated circuits. Moreover, when eight different circuits are mapped, the NVMC-FPGA improves the performance by 58.5% on average and reduces the power consumption by 6.2%–63.3%, depending on the number of simultaneously activated circuits. Jeongbin Kim 0001, Yongwoon Song, Kyungseon Cho, Hyuk-Jun Lee, Hongil Yoon, Eui-Young Chung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | A High Speed Modified Dickson Charge PumpabstractThis paper proposes a high speed modified Dickson charge pump. The conventional circuit has an auxiliary stage to control the gate of the transistor to the output load. However, the auxiliary stage's voltage level has a problem with a threshold voltage loss due to its diode- connected n-channel metal oxide semiconductor (NMOS) transistor configuration. The proposed circuit demonstrates a scheme with p-channel metal oxide semiconductor (PMOS) transistor configuration instead of the conventional NMOS diode connection in the auxiliary stage. Using this method, the threshold voltage loss is eliminated in the voltage level of the auxiliary stage. The gates of the main transfer transistors are controlled more effectively, leading to the proposed circuit's better performances. The bulk connection of PMOS transistor in the auxiliary stage is also modified to suppress the parasitic bipolar effect. The HSPICE simulations with the TSMC 0.18 μm process technology indicate and verify that the proposed circuit achieves better performances in pump-up speed, output ripple peak-to-peak voltage, and power efficiency compared to those obtained for the conventional schemes. Taegun Yim, Choong Keun Lee, Hongil Yoon |
ISCAS | 3 |
| 2021 | ASAP: Fast Mobile Application Switch via Adaptive Prepaging
Sam Son, Seung Yul Lee, Yunho Jin, Jonghyun Bae, Jinkyu Jeong, Tae Jun Ham, Jae W. Lee, Hongil Yoon |
USENIX ATC | 8 |
| 2018 | Filtering Translation Bandwidth with Virtual CachingabstractHeterogeneous computing with GPUs integrated on the same chip as CPUs is ubiquitous, and to increase programmability many of these systems support virtual address accesses from GPU hardware. However, this entails address translation on every memory access. We observe that future GPUs and workloads show very high bandwidth demands (up to 4 accesses per cycle in some cases) for shared address translation hardware due to frequent private TLB misses. This greatly impacts performance (32% average performance degradation relative to an ideal MMU). To mitigate this overhead, we propose a software-agnostic, practical, GPU virtual cache hierarchy. We use the virtual cache hierarchy as an effective address translation bandwidth filter. We observe many requests that miss in private TLBs find corresponding valid data in the GPU cache hierarchy. With a GPU virtual cache hierarchy, these TLB misses can be filtered (i.e., virtual cache hits), significantly reducing bandwidth demands for the shared address translation hardware. In addition, accelerator-specific attributes (e.g., less likelihood of synonyms) of GPUs reduce the design complexity of virtual caches, making a whole virtual cache hierarchy (including a shared L2 cache) practical for GPUs. Our evaluation shows that the entire GPU virtual cache hierarchy effectively filters the high address translation bandwidth, achieving almost the same performance as an ideal MMU. We also evaluate L1-only virtual cache designs and show that using a whole virtual cache hierarchy obtains additional performance benefits (1.31× speedup on average). Hongil Yoon, Jason Lowe-Power, Gurindar S. Sohi |
ASPLOS | 1 |
| 2018 | A New Charge Pump Switch Parallel to Series ConnectionabstractThis paper proposed increase 2VIN+VCLKvoltage series charge pump. Recently, trend of low power circuit designs a lower input voltage. But several circuit require high voltage. In this case, charge pump solves a problem with small area and cheap cost. Many Conventional charge pumps are connected parallel each capacitor. So many capacitors are required for high voltage. However proposed charge pump can supply high voltage easily and design less capacitors. Proposed circuit are comprised 2-stages with 5 capacitors. To compare with conventional charge pumps, proposed circuit design same component units. SeoungMin Lee, Choong Keun Lee, Hongil Yoon |
TENCON | 3 |
| 2018 | Bit-line Sense Amplifier Using PMOS Charge Transfer Pre-amplifier for Low-Voltage DRAMabstractA bit-line sense amplifier using PMOS charge transfer pre-sensing (CTPS) circuit is proposed. The CTPS circuit using charge sharing operation between lines which have different capacitance is useful for increasing small bit-line voltage differences. The PMOS latch transistors of CTPS circuit are used for pre-sensing operation and pull-up latching operation. Because the PMOS latch transistors are used mutually, area overhead of the CTPS is minimized. The PMOS CTPS circuit increases the small bit-line voltage difference and the output voltage difference of the CTPS circuit becomes the input voltage difference of the latch sense amplifier. Using the PMOS CTPS circuit, the speed of read operation and the reliability of the read operation is improved. The performance of the proposed scheme is verified by the simulation using a 45 nm process. Choong Keun Lee, Taegun Yim, Hongil Yoon |
TENCON | 3 |
| 2018 | Low-Energy Consumption Write Circuit using Comparing Operation in STT-MRAMabstractIn this paper, a method is proposed to reduce the write energy consumption in order to reduce the energy consumption of the circuit composed of spin-torque transfer magnetic RAM(STT-MRAM). Of the total energy consumption of the circuit, the write energy consumption occupies a large proportion. Omission of unnecessary write operations increases efficiency in terms of energy consumption. Compare the data to be written with the data to be written before proceeding with the write operation of the circuit to omit the unnecessary write operation and selectively execute the necessary write operation. As the number of MTJ increases, the gain in side of energy can be increased, so the simulation had proceeded by increasing the number of MTJ to 512. Simulation results show that the worst case energy consumption consumes 16.5% more energy than the conventional one, but it can reduce energy consumption by up to 90.6% at best. Reliability also increases through the process of comparing write data (WD) with data stored in MTJ. Taegun Yim, Hongil Yoon |
TENCON | 3 |
| 2018 | A Low-Voltage Charge Pump with High Pumping EfficiencyabstractA low voltage charge pump with high pumping efficiency is proposed in this paper. To get high pumping efficiency, the proposed circuit uses 2 branches cross coupled structure. To eliminate the body effect and threshold voltage drops, PMOS transistors are used as Charge Transfer Switch (CTS). The inverters below and above the CTS PMOS for controlling the gate of CTS PMOS properly. Undesired charge transfer flows higher voltage node voltage to lower voltage node is removed as the NMOS transistor and PMOS transistor below and above the CTS PMOS turn off during the clock transition. Furthermore, the time when those NMOS transistor and PMOS transistor are both off is short and due to the cross coupled wiring, the fast gate control would be able to made by present and previous stage. So that the cross coupled CTS charge pump has higher pumping efficiency than conventional circuits. The simulation results show that the proposed circuit's performance is better than other circuits. Taegun Yim, Choong Keun Lee, Hongil Yoon |
TENCON | 4 |
| 2016 | Revisiting virtual L1 caches: A practical design using dynamic synonym remappingabstractVirtual caches have potentially lower access latency and energy consumption than physical caches because they do not consult the TLB prior to cache access. However, they have not been popular in commercial designs. The crux of the problem is the possibility of synonyms. This paper makes several empirical observations about the temporal characteristics of synonyms, especially in caches of sizes that are typical of L1 caches. By leveraging these observations, the paper proposes a practical design of an L1 virtual cache that (1) dynamically decides a unique virtual page number for all the synonymous virtual pages that map to the same physical page and (2) uses this unique page number to place and look up data in the virtual caches. Accesses to this unique page number proceed without any intervention. Accesses to other synonymous pages are dynamically detected, and remapped to the corresponding unique virtual page number to correctly access data in the cache. Such remapping operations are rare, due to the temporal properties of synonyms, allowing a Virtual Cache with Dynamic Synonym Remapping (VC-DSR) to achieve most of the benefits of virtual caches but without software involvement. Experimental results based on real world applications show that VC-DSR can achieve about 92% of the dynamic energy savings for TLB lookups, and 99.4% of the latency benefits of ideal (but impractical) virtual caches for the configurations considered. Hongil Yoon, Gurindar S. Sohi |
HPCA | 1 |
| 2007 | High Speed, Minimal Area, and Low Power SEC Code for DRAMs with Large I/O Data WidthsabstractMany ECCs have been proposed to enhance the reliability of DRAMs, but most of them lack in the aspects of practical feasibility. We prioritize on how well and efficiently code could be actually implemented and propose high speed, minimal area, and low power SEC code for DRAMs with large I/O data widths. The proposed code minimizes the column weight, row weight and total weight of the H-matrix. Consequently, the area overhead and power consumption of the check bit generator are reduced by 17.7% at the most. The propagation delay is also decreased by reducing the level of XOR trees. Moreover, maximal power reduction is possible as the optimal H-matrix variants can be formulated to specifically tailor for various system applications with different spatial and temporal data correlations. Sang-uhn Cha, Hongil Yoon |
ISCAS | 2 |
| 2004 | An In-Order SMT Architecture with Static Resource Partitioning for Consumer Applications
Byung In Moon, Hongil Yoon, Ilgu Yun, Sungho Kang 0001 |
PDCAT | 2 |