VLDB 2026 Research / reviewers in the wild / expert
Woo-Chan Park
dblp:05/6890
· DBLP profile ↗
20ranked-venue papers
5as first author
3since 2021 · last 2023
0000-0002-9249-2887ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | An Architecture and Implementation of Real-Time Sound Propagation Hardware for Mobile DevicesabstractThis paper presents a high-performance and low-power hardware architecture for real-time sound rendering on mobile devices. Traditional sound rendering algorithms require high-performance CPUs or GPUs for processing because of its high computational complexities to realize ultra-realistic 3D audio. Thus, it has been hard to achieve real-time rates on low-power mobile devices. To overcome this limitation, we propose a hardware architecture that adopts hardware-friendly sound-propagation-path calculation algorithms. We verified the function and performance of our architecture through its implementation on an FPGA board. According to ASIC evaluation with the 8-nm process technology, it achieves high performance with 120 FPS, low power consumption with 50 mW, and a small silicon area with 0.31 mm2, allowing real-time sound rendering on mobile devices. Eunjae Kim, Sukwon Choi, Jae-Ho Nah, Woonam Jung, Tae-Hyeong Lee, Yeon-Kug Moon, Woo-Chan Park |
SIGGRAPH Asia | 8 |
| 2022 | Effective Algorithm to Control Depth Level for Performance Improvement of Sound TracingabstractSound tracing, a 3D sound rendering technology based on ray tracing, is a very costly method for calculating sound propagation. To reduce its expense, we propose an algorithm for adjusting the depth based on frame coherence and spatial characteristics. The results of the experiment indicate that when the sound source and listener were indoors, the reflection path loss rate was 3%, the diffraction path loss rate was 15.4%, and the total frame rate increased by 6.25%. When the listener was outdoors and the sound source was indoors, the reflection path and diffraction path loss rate were 0%, and the total frame rate was increased by 33.33 compared to the conventional method. Thus, the proposed algorithm can improve rendering performance while minimizing path loss rate. Eunjae Kim, Juwon Yun, Woo-Nam Chung, Jae-Ho Nah, Youngsik Kim, Cheoung Ghil Kim, Woo-Chan Park |
J. Web Eng. | 7 |
| 2021 | Lossless Compression Algorithm and Architecture for Reduced Memory Bandwidth Requirement with Improved Prediction Based on the Multiple DPCM Golomb-Rice AlgorithmabstractIn a computing environment, higher resolutions generally require more memory bandwidth, which inevitably leads to the consumption more power. This may become critical for the overall performance of mobile devices and graphic processor units with increased amounts of memory access and memory bandwidth. This paper proposes a lossless compression algorithm with a multiple differential pulse-code modulation variable sign code Golomb-Rice to reduce the memory bandwidth requirement. The efficiency of the proposed multiple differential pulse-code modulation is enhanced by selecting the optimal differential pulse code modulation mode. The experimental results show compression ratio of 1.99 for high-efficiency video coding image sequences and that the proposed lossless compression hardware can reduce the bus bandwidth requirement. Imjae Hwang, Juwon Yun, Woo-Nam Chung, Jaeshin Lee, Cheong-Ghil Kim, Youngsik Kim, Woo-Chan Park |
J. Web Eng. | 7 |
| 2017 | Real-time sound propagation hardware accelerator for immersive virtual reality 3D audioabstractIn order to support the realistic virtual reality environment, it is necessary to reproduce virtual space and virtual acoustic space. The multi-channel audio system or a 3D sound technology using head related transfer function (HRTF) is used to reproduce the virtual acoustic space [Vorländer 2010]. Dukki Hong, Tae-Hyoung Lee, Yejong Joo, Woo-Chan Park |
I3D | 4 |
| 2016 | Geometry transition method to improve ray-tracing precision
Dong-Seok Kim, Jae-Ho Nah, Woo-Chan Park |
Multim. Tools Appl. | 3 |
| 2015 | HART: A Hybrid Architecture for Ray Tracing Animated ScenesabstractWe present a hybrid architecture, inspired by asynchronous BVH construction [1], for ray tracing animated scenes. Our hybrid architecture utilizes heterogeneous hardware resources: dedicated ray-tracing hardware for BVH updates and ray traversal and a CPU for BVH reconstruction. We also present a traversal scheme using a primitive's axis-aligned bounding box (PrimAABB). This scheme reduces ray-primitive intersection tests by reusing existing BVH traversal units and the primAABB data for tree updates; it enables the use of shallow trees to reduce tree build times, tree sizes, and bus bandwidth requirements. Furthermore, we present a cache scheme that exploits consecutive memory access by reusing data in an L1 cache block. We perform cycle-accurate simulations to verify our architecture, and the simulation results indicate that the proposed architecture can achieve real-time Whitted ray tracing animated scenes at 1,920 × 1,200 resolution. This result comes from our high-performance hardware architecture and minimized resource requirements for tree updates. Jae-Ho Nah, Jin-Woo Kim 0004, Won-Jong Lee, Jeong-Soo Park 0004, Seokyoon Jung, Woo-Chan Park, Dinesh Manocha, Tack-Don Han |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2014 | RayChip®: Real-time ray-tracing chip for embedded applicationsabstractThis article consists of a collection of slides from the author's conference presentation on the special features, system design and architectures, processing capabilities, and targeted markets for SiliconArts' RayChip, the world's first commercialized chip targeted to realize real-time ray tracing for embedded applications such as TV, media box and game console. Woo-Chan Park, Hee-Jin Shin, Byoungok Lee, Hyung-Min Yoon, Tack-Don Han |
Hot Chips Symposium | 1 |
| 2014 | Effective traversal algorithms and hardware architecture for pyramidal inverse displacement mapping
Hyuck-Joo Kwon, Jae-Ho Nah, Dinesh Manocha, Woo-Chan Park |
Comput. Graph. | 4 |
| 2014 | RayCore: A Ray-Tracing Hardware Architecture for Mobile DevicesabstractWe present RayCore, a mobile ray-tracing hardware architecture. RayCore facilitates high-quality rendering effects, such as reflection, refraction, and shadows, on mobile devices by performing real-time Whitted ray tracing. RayCore consists of two major components: ray-tracing units (RTUs) based on a unified traversal and intersection pipeline and a tree-building unit (TBU) for dynamic scenes. The overall RayCore architecture offers considerable benefits in terms of die area, memory access, and power consumption. We have evaluated our architecture based on FPGA and ASIC evaluations and demonstrate its performance on different benchmarks. According to the results, our architecture demonstrates high performance per unit area and unit energy, making it highly suitable for use in mobile devices. Jae-Ho Nah, Hyuck-Joo Kwon, Dong-Seok Kim, Cheol-Ho Jeong, Jin-Hong Park, Tack-Don Han, Dinesh Manocha, Woo-Chan Park |
ACM Trans. Graph. | 8 |
| 2013 | gkDtree: A group-based parallel update kd-tree for interactive ray tracing
Yoon-Sig Kang, Jae-Ho Nah, Woo-Chan Park, Sung-Bong Yang |
J. Syst. Archit. | 3 |
| 2011 | A Lossless Color Image Compression Architecture Using a Parallel Golomb-Rice Hardware CODECabstractIn this paper, a high performance lossless color image compression and decompression architecture to reduce both memory requirement and bandwidth is proposed. The proposed architecture consists of differential-differential pulse coded modulation (DDPCM) and Golomb-Rice coding. The original image frame is organized as m by n sub-window arrays, to which DDPCM is applied to produce one seed and m × n - 1 pieces of differential data. Then the differential data are encoded using the Golomb-Rice algorithm to produce losslessly compressed data. According to the experimental results on benchmark images, the proposed architecture can guarantee high enough compression rate and throughput to perform real-time lossless CODEC operations with a reasonable hardware area. Hong-Sik Kim, Joohong Lee, Sungho Kang 0001, Woo-Chan Park |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | T&I engine: traversal and intersection engine for hardware accelerated ray tracingabstractRay tracing naturally supports high-quality global illumination effects, but it is computationally costly. Traversal and intersection operations dominate the computation of ray tracing. To accelerate these two operations, we propose a hardware architecture integrating three novel approaches. First, we present an ordered depth-first layout and a traversal architecture using this layout to reduce the required memory bandwidth. Second, we propose a three-phase ray-triangle intersection architecture that takes advantage of early exit. Third, we propose a latency hiding architecture defined as the ray accumulation unit. Cycle-accurate simulation results indicate our architecture can achieve interactive distributed ray tracing. Jae-Ho Nah, Jeong-Soo Park 0004, Chanmin Park, Jin-Woo Kim 0004, Yun-Hye Jung, Woo-Chan Park, Tack-Don Han |
ACM Trans. Graph. | 6 |
| 2007 | A consistency-free memory architecture for sort-last parallel rendering processors
Woo-Chan Park, Cheong-Ghil Kim, Duk-Ki Yoon, Kil-Whan Lee, Il-San Kim, Tack-Don Han |
J. Syst. Archit. | 1 |
| 2006 | An Effective Visibility Culling Method Based on Cache BlockabstractAs the complexity of 3D scenes is on the increase, the search for an effective visibility culling method has become one of the most important issues to be addressed in the design of 3D rendering processors. Here, we propose a new rasterization pipeline with visibility culling; the proposed architecture performs the visibility culling at an early stage of the rasterization pipeline (especially at the traversal stage) by retrieving data in a pixel cache without any significant hardware logics such as the hierarchical z-buffer. If the data to be retrieved does not exist in the pixel cache, the proposed architecture performs a prefetch operation in order to reduce the miss penalty of the pixel cache. That is, the cache miss penalty can be reduced as the transfer of a missed cache block from the frame memory into the pixel cache can be handled simultaneously with the rasterization pipeline executions. Simulation results show that the proposed architecture can achieve a performance gain of about 32% compared with the conventional pretexturing architecture and about 7% compared to the hierarchical z-buffer visibility scheme. Moon-Hee Choi, Woo-Chan Park, Francis Neelamkavil, Tack-Don Han, Shin-Dug Kim |
IEEE Trans. Computers | 2 |
| 2004 | A Bandwidth Reduction Scheme for 3D Texture-Based Volume Rendering on Commodity Graphics Hardware
Won-Jong Lee, Woo-Chan Park, Tack-Don Han, Sung-Bong Yang, Francis Neelamkavil |
ICCSA (2) | 2 |
| 2004 | A Cost-Effective Pipelined Divider with a Small Lookup TableabstractCurrent pipelinable dividers require very large lookup tables. We propose a cost-effective pipelinable divider that uses a modified Taylor-series expansion and has a smaller lookup table than other pipelinable dividers. The proposed divider requires about 27 percent less area than the pipelinable divider based on normal Taylor-series expansion in single precision. Jong-Chul Jeong, Woo-Chan Park, Woong Jeong, Tack-Don Han |
IEEE Trans. Computers | 2 |
| 2003 | An Effective Pixel Rasterization Pipeline Architecture for 3D Rendering ProcessorsabstractAs a 3D scene becomes increasingly complex and the screen resolution increases, the design of an effective memory architecture is one of the most important issues for 3D rendering processors. We propose a pixel rasterization architecture that performs the depth test twice, before and after texture mapping. The proposed architecture eliminates memory bandwidth waste due to fetching unnecessary obscured texture data by performing the depth test before texture mapping. It also reduces the miss penalties of the pixel cache by using a prefetch scheme-that is, a frame memory access, due to a cache miss at the first depth test, is done simultaneously with texture mapping. We have built a trace-driven simulator for the proposed architecture. To validate the proposed architecture, the results of various simulations are provided. The proposed pixel rasterization architecture achieves memory bandwidth effectiveness and reduces power consumption while producing high-performance gains. Woo-Chan Park, Kil-Whan Lee, Il-San Kim, Tack-Don Han, Sung-Bong Yang |
IEEE Trans. Computers | 1 |
| 2002 | A Mid-Texturing Pixel Rasterization Pipeline Architecture for 3D Rendering ProcessorsabstractAs a 3D scene becomes increasingly complex and the screen resolution increases, the design of effective memory architecture is one of the most important issues for 3D rendering processors. We propose a pixel rasterization architecture, which performs a depth test operation twice, before and after texture mapping. The proposed architecture eliminates memory bandwidth waste caused by fetching unnecessary obscured texture data, by performing the depth test before texture mapping. The proposed architecture reduces the miss penalties of the pixel cache by using a pre-fetch scheme - that is, a frame memory access, due to a cache miss at the first depth test, is done simultaneously with texture mapping. The proposed pixel rasterization architecture achieves memory bandwidth effectiveness and reduces power consumption, producing high-performance gains. Woo-Chan Park, Kil-Whan Lee, Il-San Kim, Tack-Don Han, Sung-Bong Yang |
ASAP | 1 |
| 2001 | In-Order Issue Out-of-Order Execution Floating-Point Coprocessor for CalmRISC32abstractThe CalmRISC32 FPU (Floating-Point Unit) is a RISC coprocessor for embedded system applications. It supports IEEE-754 standard single precision floating-point addition, floating-point subtraction, floating-point multiplication, floating-point division, format conversion, comparison, rounding, load, store, etc. It also supports four rounding modes, and precise exception. It can execute and complete instructions out of order, if constraints such as data dependency, resource conflict, and exception prediction are resolved. Standard cell-base design techniques were used to reduce design time and expense. The first prototype operated at approximately 70 MHz with the worst-case delay in gate level simulation. Cheol-Ho Jeong, Woo-Chan Park, Tack-Don Han |
IEEE Symposium on Computer Arithmetic | 2 |
| 1999 | A floating point multiplier performing IEEE rounding and addition in parallel
Woo-Chan Park, Tack-Don Han, Shin-Dug Kim, Sung-Bong Yang |
J. Syst. Archit. | 1 |