EDBT 2026 Demo / reviewers in the wild / expert
Hyun Woo Oh
dblp:40/8862
· DBLP profile ↗
7ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0001-9608-4453ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multimodal AI Acceleration with Dynamic Pruning and Run-Time ConfigurationabstractThe computational diversity of multimodal AI workloads-spanning vision transformers (ViTs), graph neural networks (GNNs), CNNs, and transformer-based NLP—poses a fundamental challenge to embedded acceleration platforms. We propose a fully integrated FPGA-based acceleration framework that addresses this heterogeneity via compile-time and run-time configurability. Our system introduces a reconfigurable processing unit (RPU) capable of executing dense and sparse matrix operations (DDMM, SpMM, SDDMM), a scalable top-k pruning engine for ViTs, and a domain-specific compiler for hardware-software co-design. The architecture supports real-time configuration without reloading bitstreams, enabling unified deployment across tasks. Implementations on Xilinx U50 and ZCU104 demonstrate up to 22.57× and 6.86× latency reductions versus RTX 4090 and Jetson Orin Nano, respectively, validating the design's efficiency for real-time, resource-limited environments. Hyun Woo Oh, Hanning Chen, Sanggeon Yun, Yang Ni 0001, Behnam Khaleghi, Fei Wen 0003, Mohsen Imani |
FCCM | 1 |
| 2024 | A Compact Real-Time Thermal Imaging System Based on Heterogeneous System-on-ChipabstractThis paper presents a real-time embedded thermal imaging system architecture for compact, energy-efficient, high-quality imaging utilizing heterogeneous system-on-chip (SoC) and uncooled infrared focal plane arrays (IRFPAs). Unlike previous systems that organized separate devices for complex image processing, our system provides integrated image processing support for robust sensor-to-surveillance. The image processing organizes two algorithm stacks: a non-uniformity correction stack to mitigate the distinctive noise vulnerabilities of uncooled IRFPAs, and an image enhancement stack including contrast enhancement and temporal noise filters. We optimized these algorithms for domain-specific factors, including asymmetric multiprocessing (AMP), cache organization, single instruction multiple data (SIMD) instructions, and very long instruction word (VLIW) architectures. The implementation on the TI TDA3x SoC demonstrates that our system can process 640×480, 60 frames per second (FPS) videos at a peak core load of 57.5% while consuming power less than 2.2 W for the entire system, denoting the possibility of processing the 1280×1024, 30 FPS videos from the cutting-edge uncooled IRFPAs. Additionally, our system improves power efficiency by 9.42% and 9.96% at 30 and 60 FPS, respectively, compared to the state-of-the-art when executing similar image processing algorithms. Hyun Woo Oh, Cheol-Ho Choi, Jeongwoo Cha, Hyunmin Choi, Jungho Shin, Joonhwan Han |
RTCSA | 1 |
| 2023 | Disparity Refinement Processor Architecture Utilizing Horizontal and Vertical Characteristics for Stereo Vision SystemsabstractIn embedded stereo vision systems based on semi-global matching, the matching accuracy of the initial disparity map can be degraded because of various factors. To solve this problem, weighted median-based disparity refinement hardware architectures are utilized to improve the matching accuracy. However, for the conventional hardware architectures, there is a trade-off between hardware resource utilization and re-finement performance when they are implemented on a field programmable gate array (FPGA). Therefore, in this paper, we propose a hybrid max-median filter and its hardware architecture to improve the refinement performance and reduce hardware resource utilization. To evaluate the refinement performance, we used two public stereo datasets. When using the various window sizes for KITTI 2012 and 2015 stereo benchmark datasets, the proposed hardware architecture showed better matching accuracy performance compared with the conventional hardware architectures. In terms of the hardware resource utilization, when implemented on an FPGA, the proposed hardware architecture has low requirements for all types of hardware resources. That is, the proposed hardware architecture overcomes the trade-off between hardware resource utilization and refinement performance. Cheol-Ho Choi, Hyun Woo Oh |
DSD | 2 |
| 2023 | An SoC FPGA-based Integrated Real-time Image Processor for Uncooled Infrared Focal Plane ArrayabstractThis paper presents an integrated image processor architecture designed for realtime interfacing and processing of high-resolution thermal video obtained from an uncooled infrared focal plane array (IRFPA) utilizing a modern system-on-chip field-programmable gate array (SoC FPGA). Our processor provides a one-chip solution for incorporating non-uniformity correction (NUC) algorithms and contrast enhancement methods (CEM) to be performed seamlessly. We have employed NUC algorithms that utilize multiple coefficients to ensure robust image quality, free from ghosting effects and blurring. These algorithms include polynomial modeling-based thermal drift compensation (TDC), two-point correction (TPC), and runtime discrete flat field correction (FFC). To address the memory bottlenecks originating from the parallel execution of NUC algorithms in realtime, we designed accelerators and parallel caching modules for pixel-wise algorithms based on a multi-parameter polynomial expression. Furthermore, we designed a specialized accelerator architecture to minimize the interrupted time for runtime FFC. The implementation on the XC7Z020CLG400 SoC FPGA with the QuantumRed VR thermal module demonstrates that our image processing module achieves a throughput of 60 frames per second (FPS) when processing 14-bit 640×480 resolution infrared video acquired from an uncooled IRFPA. Hyun Woo Oh, Cheol-Ho Choi, Jeongwoo Cha, Hyunmin Choi, Joonhwan Han, Jungho Shin |
DSD | 1 |
| 2023 | RF2P: A Lightweight RISC Processor Optimized for Rapid Migration from IEEE-754 to PositabstractThis paper presents a lightweight processor and evaluation platform for migrating from IEEE-754 to posit arithmetic, with an optimized posit arithmetic unit (PAU) supporting existing floating-point instructions. The PAU features a reconfigurable divider architecture for diverse operating conditions and lightweight square root logic. The platform includes a posit-optimized compiler, divider generator, JTAG environment builder, and programmable logic controller. The experimental results demonstrate the successful execution of legacy IEEE-754 code with a small additional workload and up to 60.09 times the performance improvement through hardware acceleration. Additionally, the PAU and divider consume 11.00% and 57.87% fewer LUTs, respectively, compared to the best prior works. Hyun Woo Oh, Seongmo An, Won Sik Jeong |
ISLPED | 1 |
| 2013 | Aggregator system of real-sense acquisition for 4D media authoring based on MPEG-VabstractIn this paper, we propose a system and method for real-sense acquisition; and, more particularly, to a system for providing a real-sense effect by sensing ambient environment information through a sensor at a time when a camera photographs an image, extracting an effective data from the sensed information and creating real-sense effect metadata based on the extracted effective data, and a method for real-sense acquisition using the system. Hyun Woo Oh, Jonghyun Jang, Kwang-Roh Park |
WCNC | 1 |
| 2010 | An explicit disjoint multipath algorithm for Cost efficiency in wireless sensor networksabstractIn this paper, we investigate the use of explicit disjoint for the multipath routing to achieve cost-efficient operation of wireless sensor networks. The focus is on constructing completely disjointed multiple paths between the source of sensing stimulus and the destination of gathered sensor data. We propose simple schemes for multipath construction based on explicit multiple anchor nodes. For the fully disjointed multipath construction, we exploit the logical and physical schemes. The logical method classifies a multipath through the pipeline of the area where there is at least one sensor. The physical approach comprises completely disjoint multiple paths to be comprised at the source level based on the geographic routing, intermediate node side, and anchor node side and destination node side. We propose explicit multiple paths to comprise a disjoint multipath so that we removes the flooding overhead used in the multipath discovery process. The proposed scheme constructs the distributed multiple paths with data transmission through the localized information exchange between neighbor nodes. Our simulation results show the fundamental importance of an explicitly disjointed multipath proposed in this paper. For example, one of our results shows that the proposed scheme significantly reduces the signaling overhead compared to the existing multipath routing methods when there are many nodes in WSN. The simulation result also shows that the proposed scheme, compared to the explicit disjoint multiple paths, is more cost efficient than an existing multipath scheme. Hyun Woo Oh, Jonghyun Jang, Kyeong-Deok Moon, Soochang Park, Euisin Lee, Sang-Ha Kim 0001 |
PIMRC | 1 |