EDBT 2026 Demo / reviewers in the wild / expert
Tae Sung Kim
dblp:18/5905
· DBLP profile ↗
11ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-3086-0055ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Live Demonstration: DVS-CIS Sensor Fusion System for Real-Time DNN-Based Object DetectionabstractThis demonstration presents a high-speed, energy-efficient sensor fusion system integrating CMOS Image Sensors (CIS) and Dynamic Vision Sensors (DVS) for advanced image recognition. Using CIS for high-res imaging and DVS for rapid event-driven capture, the FPGA-implemented architecture with an NPU running YOLOv3-Tiny achieves 18 ms inference latency with a minimal 2.78% mAP loss. Selective NPU activation based on DVS-detected regions yielded 31.5% power savings, while a custom receiver module efficiently fused DVS (13,900 fps) and CIS (60 fps) data. The system uses a power of 6.977 W on a Xilinx Zynq+ ZCU106 board. Mincheol Cha, Keehyuk Lee, Bobaro Chang, Soosung Kim 0003, Xuan Truong Nguyen, Tae Sung Kim, Hyunsurk Ryu |
ISCAS | 7 |
| 2025 | A DVS-CIS Sensor Data Receiver on FPGA with a 10 Gbps MIPI ControllerabstractFusing a dynamic vision sensor (DVS) and a CMOS image sensor (CIS) is promising in real-time vision applications. However, unlike common CIS, DVS typically come with a custom data format due to their naturally sparse data, which becomes a challenge to fuse DVS and CIS data streams on a general-purpose CPU. To address this problem, this work proposes a DVS-CIS sensor stream receiver on FPGA. The proposed receiver incorporates a cost-effective address decoder and an inline transpose to receive and store a DVS stream on DRAM effectively. At a system level, a host PC can stream the DVS-CIS stream from FPGA via PCIe and display streams on a monitor. Experimental results demonstrate that our architecture can decode up to 13,900fps of DVS frames without incurring any frame drops while concurrently streaming frames at 60fps from a CIS. The design only uses 135 BRAM, 38 DSPs, 69489 LUTs, and 86626 FFs on a Xilinx Zynq+ ZCU106 FPGA board and consumes a power of 6.977 W. Mincheol Cha, Keehyuk Lee, Bobaro Chang, Soosung Kim 0003, Xuan Truong Nguyen, Tae Sung Kim, Hyunsurk Ryu |
ISCAS | 7 |
| 2025 | An Energy-Efficient Daily Surveillance System with DVS-CIS Sensor Fusion and Event-based NPU TriggeringabstractThis study presents a daily surveillance system based on dynamic vision sensors (DVS) and CMOS image sensors (CIS) to enable real-time image recognition with low energy consumption. In such a system, a neural processing unit (NPU) - which executes a DNN model to detect objects on a given CIS image - may consume a lot of energy when always-on. To address the problem, this work introduces a system with a DVS-based region of interest (ROI) detector and an event-based NPU trigger for energy savings. Based on DVS, the ROI detector effectively recognizes scene changes in dynamic environments, e.g., low-light scenes at midnight, which serves as a trigger to invoke the NPU for object detection. Our system prototype was built on a host PC and two Xilinx Zynq+ ZCU106 FPGA boards, one for the DVS-CIS receiver and the other for our NPU. The experimental results demonstrated that Over a 24-hour testing period, our system achieved a 31.5% reduction in energy usage. Operating a YOLOv3-Tiny object detector at 200 MHz, our NPU achieves a latency of just 18 ms, enabling seamless real-time monitoring capabilities. Mincheol Cha, Keehyuk Lee, Bobaro Chang, Soosung Kim 0003, Daniel Moon, Tae Sung Kim, Hyunsurk Ryu, Xuan Truong Nguyen |
ISCAS | 7 |
| 2024 | Toward Real-World Multi-View Object Classification: Dataset, Benchmark, and AnalysisabstractAggregating information from multiple views is essential to accurately identifying similar objects. Nevertheless, existing datasets have limitations that hinder the development of practical multi-view object classification methods for real-world scenarios. The limitations include synthetic and coarse-grained objects in the datasets and the absence of a validation split to enable standard hyperparameter tuning. This paper proposes a new dataset, MVP-N (Multi-View, Retail Products, Label Noise), which contains 16k real captured views and 9k multi-view sets collected from 44 retail products. In MVP-N, each view is annotated with a human-perceived information quantity (HPIQ) for analyzing how views are utilized in information aggregation. Moreover, the fine-grained categorization of objects provides the inter-class view similarity and intra-class view variance, enabling the research on learning from noisy labels of the multi-view images. Finally, a new soft label scheme, HS-HPIQ, is proposed considering the hidden stratification phenomenon in the multi-view images and achieves superior performance. To assess the effectiveness of MVP-N and the proposed HS-HPIQ, this study overviews 50 recent multi-view-based methods regarding their practicality in real-world scenarios. Six feature aggregation methods and twelve soft label methods are benchmarked on MVP-N with a deep analysis. The dataset and code are publicly available at https://github.com/SMNUResearch/MVP-N. Ren Wang 0012, Tae Sung Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | MVP-N: A Dataset and Benchmark for Real-World Multi-View Object ClassificationabstractCombining information from multiple views is essential for discriminating similar objects. However, existing datasets for multi-view object classification have several limitations, such as synthetic and coarse-grained objects, no validation split for hyperparameter tuning, and a lack of view-level information quantity annotations for analyzing multi-view-based methods. To address this issue, this study proposes a new dataset, MVP-N, which contains 44 retail products, 16k real captured views with human-perceived information quantity annotations, and 9k multi-view sets. The fine-grained categorization of objects naturally generates multi-view label noise owing to the inter-class view similarity, allowing the study of learning from noisy labels in the multi-view case. Moreover, this study benchmarks four multi-view-based feature aggregation methods and twelve soft label methods on MVP-N. Experimental results show that MVP-N will be a valuable resource for facilitating the development of real-world multi-view object classification methods. The dataset and code are publicly available at https://github.com/SMNUResearch/MVP-N. Tae Sung Kim |
NeurIPS | 3 |
| 2020 | Fast Hardware-Based IME With an Idle Cycle and Computational Redundancy ReductionabstractExtensive efforts have been made to design hardware-based integer motion estimation (IME) that is much faster than software-based IME but suffers from the degradation in the coding efficiency. This is because the strategy for previous efforts was a simple algorithmic modification of the fast IME to facilitate the given hardware design at the expense of coding efficiency. This paper proposes a novel hardware design of the IME that not only offers real-time processing capability but also provides a flexible tradeoff between computational complexity and coding efficiency. First, a prediction unit (PU) loop unrolling scheme is proposed to solve the pipeline stall problem owing to the nature of fast IME algorithms such as the test zone search (TZS). It reduces idle cycles by 89.24%. Next, to further reduce the computational complexity of the TZS algorithm, a computational redundancy among PUs within a coding unit is reduced through a search step synchronization and search point sharing scheme. Thus, the computational complexity is reduced by 72.25%. The proposed schemes eliminate the inefficiency of hardware design; thus, they do not suffer from serious degradation in the coding efficiency. Consequently, the proposed hardware-based IME processes 7680 × 4320 videos at 30 frames per second while increasing the Bjøntegaard delta bitrate by only 0.90% on average. The hardware design is synthesized using a 65 nm general purpose CMOS technology, and its gate count is 268.5K at an operating clock frequency of 500 MHz. Tae Sung Kim, Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Fast Integer Motion Estimation With Bottom-Up Motion Vector Prediction for an HEVC EncoderabstractAlthough advanced motion vector prediction (AMVP) modes based on motion estimation (ME) are selected significantly less due to the merge mode newly adopted in the high-efficiency video coding (HEVC), integer ME (IME) still occupies a large amount of computation in HEVC because the HEVC supports a highly flexible block partitioning structure. Introduction of the merge mode in HEVC substantially affects the optimal search algorithm of IME. Therefore, this gives a chance to reduce the computational complexity of ME with marginal drop of the compression efficiency. This paper proposes a new IME algorithm that significantly reduces its search ranges. Computational complexity of the proposed algorithm is reduced without serious degradation in coding efficiency by obtaining additional accurate motion vector predictors (MVPs) in a bottom-up order and searching narrow regions around multiple MVPs. These bottom-up MVPs are obtained from the prediction units (PUs) in the coding units (CUs) in the hierarchically lower level, which share the pixel area either completely or partially with the current PU. Although the encoder should perform the proposed IME in a bottom-up order from the smallest CUs to the larger CUs, several other processes, such as fractional ME and merges mode are performed in a top-down order to exploit several fast algorithms already adopted in the HEVC test model encoder. To keep compatibility with HEVC standard, the bottom-up MVPs are utilized for only IME. The bitstream is generated using standard AMVP. According to the simulation results, it is confirmed that the proposed algorithm improves the accuracy of MVPs, which leads to the reduction of IME computational complexity up to 81.95% on average with the Bjøntegaard delta bitrate of 0.39%. Tae Sung Kim, Chae-Eun Rhee, Soo-Ik Chae |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Merge Mode Estimation for a Hardware-Based HEVC EncoderabstractHigh Efficiency Video Coding (HEVC) is a video coding standard that offers higher performance than previous video coding standards such as H.264/AVC. Merge mode is one of the new tools adopted in HEVC to improve the inter-frame coding efficiency. Merge mode saves the bits for the motion vector (MV) by sharing the MV with neighboring blocks. Merge mode estimation (MME) is the process of finding a merge mode candidate, which requires extensive computations and memory accesses due to the associated motion compensation. Although MME is very similar to motion estimation (ME) in many ways, previous research on ME cannot be directly applied to solve many difficulties in designing MME hardware. In this paper, the characteristics of and the computational complexity involved in MME are discussed. To improve the throughput of the MME hardware, partially increased parallelism is efficiently exploited. Furthermore, the M-of-N-pixel combination and flexible memory access schemes are proposed to maximize the scalability to support various block sizes of HEVC and to reduce the time for fetching reference data. The proposed schemes are applied to the MME hardware design in this paper. The proposed hardware can process 56074 of 64 × 64 coding tree units per second with a clock frequency of 366 MHz, and its gate count is 585.4k with 2 kB of dual-port static RAM. Tae Sung Kim, Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | An H.264 High-Profile Intra-Prediction with Adaptive Selection Between the Parallel and Pipelined Executions of Prediction ModesabstractA high-profile H.264 intra-frame encoder is suitable for low-cost and low-power applications and capable of providing enhanced compression efficiency. The high-profile is targeting the high-resolution videos. Thus, the encoding speed should be faster than or comparable to the baseline-profile. In previous work related to a hardware-based baseline-profile intra-frame encoder, a speed-up is achieved by the early termination of the intra modes and by an increase in the rate of hardware utilization only under one of the serialized and parallel schedules. This paper proposes a novel pipeline schedule for a hardware-based high-profile intra-prediction scheme in which the 8 × 8 prediction is performed in Stage 1 and 4 × 4, 16 ×16 and chroma predictions are executed during Stage 2. The processing time of Stage 2 is efficiently accelerated based on the result of the 8 × 8 prediction in Stage 1. According to the distribution of each mode, the schedule is adaptively selected between parallel and pipeline schedules. To increase the hardware utilization of the 8 × 8 prediction, the order of prediction modes and the inverse vertical transform is adaptively adjusted. In addition, early termination of the prediction modes is employed for a fast 8 × 8 prediction. The proposed 8 × 8 intra-prediction is implemented and verified as an entire intra-frame encoder. Experimental results show that the average number of cycles necessary to process one MB for videos with resolutions of 1920 ×1080 and 3840 × 2160 are only 269 and 253 cycles, respectively. Compared to JM13.2, the bitrate is increased by 1.13% on average with a small PSNR degradation of 0.06 dB. The difference in the rate-distortion performance between the proposed high-profile intra-prediction scheme and JM 13.2 is not significant, whereas the achieved speed-up due to the proposed schemes is considerable compared to the conventional hardware-based intra-prediction encoders. Chae-Eun Rhee, Tae Sung Kim |
IEEE Trans. Multim. | 2 |
| 2006 | An Interfererence Minimization and Predictive Location Based Relaying for Ad-Hoc Cellular SystemsabstractWe propose a predictive location based relaying for hybrid cellular and ad-hoc systems that can provide interference minimization to uplink transmission. Location Information Server (LIS) managed the mobile's routing information with centralized scheme. To maintain the routing correctness, mobile must send location information within a relatively short time period. However the tradeoff as the location-update period decreases is the increase in the overall cellular overhead. To overcome this problem, a predictive location-based routing scheme is proposed. Simulation results show that interference is greatly improved with very low cellular overheads. Tae Sung Kim, Young Yong Kim |
INFOCOM | 1 |
| 2006 | Adaptive Admission Control Algorithm for Multiuser OFDMA Wireless NetworksabstractWe propose a new adaptive admission control scheme for multiuser OFDMA systems to allocate wireless resource with low complexity. In this adaptive admission control scheme, the number of users is estimated adaptively by using the required data rate, bit error rate (BER) and average channel gain of each user at frame by frame. The resource allocation complexity becomes lower since the proposed method estimates the approximate number of users to be supported and reduce the repetition of actual resource allocation. Simulation result shows that the proposed method has high accuracy when ACG (Adaptive Craving Greedy) is used as the resource allocation scheme. Kyung-Ho Sohn, Jongkyung Kim, Tae Sung Kim, Young Yong Kim, Jongsoo Seo |
INFOCOM | 3 |