VLDB 2026 Research / reviewers in the wild / expert
Chae-Eun Rhee
dblp:66/4905
· DBLP profile ↗
33ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-7851-1703ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 8 since 2021Systems, architecture and hardware · 15 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BRAM-Free ML-KEM NTT with SRL-Based Reordering Unit
Gijung Kim, Yongmin Park, Chae-Eun Rhee |
ISCAS | 3 |
| 2025 | LC-Mamba: Local and Continuous Mamba with Shifted Windows for Frame InterpolationabstractIn this paper, we propose LC-Mamba, a Mamba-based model that captures fine-grained spatiotemporal information in video frames, addressing limitations in current interpolation methods and enhancing performance. The main contributions are as follows: First, we apply a shifted local window technique to reduce historical decay and enhance local spatial features, allowing multi-scale capture of detailed motion between frames. Second, we introduce a Hilbert curve-based selective state scan to maintain continuity across window boundaries, preserving spatial correlations both within and between windows. Third, we extend the Hilbert curve to enable voxel-level scanning to effectively capture spatiotemporal characteristics between frames. The proposed LC-Mamba achieves competitive results, with a PSNR of 36.53 dB on Vimeo-90k, outperforming prior models by +0.03 dB. The code and models are publicly available at https://github.com/Miinuuu/LCMamba.git Min-Wu Jeong, Chae-Eun Rhee |
CVPR | 2 |
| 2025 | Hybrid Embedding Framework for Memory-Efficient Recommendation SystemsabstractThis study introduces a memory-efficient mixed representation for deep learning recommendation models (DLRM), addressing the embedding table memory bottleneck from growing data scale. By distinguishing between frequently accessed (hot) and infrequently accessed (cold) embeddings, we store hot embeddings in a compact table while representing cold embeddings using a deep hash embedding (DHE) network, significantly reducing memory usage. This hybrid approach performs table lookups for hot embeddings and parallelized computations for cold embeddings, minimizing training time while maintaining accuracy. Experimental results demonstrate that our method outperforms other embedding reduction techniques in memory efficiency, accuracy, and training speed in CPU-GPU hybrid environments. Seung Jin Yang, Chae-Eun Rhee |
DAC | 3 |
| 2025 | Viewpoint-Adaptive Collage-based Streaming for 4K Light Field VideoabstractThis paper presents a viewpoint-adaptive, collage-based light field (LF) video streaming method to optimize data transmission for real-time rendering. Traditional LF video rendering encounters bottlenecks due to the high data volume needed for real-time viewpoint tracking, making storage and GPU memory constraints a significant challenge. To address this, we propose a collage-based approach that transmits only the necessary data based on the user’s current viewpoint, reducing data load by converting 4-dimensional (4D) LF frames into 2-dimensional (2D) collage frames. Additionally, we introduce a collage-based LF stream-switching technique at the group-of-pictures (GOP) level, leveraging I-frames for efficient switching and minimizing latency. Experiments demonstrate the effectiveness of our method, achieving stable transmission and optimized memory usage for 4K LF videos without sacrificing real-time performance. The results confirm that small tile sizes and avoiding partitioning provide the best performance for real-time LF video streaming. Kyungdae Park, Chae-Eun Rhee |
ISCAS | 2 |
| 2025 | See Through the Occlusions: Few-Shot Gaussian Splatting with Layered Amodal Supervision
Gwon-Jung Kim, Du Yeol Lee, Jae Hong Yang, Chae-Eun Rhee |
ACM Multimedia | 4 |
| 2024 | A Resource-Constrained Spatio-Temporal Super Resolution ModelabstractThis paper proposes a resource-constrained spatio-temporal super resolution (SR) model, which effectively enhances both the spatial resolution and frame rate of input videos. Replacing the entire deep learning model for spatio-temporal SR on devices that already have spatial SR capability is a challenging task. This is especially true for compact devices like image sensors that are composed of hardware modules. There is a need to enable spatio-temporal SR with minimal hardware overhead on devices that already have the SR module. The proposed model demonstrates an example of hardware implementation by combining independent hardware-friendly spatial SR and frame interpolation (FI). This configuration allows for seamless support of spatial SR, temporal SR, and spatio-temporal SR functionalities through data flow reconfiguration. Moreover, we propose schemes that leverage the flow estimation module to further reduce the computational burden of spatial SR. The experimental results show that the proposed model achieves competitive quality with state-of-the-art (SOTA) methods, while utilizing very limited computational resources. Da Hyeon Jung, Min-Wu Jeong, Xuan Truong Nguyen, Chae-Eun Rhee |
ISCAS | 4 |
| 2024 | Accelerating Large-Scale DLRM Inference through Dynamic Hot Data RearrangementabstractDeep learning recommendation systems, such as Facebook’s DLRM, enhance user experiences by providing personalized recommendations on social platforms. The use of CXL-based memory extension is gaining attention while existing server DRAM capacity is not sufficient for huge memory requirements. Typically, frequently accessed hot embedding data is stored in local memory, whereas occasionally accessed cold embedding data resides in CXL memory. The distinction between hot and cold data is based on the training results. However, the characteristics of hot and cold embedding vectors can change between training sessions, posing challenges for consistent inference latency with increasing model sizes. This study explores techniques for accelerating large-scale DLRM inference through dynamic hot data rearrangement. The proposed hotness score-based page promotion involves periodic page promotion and demotion based on the changing hotness of embedding data. Additionally, a prioritizing cache prefetch based on hotness improves cache temporal locality, especially in multi-user scenarios. Simulation results demonstrate that proposed approaches is able to enhance DLRM inference speed by up to 8.65% compared to existing techniques. Taehyung Park, Seungjin Yang, Jongmin Seok, Ju-Hyun Kim, Chae-Eun Rhee |
ISCAS | 6 |
| 2024 | Adversarial Mixture Density Network and Uncertainty-Based Joint Learning for 360$^\circ$ Monocular Depth EstimationabstractDue to the increased demand of 360$^{\circ }$images (e.g.virtual reality), estimating 360$^{\circ }$depths via deep learning has drawn attention recently. However, all previous studies share the same fundamental limitation: a lack of data. To address the issue of data insufficiency, self-supervised learning and uncertainty-aware learning based on a mixture density network (MDN) have been actively studied and have achieved great success on various tasks. Unfortunately, under the harsh training environment of 360$^{\circ }$depth estimation tasks (e.g.a large field-of-view, distortion), we observe that the practical difficulties of self-supervised learning and MDN-based uncertainty-aware learning become a critical factor degrading the depth results. In this paper, we propose anadversarial mixture density network (AMDN)anduncertainty-based joint learningto improve the depth qualities by addressing the data insufficiency issue properly. For the AMDN, architectures and objective functions of the MDN are redesigned in an adversarial manner. For uncertainty-based joint supervised and self-supervised learning, the negative effects of the self-supervised learning of 360$^{\circ }$depths are filtered out based on the epistemic uncertainties of the AMDN. Therefore, only the positive effects of self-supervised learning can be realized. Through extensive experiments, we demonstrate that the proposed approaches achieve much more accurate depths as compared with very recent studies for various datasets. Moreover, the proposed approaches also yield sophisticated uncertainties in a single forward path, in which previous studies could not. Ilwi Yun, Chae-Eun Rhee |
IEEE Trans. Multim. | 3 |
| 2023 | EGformer: Equirectangular Geometry-biased Transformer for 360 Depth EstimationabstractEstimating the depths of equirectangular (i.e., 360°) images (EIs) is challenging given the distorted 180° × 360° field-of-view, which is hard to be addressed via convolutional neural network (CNN). Although a transformer with global attention achieves significant improvements over CNN for EI depth estimation task, it is computationally inefficient, which raises the need for transformer with local attention. However, to apply local attention successfully for EIs, a specific strategy, which addresses distorted equirectangular geometry and limited receptive field simultaneously, is required. Prior works have only cared either of them, resulting in unsatisfactory depths occasionally. In this paper, we propose an equirectangular geometry-biased transformer termed EGformer. While limiting the computational cost and the number of network parameters, EGformer enables the extraction of the equirectangular geometry-aware local attention with a large receptive field. To achieve this, we actively utilize the equirectangular geometry as the bias for the local attention instead of struggling to reduce the distortion of EIs. As compared to the most recent EI depth estimation studies, the proposed approach yields the best depth outcomes overall with the lowest computational cost and the fewest parameters, demonstrating the effectiveness of the proposed methods. Ilwi Yun, Chanyong Shin, Hyunku Lee, Chae-Eun Rhee |
ICCV | 5 |
| 2023 | Improving the Compression Efficiency of Displacement using Morton-ordered Micro-Image in Video-based Dynamic Mesh CodingabstractFor immersive media services such as volumetric video, 3-dimensional (3D) data such as point cloud or mesh is commonly used. The size of such data is significantly larger than that of 2-dimensional (2D) video, and its characteristics are also different from those of conventional video. Recently, MPEG standardized video-based point cloud (V-PCC) and is now developing for video-based dynamic mesh coding (V-DMC), which changes connectivity information of vertices over time. Encoding of mesh parts is divided into generation of decimated mesh and compression of displacement for mesh reconstruction. The displacement data is being compressed using a conventional video CODEC in the form of YUV. However, in YUV of displacement, there is no spatial coherence which is commonly observed in general video. Therefore, currently adopted intra-prediction based video encoding in V-DMC, where prediction is performed in unit of blocks within one frame, is not able to exploit the characteristic of displacement data and brings low compression. This paper proposes a conversion technique for displacement data to increase compression efficiency. Displacement YUVs in a group-of-frames (GOF) are changed into a single micro image (MI) to secure better coherence. In order to maximize compression efficiency, pixels in MI are reordered in a Morton-order. This Morton-ordered MI (MoMI) performs decomposition of frequencies, and ensures a smooth area composed of zero pixels as much as possible. Experimental results show that the proposed MoMI gives better spatial coherency and improves the compression efficiency up to 22.64% Yongwook Seo, Gwangcheol Ryu, Chae-Eun Rhee, Hyunmin Jung, Dayun Nam, Hyuncheol Kim, Seongyong Lim |
ISCAS | 3 |
| 2022 | Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-Supervised LearningabstractDue to difficulties in acquiring ground truth depth of equirectangular (360) images, the quality and quantity of equirectangular depth data today is insufficient to represent the various scenes in the world. Therefore, 360 depth estimation studies, which relied solely on supervised learning, are destined to produce unsatisfactory results. Although self-supervised learning methods focusing on equirectangular images (EIs) are introduced, they often have incorrect or non-unique solutions, causing unstable performance. In this paper, we propose 360 monocular depth estimation methods which improve on the areas that limited previous studies. First, we introduce a self-supervised 360 depth learning method that only utilizes gravity-aligned videos, which has the potential to eliminate the needs for depth data during the training procedure. Second, we propose a joint learning scheme realized by combining supervised and self-supervised learning. The weakness of each learning is compensated, thus leading to more accurate depth estimation. Third, we propose a non-local fusion block, which can further retain the global information encoded by vision transformer when reconstructing the depths. With the proposed methods, we successfully apply the transformer to 360 depth estimations, to the best of our knowledge, which has not been tried before. On several benchmarks, our approach achieves significant improvements over previous works and establishes a state of the art. Ilwi Yun, Chae-Eun Rhee |
AAAI | 3 |
| 2022 | Re-ordered Micro Image based High Efficient Residual Coding in Light Field CompressionabstractLight field (LF), a new approach in three-dimensional image processing, has been actively used in various applications in recent years. LF is based on a large amount of data and this always leads to LF compression (LFC) issues. Pseudo-sequence (PS)-based LFC converts a LF into a video sequence and compresses it through a video codec, whereas synthesis-based LFC (SYN-LFC) synthesizes the rest from some of the LF to reduce the number of bits. SYN-LFC is superior to PS-based LFC at low bitrates. However, its competitiveness decreases at high bitrates due to the inefficient compression of residuals. This paper maximizes the advantages of SYN-LFC by increasing the compression efficiency of residuals. To exploit the characteristic of the residual in favor of compression, this paper compresses the residual in the form of a micro image (MI). The conversion of residuals to MI has the effect of gathering similar residuals of each viewpoint, which increases the spatial coherence. However, the conventional MI conversion does not reflect the geometric characteristics of LF at all. To tackle this problem, this paper proposes the re-ordered micro image (RoMI), which is a novel MI conversion that takes advantage of the geometric characteristics of LF, thereby maximizing the spatial coherence and compression efficiency. To compress MI-type residuals, JPEG2000, an image-level codec, is used. It is highly suitable for RoMI with spatial coherence beyond the block level. In the experimental results, the proposed RoMI shows average improvements of 30.29% and 14.05% in the compression efficiency compared to the existing PS-based LFC and SYN-LFC methods, respectively. Hyunmin Jung, Chae-Eun Rhee |
ACM Multimedia | 3 |
| 2022 | A Lightweight and Efficient GPU for NDP Utilizing Data Access Pattern of Image ProcessingabstractAs the demand for image applications with high resolution increases, the importance of the system for image processing is growing. Graphics processing units (GPUs) can increase computational capacity with massive parallelism, but are still subject to limited memory bandwidth. Near-data-processing (NDP) is expected to mitigate the performance and energy overhead caused as a result of data transfer by performing computations on the logic die of 3D-stacked memory. Although prior studies have demonstrated the advantages of NDP, a NDP solution focused on image processing has not yet been developed. This article proposes a GPU-based NDP architecture and well-matched optimization strategies considering both the characteristics of image applications and NDP constraints. First, data allocation to the processing unit is addressed to maintain the data locality and data access pattern. Second, a lightweight yet efficient NDP GPU architecture is proposed. By applying a prefetcher that leverages the pattern-aware data allocation, the number of active warps and the on-chip SRAM size of the NDP are significantly reduced. This enables the NDP constraints to be satisfied and a greater number of processing units to be integrated on a logic die. The evaluation results show that the proposed NDP GPU improves the performance by 1.85× and consumes 82.7 percent energy compared to the baseline NDP GPU. Jungwoo Choi, Boyeal Kim, Ji-Ye Jeon, Eui-Cheol Lim, Chae-Eun Rhee |
IEEE Trans. Computers | 6 |
| 2022 | Data Orchestration for Accelerating GPU-Based Light Field Rendering Aiming at a Wide Virtual SpaceabstractRecently, research on six-degree-of-freedom virtual reality (VR) systems based on a light field (LF) has been actively conducted. The LF-based approach is photorealistic and less prone to errors than the existing three-dimensional (3D) modeling-based approach. On the other hand, because the amount of data is very large for LF-based immersive virtual reality, it is important to manage the data efficiently. This paper makes two major contributions to data management for rendering acceleration in this context. First, the GPUs commonly used for high-speed rendering require optimization of the data transfer process between the host memory and the device memory. This paper proposes a simple but effective way to maximize data reuse by finding the proper size of a transfer unit depending on the target applications and the system capability constraints. Also, view synthesis and data transfer are performed in a pipelined manner to hide the data movement time. Second, for the LF rendering of a wide space, data management in the storage-host-GPU hierarchy is attempted for the first time. As the size of the virtual space increases, the required amount of LF data grows rapidly and exceeds the allocatable host memory capacity. Because the loading of data stored in storage is a very slow process, a system-level LF rendering structure that carefully considers the characteristics of the storage-host-GPU memory hierarchy should be designed. In this paper, progressive LF updating is proposed, and it prevents rendering stalls caused by slow storage data loading. Experimental results show that a 360-degree view with a$9000 \times 2048$resolution can be rendered in an average of 2.57 ms from LF data with a$4096 \times 2048$resolution, achieving performance close to 400 frames per second. Hyunmin Jung, Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | AAGAN: Accuracy-Aware Generative Adversarial Network for Supervised TasksabstractOwing to their great success in unsupervised tasks, generative adversarial networks (GANs) are widely adopted for supervised (conditional) image-generation tasks such as in-painting. Unfortunately, designing an objective function using GANs is not trivial. For supervised tasks, the generator trained only with GAN loss does not yield the corresponding target for the given input because GAN is trained to match the data distribution, not to find the exact answer. Therefore, the loss function of a generator is often formulated by linearly combining supervised loss and GAN loss as similar to multi-task learning, expecting that each loss’s weakness is complemented. Contrary to expectation, both losses cause a conflict in practice due to different optimum of each loss, yielding low objective and subjective image qualities. To address this problem, we empirically investigated the conflict caused by using a conventional GAN with pixel-wise losses; we then propose a novel (relativistic) accuracy-aware discriminator. Based on the proposed discriminator, we developed an accuracy-aware GAN (AAGAN) and proved its optimality under an ideal assumption. We then propose a relativistic accuracy-aware GAN (RAAGAN) by considering practical assumptions. Experimental results on supervised tasks demonstrated that the proposed schemes alleviated the competition between losses and outperformed conventional GANs in terms of both objective and subjective qualities. Ilwi Yun, Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Fast Hardware-Based IME With an Idle Cycle and Computational Redundancy ReductionabstractExtensive efforts have been made to design hardware-based integer motion estimation (IME) that is much faster than software-based IME but suffers from the degradation in the coding efficiency. This is because the strategy for previous efforts was a simple algorithmic modification of the fast IME to facilitate the given hardware design at the expense of coding efficiency. This paper proposes a novel hardware design of the IME that not only offers real-time processing capability but also provides a flexible tradeoff between computational complexity and coding efficiency. First, a prediction unit (PU) loop unrolling scheme is proposed to solve the pipeline stall problem owing to the nature of fast IME algorithms such as the test zone search (TZS). It reduces idle cycles by 89.24%. Next, to further reduce the computational complexity of the TZS algorithm, a computational redundancy among PUs within a coding unit is reduced through a search step synchronization and search point sharing scheme. Thus, the computational complexity is reduced by 72.25%. The proposed schemes eliminate the inefficiency of hardware design; thus, they do not suffer from serious degradation in the coding efficiency. Consequently, the proposed hardware-based IME processes 7680 × 4320 videos at 30 frames per second while increasing the Bjøntegaard delta bitrate by only 0.90% on average. The hardware design is synthesized using a 65 nm general purpose CMOS technology, and its gate count is 268.5K at an operating clock frequency of 500 MHz. Tae Sung Kim, Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Flexibly Connectable Light Field System For Free View ExplorationabstractConventional captured-image-based virtual reality (VR) systems have three degrees of freedom (DoFs), where only rotational user motion is tracked for view rendering. This is a major cause of the reduced sense of reality. To increase user immersion levels akin to the real world, 3-DoF+ VR systems that support not only rotational but also translational view changes have been proposed. The light-field (LF) approach is suitable for this type of 3-DoF+ VR because it renders a view from a free view position by simply combining lights. Many previous systems have limited scalability because they assume a single LF for the acquisition and representation of light. One recent work connects multiple LFs at a physical intersection to increase the scalability. However, these fixed connection points limit the renderable view range and the layout of multiple LFs. Furthermore, in conventional single- or multiple-LF systems, the representable ranges of the view positions are highly dependent on the input field-of-view (FOV) of the camera used. In order to realize a wide view exploration range, the above-mentioned limitations must be overcome. This paper proposes a flexible connection scheme for multiple-LF systems taking advantage of the constant radiance of rays in LF theory. The proposed flexibly connectable LF system is able to widen the range of the renderable view position under reasonable conditions of the camera FOV. A light-field unit (LFU) which uses the proposed flexible connection is implemented. The LFU has a square-shaped structure and is thus easily stackable. This offers the advantage of dramatically expanding the scope of view explorations. The proposed LFU achieves 3-DoF+ VR with good quality as well as high scalability. Its cost-performance outcome is also better compared to those in previous works. Hyunmin Jung, Chae-Eun Rhee |
IEEE Trans. Multim. | 3 |
| 2019 | POSTER: GPU Based Near Data Processing for Image Processing with Pattern Aware Data Allocation and PrefetchingabstractThe following topics are dealt with: parallel processing; multiprocessing systems; graphics processing units; cache storage; storage management; data structures; shared memory systems; learning (artificial intelligence); program compilers; memory architecture. Jungwoo Choi, Boyeal Kim, Ji-Ye Jeon, Eui-Cheol Lim, Chae-Eun Rhee |
PACT | 6 |
| 2019 | Exploration of a PIM Design Configuration for Energy-Efficient Task OffloadingabstractProcessing in memory (PIM) has been proposed to overcome the structural difficulties associated with conventional types of computing architecture and to realize a breakthrough for applications with high data requirements. There have been numerous attempts to utilize the PIM concept. Among them, PIM offloading takes advantage of high memory bandwidths, and various sophisticated conditions have been proposed to offload certain jobs to PIM on an instruction, task or application basis. However, schemes thus far are complicated and difficult to use. This can hinder the widespread use of PIM. Moreover, a performance improvement is not guaranteed in all cases even with these complicated schemes. This paper focuses on the energy efficiency of PIM technology. The potential-based conditions are defined to offload as many tasks as possible. This is justified in that PIM takes absolute advantage over the host in terms of power consumption. A PIM configuration favorable to the selected tasks is then devised. Simulation results show that the combination of potential-based task selection and the associated design configuration can effectively speed up the process while also reducing the energy use in most cases. Byoung-Hak Kim, Eui-Cheol Lim, Chae-Eun Rhee |
ISCAS | 3 |
| 2018 | Encoder-friendly Global-view-depth for Free Viewpoint VideoabstractThe multi-view video plus depth (MVD) format has been used to provide a large amount of views for immersive viewing experience in various applications such as virtual reality systems. Meanwhile, the global view and depth (GVD) format is also considered as an alternative to MVD. It reduces the amount of valid data of frames based on photo consistency in the light field. However, the applicability of GVD as an input of the standard video compression has been rarely studied. In this paper, three schemes are proposed where the spatial and temporal correlation in frames are seriously taken to apply the GVD to the standard video encoder successfully. Processes for GVD construction and video coding are tightly coupled each other. This improves the coding efficiency as well as encoding speed significantly in the GVD-input video encoder. Ji Hun Jang, Chae-Eun Rhee |
ISCAS | 2 |
| 2018 | Ray-space360: An extension of ray-space for omnidirectional free viewpointabstractConventional 3DoF (Degree of Freedom) systems provide only rotational view transformation. 6DoF VR systems offer a more realistic VR by supporting view transformation for translational movement. This paper proposes ray-space 360 which supports both rotational and translational view transformation simultaneously. Ray-space360 is an extension of ray-space [1-3] to an omnidirectional view in order to support a new camera structure in a square form for ray-space360 configuration. Ray-space360 also provides a way to use the intersection to connect independent ray-spaces. Thanks to its low complexity, the proposed system can provide real-time service with a low performance processor such as a mobile processor. This paper presents the result of 4DoF ray-space360 which provides view transformation of yaw, pitch, x and z axes. The ray-space360 can be extended to 6DoF simply by stacking cameras. Hyunmin Jung, Chae-Eun Rhee |
ISCAS | 3 |
| 2018 | Fast Integer Motion Estimation With Bottom-Up Motion Vector Prediction for an HEVC EncoderabstractAlthough advanced motion vector prediction (AMVP) modes based on motion estimation (ME) are selected significantly less due to the merge mode newly adopted in the high-efficiency video coding (HEVC), integer ME (IME) still occupies a large amount of computation in HEVC because the HEVC supports a highly flexible block partitioning structure. Introduction of the merge mode in HEVC substantially affects the optimal search algorithm of IME. Therefore, this gives a chance to reduce the computational complexity of ME with marginal drop of the compression efficiency. This paper proposes a new IME algorithm that significantly reduces its search ranges. Computational complexity of the proposed algorithm is reduced without serious degradation in coding efficiency by obtaining additional accurate motion vector predictors (MVPs) in a bottom-up order and searching narrow regions around multiple MVPs. These bottom-up MVPs are obtained from the prediction units (PUs) in the coding units (CUs) in the hierarchically lower level, which share the pixel area either completely or partially with the current PU. Although the encoder should perform the proposed IME in a bottom-up order from the smallest CUs to the larger CUs, several other processes, such as fractional ME and merges mode are performed in a top-down order to exploit several fast algorithms already adopted in the HEVC test model encoder. To keep compatibility with HEVC standard, the bottom-up MVPs are utilized for only IME. The bitstream is generated using standard AMVP. According to the simulation results, it is confirmed that the proposed algorithm improves the accuracy of MVPs, which leads to the reduction of IME computational complexity up to 81.95% on average with the Bjøntegaard delta bitrate of 0.39%. Tae Sung Kim, Chae-Eun Rhee, Soo-Ik Chae |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Complexity Reduction by Modified Scale-Space Construction in SIFT Generation Optimized for a Mobile GPUabstractScale-invariant feature transform (SIFT) is one of the most widely used local features for computer vision in mobile devices. A mobile graphic processing unit (GPU) is often used to run computer-vision applications using SIFT features, but the performance in such a case is not powerful enough to generate SIFT features in real time. This paper proposes an efficient scheme to optimize the SIFT algorithm for a mobile GPU. It analyzes the conventional scale-space construction step in the SIFT generation, finding that reducing the size of the Gaussian filter and the scale-space image leads to a significant speedup with only a slight degradation of the quality of the features. Based on this observation, the SIFT algorithm is modified and implemented for real-time execution. Additional optimization techniques are employed for a further speedup by efficiently utilizing both the CPU and the GPU in a mobile processor. The proposed SIFT generation scheme achieves a processing speed of 28.30 frames/s for an image with a resolution of 1280 × 720 running on a Galaxy S5 LTE-A device, thereby gaining a speedup by the factors of 114.78 and 4.53 over CPU- and GPU-only implementations, respectively. Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Merge Mode Estimation for a Hardware-Based HEVC EncoderabstractHigh Efficiency Video Coding (HEVC) is a video coding standard that offers higher performance than previous video coding standards such as H.264/AVC. Merge mode is one of the new tools adopted in HEVC to improve the inter-frame coding efficiency. Merge mode saves the bits for the motion vector (MV) by sharing the MV with neighboring blocks. Merge mode estimation (MME) is the process of finding a merge mode candidate, which requires extensive computations and memory accesses due to the associated motion compensation. Although MME is very similar to motion estimation (ME) in many ways, previous research on ME cannot be directly applied to solve many difficulties in designing MME hardware. In this paper, the characteristics of and the computational complexity involved in MME are discussed. To improve the throughput of the MME hardware, partially increased parallelism is efficiently exploited. Furthermore, the M-of-N-pixel combination and flexible memory access schemes are proposed to maximize the scalability to support various block sizes of HEVC and to reduce the time for fetching reference data. The proposed schemes are applied to the MME hardware design in this paper. The proposed hardware can process 56074 of 64 × 64 coding tree units per second with a clock frequency of 366 MHz, and its gate count is 585.4k with 2 kB of dual-port static RAM. Tae Sung Kim, Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | A Novel Hardware Architecture of the Lucas-Kanade Optical Flow for Reduced Frame Memory AccessabstractThe Lucas-Kanade (LK) algorithm is a cost-efficient gradient-based algorithm for real-time optical flow generation. An excessive external memory access limits the LK algorithm from being broadly used in practical high-frame-rate applications. To overcome this limitation, this paper proposes a novel hardware architecture that stores the input image after the Gaussian filtering operation instead of the original input image itself. The Gaussian-filtered image is downsampled in both the horizontal and vertical directions, thus reducing the external memory access to one quarter of the original data. The downsampling operation does not cause a significant degradation of accuracy because the Gaussian filter is a low-pass filter that reduces the aliasing effect of downsampling. The downsampled pixels are selected in an interleaved manner across multiple frames to reduce the degradation of accuracy. Experimental results show that the proposed algorithm reduces the frame memory access by 61%-75% compared with the previous research. Han Seong Son, Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | A Low-Power Video Recording System With Multiple Operation Modes for H.264 and Light-Weight CompressionabstractAn increasing demand for mobile video recording systems makes it important to reduce power consumption and to increase battery lifetime. The H.264/AVC compression is widely used for many video recording systems because of its high compression efficiency; however, the complex coding structure of H.264/AVC compression requires large power consumption. A light-weight video compression (LWC), based on discrete wavelet transform and set partitioning in hierarchical trees, consumes less power than H.264/AVC compression thanks to its relatively simple coding structure, although its compression efficiency is lower than that of H.264/AVC compression. This paper proposes a low-power video recording system that combines both the H.264/AVC encoder with high compression efficiency and LWC with low power consumption. The LWC is used to compress video data for temporal storage while the H.264/AVC encoder is used for permanent storage of data when some events are detected. For further power reduction, a down-sampling operation is utilized for permanent data storage. For an effective use of the two compressions with the down-sampling operation, an appropriate scheme is selected according to the proportion of long-term to short-term storage and the target bitrate. The proposed system reduces power consumption by up to 72.5% compared to that in a conventional video recording system. Hyun Kim 0001, Chae-Eun Rhee |
IEEE Trans. Multim. | 2 |
| 2015 | An Effective Combination of Power Scaling for H.264/AVC CompressionabstractThis brief proposes a novel method to determine the best combination of operation conditions for multiple power-scaling schemes. The power saving and rate-distortion performances of individual schemes are simulated, and then, the combined effects are modeled to obtain the best operation combination. The optimized combinations are defined as a power-level table. The proposed power-aware design is tested with four popular power-saving schemes and simulations show that a power saving of ~25% is achieved at the sacrifice of .0.172 dB Bjontegaard Delta peak signal-to-noise ratio degradation. Hyun Kim 0001, Chae-Eun Rhee |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | An H.264 High-Profile Intra-Prediction with Adaptive Selection Between the Parallel and Pipelined Executions of Prediction ModesabstractA high-profile H.264 intra-frame encoder is suitable for low-cost and low-power applications and capable of providing enhanced compression efficiency. The high-profile is targeting the high-resolution videos. Thus, the encoding speed should be faster than or comparable to the baseline-profile. In previous work related to a hardware-based baseline-profile intra-frame encoder, a speed-up is achieved by the early termination of the intra modes and by an increase in the rate of hardware utilization only under one of the serialized and parallel schedules. This paper proposes a novel pipeline schedule for a hardware-based high-profile intra-prediction scheme in which the 8 × 8 prediction is performed in Stage 1 and 4 × 4, 16 ×16 and chroma predictions are executed during Stage 2. The processing time of Stage 2 is efficiently accelerated based on the result of the 8 × 8 prediction in Stage 1. According to the distribution of each mode, the schedule is adaptively selected between parallel and pipeline schedules. To increase the hardware utilization of the 8 × 8 prediction, the order of prediction modes and the inverse vertical transform is adaptively adjusted. In addition, early termination of the prediction modes is employed for a fast 8 × 8 prediction. The proposed 8 × 8 intra-prediction is implemented and verified as an entire intra-frame encoder. Experimental results show that the average number of cycles necessary to process one MB for videos with resolutions of 1920 ×1080 and 3840 × 2160 are only 269 and 253 cycles, respectively. Compared to JM13.2, the bitrate is increased by 1.13% on average with a small PSNR degradation of 0.06 dB. The difference in the rate-distortion performance between the proposed high-profile intra-prediction scheme and JM 13.2 is not significant, whereas the achieved speed-up due to the proposed schemes is considerable compared to the conventional hardware-based intra-prediction encoders. Chae-Eun Rhee, Tae Sung Kim |
IEEE Trans. Multim. | 1 |
| 2012 | Cascaded Direction Filtering for Fast Multidirectional Inter-Prediction in H.264/AVC Main and High Profile CompressionabstractEarly direct mode decision is a popular technique to improve the execution speed of inter-predictions in a B slice. However, the improvement is limited when this technique is applied to a hardware-based pipelined architecture or to the video sequences where the ratio of macroblocks (MBs) encoded as the direct mode is low. This paper proposes a novel fast inter-prediction algorithm for B slices which increases the encoding speed by early decision of the direction of the inter-prediction using spatial and temporal correlation in a video. A speed-up is achieved by selecting a prediction direction with a lower complexity and discarding a prediction direction with a higher complexity. For the case when early mode decision is not possible, additional speed-up is achieved by partially skipping unidirectional motion estimation (ME) if the estimated cost of the ME is greater than the result of the direct mode prediction that is completed much faster than ME. This proposed sequence of early decisions is referred to as cascaded direction filtering (CDF). The proposed algorithm increases the early mode decision rate, which improves the encoding speed effectively for a hardware-based pipelined architecture as well as videos with a low ratio of MBs encoded as the direct mode. The additional memory size and bandwidth overhead required for the proposed CDF are very small. Experimental results show that the proposed CDF scheme improves the encoding speed by 45% for B slices and 32% for overall sequences. The bitrate is increased by less than 1% with a small peak signal-to-noise ratio degradation of 0.03 dB on average. Chae-Eun Rhee |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Power-aware design with various low-power algorithms for an H.264/AVC encoderabstractH.264/AVC video compression standard provides high coding efficiency, but requires a considerable amount of complexity and power consumption. This paper presents advanced low-power algorithms for an H.264/AVC encoder and a power-aware design composed of low-power algorithms. Power reduction algorithms with frame memory compression and early skip mode decision are presented, and the search range for motion estimation is reduced for further power reduction. The proposed power aware design controls the power consumption depending on the remaining energy by controlling the operation condition of the proposed low-power algorithms. In order to estimate the power reduction by the proposed algorithms, the power consumed by external memory as well as the bus between an H.264 encoder and an external DRAM is considered. Simulation results show that up to 49.9% of the power consumed by bus and external memory is reduced and that the power consumption from 0% to 41.56% is achieved with a reasonably small degradation of R-D performance. Hyun Kim 0001, Chae-Eun Rhee, Sunwoong Kim |
ISCAS | 2 |
| 2010 | A Real-Time H.264/AVC Encoder With Complexity-Aware Time AllocationabstractThis paper presents a novel processing time control algorithm for a hardware-based H.264/AVC encoder. The encoder employs three complexity scaling methods partial cost evaluation for fractional motion estimation (FME), block size adjustment for FME, and search range adjustment for integer motion estimation (IME). With these methods, 12 complexity levels are defined to support tradeoffs between the processing time and compression efficiency. A speed control algorithm is proposed to select the complexity level that compresses most efficiently among those that meet the target time budget. The time budget is allocated to each macroblock based on the complexity of the macroblock and on the execution time of other macroblocks in the frame. For main profile compression, an additional complexity scaling method called direction filtering is proposed to select the prediction direction of FME by comparing the costs resulting from forward and backward IMEs. With direction filtering in addition to the three complexity scaling methods for baseline compression, 32 complexity levels are defined for main profile compression. Experimental results show that the speed control algorithm guarantees the processing time to meet the given time budget with negligible quality degradation. Various complexity levels for speed control are also used to speed up the encoding time with a slight degradation in quality and a minor reduction of the compression efficiency. Chae-Eun Rhee, Jin-Su Jung |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | An SoC Integrating an H.264 Encoder with an ISPabstractThe SoC presented in this paper integrates an H.264 encoder with an ISP (Image Signal Processor). It is currently implemented in an FPGA and processes an HD-size (1280 × 720) image at the speed of 15 fps with the operating clock frequency of 50 MHz. In the presented demo system, a Bayer input from a CMOS image is given to the FPGA and the output stream is transmitted through an USB transceiver to a PC that decodes and displays the H.264 stream. Eung Sup Kim, Seongyoon Kim, Gyoung-Hwan Hyun, Jin-Su Jung, Chae-Eun Rhee, Yongseok Jin |
ISCAS | 5 |
| 2007 | A New Frame Recompression Algorithm Integrated with H.264 Video CompressionabstractTo reduce the size and bandwidth requirement of a frame memory for video compression, a number of memory recompression algorithms have been proposed. These previous algorithms are performed independently of a video compression standard and therefore do not take advantage of the information obtained during the processing of the compression standard. This paper proposes a new recompression algorithm that makes use of the information from H.264 intra prediction results. The proposed algorithm decomposes a frame into 4times4 blocks which are then compressed into 64-bit segments. The result of 4times4 intra prediction is used to select the scan order of the 4times4 block and DPCM (differential pulse code modulation) is performed along this scan order. Then, the DPCM results are further compressed by Golomb-Rice coding. The proposed recompression algorithm is implemented in hardware and integrated with an H.264 encoder. The proposed algorithm improves the average PSNR by 2.9dB compared to the previous work in (Lee, 2003). The hardware cost for the implementation of the recompression algorithm is 28 K gates and the additional latency to read the compressed frame memory is 162 cycles per a macroblock Yongje Lee, Chae-Eun Rhee |
ISCAS | 2 |