EDBT 2026 Demo / reviewers in the wild / expert
Je-Won Kang
dblp:13/5636
· DBLP profile ↗
32ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-1637-9479ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Future object localization using multi-modal ego-centric video
Jee-Ye Yoon, Je-Won Kang |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Illumination Spectrum Estimation for Multispectral Images via Surface Reflectance Modeling and Spatial-Spectral Feature GenerationabstractMultispectral (MS) images contain richer spectral information than RGB images due to their increased number of channels and are widely used for various applications. However, achieving accurate estimation in MS images remains challenging, as previous studies have struggled with spectral diversity and the inherent entanglement between the illuminant and surface reflectance spectra. To tackle these challenges, in this paper, we propose a novel Illumination spectrum estimation technique for MS images via Surface reflectance modeling and Spatial-spectral feature generation (ISS). The proposed technique employs a learnable spectral unmixing (SU) block to enhance surface reflectance modeling, which was unattempted in the illumination spectrum estimation, and a feature mixing block to fuse spectral and spatial features of MS images with cross-attention. The features are refined iteratively and processed through a decoder to produce an illumination spectrum estimator. Experimental results demonstrate that the proposed technique achieves state-of-the-art performance in illumination spectrum estimation in various MS image datasets. The code is available at https://github.com/heyjinnii/ISS-MSI.git. Hyejin Oh, Woo-Shik Kim, Sangyoon Lee 0003, YungKyung Park, Je-Won Kang |
CVPR | 5 |
| 2025 | Lightweight Temporal Contextual Fine-Tuning Method of Large Multimodal Model for Video Moment RetrievalabstractVideo moment retrieval (VMR) tasks require a comprehensive understanding of the video-language features of an input video, based on large multimodal models (LMMs). In this paper, we introduce temporal contextual prompts and provide contextual information into the LMM model to improve the generalization ability of the VMR. We integrate temporal contextual prompts into conventional prompts and obtain integrated tokens through an embedding module. The temporal contextual prompts are transformed into temporal conditional tokens, and these are fine-tuned as the embedding representation of the temporal correlation between the video and the given query text. In addition, we quantize the model and apply LoRA to demonstrate efficient learning for limited resources. Experimental results on Charades-STA show improvements of 3.78 percent in mIoU and 0.85 percent in [email protected] over state-of-the-art performance. Semi Kwon, Ju-Hee Lee, Je-Won Kang |
ICIP | 3 |
| 2025 | Erp-Aware Text-To-360 Panorama Diffusion ModelabstractGrowing demand for immersive media and virtual reality highlights the importance of text-to-360° panoramic image (T2P) generation. While diffusion models have achieved remarkable success in image generation, generating equirectangular projection (ERP) images from text is challenging due to spherical projection distortion and the lack of sufficient ERP data to establish clear priors during training. In contrast, cube-map perspective views, a standard 2D image trained on abundant text-image pairs, provide stable generation results. In this paper, we propose a novel T2P model, consisting of an ERP-aware diffusion process in the forward direction and a perspective-aware diffusion process in the reverse direction, supported by several enabling modules to bridge the domain gap between ERP and 2D images. Experimental results demonstrate that the proposed model effectively generates visually compelling and high-quality ERP images while maintaining realistic patterns in ERP. Ju-Hee Lee, Je-Won Kang |
ICIP | 2 |
| 2025 | Illumination Spectrum Estimation for Multispectral Images Using Illuminant PriorabstractMultispectral (MS) imaging, which captures richer spectral information than traditional three-channel RGB images, enables more precise reconstruction of illuminant spectral power distribution. However, accurate illuminant spectrum estimation (ISE) using MS images remains a challenging task, as existing studies often neglect the physical characteristics of spectral images. To address these challenges, we propose a novel deep learning model that incorporates illuminant prior (IP) information, extending a Gray-World assumption commonly used in RGB color constancy. In particular, the IP enhances the accuracy of the proposed network through an IP-aware attention network. We demonstrate the superiority of our proposed method through quantitative and qualitative results across various datasets for illumination spectral estimation in multispectral images. Hyejin Oh, Je-Won Kang |
ICIP | 2 |
| 2025 | Label Space-Induced Pseudo Label Refinement for Multi-Source Black-Box Domain AdaptationabstractConventional unsupervised domain adaptation (UDA) requires access to source data and/or source model parameters, prohibiting its practical application in terms of privacy, security, and intellectual property. Recent black-box UDA (BDA) reduces such constraints by defining a pseudo label from a single encapsulated source application programming interface (API) prediction, which allows for self-training of the target model. Nonetheless, existing methods have limited consideration for multi-source settings, in which multiple source domain APIs are available to generate pseudo labels. In this work, we introduce a novel training framework for multi-source BDA (MSBDA), dubbed Label Space-Induced Pseudo Label Refinement (LPR). Specifically, LPR incorporates a Pseudo label Refinery Network (PRN) that learns the relationship among source domains conditioned by the target domain only utilizing source API's prediction. The target model is adapted by our dual phases PRN. First, a warm-up phase targets to avoid failure due to noisy samples and provide an initial pseudo-label, which is followed by a label refinement phase with domain relationship exploration. We provide theoretical support for the mechanism of the LPR. Experimental results on four benchmark datasets demonstrate that MSBDA using LPR achieves competitive performance compared to state-of-the-art approaches with different DA settings. Chae Hwa Yoo, Xiaofeng Liu 0001, Fangxu Xing, Jonghye Woo, Je-Won Kang |
IEEE Trans. Image Process. | 5 |
| 2025 | Neural Volumetric Video Coding With Hierarchical Coded Representation of Dynamic VolumeabstractThis article proposes a novel multi-view (MV) video coding technique that leverages a four-dimensional (4D) voxel-grid representation to enhance coding efficiency, particularly in novel view synthesis. Although the voxel grid approximation provides a continuous representation for dynamic scenes, its volumetric nature requires substantial storage. The compression of MV videos can be interpreted as the compression of dense features. However, the substantial size of these features poses a significant problem relative to the generation of dynamic scenes at arbitrary viewpoints. To address this challenge, this study introduces a hierarchical coded representation of dynamic volumes based on low-rank tensor decomposition of volumetric features and develops effective coding techniques based on this representation. The proposed method employs a two-level coding strategy to capture the temporal characteristics of the decomposed features. At a higher level, spatial features are encoded, representing 3D structural information, with time-invariant components over short intervals of an MV video sequence. At a lower level, temporal features are encoded to capture the dynamics of current scenes. The spatial features are shared in a group, and temporal features are encoded at each time step. The experimental results demonstrate that the proposed technique outperforms existing MV video coding standards and current state-of-the-art methods, providing superior rate-distortion performance in the novel view synthesis of MV video compression. Ju-Yeon Shin, Jung-Kyung Lee, Gun Bang, Junsik Kim 0002, Je-Won Kang |
IEEE Trans. Multim. | 5 |
| 2025 | Reference-based In-loop Filter with Robust Neural Feature Transfer for Video CodingabstractIn this article, we propose an efficient reference-based deep in-loop filtering method for video coding. Existing reference-based in-loop filters often face challenges in improving coding efficiency due to the difficulty in capturing relevant textures from the reference frames. Our method accurately predicts the texture of a reference block and uses this information to restore the current block. To achieve this, we develop a reference-to-current feature estimation module that conveys high-quality information from previously coded frames in the feature domain, thereby preventing loss of detail due to inaccurate prediction. Although a neural network is trained to restore a coded video frame to be similar to the current frame, their performance can significantly degrade when operating with various quantization parameters (QPs) and managing different levels of distortion. This problem becomes further severe in the reference-to-current feature estimation, in which QP values are applied differently to video frames. We address this problem by developing a QP-aware convolution layer with a small number of learnable parameters to generate reliable features and adapt to fine-grained adaptive QPs among consecutive frames. The proposed method is implemented into the versatile video coding (VVC) reference software, VTM version 10.0. Experimental results demonstrate that the proposed method improves coding performance significantly in VVC. Jung-Kyung Lee, Je-Won Kang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | SRTube: Video-Language Pre-Training with Action-Centric Video Tube Features and Semantic Role LabelingabstractIn recent years, large-scale video-language pre-training (VidLP) has received considerable attention for its effectiveness in relevant tasks. In this paper, we propose a novel action-centric VidLP framework that employs video tube features for temporal modeling and language features based on semantic role labeling (SRL). Our video encoder generates multiple tube features along object trajectories, identifying action-related regions within videos, to over-come the limitations of existing temporal attention mechanisms. Additionally, our text encoder incorporates high-level, action-related language knowledge, previously under-utilized in current VidLP models. The SRL captures action-verbs and related semantics among objects in sentences and enhances the ability to perform instance-level text matching, thus enriching the cross-modal (CM) alignment process. We also introduce two novel pre-training objectives and a self-supervision strategy to produce a more faithful CM representation. Experimental results demonstrate that our method outperforms existing VidLP frameworks in various downstream tasks and datasets, establishing our model a baseline in the modern VidLP framework. Ju-Hee Lee, Je-Won Kang |
CVPR | 2 |
| 2024 | SNeRV: Spectra-Preserving Neural Representation for Video
Jihoo Lee, Je-Won Kang |
ECCV (55) | 3 |
| 2024 | Dynamic Volumetric Video Coding with Tensor DecompositionabstractRecently, volumetric video coding based on neural radiance fields has gained significant attention for storing and transmitting three-dimensional (3D) scenes captured from multi-view video. Because the neural networks are trained to produce novel view synthesis of surrounding 3D scenes, compressing the model and then rendering the colors and geometry through the decompressed model can be utilized as a 3D video coding system. However, although this approach provides superior performance compared to conventional 3D video coding standards using depth video, challenges remain in reducing overall model sizes to improve coding efficiency. In this paper, we propose a novel dynamic volumetric video coding technique that employs a Group of Volume (GoV) to divide multi-view video sequences into smaller chunks, addressing complex temporal dynamics. Our method uses volumetric video features represented with 3D spatial and temporal tensor matrices and vectors and encodes them with the GoVs. The tensors are compressed by existing 2D video codec, allowing for fast rendering and easing deployment. Experimental results validate that our method not only reduces memory footprint but also maintains high-quality rendering as compared to state-of-the-art studies. Ju-Yeon Shin, Yeoneui Kim, Je-Won Kang, Gun Bang |
VCIP | 3 |
| 2022 | Relation Enhanced Vision Language Pre-TrainingabstractIn this paper, we propose a relation enhanced vision-language pre-training (VLP) method for a transformer model (TM) to improve performance in vision-language (V+L) tasks. Current VLP studies attempted to generate a multimodal representation with individual objects as input and relied on a self-attention to learn semantic representation in a brute force manner. However, the relations among objects in an image are largely ignored. To address the problem, we generate a paired visual feature (PVF) that is organized to express the relations between objects. Prior knowledge that reflects co-occurrences of paired objects and a pair-wise distance matrix adjusts the relations, and a triplet is used for sentence embedding. Experimental results demonstrate that the proposed method is efficiently used for VLP by bridging relations between objects, and thus improves performance on V+L downstream tasks. Ju-Hee Lee, Je-Won Kang |
ICIP | 2 |
| 2022 | Transferring Structured Knowledge in Unsupervised Domain Adaptation of a Sleep Staging NetworkabstractAutomatic sleep staging based on deep learning (DL) has been attracting attention for analyzing sleep quality and determining treatment effects. It is challenging to acquire long-term sleep data from numerous subjects and manually labeling them even though most DL-based models are trained using large-scale sleep data to provide state-of-the-art performance. One way to overcome this data shortage is to create a pre-trained network with an existing large-scale dataset (source domain) that is applicable to small cohorts of datasets (target domain); however, discrepancies in data distribution between the domains prevent successful refinement of this approach. In this paper, we propose an unsupervised domain adaptation method for sleep staging networks to reduce discrepancies by re-aligning the domains in the same space and producing domain-invariant features. Specifically, in addition to a classical domain discriminator, we introduce local discriminators - subject and stage - to maintain the intrinsic structure of sleep data to decrease local misalignments while using adversarial learning to play a minimax game between the feature extractor and discriminators. Moreover, we present several optimization schemes during training because the conventional adversarial learning is not effective to our training scheme. We evaluate the performance of the proposed method by examining the staging performances of a baseline network compared with direct transfer (DT) learning in various conditions. The experimental results demonstrate that the proposed domain adaptation significantly improves the performance though it needs no labeled sleep data in target domain. Chae Hwa Yoo, Hyang Woon Lee, Je-Won Kang |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Video Question Answering Using Language-Guided Deep Compressed-Domain Video FeatureabstractVideo Question Answering (Video QA) aims to give an answer to the question through semantic reasoning between visual and linguistic information. Recently, handling large amounts of multi-modal video and language information of a video is considered important in the industry. However, the current video QA models use deep features, suffered from significant computational complexity and insufficient representation capability both in training and testing. Existing features are extracted using pre-trained networks after all the frames are decoded, which is not always suitable for video QA tasks. In this paper, we develop a novel deep neural network to provide video QA features obtained from coded video bit-stream to reduce the complexity. The proposed network includes several dedicated deep modules to both the video QA and the video compression system, which is the first attempt at the video QA task. The proposed network is predominantly model-agnostic. It is integrated into the state-of-the-art networks for improved performance without any computationally expensive motion-related deep models. The experimental results demonstrate that the proposed network outperforms the previous studies at lower complexity. https://github.com/Nayoung-Kim-ICP/VQAC Seong Jong Ha, Je-Won Kang |
ICCV | 3 |
| 2021 | Dynamic Motion Estimation and Evolution Video Prediction NetworkabstractFuture video prediction provides valuable information that helps a computer machine understand the surrounding environment and make critical decisions in real-time. However, long-term video prediction remains a challenging problem due to the complicated spatiotemporal dynamics in a video. In this paper, we propose a dynamic motion estimation and evolution (DMEE) network model to generate unseen future videos from the observed videos in the past. Our primary contribution is to use trained kernels in convolutional neural network (CNN) and long short-term memory (LSTM) architectures, adapted to each time step and sample position, to efficiently manage spatiotemporal dynamics. DMEE uses the motion estimation (ME) and motion update (MU) kernels to predict the future video using an end-to-end prediction-update process. In the prediction, the ME kernel estimates the temporal changes. In the update step, the MU kernel combines the estimates with the previously generated frames as reference frames using a weighted average. The kernels are not only used for a current frame, but also are evolved to generate successive frames to enable temporally specific filtering. We perform qualitative performance analysis and quantitative performance analysis based on the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and video classification score developed for examining the visual quality of the generated video. It is demonstrated with experiments that our algorithm provides better qualitative and quantitative performance superior to the current state-of-the-art algorithms. Our source codes are available inhttps://github.com/Nayoung-Kim-ICP/Video-Generation. Na-Young Kim, Je-Won Kang |
IEEE Trans. Multim. | 2 |
| 2021 | Fast Multi-Type Tree Partitioning for Versatile Video Coding Using a Lightweight Neural NetworkabstractIn this paper, we propose a fast decision scheme using a lightweight neural network (LNN) to avoid redundant block partitioning in versatile video coding (VVC). A more versatile block structure, named the multi-type tree (MTT) structure, which includes binary trees (BTs) and ternary trees (TTs), is adopted by VCC, in addition to the traditional quadtree structure. The MTT improved the coding efficiency compared with previous video coding standards. However, the new tree structures, mainly TT, significantly increased the complexity of the VVC encoder. Although widespread application of VVC has been inhibited, this problem has not yet been investigated thoroughly in the literature. In this study, we first determine the statistical characteristics of coded parameters that exhibit correlation with the TT and develop two useful types of features—explicit VVC features(EVFs) andderived VVC features(DVFs)—to facilitate the intra coding of VVC. These features can be obtained efficiently during the intra prediction before the determination of the best block partitioning during rate-distortion optimization in VVC encoding. Our LNN model decides whether to terminate the nested TT block structures subsequent to a quadtree based on the features. The experimental results confirm that the proposed method substantially decreases the encoding complexity of VVC with a slight coding loss under the All Intra configuration. Our code, models, and dataset are available athttps://github.com/foriamweak/MTTPartitioning_LNN. Sanghyo Park 0001, Je-Won Kang |
IEEE Trans. Multim. | 2 |
| 2018 | Long-Term Video Generation with Evolving Residual Video FramesabstractIn this paper, we propose a novel long-term video generation algorithm, motivated by the recent developments of unsupervised deep learning techniques. The proposed technique learns two ingredients of internal video representation, i.e., video textures and motions to reproduce realistic pixels in the future video frames. To this aim, the proposed technique uses two encoders comprising convolutional neural networks (CNN) to extract spatiotemporal features from the original video frame and a residual video frame, respectively. The use of the residual frame facilitates the learning with fewer parameters as there are high spatiotemporal correlations in a video. Moreover, the residual frames are efficiently used for evolving pixel differences in the future frame. In a decoder, the future frame is generated by transforming the combination of two feature vectors into the original video size. Experimental results demonstrate that the proposed technique provides more robust and accurate results of long-term video generation than conventional techniques. Na-Young Kim, Je-Won Kang |
ICIP | 2 |
| 2018 | Machine Learning-Based Fast Angular Prediction Mode Decision Technique in Video CodingabstractIn this paper we propose a machine learning-based fast intra-prediction mode decision algorithm, using random forest that is an ensemble model of randomized decision trees. The random forest is used to estimate an intra-prediction mode from a prediction unit and to reduce encoding time significantly by avoiding the intensive Rate-Distortion optimization of a number of intra-prediction modes. To this aim, we develop a randomized tree model including parameterized split functions at nodes to learn directional block-based features. The feature uses only four pixels reflecting a directional property of a block, and, thus the evaluation is fast and efficient. To integrate the proposed technique into the conventional video coding standard frameworks, the intra-prediction mode derived from the proposed technique, called an inferred mode (IM), is used to shrink the pool of the candidate modes before carrying out the Rate-Distortion (R-D) optimization. The proposed technique is implemented into the High Efficiency Video Coding Test Model (HM) reference software of the state-of-the-art video coding standard and Joint Exploration Model (JEM) reference software, by integrating the random forest trained off-line into the codecs. Experimental results demonstrate that the proposed technique achieves significant encoding time reduction with only slight coding loss as compared the reference software models. Sookyung Ryu, Je-Won Kang |
IEEE Trans. Image Process. | 2 |
| 2016 | A Novel Intrusion Detection Method Using Deep Neural Network for In-Vehicle Network SecurityabstractIn this paper, we propose a novel intrusion detection technique using a deep neural network (DNN). In the proposed technique, in-vehicle network packets exchanged between electronic control units (ECU) are trained to extract low- dimensional features and used for discriminating normal and hacking packets. The features perform in high efficient and low complexity because they are generated directly from a bitstream over the network. The proposed technique monitors an exchanging packet in the vehicular network while the feature are trained off-line, and provides a real-time response to the attack with a significantly high detection ratio in our experiments. Min Ju Kang, Je-Won Kang |
VTC Spring | 2 |
| 2016 | Compressed domain video saliency detection using global and local spatiotemporal features
Se-Ho Lee, Je-Won Kang, Chang-Su Kim 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Multiview and 3D Video Compression Using Neighboring Block Based Disparity VectorsabstractCompression of the statistical redundancy among different viewpoints, i.e., inter-view redundancy, is a fundamental and critical problem in multiview and three-dimensional (3D) video coding. To exploit the inter-view redundancy, disparity vectors are required to identify pixels of the same objects within two different views; in this way, the enhancement coding tools can be efficiently employed as new modes in block-based video codecs to achieve higher compression efficiency. Although disparity can be converted from depth, it is not possible in multiview video coding since depth information is not considered. Even when depth information is coded, it breaks the so-called multiview compatibility wherein texture views can be decoded without depth information. To resolve this problem, in this paper, a neighboring block-based disparity vector derivation (NBDV) method is proposed. The basic concept of NBDV is to derive a disparity vector (DV) of a current block by utilizing the motion information of spatially and temporally neighboring blocks predicted from another view. Through extensive experiments and analysis, it is shown that the proposed NBDV method achieves efficient DV derivation in the state-of-art video codecs, and it keeps the multiview compatibility with a relatively lower complexity. The proposed method has become an essential part of the 3D video standard extensions of H.264/AVC and HEVC. Ying Chen 0011, Xin Zhao 0003, Li Zhang 0006, Je-Won Kang |
IEEE Trans. Multim. | 4 |
| 2016 | Efficient Residual DPCM Using an l1 Robust Linear Prediction in Screen Content Video CodingabstractIn this paper, a residual differential pulse code modulation (RDPCM) coding technique using a weighted linear combination of neighboring residual samples is proposed to provide coding efficiency in the screen content video coding. The RDPCM performs the sample-based prediction of residue to reduce spatial redundancies. The proposed method uses the$l_1$optimization in the weight derivation by considering the statistical characteristics of graphical components in videos in an intracoding. Specifically we use the least absolute shrinkage and selection operator to derive the weights because the solution is accurate in high variance residue. Furthermore, we enhance parallelism in a line processing by restricting the support to the row-wise prediction to above samples or the column-wise prediction to the left samples. The proposed method uses an explicit RDPCM scheme, so a coding mode determined by rate-distortion optimization is transmitted to a decoder. For coding the overhead, we develop a context design in CABAC based on correlation between an intraprediction direction and an RDPCM prediction mode. It is demonstrated with the experimental results that the proposed method provides a significant coding gain over the state-of-the-art reference codec for screen content video coding. Je-Won Kang, Su-Kyung Ryu, Na-Young Kim, Min-Joo Kang |
IEEE Trans. Multim. | 1 |
| 2014 | Low complexity Neighboring Block based Disparity Vector Derivation in 3D-HEVCabstract3D-HEVC incorporates advanced inter-view prediction techniques based on more accurately derived disparity vector to better exploit the correlation between objects in different views. The efficient disparity vector derivation method, namely, Neighboring Block based Disparity Vector Derivation (NBDV) provides disparity without accessing any depth information. The NBDV has been developed as a part of the 3D-HEVC in Joint Collaborative Team on 3D Video Coding (JCT-3V) for several meeting cycles, and adopted as a common coding tool used for all the inter-view prediction techniques for high efficient coding of texture views. This paper presents a low complexity NBDV, which is the state-of-the-art disparity vector derivation method in 3D-HEVC. Je-Won Kang, Ying Chen 0011, Li Zhang 0006, Marta Karczewicz |
ISCAS | 1 |
| 2013 | Neighboring block based disparity vector derivation for 3D-AVCabstract3D-AVC, being developed under Joint Collaborative Team on 3D Video Coding (JCT-3V), significantly outperforms the Multiview Video Coding plus Depth (MVC+D) which has no new macroblock level coding tools compared to Multiview video coding extension of H.264/AVC (MVC). However, for multiview compatible configuration, i.e., when texture views are decoded without accessing depth information, the performance of the current 3D-AVC is only marginally better than MVC+D. The problem is caused by the lack of disparity vectors which can be obtained only from the coded depth views in 3D-AVC. In this paper, a disparity vector derivation method is proposed by using the motion information of neighboring blocks and applied along with existing coding tools in 3D-AVC. The proposed method improves 3D-AVC in the multiview compatible mode substantially, resulting in about 20% bitrate reduction for texture coding. When enabling the so-called view synthesis prediction to further refine the disparity vectors, the performance of the proposed method is 31% better than MVC+D and even better than 3D-AVC under the best performing 3D-AVC configuration. Li Zhang 0006, Je-Won Kang, Xin Zhao 0003, Ying Chen 0011, Rajan L. Joshi |
VCIP | 2 |
| 2013 | Efficient HD video coding with joint first-order-residual (FOR) and second-order-residual (SOR) coding technique
Je-Won Kang, Chung-Cheng Lou, Seung-Hwan Kim 0001, C.-C. Jay Kuo |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | Sparse/DCT (S/DCT) Two-Layered Representation of Prediction Residuals for Video CodingabstractIn this paper, we propose a cascaded sparse/DCT (S/DCT) two-layer representation of prediction residuals, and implement this idea on top of the state-of-the-art high efficiency video coding (HEVC) standard. First, a dictionary is adaptively trained to contain featured patterns of residual signals so that a high portion of energy in a structured residual can be efficiently coded via sparse coding. It is observed that the sparse representation alone is less effective in the R-D performance due to the side information overhead at higher bit rates. To overcome this problem, the DCT representation is cascaded at the second stage. It is applied to the remaining signal to improve coding efficiency. The two representations successfully complement each other. It is demonstrated by experimental results that the proposed algorithm outperforms the HEVC reference codec HM5.0 in the Common Test Condition. Je-Won Kang, Moncef Gabbouj, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 1 |
| 2011 | FOR/SOR video coding with super macroblock and inter-frame stripe predictionabstractA first-order-residual/second-order-residual (FOR/SOR) based video coding algorithm that incorporates the super macroblock (SMB) and the inter-frame stripe prediction (ISP) technique is proposed for high definition (HD) video coding in this work. We first examine the limitation of the SMB in high-bit-rate coding, and show that a simple extension of the block-size is not sufficient to get a significant improvement in coding performance. Then, we introduce the FOR/SOR coding technique to resolve this issue. For the FOR coding, we apply a larger quantization stepize and motion-predictive coding to remove the temporal correlation. Then, for the SOR coding, we propose an efficient prediction method, called the inter frame stripe prediction (ISP) technique, to remove the structured residuals after the FOR coding to achieve a higher coding gain. It is demonstrated by experimental results that the proposed FOR/SOR algorithm outperforms H.264/AVC by a significant margin; namely, about 21.57% in the bit rate saving or 1.09dB in the PSNR improvement. Je-Won Kang, Seung-Hwan Kim 0001, C.-C. Jay Kuo |
ICASSP | 1 |
| 2011 | Improved FOR/SOR-based video coding and its performance analysisabstractThe FOR/SOR algorithm was initially proposed in [1] and shown to outperform H.264/AVC by a significant margin for high definition (HD) video coding when the bit rate is equal to 20Mbps or higher. In this work, we propose an improved FOR/SOR coding scheme which offers a superior coding performance in a bit rate range of practical interest, which ranges from 1-15 Mbps. To enhance the coding performance of the proposed algorithm, we incorporate several new coding tools in the basic FOR/SOR coding algorithm in [1]. Furthermore, we attempt to give an intuitive explanation of the superior performance of the FOR/SOR coding scheme using arguments based on the R-D analysis. Finally, it is shown by experimental results that the proposed algorithm outperforms H.264/AVC significantly for HD video in the bit rate range of common interest. Je-Won Kang, Chung-Cheng Lou, Seung-Hwan Kim 0001, C.-C. Jay Kuo |
ICIP | 1 |
| 2011 | Efficient dictionary based video coding with reduced side informationabstractIn this paper, we propose a novel dictionary based video coding technique with adaptive construction of over complete dictionaries and advanced coding methods tailored to sparse signal representations. A set of dictionaries is trained off-line using inter or intra predicted residual samples and is applied for encoding. New coding tools are developed so that the encoder can more compactly represent the residual signal. The same set of dictionary elements can be reused for neighboring blocks, and the optimal number of dictionary elements can be decided using rate-distortion optimization. Experimental results demonstrate that the proposed algorithm yields both improved coding performance and improved perceptual quality at low bit rates. Je-Won Kang, C.-C. Jay Kuo, Robert A. Cohen, Anthony Vetro |
ISCAS | 1 |
| 2011 | Improved H.264/AVC Lossless Intra Coding With Two-Layered Residual Coding (TRC)abstractA lossless image coding method, known as the two-layered residual coding (TRC) scheme, is proposed in this letter. After the H.264/AVC lossy intra prediction, we propose an advanced scheme for residual coding, which consists of two residual coders in cascade. The first-layer residual coder is conducted via transform and quantization with a coarser quantization parameter. The second-layer residual coder is a bit-plane coding method. It is shown experimentally that the TRC scheme outperforms the H.264/AVC lossless intra coding with an averaged bit rate saving of about 24%. Seung-Hwan Kim 0001, Je-Won Kang, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Coding Order Decision of B Frames for Rate-Distortion Performance Improvement in Single-View Video and Multiview Video CodingabstractThe coding gain that can be achieved by improving the coding order of B frames in the H.264/AVC standard is investigated in this work. We first represent the coding order of B frames and their reference frames with a binary tree. We then formulate a recursive equation to find out the binary tree that provides a suboptimal, but very efficient, coding order. The recursive equation is efficiently solved using a dynamic programming method. Furthermore, we extend the coding order improvement technique to the case of multiview video sequences, in which the quadtree representation is used instead of the binary tree representation. Simulation results demonstrate that the proposed algorithm provides significantly better R-D performance than conventional prediction structures. Je-Won Kang, Young-Yoon Lee, Chang-Su Kim 0001, Sang Uk Lee |
IEEE Trans. Image Process. | 1 |
| 2007 | Graph Theoretical Optimization of Prediction Structure in Multiview Video CodingabstractAn algorithm to construct the optimal prediction structure in multiview video coding (MVC) is proposed in this work. We employ the graph theory as a framework. By considering each frame as a vertex and the motion compensation or disparity compensation as an edge, we represent a prediction structure as a spanning tree. Then, we obtain the optimal structure by finding the minimum spanning tree using the Prim's algorithm. Simulation results demonstrate that the proposed algorithm provides about 0.2-0.4 dB better PSNR performance than the conventional prediction structure, and about 1.5 dB better performance than the simulcast. Je-Won Kang, Suk-Hee Cho, Namho Hur, Chang-Su Kim 0001, Sang Uk Lee |
ICIP (6) | 1 |