Hui Yong Kim

dblp:59/8842 · also Hui-Yong Kim · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0001-7308-133XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient dataset condensation with learnable color representation and subset matching
Linh-Tam Tran, Quang Hieu Vo, Maryam Qamar, Chaoning Zhang, Hui Yong Kim, Sung-Ho Bae
Knowl. Based Syst.5
2025 Phase Distribution Matters: On the Importance of Phase Distribution Alignment (PDA) in Holographic Applications
Seungmi Choi, TaeHwa Lee, Jun Yeong Cha, Suhyun Jo, Hyunmin Ban, Kwan-Jung Oh, Hyunsuk Ko, Hui Yong Kim
ACM Multimedia8
2025 Spatial-Channel Mixing Block for Neural Network-based Video Coding (NNVC) Tools
abstract
In this paper, we propose an efficient building block for neural network-based video coding (NNVC) tools, which are part of the post-VVC research. Our proposed Spatial-Channel Mixing (SCM) block explicitly separates spatial and channel mixing operations and applies them sequentially. The SCM block is integrated into existing NNLF (NN-based in-Loop Filter) and NNSR (NN-based Super Resolution) tools in NNVC by replacing their backbone blocks. Under the Common Test Conditions (CTC) of the Joint Video Experts Team (JVET) NNVC, the proposed NNLF and NNSR showed consistent BD-Rate improvements across Y, U, and V channels in various configurations, along with reductions in parameter count and MACs per pixel, demonstrating the effectiveness and efficiency of the proposed SCM block for NNVC tools.
HyunDong Cho, Suyong Bahk, Jong Wook Kim, Donghyun Kim 0017, Sung-Chang Lim, Hui Yong Kim
VCIP6
2024 NHVC: Neural Holographic Video Compression with Scalable Architecture
abstract
Recently, neural network-based approaches for hologram generation and compression have gained popularity as they allow for efficient inference on GPUs without the need for iterative optimization required in traditional methods. In this paper, we introduce Neural Holographic Video Compression (NHVC), an end-to-end trainable and scalable model designed for high-quality phase hologram video generation and compression. NHVC consists of an auto-encoder-based phase hologram generator, a latent coder and-two hyper-prior coders. For each input image, the latent features are extracted through the encoder part of the phase generator and then entropy coded at the shared latent coder based on the hyper-prior information. The two hyper-prior coders employ a spatial and a spatio-temporal entropy model for I-frames and P-frames, respectively. With this architecture, our NHVC can offer task-scalability, allowing a single trained model to serve as a phase hologram generator, phase hologram image compressor, or phase hologram video compressor as required.Experimental results on phase hologram video compression with UVG dataset show that our model outperforms ‘HoloNet + VVC’ by 75.6% BD-Rate reduction, with modest 2K encoding and decoding speeds (5 fps and 12 fps, respectively). For the phase hologram video generation task, our model showed much higher-quality (almost 42dB PSNR) reconstruction using the UVG dataset, while the previous neural generation model HoloNet provides at most 36dB reconstruction quality. We also provide an extensive experimental study on several important design questions such as the need for quadruple extension (QE) in the neural compression model, the feasibility of motion estimation in the phase domain, and an alternative, the need for increasing receptive field to learn better phase features, and variable rate support with a single trained model. It is noteworthy that our model is the first and best neural phase video compression model providing such high-quality reconstruction and task-scalability.
Hyunmin Ban, Seungmi Choi, Jun Yeong Cha, Yeongwoong Kim, Hui Yong Kim
VR5
2024 P-Hologen: An End-to-End Generative Framework for Phase-Only Holograms
abstract
Abstract Holography stands at the forefront of visual technology, offering immersive, three‐dimensional visualizations through the manipulation of light wave amplitude and phase. Although generative models have been extensively explored in the image domain, their application to holograms remains relatively underexplored due to the inherent complexity of phase learning. Exploiting generative models for holograms offers exciting opportunities for advancing innovation and creativity, such as semantic‐aware hologram generation and editing. Currently, the most viable approach for utilizing generative models in the hologram domain involves integrating an image‐based generative model with an image‐to‐hologram conversion model, which comes at the cost of increased computational complexity and inefficiency. To tackle this problem, we introduce P‐Hologen, the first end‐to‐end generative framework designed for phase‐only holograms (POHs). P‐Hologen employs vector quantized variational autoencoders to capture the complex distributions of POHs. It also integrates the angular spectrum method into the training process, constructing latent spaces for complex phase data using strategies from the image processing domain. Extensive experiments demonstrate that P‐Hologen achieves superior quality and computational efficiency compared to the existing methods. Furthermore, our model generates high‐quality unseen, diverse holographic content from its learned latent space without requiring pre‐existing images. Our work paves the way for new applications and methodologies in holographic content creation, opening a new era in the exploration of generative holographic content. The code for our paper is publicly available on https://github.com/james0223/P-Hologen .
JooHyun Park, Yujin Jeon, Hui Yong Kim, Seung-Hwan Baek, HyeongYeop Kang
Comput. Graph. Forum3
2024 Rate-Rendering Distortion Optimized Preprocessing for Texture Map Compression of 3D Reconstructed Scenes
abstract
Textured meshes are widely used in computer graphics to represent 3D scenes, with UV mapping playing a crucial role in establishing a bijective mapping between the 3D mesh surface and a 2D texture. This mapping not only allows for the enhancement of rendering quality but also enables the compression of mesh textures using standard 2D image or video codecs. However, when reconstructing meshes from real-world multiview images, the resulting UV texture maps often suffer from fragmentation due to geometric inaccuracies and excessive tessellation of the reconstructed surfaces, leading to decreased compression performance. In this paper, we propose a novel and effective preprocessing approach for UV texture map compression based on rate-rendering distortion (R-RD) optimization. Unlike existing methods that rely on padding or smoothing, our method iteratively updates the texture map using the gradient of a joint cost of bitrate and rendering distortion. This cost is estimated through a differentiable image encoder and a differentiable texture sampling. Experimental results with lossless compressed mesh geometry demonstrate that our preprocessing method outperforms existing texture padding methods, achieving BD-rate reductions of at least 10.23%, 15.24%, and 12.10% when combined with JPEG, HEVC, and VvC, respectively. We also validate the effectiveness of our method with lossy compressed meshes using Google Draco, showing improved compression efficiency compared to the lossless geometry scenario. Subjective evaluations further confirm that our method enhances both color and structural continuities in the texture map by automatically eliminating high-frequency components unfavorable to compression. The paper provides comprehensive experiments and analyses, including rate estimation with different choices of differentiable image encoders, texture map distortion vs. rendering distortion, and complexity comparison with existing methods.
Soowoong Kim, Jihoon Do, Jungwon Kang, Hui Yong Kim
IEEE Trans. Circuits Syst. Video Technol.4
2024 End-to-End Learnable Multi-Scale Feature Compression for VCM
abstract
The proliferation of deep learning-based machine vision applications has given rise to a new type of compression, so called video coding for machine (VCM). VCM differs from traditional video coding in that it is optimized for machine vision performance instead of human visual quality. In the feature compression track of MPEG-VCM, multi-scale features extracted from images are subject to compression. Recent feature compression works have demonstrated that the versatile video coding (VVC) standard-based approach can achieve a BD-rate reduction of up to 96% against MPEG-VCM feature anchor. However, it is still sub-optimal as VVC was not designed for extracted features but for natural images. Moreover, the high encoding complexity of VVC makes it difficult to design a lightweight encoder without sacrificing performance. To address these challenges, we propose a novel multi-scale feature compression method that enables both the end-to-end optimization on the extracted features and the design of lightweight encoders. The proposed model combines a learnable compressor with a multi-scale feature fusion network so that the redundancy in the multi-scale features is effectively removed. Instead of simply cascading the fusion network and the compression network, we integrate the fusion and encoding processes in an interleaved way. Our model first encodes a larger-scale feature to obtain a latent representation and then fuses the latent with a smaller-scale feature. This process is successively performed until the smallest-scale feature is fused and then the encoded latent at the final stage is entropy-coded for transmission. The results show that our model outperforms previous approaches by at least 52% BD-rate reduction and has$\times 5$to$\times 27$times less encoding time for object detection. It is noteworthy that our model can attain near-lossless task performance with only 0.002-0.003% of the uncompressed feature data size.
Yeongwoong Kim, Hyewon Jeong, Janghyun Yu, Younhee Kim, Jooyoung Lee 0004, Seyoon Jeong, Hui Yong Kim
IEEE Trans. Circuits Syst. Video Technol.7
2023 Towards Efficient Image Compression Without Autoregressive Models
abstract
Recently, learned image compression (LIC) has garnered increasing interest with its rapidly improving performance surpassing conventional codecs. A key ingredient of LIC is a hyperprior-based entropy model, where the underlying joint probability of the latent image features is modeled as a product of Gaussian distributions from each latent element. Since latents from the actual images are not spatially independent, autoregressive (AR) context based entropy models were proposed to handle the discrepancy between the assumed distribution and the actual distribution. Though the AR-based models have proven effective, the computational complexity is significantly increased due to the inherent sequential nature of the algorithm. In this paper, we present a novel alternative to the AR-based approach that can provide a significantly better trade-off between performance and complexity. To minimize the discrepancy, we introduce a correlation loss that forces the latents to be spatially decorrelated and better fitted to the independent probability model. Our correlation loss is proved to act as a general plug-in for the hyperprior (HP) based learned image compression methods. The performance gain from our correlation loss is ‘free’ in terms of computation complexity for both inference time and decoding time. To our knowledge, our method gives the best trade-off between the complexity and performance: combined with the Checkerboard-CM, it attains **90%** and when combined with ChARM-CM, it attains **98%** of the AR-based BD-Rate gains yet is around **50 times** and **30 times** faster than AR-based methods respectively
Muhammad Salman Ali, Yeongwoong Kim, Maryam Qamar, Sung-Chang Lim, Donghyun Kim 0017, Chaoning Zhang, Sung-Ho Bae, Hui Yong Kim
NeurIPS8
2023 MEDO: Minimizing Effective Distortions Only for Machine-Oriented Visual Feature Compression
abstract
In search for efficient feature compression technologies for machine consumption, MPEG recently issued a call for proposal (CfP) on feature compression for video coding for machine (FCVCM). One issue in feature compression is that the input feature maps generally have high redundancy in them. Various researches to reduce such redundancy have been made. For example, a recent study called L-MSFC (learnable multi-scale feature compression), which effectively combines multi-scale feature fusion and compression in an end-to-end learnable framework, showed up to 98% BD rate gain over the anchor model defined in the FCVCM CfP. Despite these advances in FCVCM, relation between distortions in feature maps and performance of vision tasks has stayed relatively unexplored. In this paper, we propose a novel loss function called MEDO (minimizing effective distortions only) based on our hypothesis that distortions below some threshold do not improve task performance. Experimental results on instance segmentation task show that our MEDO loss on top of L-MSFC improves the overall rate-mAP performance without compromising complexity. Being more practical for real-world uses, we also present an extension to L-MSFC for variable-rate support with a single model.
Curie Yoon, Dalhong Lim, Yeongwoong Kim, Hyewon Jeong, Hui Yong Kim, Jooyoung Lee 0004, Younhee Kim, Seyoon Jeong
VCIP5
2023 Visual Quality Assessment of Point Clouds Compared to Natural Reference Images
abstract
This paper proposes a point cloud (PC) visual quality assessment (VQA) framework that reflects the human visual system (HVS). The proposed framework compares natural images acquired using a digital camera and PC images generated via 2D projection in terms of appropriate objective quality evaluation metrics. Humans primarily consume natural images; thus, human knowledge is typically formed from natural images. Thus, natural images can be more reliable reference data than PC data. The proposed framework performs an image alignment process based on feature matching and image warping to use the natural images as a reference which enhances the similarities of the acquired natural and corresponding PC images. The framework facilitates identifying which objective VQA metrics can be used to reflect the HVS effectively. We constructed a database of natural images and three PC image qualities, and objective and subjective VQAs were conducted. The experimental result demonstrates that the acceptable consistency among different PC qualities appears in the metrics that compare the global structural similarity of images. We found that the SSIM, MAD, and GMSD achieved remarkable Spearman rank-order correlation coefficient scores of 0.882, 0.871, and 0.930, respectively. Thus, the proposed framework can reflect the HVS by comparing the global structural similarity between PC and natural reference images.
Aram Baek, Minseop Kim, Sohee Son, Sangwoo An, Jeongil Seo, Hui Yong Kim, Haechul Choi
J. Web Eng.6
2022 Modelling Surround-aware Contrast Sensitivity for HDR Displays
abstract
Abstract Despite advances in display technology, many existing applications rely on psychophysical datasets of human perception gathered using older, sometimes outdated displays. As a result, there exists the underlying assumption that such measurements can be carried over to the new viewing conditions of more modern technology. We have conducted a series of psychophysical experiments to explore contrast sensitivity using a state‐of‐the‐art HDR display, taking into account not only the spatial frequency and luminance of the stimuli but also their surrounding luminance levels. From our data, we have derived a novel surround‐aware contrast sensitivity function (CSF), which predicts human contrast sensitivity more accurately. We additionally provide a practical version that retains the benefits of our full model, while enabling easy backward compatibility and consistently producing good results across many existing applications that make use of CSF models. We show examples of effective HDR video compression using a transfer function derived from our CSF, tone‐mapping and improved accuracy in visual difference prediction.
Shinyoung Yi 0001, Daniel S. Jeon, Ana Serrano, Seyoon Jeong, Hui Yong Kim, Diego Gutierrez, Min H. Kim 0001
Comput. Graph. Forum5
2021 Tiny Drone Tracking Framework Using Multiple Trackers and Kalman-based Predictor
abstract
Unmanned aerial vehicles like drones are one of the key development technologies with many beneficial applications. As they have made great progress, security and privacy issues are also growing. Drone tacking with a moving camera is one of the important methods to solve these issues. There are various challenges of drone tracking. First, drones move quickly and are usually tiny. Second, images captured by a moving camera have illumination changes. Moreover, the tracking should be performed in real-time for surveillance applications. For fast and accurate drone tracking, this paper proposes a tracking framework utilizing two trackers, a predictor, and a refinement process. One tracker finds a moving target based on motion flow and the other tracker locates the region of interest (ROI) employing histogram features. The predictor estimates the trajectory of the target by using a Kalman filter. The predictor contributes to keeping track of the target even if the trackers fail. Lastly, the refinement process decides the location of the target taking advantage of ROIs from the trackers and the predictor. In experiments on our dataset containing tiny flying drones, the proposed method achieved an average success rate of 1.134 times higher than conventional tracking methods and it performed at an average run-time of 21.08 frames per second.
Sohee Son, Jeongin Kwon, Hui Yong Kim, Haechul Choi
J. Web Eng.3
2020 Edge-Preserving Reference Sample Filtering and Mode-Dependent Interpolation for Intra-Prediction
abstract
High Efficiency Video Coding is the latest video compression standard, which achieves the best coding performance up until now. Specifically, intra prediction is a tool that removes spatial redundancy in a single frame and then a predictor is generated from its neighboring reference samples based on a specific interpolation scheme. In this paper, we propose an edge-preserving intra reference sample filtering method using a bilateral filter, which is implemented as hardware-friendly. Two parameters of the bilateral filter are modeled by block size and mean amplitude of the pixel intensity. In addition, a mode-dependent interpolation scheme is proposed, which takes the directionality of angular predictions into account. The experimental results show that a BD rate-reduction of 0.63% can be achieved for all intra configurations by combining the two methods. We also demonstrate that the subjective quality of the reconstructed frames can be improved.
Hyunsuk Ko, Jungwon Kang, Hui Yong Kim
IEEE Trans. Circuits Syst. Video Technol.4
2019 Transform with residual rearrangement for HEVC intra coding
abstract
In video compression, transformation plays a significant role in the energy compaction of spatial domain data into frequency domain data. In the HEVC intra prediction, Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST) are used to concentrate the spatial residual signals into low-frequency components. DCT has shown good compression performance in both intra and inter residual coding, but when the spatial residual signals are not uniformly distributed, its coding efficiency decreases. This paper proposes a transform method that applies DCT or residual-rearranged DST to improve the coding efficiency in HEVC intra coding. The proposed method selects the best transform in terms of coding efficiency between the DCT and residual-rearranged DST for all block sizes. The experimental results show that, compared with the HEVC intra coding, the proposed method reduces the luma Bjontegaard Delta (BD) rates by 2.6%.
NamUk Kim, Sung-Chang Lim, Jungwon Kang, Hui Yong Kim, Yung Lyul Lee
Signal Process. Image Commun.4
2019 A New No-Reference Method for Judder Artifact Assessment
abstract
This paper proposes a new metric to measure judder artifacts of video sequences. The judder artifacts appear as non-smooth motions in hold-type displays when the frame rate is low and object motion is fast. To analyze the judder artifacts in video sequences, the proposed judder metric is defined by analyzing the effects and cross-relations of judder features in the video sequences. The judder features include spatial features of image gradients and temporal features of motion vectors and the frame rate. In addition, sensitivity of the human visual system (HVS) is considered to determine the perceptual judder artifacts because it has special characteristics of sensitivity to image brightness and contrast. Therefore, a sensitivity map of the HVS is used to mask the judder artifacts. Then, the degree of perceptual judder artifacts is estimated using the judder features and a regression model. The experimental results demonstrate that the proposed judder metric is highly correlated with the subjective assessment results.
Se Ri Oh, Seyoon Jeong, Pyeong Gang Heo, Hui Yong Kim, Hyun Wook Park
IEEE Trans. Circuits Syst. Video Technol.5
2018 Low complexity based ultra-high quality video compression method for multimedia-centric internet of things (IoT) services
Dong-San Jun, Hui Yong Kim
Multim. Tools Appl.2
2018 Understanding and Removal of False Contour in HEVC Compressed Images
abstract
A contour-like artifact called false contour is often observed in large smooth areas of decoded images and video. Without loss of generality, we focus on detection and removal of false contours resulting from the state-of-the-art High Efficiency Video Coding codec. First, we identify the cause of false contours by explaining the human perceptual experiences on them with specific experiments. Next, we propose a precise pixel-based false contour detection method based on the evolution of a false contour candidate (FCC) map. The number of points in the FCC map becomes fewer by imposing more constraints step by step. Special attention is paid to separating false contours from real contours such as edges and textures in the video source. Then, a decontour method is designed to remove false contours in the exact contour position while preserving edge/texture details. Extensive experimental results are provided to demonstrate the superior performance of the proposed false contour detection and removal method in both compressed images and videos.
Qin Huang 0006, Hui Yong Kim, Wen-Jiin Tsai, Seyoon Jeong, Jin Soo Choi, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.2
2018 GPU-based real-time super-resolution system for high-quality UHD video up-conversion
Dae Yeol Lee, Jooyoung Lee 0004, Ji-Hoon Choi, Jong-Ok Kim, Hui Yong Kim, Jin Soo Choi
J. Supercomput.5
2017 Measure and Prediction of HEVC Perceptually Lossy/Lossless Boundary QP Values
abstract
Evaluation of coding efficiency is traditionally modeled as a continuous rate-distortion (R-D) function, where the peak signal-to-noise ratio (PSNR) is adopted as the quality measure. Although the PSNR-versus-bitrate curve offers some useful tradeoff information between video quality and coding bit-rates, it does not take human perceptual experience into account. In this work, by following the recent image/video quality assessment framework based on the just-noticeable-difference (JND) notion, we conduct a subjective test for HEVC (High Efficiency Video Codec) video to measure the QP value that lies in the boundary of perceptually lossless and lossy coded bit streams for each human subject. This is also known as the first JND point. It is observed that the statistics of the first JND points of 30 subjects follows the normal distribution for a great majority of test sequences. Finally, a machine-learning approach is proposed to predict the mean of the group-based JND distribution based on extracted video features. It is shown by experimental results that the mean JND point can be predicted accurately.
Qin Huang 0006, Haiqiang Wang, Sung-Chang Lim, Hui Yong Kim, Seyoon Jeong, C.-C. Jay Kuo
DCC4
2017 Development of an ultra-HD HEVC encoder using SIMD implementation and fast encoding schemes for smart surveillance system
Dong-San Jun, Sung-Chang Lim, Hahyun Lee, Jungwon Kang, Jinwook Seok, Younhee Kim, Soon-Heung Jung, Hui Yong Kim, Jin Soo Choi
J. Supercomput.10
2015 Efficient In-Loop Filtering Across Tile Boundaries for Multi-Core HEVC Hardware Decoders With 4 K/8 K-UHD Video Applications
abstract
HEVC is a next generation video coding standard designed with modern coding techniques to be especially efficient for coding high-resolution video such as 4 K/8 K-ultra high- definition (UHD) video. Among the advanced coding tools of HEVC, tiles and wavefront parallel processing (WPP) have been newly adopted for parallel processing of such high-resolution (4 K/8 K-UHD) video. To realize UHD video services over portable devices with limited battery power, it is essential to implement multi-core-based and dedicated HEVC hardware decoders that support the tile- and wavefront-based parallel processing. By doing so, each frame is divided into a multiple number of picture partitions which can then be processed by multiple hardware decoder cores in parallel. However, in-loop filtering (ILF) at tile boundaries cannot be easily parallelized by a multi-core HEVC hardware decoder because of the data dependency between samples in different tiles. In this paper, an efficient control method for ILF across tile boundaries is proposed for multi-core HEVC hardware decoders. The proposed method does not require additional in-loop filters for ILF across the tile boundaries and it allows a decoder core to continue to process the next coding tree unit (CTU) without waiting for other decoders until they finish their ILF processing for the neighboring CTUs in other tiles. From experiments, we show the effectiveness of our ILF control method via a quad-core HEVC decoder for 4 K-UHD video implemented on a prototyping FPGA board.
Seunghyun Cho, Hyunmi Kim, Hui Yong Kim, Munchurl Kim
IEEE Trans. Multim.3
2014 Performance analysis of hierarchical transform coding with a large kernel for video codecs
abstract
In this study, the performance of hierarchical transform coding is analysed with design of an order‐16 integer transform kernel. The proposed hierarchical transform‐coding structure is constructed with a set of 4 × 4, 8 × 8 and 16 × 16 integer transforms of variable transform block sizes, which takes the advantages of both lower and higher transform kernels by flexibly adapting to varying image characteristics of video sequences with homogeneous and complex regions. The proposed hierarchical transform‐coding structure is implemented as an extension to H.264/advanced video coding joint model. The authors show the effectiveness of the hierarchical variable‐sized block transform scheme by analysing the quantisation effects and the correlation among neighbouring pixels in video sequences of different spatial resolutions. The experimental results show that: (i) the variable‐sized block transform scheme with the hierarchical structure is advantageous to the texture regions with strong local edges and (ii) the higher‐order‐16 integer transform kernel itself is more effective for the homogeneous texture regions, which are often encountered in higher resolution sequences. Therefore these two features can complementarily work in an rate‐distortion (RD) optimised manner for various characteristics of the input signals.
Bumshik Lee, Munchurl Kim, Hui Yong Kim, Jin Soo Choi
IET Image Process.3
2014 MC Complexity Reduction for Generalized P and B Pictures in HEVC
abstract
Motion compensation (MC) is a critical component in terms of computational complexity and memory bandwidth. The MC complexity of High Efficiency Video Coding (HEVC) for UHD contents significantly increased more than that of AVC/H.264. This paper reveals and analyzes a feature of generalized P and B pictures in HEVC and introduces a simple and effective method for MC complexity reduction that can be exploited at both the encoder and decoder without affecting the compression performance. The proposed method bypasses the \(L1\) interpolation process when the \(L0\) and \(L1\) motion information of a bipredicted block are identical. The simulation results show that the time reductions of 14.5% and 6.4% for the encoder and decoder, respectively, were achieved for LD-B configuration without any changes in coding results. The proposed method was adopted in the HEVC test model as a non-normative complexity reduction tool.
Kyung-Yong Kim, Hui Yong Kim, Jin Soo Choi, Gwang Hoon Park
IEEE Trans. Circuits Syst. Video Technol.2
2010 Fast block mode decision scheme for B-picture coding in H.264/AVC
abstract
The recent H.264/AVC video coding standard provides a higher coding efficiency than previous standards. H.264/AVC achieves a bit rate saving of more than 50 % with many new technologies, but it is computationally complex. Most of fast mode decision algorithms have focused on Baseline profile of H.264/AVC which does not consider B-picture coding. In this paper, a fast block mode decision scheme for B-pictures in High profile and Main profile is proposed to reduce the computational complexity for H.264/AVC. To reduce the block mode decision complexity in B-pictures of High profile, we use the SAD value after 16×16 block motion estimation. This SAD value is used for the classification feature to divide all block modes into some proper candidate search block modes. A differential mode allocation method is also used for the list (list 0, list 1) of B-slices based on the SAD value of the 16 × 16 block mode. The proposed algorithm shows the average speed-up factors of 41.9 ∼ 58.57% for IBBPBB sequences with a negligible bit increment and a minimal loss of image quality.
Jong-Ho Kim, Hyo-Sung Kim, Byung-Gyu Kim, Hui Yong Kim, Seyoon Jeong, Jin Soo Choi
ICIP4
2010 A hierarchical variable-sized block transform coding scheme for coding efficiency improvement on H.264/AVC
abstract
In this paper, a rate-distortion optimized variable block transform coding scheme is proposed based on a hierarchical structured transform for macroblock (MB) coding with a set of the order-4 and −8 integer cosine transform (ICT) kernels of H.264/AVC as well as a new order-16 ICT kernel. The set of order-4, −8 and −16 ICT kernels are applied for inter-predictive coding in square (4×4, 8×8 or 16×16) or non-square (16×8 or 8×16) transform for each MB in a hierarchical structured manner. The proposed hierarchical variable-sized block transform scheme using the order-16 ICT kernel achieves significant bitrate reduction up to 15%, compared to the High profile of H.264/AVC. Even if the number of candidates for the transform types increases, the encoding time can be reduced to average 4–6% over the H.264/AVC
Bumshik Lee, Jaeil Kim, Sangsoo Ahn, Munchurl Kim, Hui Yong Kim, Jong-Ho Kim, Jin Soo Choi
PCS5