Kwangpyo Choi

dblp:225/1469 · also Kwang Pyo Choi, Kwang-Pyo Choi · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
16since 2021 · last 2025
0000-0003-1638-5446ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 16 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Guided Detail Filter for AVM
abstract
This paper proposes an additional filter — that aims at enhancing the fidelity of local details in reconstructed frames and consequently improving coding efficiency — to the in-loop filtering process in the AOM next generation video codec (AVM). This involves correcting each sample through a rectification (i.e., scaling and clipping) process applied to the expected coding error obtained via a multidimensional table look-up with respect to a query vector that is derived by passing the sample’s spatial neighborhood and its intensity gradients through a multichannel filter. Guiding information comprising (multichannel) linear kernels and the table of expected coding errors is chosen according to edge classification and coding settings for optimal adaptation to local sample activity and coding context. Experimental results demonstrate that our approach achieves significant improvements over AVM v7.0.1, with PSNR-Y bitrate saving of -0.57% and -0.52% for All Intra and Random Access configurations respectively.
Khanh Quoc Dinh, Yangwoo Kim, Madhukar Budagavi, Rajan Joshi, Kwangpyo Choi
ICIP5
2025 SPCODEC: Split and Prediction for Neural Speech Codec
Liang Wen, Lizhong Wang, Yuxing Zheng, Weijing Shi, Kwangpyo Choi
INTERSPEECH5
2025 OpenAPV: Open Collaborative Innovation in Professional Video Ecosystem
abstract
Over the past few decades, significant resources have been invested in developing video codecs for the storage and delivery of video data. International standards development organizations, such as MPEG (Moving Picture Experts Group) and AOM (Alliance for Open Media), have fostered large-scale competitive collaboration among industry players. While open-source software developed alongside these video codec standards has accelerated verification of the standard and enabled rapid adoption, as it is used as a starting point for implementation and future research, this practice has not been fully applied to the ecosystem of mezzanine video codecs used in high-quality video capture and post-production. To address this gap, the industry's first open source project for collaborative innovation in professional video codecs, OpenAPV, has been established. The project aims to facilitate collaborative research and development of a royalty free, open source, and open standard video codec for professional use. This paper outlines the technical aspects of the video codec, together with the open source software implementation. The codec has been successfully tested across various platforms and shows excellent R-D (Rate-Distortion) performance and coding speed, even without hardware acceleration. Furthermore, it demonstrates robust quality resilience against multiple rounds of encoding and decoding cycles, which are common in video post-production. The open source project is hosted at https://github.com/AcademySoftwareFoundation/openapv.
Minsoo Park, Youngkwon Lim, Yangwoo Kim, Sam Richards, Kwangpyo Choi
ACM Multimedia6
2024 FT-CSR: Cascaded Frequency-Time Method for Coded Speech Restoration
abstract
Lossy speech codecs often introduce coding distortions such as coding noise and constrained bandwidth, which can affect the quality of the decoded speech. This paper proposes a method called FT-CSR, which is used for coded speech restoration. FT-CSR reduces coding noise and recovers missing frequencies sequentially using a cascaded frequency-time domain model. In experiments using the Opus codec, FT-CSR was found to be effective across bitrates ranging from 8 to 16 kbps and outperformed the baseline on both objective and subjective measurements. FT-CSR achieves a MOS-POLQA score of 3.6 or higher and improves MOS-POLQA by more than 0.23 when compared to decoded speech. The results of the subjective test show that FT-CSR can improve MOS by over 0.85 for decoded speech.
Liang Wen, Lizhong Wang, Yuxing Zheng, Weijing Shi, Kwangpyo Choi
ICME5
2023 Context-Based Trit-Plane Coding for Progressive Image Compression
abstract
Trit-plane coding enables deep progressive image compression, but it cannot use autoregressive context models. In this paper, we propose the context-based trit-plane coding (CTC) algorithm to achieve progressive compression more compactly. First, we develop the context-based rate reduction module to estimate trit probabilities of latent elements accurately and thus encode the trit-planes compactly. Second, we develop the context-based distortion reduction module to refine partial latent tensors from the trit-planes and improve the reconstructed image quality. Third, we propose a retraining scheme for the decoder to attain better rate-distortion tradeoffs. Extensive experiments show that CTC outperforms the baseline trit-plane codec significantly, e.g. by -14.84% in BD-rate on the Kodak loss less dataset, while increasing the time complexity only marginally. The source codes are available at https://github.com/seungminjeon-github/CTC.
Seungmin Jeon, Kwangpyo Choi, Youngo Park, Chang-Su Kim 0001
CVPR2
2023 Learned Video Coding with Motion Compensation Mixture Model
abstract
Learned video coding employs explicit motion compensation (MC) with neural networks to predict the original frame from its reference frame and to compress its residual from the predicted frame, where neural networks are optimized with rate-distortion trade-offs. However, good predictions are hard to find or even do not exist due to fast motions, dis-occlusions, and coding errors of the reference frame. To avoid the problem of carrying false edges/details caused by inaccurate optical flow in the predicted frame to the residual, we propose a dynamic mixture of explicit and implicit motion compensations, where implicitness means that the encoding and decoding of the original frame are conditioned on the predicted frame in pixel and latent domains, respectively. The proposed mixture model saves up to 30% bitrate over the baseline and achieves state-of-the-art performance.
Khanh Quoc Dinh, Kwangpyo Choi
ICASSP2
2023 Distortion-Aware Convolutional Neural Network-Based Interpolation Filter for AVS3
abstract
Motion compensation is a key technology in video coding for removing the temporal redundancy between video frames. Considering the incompatibility between traditional interpolation filters and diversified video content, the inter prediction method still has considerable room for improvement. This paper proposed a distortion-aware convolutional neural network-based interpolation filter (DA-NNIF) to further improve the interpolation prediction accuracy of sub-pixels with one model. Distortion parameters are introduced into the proposed network to reflect the quantization noise of reference frames. The experimental result shows that the proposed method achieves on average 1.47 % BD-rate reduction on Y component for ClassB, ClassC and ClassD sequences under the random access configuration of AVS3.
Liang Wen, Lizhong Wang, Yinji Piao, Weijing Shi, Kwangpyo Choi
ICASSP6
2023 End-to-End Single-Frame Image Signal Processing for High Dynamic Range Scenes
abstract
This paper considers photography of high dynamic range scenes containing mixtures of shadows and highlights on mobile phones. Multi-frame merging constructs a high-quality image at the cost of capturing multiple frames of the same scene. Contrarily, end-to-end optimized image signal processing (E2EISP) produces an enhanced image from a single-frame Bayer array. This paper combines the merits of the two approaches by using labels of high-quality multi-frame merged images to train E2EISP with a novel neural network architecture composed of a multi-head mixture of brightness enhancement for accurately processing shadows/highlights and a multi-head mixture of image processing featured camera settings of white balance and color correction for a proper color generation. We also proposed a combination of supervised, unsupervised, and generative adversarial losses for brightness, edge, and detail enhancement. Experimental results show that the proposed single-frame ISP produces enhanced images and outperforms state-of-the-art methods.
Khanh Quoc Dinh, Kwangpyo Choi
WACV2
2022 DPICT: Deep Progressive Image Compression Using Trit-Planes
abstract
We propose the deep progressive image compression using trit-planes (DPICT) algorithm, which is the first learning-based codec supporting fine granular scalability (FGS). First, we transform an image into a latent tensor using an analysis network. Then, we represent the latent tensor in ternary digits (trits) and encode it into a compressed bitstream trit-plane by trit-plane in the decreasing order of significance. Moreover, within each trit-plane, we sort the trits according to their rate-distortion priorities and transmit more important information first. Since the compression network is less optimized for the cases of using fewer tritplanes, we develop a postprocessing network for refining reconstructed images at low rates. Experimental results show that DPICT outperforms conventional progressive codecs significantly, while enabling FGS transmission. Codes are available at https://github.com/jaehanlee-mcl/DPICT.
Jae-Han Lee, Seungmin Jeon, Kwangpyo Choi, Youngo Park, Chang-Su Kim 0001
CVPR3
2022 Learning Local Implicit Fourier Representation for Image Warping
Kwangpyo Choi, Kyong Hwan Jin
ECCV (18)2
2022 Low-Complexity Scaler Based on Convolutional Neural Networks for Adaptive Video Streaming
abstract
Adaptive streaming service nowadays, became an essential key technology for delivering videos over internet protocol network. Due to limitations and fluctuations in internet bandwidth, scaler has become essential for streaming service. Recently, learning-based scalers have greatly improved performance compared to conventional methods in visual quality. However, performance is not guaranteed for compatibility of conventional scaler and real-time processing is difficult due to high complexity. In this paper, we present a low complexity scaler based on convolutional neural networks called Video Scale Network (VSN). The proposed method has a simple structure with real-time processing and a loss function compatible of conventional scaler. Furthermore, we propose a learning method to improve performance using only a single network. Experimental results on the AOM test sequences reveal improved performance by using the proposed method compared to the conventional method and real-time processing in the decoder of AV1 codec is also able to be done.
Dongkyu Kim, Chaeeun Lee, Youngo Park, Kwangpyo Choi
ICIP6
2022 Multi-Stage Progressive Audio Bandwidth Extension
abstract
Audio bandwidth extension can enhance subjective sound quality by increasing bandwidth of audio signal. This paper presents a novel multi-stage progressive method for time domain causal bandwidth extension. Each stage of the progressive model contains a light weight scale-up module to generate high frequency signal and a supervised attention module to guide features propagating between stages. Time-frequency two-step training method with weighted loss for progressive output is adopted to supervise bandwidth extension performance improves along stages. Test results show that multi-stage model can improve both objective results and perceptual quality progressively. The multi-stage progressive model makes bandwidth extension performance adjustable according to energy consumption, computing capacity and user preferences.
Liang Wen, Lizhong Wang, Kwangpyo Choi
SLT4
2022 Depth-Wise Split Unit Coding Order for Video Compression
abstract
In this paper, we propose a depth-wise flexible block processing method called split unit coding order (SUCO) for video coding. Conventionally, block-based image and video compression frameworks always apply raster scans to process blocks in order from left to right. Owing to the fixed coding order, the available information for prediction in the coding blocks is limited to adjacent blocks on the left and top. To address the limitations of block-based images and video compression frameworks, the proposed SUCO provides more flexibility in handling coding block sequences than predicted on the left and top, and thus coding blocks can take advantage of adjacent right information, such as reconstructed pixels and motion information. The flexibility is achieved by depth-wise signaling of preferred coding order for the given partitions. The experiment results demonstrate that the proposed SUCO can effectively improve the coding efficiency of both intra and inter prediction in the latest video coding standards.
Yinji Piao, Kiho Choi, Kwangpyo Choi, Minsoo Park
IEEE Trans. Circuits Syst. Video Technol.3
2021 High-Frequency Preserving Image Downscaler
abstract
The goal of downscaling an image is to reduce its resolution to a lower resolution, while maintaining the visual characteristics of the original image. Recent learning-based algorithms have shown great improvement over conventional methods in preserving high-frequency information. However, they continue to suffer from various artifacts due to the lack of true high-low resolution image pairs in their training datasets. In this paper, an unsupervised image downscaler that preserves the high frequency content of the original image based on an autoencoder is presented. Specifically, the image downscaler is obtained by extracting the decoder of the developed autoencoder. Furthermore, we propose a preprocessing step that enables image downscaler to any arbitrary scales. Experimental results on five benchmark datasets reveal the qualitative and quantitative superiority of the proposed method at various scales compared to other methods.
Soo Min Kang, Kwangpyo Choi, Youngo Park, Chaeeun Lee, Jongseok Lee
ICIP3
2021 Split Unit Coding Order for Video Coding
abstract
This paper presents a flexible block processing order, called split unit coding order (SUCO), for video coding. In the conventional block-based image and video compression, the largest blocks are processed in a raster scan order while further partitioned blocks are processed in a z-scan order. Due to the fixed coding order, the information that can be exploited for prediction at coding block is limited to the left and above neighbors. Beyond the traditional prediction from left and above, the proposed SUCO allows more flexible processing order for the partitioned blocks so that the coding blocks could also utilize right neighboring information, such as reconstructed pixels and motion information, more adaptively. The impact of proposed coding order has been verified in the several video coding platforms: 2.1% and 2.1% BD-rate reduction on average in AI and RA configuration over HM12.1; 1.6% BD-rate reduction on average in RA over JEM3.1; 1.0% BD-rate reduction on average in RA over ETM4.1.
Yinji Piao, Kiho Choi, Minsoo Park, Kwangpyo Choi
ICME5
2021 X-net: A Joint Scale Down and Scale Up Method for Voice Call
Liang Wen, Lizhong Wang, Yuxing Zheng, Youngo Park, Kwangpyo Choi
Interspeech6
2020 Video Codec Using Flexible Block Partitioning and Advanced Prediction, Transform and Loop Filtering Technologies
abstract
This paper describes a joint response to the Call for Proposals by Samsung, Huawei, GoPro, and HiSilicon on Video Compression with Capability beyond HEVC/H.265, jointly issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11 (MPEG). In the proposed codec, the coding framework supports hierarchical splitting with binary and ternary trees and flexible coding order representations. Additionally, novel compression tools on inter/intra prediction, in-loop filtering, and entropy coding have been proposed. The proposed compression scheme provides significantly higher compression capability than the state-of-the-art HEVC/H.265 standard for SDR (Standard Dynamic Range) category while maintaining complexity acceptable for emerging applications. When all the proposed algorithmic tools are used, the proposed video codec achieves approximately 40% bit-saving for the SDR cetegory on average compared to HEVC/H.265 anchor.
Kiho Choi, Jianle Chen, Haitao Yang 0001, Woongil Choi, Sergey Ikonin, Yinji Piao, Semih Esenlik, Minsoo Park, Ye-Kui Wang, Narae Choi, Yin Zhao, Seungsoo Jeong, Anish Tamse, Alexey Filippov, Heechul Yang, Junghye Min, Roman Chernyak, Bora Jin, Anand Meher Kotra, Sunil Lee, Han Gao 0001, Chanyul Kim, Timofey Solovyev, Kwangpyo Choi, Vasily Rufitskiy, Maxim Sychev, Jeonghoon Park
IEEE Trans. Circuits Syst. Video Technol.27
2018 Only-Reference Video Quality Assessment for Video Coding Using Convolutional Neural Network
abstract
Conventional video quality assessment methods are either full-, reduced-, or no-reference methods that need to access decoded videos. Hence, to calculate quality of decoded video in video coding regarding an image/video quality metric, complete encoding and decoding have to executed, which is computationally expensive. To address this problem, we propose to estimate quality of decoded videos from the original video only (i.e., only-reference) using convolutional neural network, as if the original video is encoded using a range of quantization parameter. The proposed network is shallow and can be trained to estimate various video quality metrics. Furthermore, among potential rate control applications using the proposed network, we demonstrate achieving a targeted decoded-video quality by selecting a proper quantization parameter before actually encoding.
Khanh Quoc Dinh, Jongseok Lee, Youngo Park, Kwangpyo Choi, Jeonghoon Park
ICIP5
2017 Omnidirectional Video Quality Metrics and Evaluation Process
abstract
Widespread of virtual reality technologies across entertainment formats has created a diverse infrastructure of related technologies as head-mount displays, dome screens and virtual reality multi-camera platforms. As omnidirectional content is processing pipeline is completely different form conventional planar video and involves multiple conversion steps which affect quality in a different way. As a result of our research we propose objective quality estimation methodology and a set of tools to evaluate different projection methods and coding tools for omnidirectional video content.
Vladyslav Zakharchenko, Kwangpyo Choi, Elena Alshina, Jeonghoon Park
DCC2
2014 Video saliency detection based on spatiotemporal feature learning
abstract
A video saliency detection algorithm based on feature learning, called ROCT, is proposed in this work. To detect salient regions, we design multiple spatiotemporal features and combine those features using a support vector machine (SVM). We extract the spatial features of rarity, compactness, and center prior by analyzing the color distribution in each image frame. Also, we obtain the temporal features of motion intensity and motion contrast to identify visually important motions. We train an SVM classifier using the spatiotemporal features extracted from training video sequences. Finally, we compute the visual saliency of each patch in an input sequence using the trained classifier. Experimental results demonstrate that the proposed algorithm provides more accurate and reliable results of saliency detection than conventional algorithms.
Se-Ho Lee, Jin-Hwan Kim, Kwangpyo Choi, Jae-Young Sim, Chang-Su Kim 0001
ICIP3