Wen-Hsiao Peng

dblp:62/2384 · DBLP profile ↗
← Back
101ranked-venue papers
8as first author
57since 2021 · last 2026
0000-0002-4421-8031ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 77 · 6 first-author · 43 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 12 since 2021Systems, architecture and hardware · 16 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Energy-Efficient Neural Video Coding via a High-Throughput Hardware Design for Pointwise Convolution and LUT-based WSiLU
Denis Maass, Vanessa Aldrighi, Ruhan A. Conceição, Wen-Hsiao Peng, Luciano Volcan Agostini, Marcelo Schiavon Porto
ISCAS4
2026 A Rate-Distortion-Complexity Analysis of Neural Video CODECs
Ricardo L. de Queiroz, Diogo C. Garcia, Yi-Hsin Chen, Ruhan Conceição, Wen-Hsiao Peng, Luciano V. Agostini
ISCAS5
2026 TED-4DGS: Temporally Activated and Embedding-based Deformation for 4DGS Compression
abstract
Building on the success of 3D Gaussian Splatting (3DGS) in static 3D scene representation, its extension to dynamic scenes-commonly referred to as 4DGS or dynamic 3DGS- has attracted increasing attention. However, designing more compact, efficient deformation schemes together with rate-distortion-optimized compression strategies for dynamic 3DGS representations remains an underexplored area. Prior methods either rely on space-time 4DGS with overspecified, short-lived Gaussian primitives or on canonical 3DGS with deformation that lacks explicit temporal control. To address this, we present TED-4DGS, a temporally activated and embedding-based deformation scheme for rate-distortion- optimized 4DGS compression that unifies the strengths of both families. TED-4DGS is built on a sparse anchor-based 3DGS representation. Each canonical anchor is assigned with learnable temporal-activation parameters to specify its appearance and disappearance transitions over time, while a lightweight per-anchor temporal embedding queries a shared deformation bank to produce anchor-specific deformation. For rate-distortion compression, we incorporate an implicit neural representation (INR)-based hyperprior to model anchor attribute distributions, along with a channelwise autoregressive model to capture intra-anchor correlations. With these novel elements, our scheme achieves the state-of-the-art rate-distortion performance on several commonly used real-world datasets. To the best of our knowledge, this work represents one of the first attempts to pursue a rate-distortion-optimized compression framework for dynamic 3DGS representations.
Cheng-Yuan Ho, Hebi Yang, Jui-Chiu Chiang, Yu-Lun Liu 0001, Wen-Hsiao Peng
WACV5
2026 MEGA-PCC: A Mamba-based Efficient Approach for Joint Geometry and Attribute Point Cloud Compression
abstract
Joint compression of point cloud geometry and attributes is essential for efficient 3D data representation. Existing methods often rely on post-hoc recoloring procedures and manually tuned bitrate allocation between geometry and attribute bitstreams in inference, which hinders end-to-end optimization and increases system complexity. To overcome these limitations, we propose MEGA-PCC, a fully end-to-end, learning-based framework featuring two specialized models for joint compression. The main compression model employs a shared encoder that encodes both geometry and attribute information into a unified latent representation, followed by dual decoders that sequentially reconstruct geometry and then attributes. Complementing this, the Mamba-based Entropy Model (MEM) enhances entropy coding by capturing spatial and channel-wise correlations to improve probability estimation. Both models are built on the Mamba architecture to effectively model long-range dependencies and rich contextual features. By eliminating the need for recoloring and heuristic bitrate tuning, MEGA-PCC enables data-driven bitrate allocation during training and simplifies the overall pipeline. Extensive experiments demonstrate that MEGA-PCC achieves superior rate-distortion performance and runtime efficiency compared to both traditional and learning-based baselines, offering a powerful solution for AI-driven point cloud compression.
Kai Hsiang Hsieh, Monyneath Yim, Wen-Hsiao Peng, Jui-Chiu Chiang
WACV3
2026 milliMamba: Specular-Aware Human Pose Estimation via Dual mmWave Radar with Multi-Frame Mamba Fusion
abstract
Millimeter-wave radar offers a privacy-preserving and lighting-invariant alternative to RGB sensors for Human Pose Estimation (HPE) task. However, the radar signals are often sparse due to specular reflection, making the extraction of robust features from radar signals highly challenging. To address this, we present milliMamba, a radar-based 2D human pose estimation framework that jointly models spatio-temporal dependencies across both the feature extraction and decoding stages. Specifically, given the high dimensionality of radar inputs, we adopt a Cross-View Fusion Mamba encoder to efficiently extract spatio-temporal features from longer sequences with linear complexity. A Spatio-Temporal-Cross Attention decoder then predicts joint coordinates across multiple frames. Together, this spatio-temporal modeling pipeline enables the model to leverage contextual cues from neighboring frames and joints to infer missing joints caused by specular reflections. To reinforce motion smoothness, we incorporate a velocity loss alongside the standard keypoint loss during training. Experiments on the TransHuPR and HuPR datasets demonstrate that our method achieves significant performance improvements, exceeding the baselines by 11.0 AP and 14.6 AP, respectively, while maintaining reasonable complexity. Code: https://github.com/NYCU-MAPL/milliMamba
Niraj Prakash Kini, Shiau-Rung Tsai, Guan-Hsun Lin, Wen-Hsiao Peng, Ching-Wen Ma, Jenq-Neng Hwang
WACV4
2026 CSGaussian: Progressive Rate-Distortion Compression and Segmentation for 3D Gaussian Splatting
abstract
We present the first unified framework for rate-distortion-optimized compression and segmentation of 3D Gaussian Splatting (3DGS). While 3DGS has proven effective for both real-time rendering and semantic scene understanding, prior works have largely treated these tasks independently, leaving their joint consideration unexplored. Inspired by recent advances in rate-distortion-optimized 3DGS compression, this work integrates semantic learning into the compression pipeline to support decoder-side applications–such as scene editing and manipulation–that extend beyond traditional scene reconstruction and view synthesis. Our scheme features a lightweight implicit neural representation-based hyperprior, enabling efficient entropy coding of both color and semantic attributes while avoiding costly grid-based hyperprior as seen in many prior works. To facilitate compression and segmentation, we further develop compression-guided segmentation learning, consisting of quantization-aware training to enhance feature separability and a quality-aware weighting mechanism to suppress unreliable Gaussian primitives. Extensive experiments on the LERF and 3D-OVS datasets demonstrate that our approach significantly reduces transmission cost while preserving high rendering quality and strong segmentation performance.
Yu-Jen Tseng, Chia-Hao Kao, Jing-Zhong Chen, Alessandro Gnutti, Shao-Yuan Lo, Yen-Yu Lin, Wen-Hsiao Peng
WACV7
2026 ExReg: Wide-range Photo Exposure Correction via a Multi-dimensional Regressor with Attention
abstract
Photo exposure correction is widely investigated, but fewer studies focus on correcting under- and over-exposed images simultaneously. Three issues remain open to handle and correct both under- and over-exposed images in a unified way. First, a locally adaptive exposure adjustment may be more flexible instead of learning a global mapping. Second, it is an ill-posed problem to determine the suitable exposure values locally. Third, photos with the same content but different exposures may not reach consistent adjustment results. To this end, we proposed a novel exposure correction network, ExReg, to address the challenges by formulating exposure correction as a multi-dimensional regression process. Given an input image, a compact multi-exposure generation network is introduced to generate images with different exposure conditions for multi-dimensional regression and exposure correction in the next stage. An auxiliary module is designed to predict the region-wise exposure values, guiding the proposed Encoder–Decoder ANP (Attentive Neural Process) to regress the final corrected image. The experimental results show that ExReg can generate well-exposed results and outperform the SOTA method in PSNR for extensive exposure problems. Furthermore, the processing speed, with 0.05 seconds per image on an RTX 3090, is efficient. When tested on the same image under various exposure levels, ExReg also yields results that are visually consistent and physically accurate.
Huu-Phu Do, Hao-Chien Hsueh, Tzu-Hao Chiang, Chi Han Chen, Wen-Hsiao Peng
ACM Trans. Intell. Syst. Technol.5
2025 Learning Optimal Linear Block Transform by Rate Distortion Minimization
abstract
The rise of deep learning has spurred advancements in image compression, with end-to-end learned systems gaining traction. However, their adoption in standard frameworks is limited, as they require a major overhaul of existing hardware designed for traditional methods. Moreover, their computational complexity, especially on the decoder side, remains significantly higher than conventional codecs. Consequently, optimizing traditional codecs remains a key research focus.
Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi
DCC3
2025 HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video Coding
abstract
Most frame-based learned video codecs can be interpreted as recurrent neural networks (RNNs) propagating reference information along the temporal dimension. This work revisits the limitations of the current approaches from an RNN perspective. The output-recurrence methods, which propagate decoded frames, are intuitive but impose dual constraints on the output decoded frames, leading to suboptimal rate-distortion performance. In contrast, the hidden-to-hidden connection approaches, which propagate latent features within the RNN, offer greater flexibility but require large buffer sizes. To address these issues, we propose HyTIP, a learned video coding framework that combines both mechanisms. Our hybrid buffering strategy uses explicit decoded frames and a small number of implicit latent features to achieve competitive coding performance. Experimental results show that our HyTIP outperforms the sole use of either output-recurrence or hidden-to-hidden approaches. Furthermore, it achieves comparable performance to state-of-the-art methods but with a much smaller buffer size, and outperforms VTM 17.0 (Low-delay B) in terms of PSNR-RGB and MS-SSIM-RGB. The source code of HyTIP is available at https://github.com/NYCU-MAPL/HyTIP.
Yi-Hsin Chen, Yi-Chen Yao, Kuan-Wei Ho, Chun-Hung Wu, Huu-Tai Phung, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng
ICCV8
2025 MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding
Huu-Tai Phung, Zong-Lin Gao, Yi-Chen Yao, Kuan-Wei Ho, Yi-Hsin Chen, Yu-Hsiang Lin, Alessandro Gnutti, Wen-Hsiao Peng
ICCV8
2025 Learned Hybrid Video Coding for Human Perception and Multiple Machine Vision Tasks
abstract
In this work, we present a learned multi-task video codec that is optimized for human and machine vision. The codec consists of an encoder that maps images from the pixel domain to a latent representation and multiple decoders that map the latent to either an image for human consumption or multiple task-specific features for different machine vision tasks. This allows a single bitstream to be used for multiple tasks while also reducing the decoder complexity for machine vision tasks. Unlike most learned codecs, our method performs inter-coding at the latent level instead of the pixel domain. Experiments show that the proposed method achieves a compression performance for machine vision tasks comparable to other multi-task codecs designed for machine vision only, while also providing video reconstruction. The code is available at https://github.com/GreenAutoML4FAS/HybridMultiTaskCoding.
Martin Benjak, Saifullah Khan, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann
ICIP4
2025 Warm Diffusion: Recipe for Blur-Noise Mixture Diffusion Models
abstract
Diffusion probabilistic models have achieved remarkable success in generative tasks across diverse data types. While recent studies have explored alternative degradation processes beyond Gaussian noise, this paper bridges two key diffusion paradigms: hot diffusion, which relies entirely on noise, and cold diffusion, which uses only blurring without noise. We argue that hot diffusion fails to exploit the strong correlation between high-frequency image detail and low-frequency structures, leading to random behaviors in the early steps of generation. Conversely, while cold diffusion leverages image correlations for prediction, it neglects the role of noise (randomness) in shaping the data manifold, resulting in out-of-manifold issues and partially explaining its performance drop. To integrate both strengths, we propose Warm Diffusion, a unified Blur-Noise Mixture Diffusion Model (BNMD), to control blurring and noise jointly. Our divide-and-conquer strategy exploits the spectral dependency in images, simplifying score model estimation by disentangling the denoising and deblurring processes. We further analyze the Blur-to-Noise Ratio (BNR) using spectral analysis to investigate the trade-off between model learning dynamics and changes in the data manifold. Extensive experiments across benchmarks validate the effectiveness of our approach for image generation.
Hao-Chien Hsueh, Wen-Hsiao Peng
ICLR2
2025 Bridging Compressed Image Latents and Multimodal Large Language Models
abstract
This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hinders deployment on resource-constrained end devices. While cloud-hosted MLLMs could be available, transmitting raw, uncompressed images captured by end devices to the cloud requires an efficient image compression system. To address this, we focus on emerging neural image compression and propose a novel framework with a lightweight transform-neck and a surrogate loss to adapt compressed image latents for MLLM-based vision tasks. Given the huge scale of MLLMs, our framework excludes the entire downstream MLLM except part of its visual encoder from training our system. This stands out from most existing coding for machine approaches that involve downstream networks in training and thus could be impractical when the networks are MLLMs. The proposed framework is general in that it is applicable to various MLLMs, neural image codecs, and multiple application scenarios, where the neural image codec can be (1) pre-trained for human perception without updating, (2) fully updated for joint human and machine perception, or (3) fully updated for only machine perception. Extensive experiments on different neural image codecs and various MLLMs show that our method achieves great rate-accuracy performance with much less complexity.
Chia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen, Alessandro Gnutti, Shao-Yuan Lo, Wen-Hsiao Peng, Riccardo Leonardi
ICLR7
2025 CAT-3DGS: A Context-Adaptive Triplane Approach to Rate-Distortion-Optimized 3DGS Compression
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising 3D representation. Much research has been focused on reducing its storage requirements and memory footprint. However, the needs to compress and transmit the 3DGS representation to the remote side are overlooked. This new application calls for rate-distortion-optimized 3DGS compression. How to quantize and entropy encode sparse Gaussian primitives in the 3D space remains largely unexplored. Few early attempts resort to the hyperprior framework from learned image compression. But, they fail to utilize fully the inter and intra correlation inherent in Gaussian primitives. Built on ScaffoldGS, this work, termed CAT-3DGS, introduces a context-adaptive triplane approach to their rate-distortion-optimized coding. It features multi-scale triplanes, oriented according to the principal axes of Gaussian primitives in the 3D space, to capture their inter correlation (i.e. spatial correlation) for spatial autoregressive coding in the projected 2D planes. With these triplanes serving as the hyperprior, we further perform channel-wise autoregressive coding to leverage the intra correlation within each individual Gaussian primitive. Our CAT-3DGS incorporates a view frequency-aware masking mechanism. It actively skips from coding those Gaussian primitives that potentially have little impact on the rendering quality. When trained end-to-end to strike a good rate-distortion trade-off, our CAT-3DGS achieves the state-of-the-art compression performance on the commonly used real-world datasets.
Yu-Ting Zhan, Cheng-Yuan Ho, Hebi Yang, Yi-Hsin Chen, Jui-Chiu Chiang, Yu-Lun Liu 0001, Wen-Hsiao Peng
ICLR7
2025 Conditional Residual Coding with Explicit-Implicit Temporal Buffering for Learned Video Compression
abstract
This work proposes a hybrid, explicit-implicit temporal buffering scheme for conditional residual video coding. Recent conditional coding methods propagate implicit temporal information for inter-frame coding, demonstrating superior coding performance to those relying exclusively on previously decoded frames (i.e. the explicit temporal information). However, these methods require substantial memory to store a large number of implicit features. This work presents a hybrid buffering strategy. For inter-frame coding, it buffers one previously decoded frame as the explicit temporal reference and a small number of learned features as implicit temporal reference. Our hybrid buffering scheme for conditional residual coding outperforms the single use of explicit or implicit information. Moreover, it allows the total buffer size to be reduced to the equivalent of two video frames with a negligible performance drop on 2K video sequences. The ablation experiment further sheds light on how these two types of temporal references impact the coding performance.
Yi-Hsin Chen, Kuan-Wei Ho, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng
ICME5
2025 Cross-Platform Neural Video Coding: A Case Study
abstract
In this paper, we first show that current learning-based video codecs, specifically the SSF codec, are not suitable for real-world applications due to the mismatch between the encoder and decoder caused by floating-point round-off errors. To address this issue, we propose the static quantization of the hyper prior decoding path. The quantization parameters are determined through an exhaustive search of all possible combinations of observers and quantization schemes from PyTorch. For the SSF codec, when encoding and decoding on different machines, the proposed solution effectively mitigates the mismatch issue and enhances compression efficiency results by preventing severe image quality degradation. When encoding and decoding are performed on the same machine, it constrains the average BD-rate increase to 9.93% and 9.02% for UVG and HEVC-B sequences, respectively.
Ruhan A. Conceição, Marcelo Schiavon Porto, Wen-Hsiao Peng, Luciano Volcan Agostini
ISCAS3
2025 MaskCRT-B: Masked Conditional Residual Transformer for Learned B-frame Coding
abstract
This paper proposes a learned hierarchical B-frame coding scheme in response to the Grand Challenge on Neural Network-based Video Coding at ISCAS 2025. Recently, masked conditional residual coding emerged as an attractive alternative to the existing inter-frame coding frameworks, including residual coding, conditional coding, and conditional residual coding. In this work, we propose masked conditional residual B-frame coding, termed MaskCRT-B, for YUV420 videos. It features an asymmetric codec architecture that includes one joint YUV encoder and two separate Y and UV decoders. Moreover, it incorporates a bi-directional adaptive fusion module that refines the bi-directional feature maps to better tackle the prediction of the occluded and dis-occluded regions within the input video. MaskCRT-B presents a significant advancement in learned B-frame coding, outperforming the state-of-the-art conditional B-frame codec from the Grand Challenge at ISCAS 2024.
Zong-Lin Gao, Yi-Chen Yao, Kuan-Wei Ho, Yi-Hsin Chen, Wen-Hsiao Peng
ISCAS5
2025 Scalable COOL-CHIC: Dual-Resolution Images from a Single Bitstream
Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann
PCS3
2025 Rate Distortion Learned Transform For Image Compression
Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi
PCS3
2025 A Cross-Framework Study of Temporal Information Buffering Strategies for Learned Video Compression
Kuan-Wei Ho, Yi-Hsin Chen, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng
PCS5
2025 Exploring Autoregressive Vision Foundation Models for Image Compression
Huu-Tai Phung, Yu-Hsiang Lin, Yen-Kuan Ho, Wen-Hsiao Peng
PCS4
2025 Progressive COOL-CHIC: Efficient Decoding for Dual-Resolution Images
abstract
In this work, we propose Progressive Cool-Chic (PCC), a scalable overfitted neural image codec that can decode an image at two different resolutions from a single bitstream. Experiments show that our method reduces the necessary bitrate to encode two representations of the same image by up to 31.54% in terms of BD-rate compared to encoding both representations independently using Cool-Chic while also decreasing the necessary decoding time. The bitstream is structured in a way that the low-resolution image can already be decoded, when only a part of the bitstream has been received. The code is available at https://github.com/mbenjak/progressive-CC.
Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann
VCIP3
2024 TransHuPR: Cross-View Fusion Transformer for Human Pose Estimation Using mmWave Radar
Niraj Prakash Kini, Ruey-Horng Shiue, Ryan Chandra, Wen-Hsiao Peng, Ching-Wen Ma, Jenq-Neng Hwang
BMVC4
2024 Omra: Online Motion Resolution Adaptation To Remedy Domain Shift in Learned Hierarchical B-Frame Coding
abstract
Learned hierarchical B-frame coding aims to leverage bidirectional reference frames for better coding efficiency. However, the domain shift between training and test scenarios due to dataset limitations poses a challenge. This issue arises from training the codec with small groups of pictures (GOP) but testing it on large GOPs. Specifically, the motion estimation network, when trained on small GOPs, is unable to handle large motion at test time, incurring a negative impact on compression performance. To mitigate the domain shift, we present an online motion resolution adaptation (OMRA) method. It adapts the spatial resolution of video frames on a per-frame basis to suit the capability of the motion estimation network in a pre-trained B-frame codec. Our OMRA is an online, inference technique. It need not re-train the codec and is readily applicable to existing B-frame codecs that adopt hierarchical bi-directional prediction. Experimental results show that OMRA significantly enhances the compression performance of two state-of-the-art learned B-frame codecs on commonly used datasets.
Zong-Lin Gao, Sang NguyenQuang, Wen-Hsiao Peng, Xiem HoangVan
ICIP3
2024 Lidar Depth Map Guided Image Compression Model
abstract
The incorporation of LiDAR technology into some high-end smartphones has unlocked numerous possibilities across various applications, including photography, image restoration, augmented reality, and more. In this paper, we introduce a novel direction that harnesses LiDAR depth maps to enhance the compression of the corresponding RGB camera images. To the best of our knowledge, this represents the initial exploration in this particular research direction. Specifically, we propose a Transformer-based learned image compression system capable of achieving variable-rate compression using a single model while utilizing the LiDAR depth map as supplementary information for both the encoding and decoding processes. Experimental results demonstrate that integrating LiDAR yields an average PSNR gain of 0.83 dB and an average bitrate reduction of 16% as compared to its absence.
Alessandro Gnutti, Stefano Della Fiore, Mattia Savardi, Yi-Hsin Chen, Riccardo Leonardi, Wen-Hsiao Peng
ICIP6
2024 Conditional Variational Autoencoders for Hierarchical B-frame Coding
abstract
In response to the Grand Challenge on Neural Network-based Video Coding at ISCAS 2024, this paper proposes a learned hierarchical B-frame coding scheme. Most learned video codecs concentrate on P-frame coding for the RGB content, while B-frame coding for the YUV420 content remains largely under-explored. Some early works explore Conditional Augmented Normalizing Flows (CANF) for B-frame coding. However, they suffer from high computational complexity because of stacking multiple variational autoencoders (VAE) and using separate Y and UV codecs. This work aims to develop a lightweight VAE-based B-frame codec in a conditional coding framework. It features (1) extracting multi-scale features for conditional motion and inter-frame coding, (2) performing frame-type adaptive coding for better bit allocation, and (3) a lightweight conditional VAE backbone that encodes YUV420 content by a simple conversion into YUV444 content for joint Y and UV coding. Experimental results confirms its superior compression performance to the CANF-based B-frame codec from the last year’s challenge while having much reduced complexity.
Zong-Lin Gao, Cheng-Wei Chen, Yi-Chen Yao, Cheng-Yuan Ho, Wen-Hsiao Peng
ISCAS5
2024 Learning-Based Conditional Image Compression
abstract
In recent years, deep learning-based image compression has achieved significant success. Most schemes adopt an end-to-end trained compression network with a specifically designed entropy model. Inspired by recent advances in conditional video coding, in this work, we propose a novel transformer-based conditional coding paradigm for learned image compression. Our approach first compresses a low-resolution version of the target image and up-scales the decoded image using an off-the-shelf super-resolution model. The super-resolved image then serves as the condition to compress and decompress the target high-resolution image. Experiments demonstrate the superior rate-distortion performance of our approach compared to existing methods.
Tianma Shen, Wen-Hsiao Peng, Huang-Chia Shih
ISCAS2
2024 On the Rate-Distortion-Complexity Trade-Offs of Neural Video Coding
abstract
This paper aims to delve into the rate-distortion-complexity trade-offs of modern neural video coding. Recent years have witnessed much research effort being focused on exploring the full potential of neural video coding. Conditional auto encoders have emerged as the mainstream approach to efficient neural video coding. The central theme of conditional auto encoders is to leverage both spatial and temporal information for better conditional coding. However, a recent study indicates that conditional coding may suffer from information bottlenecks, potentially performing worse than traditional residual coding. To address this issue, recent conditional coding methods incorporate a large number of high-resolution features as the condition signal, leading to a considerable increase in the number of multiply-accumulate operations, memory footprint, and model size. Taking DCVC as the common code base, we investigate how the newly proposed conditional residual coding, an emerging new school of thought, and its variants may strike a better balance among rate, distortion, and complexity.
Yi-Hsin Chen, Kuan-Wei Ho, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng
MMSP5
2024 Transformer-Based Learned Image Compression for Joint Decoding and Denoising
abstract
This work introduces a Transformer-based image compression system. It has the flexibility to switch between the standard image reconstruction and the denoising reconstruction from a single compressed bitstream. Instead of training separate decoders for these tasks, we incorporate two add-on modules to adapt a pre-trained image decoder from performing the standard image reconstruction to joint decoding and denoising. Our scheme adopts a two-pronged approach. It features a latent refinement module to refine the latent representation of a noisy input image for reconstructing a noise-free image. Additionally, it incorporates an instance-specific prompt generator that adapts the decoding process to improve on the latent refinement. Experimental results show that our method achieves a similar level of denoising quality to training a separate decoder for joint decoding and denoising at the expense of only a modest increase in the decoder's model size and computational complexity.
Yi-Hsin Chen, Kuan-Wei Ho, Shiau-Rung Tsai, Guan-Hsun Lin, Alessandro Gnutti, Wen-Hsiao Peng, Riccardo Leonardi
PCS6
2024 Indirect: invertible and discrete noisy image rescaling with enhancement from case-dependent textures
abstract
Abstract Rescaling digital images for display on various devices, while simultaneously removing noise, has increasingly become a focus of attention. However, limited research has been done on a unified framework that can efficiently perform both tasks. In response, we propose INDIRECT (INvertible and Discrete noisy Image Rescaling with Enhancement from Case-dependent Textures), a novel method designed to address image denoising and rescaling jointly. INDIRECT leverages a jointly optimized framework to produce clean and visually appealing images using a lightweight model. It employs a discrete invertible network, DDR-Net, to perform rescaling and denoising through its reversible operations, efficiently mitigating the quantization errors typically encountered during downscaling. Subsequently, the Case-dependent Texture Module (CTM) is introduced to estimate missing high-frequency information, thereby recovering a clean and high-resolution image. Experimental results demonstrate that our method achieves competitive performance across three tasks: noisy image rescaling, image rescaling, and denoising, all while maintaining a relatively small model size.
Huu-Phu Do, Yan-An Chen, Nhat-Tuong Do-Tran, Kai-Lung Hua, Wen-Hsiao Peng
Multim. Syst.5
2024 B-CANF: Adaptive B-Frame Coding With Conditional Augmented Normalizing Flows
abstract
Over the past few years, learning-based video compression has become an active research area. However, most works focus on P-frame coding. Learned B-frame coding is under-explored and more challenging. This work introduces a novel B-frame coding framework, termed B-CANF, that exploits conditional augmented normalizing flows for B-frame coding. B-CANF additionally features two novel elements: frame-type adaptive coding and B*-frames. Our frame-type adaptive coding learns better bit allocation for hierarchical B-frame coding by dynamically adapting the feature distributions according to the B-frame type. Our B*-frames allow greater flexibility in specifying the group-of-pictures (GOP) structure by reusing the B-frame codec to mimic P-frame coding, without the need for an additional, separate P-frame codec. On commonly used datasets, B-CANF achieves the state-of-the-art compression performance as compared to the other learned B-frame codecs and shows comparable BD-rate results to HM-16.23 under the random access configuration in terms of PSNR. When evaluated on different GOP structures, our B*-frames achieve similar performance to the additional use of a separate P-frame codec.
Mu-Jung Chen, Yi-Hsin Chen, Wen-Hsiao Peng
IEEE Trans. Circuits Syst. Video Technol.3
2024 MaskCRT: Masked Conditional Residual Transformer for Learned Video Compression
abstract
Conditional coding has lately emerged as the mainstream approach to learned video compression. However, a recent study shows that it may perform worse than residual coding when the information bottleneck arises. Conditional residual coding was thus proposed, creating a new school of thought to improve on conditional coding. Notably, conditional residual coding relies heavily on the assumption that the residual frame has a lower entropy rate than that of the intra frame. Recognizing that this assumption is not always true due to dis-occlusion phenomena or unreliable motion estimates, we propose a masked conditional residual coding scheme. It learns a soft mask to form a hybrid of conditional coding and conditional residual coding in a pixel adaptive manner. We introduce a Transformer-based conditional autoencoder. Several strategies are investigated with regard to how to condition a Transformer-based autoencoder for inter-frame coding, a topic that is largely under-explored. Additionally, we propose a channel transform module (CTM) to decorrelate the image latents along the channel dimension, with the aim of using the simple hyperprior to approach similar compression performance to the channel-wise autoregressive model. Experimental results confirm the superiority of our masked conditional residual transformer (termed MaskCRT) to both conditional coding and conditional residual coding. On commonly used datasets, MaskCRT shows comparable BD-rate results to VTM-17.0 under the low delay P configuration in terms of PSNR-RGB and outperforms VTM-17.0 in terms of MS-SSIM-RGB. It also opens up a new research direction for advancing learned video compression.
Yi-Hsin Chen, Hong-Sheng Xie, Cheng-Wei Chen, Zong-Lin Gao, Martin Benjak, Wen-Hsiao Peng, Jörn Ostermann
IEEE Trans. Circuits Syst. Video Technol.6
2023 Hierarchical B-Frame Video Coding Using Two-Layer CANF Without Motion Coding
abstract
Typical video compression systems consist of two main modules: motion coding and residual coding. This general architecture is adopted by classical coding schemes (such as international standards H.265 and H.266) and deep learning-based coding schemes. We propose a novel B-frame coding architecture based on two-layer Conditional Augmented Normalization Flows (CANF). It has the striking feature of not transmitting any motion information. Our proposed idea of video compression without motion coding offers a new direction for learned video coding. Our base layer is a low-resolution image compressor that replaces the full-resolution motion compressor. The low-resolution coded image is merged with the warped high-resolution images to generate a high-quality image as a conditioning signal for the enhancement-layer image coding in full resolution. One advantage of this architecture is significantly reduced computational complexity due to eliminating the motion information compressor. In addition, we adopt a skip-mode coding technique to reduce the transmitted latent samples. The rate-distortion performance of our scheme is slightly lower than that of the state-of-the-art learned B-frame coding scheme, B-CANF, but outperforms other learned B-frame coding schemes. However, compared to B-CANF, our scheme saves 45% of multiply-accumulate operations (MACs) for encoding and 27% of MACs for decoding. The code is available at https://nycu-clab.github.io.
David Alexandre, Hsueh-Ming Hang, Wen-Hsiao Peng
CVPR3
2023 MoTIF: Learning Motion Trajectories with Local Implicit Neural Functions for Continuous Space-Time Video Super-Resolution
abstract
This work addresses continuous space-time video super-resolution (C-STVSR) that aims to up-scale an input video both spatially and temporally by any scaling factors. One key challenge of C-STVSR is to propagate information temporally among the input video frames. To this end, we introduce a space-time local implicit neural function. It has the striking feature of learning forward motion for a continuum of pixels. We motivate the use of forward motion from the perspective of learning individual motion trajectories, as opposed to learning a mixture of motion trajectories with backward motion. To ease motion interpolation, we encode sparsely sampled forward motion extracted from the input video as the contextual input. Along with a reliability-aware splatting and decoding scheme, our framework, termed MoTIF, achieves the state-of-the-art performance on C-STVSR. The source code of MoTIF is available at https://github.com/sichun233746/MoTIF.
Si-Cun Chen, Yi-Hsin Chen, Yen-Yu Lin, Wen-Hsiao Peng
ICCV5
2023 TransTIC: Transferring Transformer-based Image Compression from Human Perception to Machine Perception
abstract
This work aims for transferring a Transformer-based image compression codec from human perception to machine perception without fine-tuning the codec. We propose a transferable Transformer-based image compression framework, termed TransTIC. Inspired by visual prompt tuning, TransTIC adopts an instance-specific prompt generator to inject instance-specific prompts to the encoder and task-specific prompts to the decoder. Extensive experiments show that our proposed method is capable of transferring the base codec to various machine tasks and outperforms the competing methods significantly. To our best knowledge, this work is the first attempt to utilize prompting on the low-level image compression task.
Yi-Hsin Chen, Ying-Chieh Weng, Chia-Hao Kao, Cheng Chien, Walon Wei-Chen Chiu, Wen-Hsiao Peng
ICCV6
2023 Learning Continuous Exposure Value Representations for Single-Image HDR Reconstruction
abstract
Deep learning is commonly used to reconstruct HDR images from LDR images. LDR stack-based methods are used for single-image HDR reconstruction, generating an HDR image from a deep learning-generated LDR stack. However, current methods generate the stack with predetermined exposure values (EVs), which may limit the quality of HDR reconstruction. To address this, we propose the continuous exposure value representation (CEVR), which uses an implicit function to generate LDR images with arbitrary EVs, including those unseen during training. Our approach generates a continuous stack with more images containing diverse EVs, significantly improving HDR reconstruction. We use a cycle training strategy to supervise the model in generating continuous EV LDR images without corresponding ground truths. Our CEVR model outperforms existing methods, as demonstrated by experimental results.
Su-Kai Chen, Hung-Lin Yen, Yu-Lun Liu 0001, Min-Hung Chen, Hou-Ning Hu, Wen-Hsiao Peng, Yen-Yu Lin
ICCV6
2023 Transformer-Based Variable-Rate Image Compression with Region-of-Interest Control
abstract
This paper proposes a transformer-based learned image compression system. It is capable of achieving variable-rate compression with a single model while supporting the region-of-interest (ROI) functionality. Inspired by prompt tuning, we introduce prompt generation networks to condition the transformer-based autoencoder of compression. Our prompt generation networks generate content-adaptive tokens according to the input image, an ROI mask, and a rate parameter. The separation of the ROI mask and the rate parameter allows an intuitive way to achieve variable-rate and ROI coding simultaneously. Extensive experiments validate the effectiveness of our proposed method and confirm its superiority over the other competing methods.
Chia-Hao Kao, Ying-Chieh Weng, Yi-Hsin Chen, Walon Wei-Chen Chiu, Wen-Hsiao Peng
ICIP5
2023 Learned Hierarchical B-frame Coding with Adaptive Feature Modulation for YUV 4: 2: 0 Content
abstract
This paper introduces a learned hierarchical B-frame coding scheme in response to the Grand Challenge on Neural Network-based Video Coding at ISCAS 2023. We address specifically three issues, including (1) B-frame coding, (2) YUV 4:2:0 coding, and (3) content-adaptive variable-rate coding with only one single model. Most learned video codecs operate internally in the RGB domain for P-frame coding. B-frame coding for YUV 4:2:0 content is largely under-explored. In addition, while there have been prior works on variable-rate coding with conditional convolution, most of them fail to consider the content information. We build our scheme on conditional augmented normalized flows (CANF). It features conditional motion and inter-frame codecs for efficient B-frame coding. To cope with YUV 4:2:0 content, two conditional inter-frame codecs are used to process the Y and UV components separately, with the coding of the UV components conditioned additionally on the Y component. Moreover, we introduce adaptive feature modulation in every convolutional layer, taking into account both the content information and the coding levels of B-frames to achieve content-adaptive variable-rate coding. Experimental results show that our model outperforms x265 and the winner of last year's challenge on commonly used datasets in terms of PSNR-YUV.
Mu-Jung Chen, Hong-Sheng Xie, Cheng Chien, Wen-Hsiao Peng, Hsueh-Ming Hang
ISCAS4
2023 Continually-Adapted Margin and Multi-Anchor Distillation for Class-Incremental Learning
abstract
This paper addresses the problem of class-incremental learning. The model is trained to recognize the classes added incrementally. It thus suffers from the challenging issue of catastrophic forgetting. Stemming from the knowledge distillation idea of attempting to retain the model's knowledge on seen classes while learning the newly-added ones, we advance to further alleviate the catastrophic forgetting via our proposed multi-anchor distillation objective, which is realized by constraining the spatial relationship between the input data and the multiple class embeddings of each seen class in the feature space while training the model. Moreover, since the knowledge distillation for incremental learning generally relies on keeping a replay buffer to store the samples of seen classes, the buffer of limited size brings another issue of class imbalance: the number of samples from each seen class decreases gradually, thus being much smaller than the number of samples from each new class. We therefore propose to introduce the continually-adapted margin into the classification objective for tackling the prediction bias towards new classes caused by the class imbalance. Experiments are conducted on various datasets and settings to demonstrate the effectiveness and superior performance of our proposed techniques in comparison to several state-of-the-art baselines.
Yi-Hsin Chen, Dian-Shan Chen, Ying-Chieh Weng, Wen-Hsiao Peng, Walon Wei-Chen Chiu
SMC4
2023 Learning-Based Scalable Video Coding with Spatial and Temporal Prediction
abstract
In this work, we propose a hybrid learning-based method for layered spatial scalability. Our framework consists of a base layer (BL), which encodes a spatially downsampled representation of the input video using Versatile Video Coding (VVC), and a learning-based enhancement layer (EL), which conditionally encodes the original video signal. The EL is conditioned by two fused prediction signals: a spatial inter-layer prediction signal, that is generated by spatially upsampling the output of the BL using super-resolution, and a temporal inter-frame prediction signal, that is generated by decoder-side motion compensation without signaling any motion vectors. We show that our method outperforms LCEVC and has comparable performance to full-resolution VVC for high-resolution content, while still offering scalability.
Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann
VCIP3
2023 Rate Adaptation for Learned Two-layer B-frame Coding without Signaling Motion Information
abstract
This paper explores the potential of a learned two-layer B-frame codec, known as TLZMC. TLZMC is one of the few early attempts that deviate from the hybrid-based coding architecture by skipping motion coding. With TLZMC, a low-resolution base layer is utilized to encode temporally unpredictable information. We address the question of whether adapting the base-layer bitrate can achieve better rate-distortion performance. We apply the feature map modulation technique to enable per-frame bitrate adaptation of the base layer. We then propose and compare three online search strategies for determining the base-layer rate parameter: per-level brute-force search, per-level greedy search, and per-frame greedy search. Experimental results show that our top-performing search strategy achieves 0.6%-15.8% Bjøntegaard-Delta rate savings over TLZMC.
Hong-Sheng Xie, Yi-Hsin Chen, Wen-Hsiao Peng, Martin Benjak, Jörn Ostermann
VCIP3
2023 HuPR: A Benchmark for Human Pose Estimation Using Millimeter Wave Radar
abstract
This paper introduces a novel human pose estimation benchmark, Human Pose with Millimeter Wave Radar (HuPR), that includes synchronized vision and radio signal components. This dataset is created using cross-calibrated mmWave radar sensors and a monocular RGB camera for cross-modality training of radar-based human pose estimation. There are two advantages of using mmWave radar to perform human pose estimation. First, it is robust to dark and low-light conditions. Second, it is not visually perceivable by humans and thus, can be widely applied to applications with privacy concerns, e.g., surveillance systems in patient rooms. In addition to the benchmark, we propose a cross-modality training framework that leverages the ground-truth 2D keypoints representing human body joints for training, which are systematically generated from the pre-trained 2D pose estimation network based on a monocular camera input image, avoiding laborious manual label annotation efforts. The framework consists of a new radar pre-processing method that better extracts the velocity information from radar data, Cross- and Self-Attention Module (CSAM), to fuse multi-scale radar features, and Pose Refinement Graph Convolutional Networks (PRGCN), to refine the predicted keypoint confidence heatmaps. Our intensive experiments on the HuPR benchmark show that the proposed scheme achieves better human pose estimation performance with only radar data, as compared to traditional pre-processing solutions and previous radiofrequency-based methods. Our code is available at here1
Shih-Po Lee, Niraj Prakash Kini, Wen-Hsiao Peng, Ching-Wen Ma, Jenq-Neng Hwang
WACV3
2022 CANF-VC: Conditional Augmented Normalizing Flows for Video Compression
Yung-Han Ho, Chih-Peng Chang, Alessandro Gnutti, Wen-Hsiao Peng
ECCV (16)5
2022 Deep Incremental Optical Flow Coding For Learned Video Compression
abstract
This work addresses motion coding in end-to-end learned video compression. The efficiency of motion coding is critical at low bit rates, at which a large portion of the bitstream signals motion information. Most end-to-end learned video codecs adopt an intra-coding approach to coding motion information as individual optical flow maps. Some recent studies introduce predictive motion coding to encode optical flow map residuals. Still, motion coding remains an active research area for learned video compression. We present an incremental optical flow coding scheme. It first leverages an extrapolated flow together with the reference frame in estimating an incremental flow between the reference and the target frames for efficient motion coding. It then derives the final flow map for motion compensation by integrating the incremental and the extrapolated flows in a double-warping scheme. Experimental results on commonly used datasets show the superiority of our method over predictive motion coding and other advanced schemes.
Chih-Peng Chang, Yung-Han Ho, Wen-Hsiao Peng
ICIP4
2022 Learned Video Compression for YUV 4: 2: 0 Content Using Flow-based Conditional Inter-frame Coding
abstract
This paper proposes a learning-based video compression framework for variable-rate coding on YUV 4:2:0 content. Most existing learning-based video compression models adopt the traditional hybrid-based coding architecture, which involves temporal prediction followed by residual coding. However, recent studies have shown that residual coding is suboptimal from the information-theoretic perspective. In addition, most existing models are optimized with respect to RGB content. Furthermore, they require separate models for variable-rate coding. To address these issues, this work presents an attempt to incorporate the conditional inter-frame coding for YUV 4:2:0 content. We introduce a conditional flow-based inter-frame coder to improve the inter-frame coding efficiency. To adapt our codec to YUV 4:2:0 content, we adopt a simple strategy of using space-to-depth and depth-to-space conversions. Lastly, we employ a rate-adaption net to achieve variable-rate coding without training multiple models. Experimental results show that our model performs better than x265 on UVG and MCL-JCV datasets in terms of PSNR-YUV. However, on the more challenging datasets from ISCAS’22 GC, there is still ample room for improvement. This insufficient performance is due to the lack of inter-frame coding capability at a large GOP size and can be mitigated by increasing the model capacity and applying an error propagationaware training strategy.
Yung-Han Ho, Chih-Hsuan Lin, Mu-Jung Chen, Chih-Peng Chang, Wen-Hsiao Peng, Hsueh-Ming Hang
ISCAS6
2022 Two-Layer Learning-Based P-Frame Coding with Super-Resolution and Content-Adaptive Conditional ANF
abstract
Deep-learning-based video compression technique has been rapidly growing in recent years. This paper adopts the Conditional Augmented Normalizing Flow video codec (CANF-VC) [8] as our basic system. To improve the quality of the condition signal (image) for CANF, we propose a two-layer structure learning-based video codec. At low cost of extra bit rate, the low-resolution base layer provides side information to improve the quality of motion-compensated reference frame through a super-resolution module with a merge-net. In addition, the base layer also provides information to the skip-mask generator. The skip-mask guides the coding mechanism to reduce the transmitted samples for the high-resolution enhancement layer. The experiment results indicate that the proposed two-layer coding scheme can provide 22.19% PSNR BD-Rate saving and 49.59% MS-SSIM BD-Rate saving over H.265 (HM 16.20) on the UVG test sequences.
David Alexandre, Hsueh-Ming Hang, Wen-Hsiao Peng
MMAsia3
2022 Content-Adaptive Motion Rate Adaption For Learned Video Compression
abstract
This paper introduces an online motion rate adaptation scheme for learned video compression, with the aim of achieving content-adaptive coding on individual test sequences to mitigate the domain gap between training and test data. It features a patch-level bit allocation map, termed the $\alpha-$map, to trade off between the bit rates for motion and inter-frame coding in a spatially-adaptive manner. We optimize the $\alpha-$map through an online back-propagation scheme at inference time. Moreover, we incorporate a look-ahead mechanism to consider its impact on future frames. Extensive experimental results confirm that the proposed scheme, when integrated into a conditional learned video codec, is able to adapt motion bit rate effectively, showing much improved rate-distortion performance particularly on test sequences with complicated motion characteristics.
Chih Hsuan Lin, Yi-Hsin Chen, Wen-Hsiao Peng
PCS3
2022 Neural Frank-Wolfe Policy Optimization for Region-of-Interest Intra-Frame Coding with HEVC/H.265
abstract
This paper presents a reinforcement learning (RL) framework that utilizes Frank-Wolfe policy optimization to solve Coding- Tree-Unit (CTU) bit allocation for Region-of-Interest (ROI) intra-frame coding. Most previous RL-based methods employ the single-critic design, where the rewards for distortion minimization and rate regularization are weighted by an empirically chosen hyper-parameter. Recently, the dual-critic design is proposed to update the actor by alternating the rate and distortion critics. However, its convergence is not guaranteed. To address these issues, we introduce Neural Frank-Wolfe Policy Op-timization (NFWPO) in formulating the CTU-level bit allocation as an action-constrained RL problem. In this new framework, we exploit a rate critic to predict a feasible set of actions. With this feasible set, a distortion critic is invoked to update the actor to maximize the ROI-weighted image quality subject to a rate constraint. Experimental results produced with x265 confirm the superiority of the proposed method to the other baselines.
Yung-Han Ho, Chia-Hao Kao, Wen-Hsiao Peng, Ping-Chun Hsieh
VCIP3
2022 Augmented Normalizing Flow for Point Cloud Geometry Coding
abstract
With the increased popularity of immersive media, point clouds have become one of the popular data representations for presenting 3D scenes. The huge amount of point cloud data poses a great challenge on their storage and real-time transmission, which calls for efficient point cloud compression. This paper presents a novel point cloud geometry compression technique based on learning end-to-end an augmented normalizing flow (ANF) model to represent the occupancy status of voxelized data points. The higher expressive power of ANF than variational autoencoders (V AE) is leveraged for the first time to represent binary occupancy status. Compared to two coding standards developed by MPEG, namely G-PCC (geometry-based point cloud compression) and V-PCC (video-based point cloud compression), our method achieves more than 80% and 30% bitrate reduction, respectively. Compared to several learning-based methods, our method also yields better performance.
Siao-Yu Li, Ji-Jin Chiu, Jui-Chiu Chiang, Wen-Hsiao Peng, Wen-Nung Lie
VCIP4
2022 Object Rearrangement Through Planar Pushing: A Theoretical Analysis and Validation
abstract
In this article, we focus on rearranging an object by pushing it to any target planar pose. We identify the essential elements to guarantee that the target pose can be reached at an acceptable precision. We present a simple rearrangement algorithm that relies on only a few known straight-line pushes for some novel object and requires no analytical models, force sensors, or large training datasets. We derive the step upper bound, which relates the initial pose of the object, stopping criterion, and quality of the set of pushes, to facilitate the estimation of the maximum number of required steps without the need to perform a task. We experimentally verified the performance of our algorithm at different noise levels, stopping criteria, and task difficulties on datasets containing several types of objects. By applying combinations of only nine known pushes, our simple algorithm performed successfully in real-world experiments with challenging objects, including partially deformable objects that are difficult to model analytically; it achieved the precise stopping criterion (7.5 mm, 5$^\circ$) in various rearrangement tasks.
Chun-Yu Chai, Wen-Hsiao Peng, Shiao-Li Tsao
IEEE Trans. Robotics2
2021 Video Rescaling Networks With Joint Optimization Strategies for Downscaling and Upscaling
abstract
This paper addresses the video rescaling task, which arises from the needs of adapting the video spatial resolution to suit individual viewing devices. We aim to jointly optimize video downscaling and upscaling as a combined task. Most recent studies focus on image-based solutions, which do not consider temporal information. We present two joint optimization approaches based on invertible neural networks with coupling layers. Our Long Short-Term Memory Video Rescaling Network (LSTM-VRN) leverages temporal information in the low-resolution video to form an explicit prediction of the missing high-frequency information for upscaling. Our Multi-input Multi-output Video Rescaling Network (MIMO-VRN) proposes a new strategy for downscaling and upscaling a group of video frames simultaneously. Not only do they outperform the image-based invertible model in terms of quantitative and qualitative results, but also show much improved upscaling quality than the video rescaling methods without joint optimization. To our best knowledge, this work is the first attempt at the joint optimization of video downscaling and upscaling.
Yan-Cheng Huang, Yi-Hsin Chen, Cheng-You Lu, Hui-Po Wang, Wen-Hsiao Peng
CVPR5
2021 A Dual-Critic Reinforcement Learning Framework for Frame-Level Bit Allocation in HEVC/H.265
abstract
This paper introduces a dual-critic reinforcement learning (RL) framework to address the problem of frame-level bit allocation in HEVC/H.265. The objective is to minimize the distortion of a group of pictures (GOP) under a rate constraint. Previous RL-based methods tackle such a constrained optimization problem by maximizing a single reward function that often combines a distortion and a rate reward. However, the way how these rewards are combined is usually ad hoc and may not generalize well to various coding conditions and video sequences. To overcome this issue, we adapt the deep deterministic policy gradient (DDPG) reinforcement learning algorithm for use with two critics, with one learning to predict the distortion reward and the other the rate reward. In particular, the distortion critic works to update the agent when the rate constraint is satisfied. By contrast, the rate critic makes the rate constraint a priority when the agent goes over the bit budget. Experimental results on commonly used datasets show that our method outperforms the bit allocation scheme in x265 and the single-critic baseline by a significant margin in terms of rate-distortion performance while offering fairly precise rate control.
Yung-Han Ho, Guo-Lun Jin, Yun Liang 0015, Wen-Hsiao Peng
DCC4
2021 Deep Video Compression for Interframe Coding
abstract
A typical learning-based video compression scheme consists of motion coding and residual coding. In this paper, our deep video compression features a motion predictor and refinement networks for interframe coding. To save the bits for transmitting motion information, our scheme performs local motion prediction and sends only the differential motion vectors to the decoder. In the residual coding, we couple the residual decoder with the refine-net to reduce residual signal bits. The experiments show that our work can produce a very competitive coding performance compared to the other learning-based predictive video codecs.
David Alexandre, Hsueh-Ming Hang, Wen-Hsiao Peng, Marek Domanski
ICIP3
2021 GSVNET: Guided Spatially-Varying Convolution for Fast Semantic Segmentation on Video
abstract
This paper addresses fast semantic segmentation on video. Video segmentation often calls for real-time, or even faster than real-time, processing. One common recipe for conserving computation arising from feature extraction is to propagate features of few selected keyframes. However, recent advances in fast image segmentation make these solutions less attractive. To leverage fast image segmentation for furthering video segmentation, we propose a simple yet efficient propagation framework. Specifically, we perform lightweight flow estimation in 1/8-downscaled image space for temporal warping in segmentation outpace space. Moreover, we introduce a guided spatially-varying convolution for fusing segmentations derived from the previous and current frames, to mitigate propagation error and enable lightweight feature extraction on non-keyframes. Experimental results on Cityscapes and CamVid show that our scheme achieves the state-of-the-art accuracy-throughput trade-off on video segmentation.
Shih-Po Lee, Si-Cun Chen, Wen-Hsiao Peng
ICME3
2021 Weakly-Supervised Image Semantic Segmentation Using Graph Convolutional Networks
abstract
This work addresses weakly-supervised image semantic segmentation based on image-level class labels. One common approach to this task is to propagate the activation scores of Class Activation Maps (CAMs) using a random-walk mechanism in order to arrive at complete pseudo labels for training a semantic segmentation network in a fully-supervised manner. However, the feed-forward nature of the random walk imposes no regularization on the quality of the resulting complete pseudo labels. To overcome this issue, we propose a Graph Convolutional Network (GCN)-based feature propagation framework. We formulate the generation of complete pseudo labels as a semi-supervised learning task and learn a 2-layer GCN separately for every training image by back-propagating a Laplacian and an entropy regularization loss. Experimental results on the PASCAL VOC 2012 dataset confirm the superiority of our scheme to several state-of-the-art baselines. Our code is available at https: //github.com/Xavier-Pan/WSGCN.
Shun-Yi Pan, Cheng-You Lu, Shih-Po Lee, Wen-Hsiao Peng
ICME4
2021 DIRECT: Discrete Image Rescaling with Enhancement from Case-specific Textures
abstract
This paper addresses image rescaling, the task of which is to downscale an input image followed by upscaling for the purposes of transmission, storage, or playback on heterogeneous devices. The state-of-the-art image rescaling network (known as IRN) tackles image downscaling and upscaling as mutually invertible tasks using invertible affine coupling layers. In particular, for upscaling, IRN models the missing high-frequency component by an input-independent (case-agnostic) Gaussian noise. In this work, we take one step further to predict a case-specific high-frequency component from textures embedded in the downscaled image. Moreover, we adopt integer coupling layers to avoid quantizing the downscaled image. When tested on commonly used datasets, the proposed method, termed DIRECT, improves high-resolution reconstruction quality both subjectively and objectively, while maintaining visually pleasing downscaled images.
Yan-An Chen, Ching-Chun Hsiao, Wen-Hsiao Peng
VCIP3
2021 Learning to Fly with a Video Generator
abstract
This paper demonstrates a model-based reinforcement learning framework for training a self-flying drone. We implement the Dreamer proposed in a prior work as an environment model that responds to the action taken by the drone by predicting the next video frame as a new state signal. The Dreamer is a conditional video sequence generator. This model-based environment avoids the time-consuming interactions between the agent and the environment, speeding up largely the training process. This demonstration showcases for the first time the application of the Dreamer to train an agent that can finish the racing task in the Airsim simulator.
Chia-Chun Chung, Wen-Hsiao Peng, Teng-Hu Cheng, Chia-Hau Yu
VCIP2
2020 Class-Incremental Learning with Rectified Feature-Graph Preservation
Cheng-Hsun Lei, Yi-Hsin Chen, Wen-Hsiao Peng, Walon Wei-Chen Chiu
ACCV (6)3
2020 Deep Video Prediction Through Sparse Motion Regularization
abstract
This paper introduces data-dependent sparse motion regularization for dense flow-based video prediction. To achieve video prediction (a form of extrapolation from past frames), the dense flow-based model estimates a motion vector for every pixel in a target frame for backward warping. Due to the sheer amount of motion vectors to be estimated, the model tends to be complex, thereby calling for proper regularization to avoid over-fitting. Most flow-based models adopt smoothness regularization. However, the smoothness requirement is detrimental to preserving the discontinuity of the motion field, which often appears in videos with distinct object motion. To address this issue, our sparse motion regularization discovers distinct sparse motion via weighted K-means clustering and regularizes the model based on minimizing clustering errors in the predicted motion field. When incorporated in an end-to-end trainable deep video prediction model, our scheme outperforms smoothness regularization, achieving superiority over direct generation-based video prediction on UCF-101 and Common Intermediate Format (CIF) datasets.
Yung-Han Ho, Chih-Chun Chan, Wen-Hsiao Peng
ICIP3
2020 Adaptive Unknown Object Rearrangement Using Low-Cost Tabletop Robot
abstract
Studies on object rearrangement planning typically consider known objects. Some learning-based methods can predict the movement of an unknown object after single-step interaction, but require intermediate targets, which are generated manually, to achieve the rearrangement task. In this work, we propose a framework for unknown object rearrangement. Our system first models an object through a small-amount of identification actions and adjust the model parameters during task execution. We implement the proposed framework based on a low-cost tabletop robot (under 180 USD) to demonstrate the advantages of using a physics engine to assist action prediction. Experimental results reveal that after running our adaptive learning procedure, the robot can successfully arrange a novel object using an average of five discrete pushes on our tabletop environment and satisfy a precise 3.5 cm translation and 5° rotation criterion.
Chun-Yu Chai, Wen-Hsiao Peng, Shiao-Li Tsao
ICRA2
2020 Semantic Segmentation on Compressed Video using Block Motion Compensation and Guided Inpainting
abstract
This paper addresses the problem of fast semantic segmentation on compressed video. Unlike most prior works for video segmentation, which perform feature propagation based on optical flow estimates or sophisticated warping techniques, ours takes advantage of block motion vectors in the compressed bitstream to propagate the segmentation of a keyframe to subsequent non-keyframes. This approach, however, needs to respect the inter-frame prediction structure, which often suggests recursive, multi-step prediction with error propagation and accumulation in the temporal dimension. To tackle the issue, we refine the motion-compensated segmentation using inpainting. Our inpainting network incorporates guided non-local attention for long-range reference and pixel-adaptive convolution for ensuring the local coherence of the segmentation. A fusion step then follows to combine both the motion-compensated and inpainted segmentations. Experimental results show that our method outperforms the state-of-the-art baselines in terms of segmentation accuracy. Moreover, it introduces the least amount of network parameters and multiply-add operations for non-keyframe segmentation.
Stefanie Tanujaya, Tieh Chu, Wen-Hsiao Peng
ISCAS4
2020 A Hybrid Layered Image Compressor with Deep-Learning Technique
abstract
This paper presents a detailed description of NCTU's proposal for learning-based image compression, in response to the JPEG AI Call for Evidence Challenge. The proposed compression system features a VVC intra codec as the base layer and a learning-based residual codec as the enhancement layer. The latter aims to refine the quality of the base layer via sending a latent residual signal. In particular, a base-layer-guided attention module is employed to focus the residual extraction on critical high-frequency areas. To reconstruct the image, this latent residual signal is combined with the base-layer output in a non-linear fashion by a neural-network-based synthesizer. The proposed method shows comparable rate-distortion performance to single-layer VVC intra in terms of common objective metrics, but presents better subjective quality particularly at high compression ratios in some cases. It consistently outperforms HEVC intra, JPEG 2000, and JPEG. The proposed system incurs 18M network parameters in 16-bit floating-point format. On average, the encoding of an image on Intel Xeon Gold 6154 takes about 13.5 minutes, with the VVC base layer dominating the encoding runtime. On the contrary, the decoding is dominated by the residual decoder and the synthesizer, requiring 31 seconds per image.
Wei-Cheng Lee, Chih-Peng Chang, Wen-Hsiao Peng, Hsueh-Ming Hang
MMSP3
2020 Recent Advances in End-to-End Learned Image and Video Compression
abstract
The DCT-based transform coding technique was adopted by the international standards (ISO JPEG, ITU H.261/264/265, ISO MPEG-2/4/H, and many others) for nearly 30 years. Although researchers are still trying to improve its efficiency by fine-tuning its components and parameters, the basic structure has not changed in the past two decades.The deep learning technology recently developed may provide a new direction for constructing a high-compression image/video coding system. Recent results, particularly from the Challenge on Learned Image Compression (CLIC) at CVPR, indicate that this new type of schemes (often trained end-to-end) may have good potential for further improving compression efficiency.In the first part of this tutorial, we shall (1) summarize briefly the progress of this topic in the past 3 or so years, including an overview of CLIC results and JPEG AI Call-for-Evidence Challenge on Learning-based Image Coding (issued in early 2020). Because Deep Neural Network (DNN)-based image compression is a new area, several techniques and structures have been tested. The recently published autoencoder-based schemes can achieve similar PSNR to BPG (Better Portable Graphics, H.265 still image standard) and has superior subject quality (e.g., MSSSIM), especially at the very low bit rates. In the second part, we shall (2) address the detailed design concepts of image compression algorithms using the autoencoder structure. In the third part, we shall switch gears to (3) explore the emerging area of DNN-based video compression. Recent publications in this area have indicated that end-to-end trained video compression can achieve comparable or superior rate-distortion performance to HEVC/H.265. The CLIC at CVPR 2020 also created for the first time a new track dedicated to P-frame coding.
Wen-Hsiao Peng, Hsueh-Ming Hang
VCIP1
2020 Guest Editorial Introduction to Special Section on Learning-Based Image and Video Compression
abstract
Video is being watched more than ever before. It is estimated that in 2020, 82% of global IP traffic and 79% of global Internet traffic will come from video; globally 3 trillion minutes (5 million years) of video content will cross the Internet each month. According to the Cisco 2020 Forecast, that is one million minutes of video streamed or downloaded every second[1]. The rapidly increasing consumption of storage capacity and transmission bandwidth from video, especially HD and UHD video content, has made video compression a critical stage to guarantee the quality of delivery and playback.
Shan Liu 0001, Wen-Hsiao Peng, Lu Yu 0003
IEEE Trans. Circuits Syst. Video Technol.2
2019 All About Structure: Adapting Structural Information Across Domains for Boosting Semantic Segmentation
abstract
In this paper we tackle the problem of unsupervised domain adaptation for the task of semantic segmentation, where we attempt to transfer the knowledge learned upon synthetic datasets with ground-truth labels to real-world images without any annotation. With the hypothesis that the structural content of images is the most informative and decisive factor to semantic segmentation and can be readily shared across domains, we propose a Domain Invariant Structure Extraction (DISE) framework to disentangle images into domain-invariant structure and domain-specific texture representations, which can further realize image-translation across domains and enable label transfer to improve segmentation performance. Extensive experiments verify the effectiveness of our proposed DISE model and demonstrate its superiority over several state-of-the-art approaches.
Wei-Lun Chang, Hui-Po Wang, Wen-Hsiao Peng, Walon Wei-Chen Chiu
CVPR3
2019 SME-Net: Sparse Motion Estimation for Parametric Video Prediction Through Reinforcement Learning
abstract
This paper leverages a classic prediction technique, known as parametric overlapped block motion compensation (POBMC), in a reinforcement learning framework for video prediction. Learning-based prediction methods with explicit motion models often suffer from having to estimate large numbers of motion parameters with artificial regularization. Inspired by the success of sparse motion-based prediction for video compression, we propose a parametric video prediction on a sparse motion field composed of few critical pixels and their motion vectors. The prediction is achieved by gradually refining the estimate of a future frame in iterative, discrete steps. Along the way, the identification of critical pixels and their motion estimation are addressed by two neural networks trained under a reinforcement learning setting. Our model achieves the state-of-the-art performance on CaltchPed, UCF101 and CIF datasets in one-step and multi-step prediction tests. It shows good generalization results and is able to learn well on small training data.
Yung-Han Ho, Chuan-Yuan Cho, Guo-Lun Jin, Wen-Hsiao Peng
ICCV4
2019 Learned Image Compression with Soft Bit-Based Rate-Distortion Optimization
abstract
This paper introduces the notion of soft bits to address the rate-distortion optimization for learning-based image compression. Recent methods for such compression train an autoencoder end-to-end with an objective to strike a balance between distortion and rate. They are faced with the zero gradient issue due to quantization and the difficulty of estimating the rate accurately. Inspired by soft quantization, we represent quantization indices of feature maps with differentiable soft bits. This allows us to couple tightly the rate estimation with context-adaptive binary arithmetic coding. It also provides a differentiable distortion objective function. Experimental results show that our approach achieves the state-of-the-art compression performance among the learning-based schemes in terms of MS-SSIM and PSNR.
David Alexandre, Chih-Peng Chang, Wen-Hsiao Peng, Hsueh-Ming Hang
ICIP3
2019 Deep Reinforcement Learning for Video Prediction
abstract
This paper introduces a hybrid video prediction scheme that combines the classic parametric overlapped block motion compensation (POBMC) technique with neural networks. Most learning-based video prediction methods rely on a black-box-like model for either direct generation of future video frames or estimation of a dense motion field. The model complexity often increases drastically with frame resolution. Departing from pure black-box approaches, this paper leverages the theoretically-grounded POBMC in a reinforcement learning framework to estimate a sparse motion field for future frame warping. Two neural networks are trained to identify critical points in the motion field for motion estimation. We train our model on 10k unlabeled frames in KITTI dataset and achieve the state-of-the-art SSIM score of 0.923 on CaltechPed and an average SSIM scroe of 0.856 on Common Intermediate Format (CIF) standard sequences.
Yung-Han Ho, Chuan-Yuan Cho, Wen-Hsiao Peng
ICIP3
2019 Learning Goal-Oriented Visual Dialog Agents: Imitating and Surpassing Analytic Experts
abstract
This paper tackles the problem of learning a questioner in the goal-oriented visual dialog task. Several previous works adopt model-free reinforcement learning. Most pretrain the model from a finite set of human-generated data. We argue that using limited demonstrations to kick-start the questioner is insufficient due to the large policy search space. Inspired by a recently proposed information theoretic approach, we develop two analytic experts to serve as a source of high-quality demonstrations for imitation learning. We then take advantage of reinforcement learning to refine the model towards the goal-oriented objective. Experimental results on the GuessWhat?! dataset show that our method has the combined merits of imitation and reinforcement learning, achieving the state-of-the-art performance.
Yen-Wei Chang, Wen-Hsiao Peng
ICME2
2018 Reinforcement Learning for HEVC/H.265 Intra-Frame Rate Control
abstract
Reinforcement learning has proven effective for solving decision making problems. However, its application to modern video codecs has yet to be seen. This paper presents an early attempt to introduce reinforcement learning to HEVC/H.265 intra-frame rate control. The task is to determine a quantization parameter value for every coding tree unit in a frame, with the objective being to minimize the frame-level distortion subject to a rate constraint. We draw an analogy between the rate control problem and the reinforcement learning problem, by considering the texture complexity of coding tree units and bit balance as the environment state, the quantization parameter value as an action that an agent needs to take, and the negative distortion of the coding tree unit as an immediate reward. We train a neural network based on Q-learning to be our agent, which observes the state to evaluate the reward for each possible action. When trained on only limited sequences, the proposed model can already perform comparably with the rate control algorithm in HM-16.15.
Jun-Hao Hu, Wen-Hsiao Peng, Chia-Hua Chung
ISCAS2
2017 Intra Line Copy for HEVC Screen Content Intra-Picture Prediction
abstract
This paper presents an intra line copy (ILC) technique for HEVC screen content coding. It shares the same origin with two other prominent techniques, intra string copy (ISC) and intra block copy (IBC), in applying the notion of string matching to intra-frame coding. This work combines their merits in one scheme with both good compression performance and high regularity. Specifically, it forms a prediction of a coding block by decomposing it into horizontal or vertical lines of pixels and performing line-based predictions based on previously coded data from the current frame. To address the massive amounts of search operations, our fast search algorithm first searches in the horizontal and vertical directions, then checks line vector candidates from spatial and temporal neighbors, and finally references lines having an identical hash value as the prediction line. The resulting line vectors are further predicted adaptively to minimize their coding overhead. Extensive experiments based on SCM-4.0, which includes IBC as an integral component, show that ILC can provide an additional 4%-7% BD-rate savings when the search area extends to the entire frame and 3%-4% improvements with a restricted local search. Compared with ISC, it achieves comparable performance without all its complications from sequential string parsing.
Chun-Chi Chen, Wen-Hsiao Peng
IEEE Trans. Circuits Syst. Video Technol.2
2015 Discriminatively-learned global image representation using CNN as a local feature extractor for image retrieval
abstract
This work introduces an image retrieval framework based on using deep convolutional neural networks (CNN) as a local feature extractor. Motivated by the great success of CNN in recognition tasks, one may be tempted to simply adopt the output of CNN as a global image representation for retrieval. This straightforward approach, however, has proved deficient, because it can be vulnerable to various image transformation attacks. To address this issue, we propose to treat CNN as a local feature extractor, and a local image patch selection mechanism is developed to extract discriminative patches by observing their objectness responses, aspect ratios, relative scales, and locations in the image. The criterion is given by a learned posterior probability indicating how likely the image patch in question will find a correspondence in another similar image. In addition, the CNN's weight parameters are specifically adapted by a contrastive loss function to suit retrieval tasks. Extensive experiments on typical retrieval datasets confirm the superiority of the proposed scheme over the state-of-the-art methods.
Wei-Lin Ku, Hung-Chun Chou, Wen-Hsiao Peng
VCIP3
2015 On comparison of intra line copy and intra string copy for HEVC screen content coding
abstract
Recently, Intra Line Copy (ILC) and Intra String Copy (ISC) have been introduced as effective means for screen content coding during the development of HEVC Screen Content Coding Extensions. Although conceptually sharing the same theoretical basis, which is essentially string matching, they differ from each other in a number of significant ways. This paper presents detailed comparisons between them with respect to their compression performance, runtime complexity and memory access bandwidth, in order to better understand their merits and faults. Experimental results indicate that they perform comparably to each other, with ILC generally performing better in coding mixed content and ISC better in coding pure screen content. When considering the impact on the memory access bandwidth and coding performance trade-offs, local search for ILC/ISC becomes more favorable than full-frame search in low-delay real-time applications.
Ru-Ling Liao, Chun-Chi Chen, Wen-Hsiao Peng
VCIP3
2014 Mode-dependent distortion modeling for H.264/SVC coarse grain SNR scalability
abstract
This paper presents a mode-dependent distortion model for H.264/SVC coarse grain SNR scalability. It estimates the base-layer and enhancement-layer's distortions with particular consideration of their prediction modes and inter-layer residual prediction. Based on a parametric signal model, the variances of the transformed prediction residual at both layers are first formulated analytically and approximated empirically. The results are then incorporated into the assumption that the transform coefficients are distributed according to the Laplacian distribution to obtain the final distortion estimates. Experimental results confirm its fairly good ability to predict the actual distortions in both the frame and macroblock levels.
Yin-An Jian, Chun-Chi Chen, Wen-Hsiao Peng
ICIP3
2014 Screen content coding using non-square intra block copy for HEVC
abstract
To achieve high coding performance for screen content, the intra block copy (IntraBC) performs block matching within a limited area of the reconstructed samples inside the current picture. We further extend its notion from the unit of square coding unit (CU) to non-square prediction unit (PU) partitions. This design is then justified by theoretical and empirical analyses which reveal the same fact that blocks coded by IntraBC mode tend to enable more at smaller partition levels. Besides, the syntax design of the proposed method is fully aligned with that of inter partition modes. Therefore the architecture-wise change in video codec design can be minimized. The experimental results justify the effectiveness of the proposed mode for coding of screen content video. In particular, up to 19.5% rate reduction (with an average of 12.7%) relative to the HM-12.0+RExt-4.1 anchor can be achieved on top of the usage of CU-based IntraBC prediction.
Chun-Chi Chen, Xiaozhong Xu, Ru-Ling Liao, Wen-Hsiao Peng, Shan Liu 0001, Shawmin Lei
ICME4
2014 Global image representation using Locality-constrained Linear Coding for large-scale image retrieval
abstract
This paper proposes a global image representation based on Locality-constrained Linear Coding (LLC), with an aim to simplify the encoding process of local descriptors so as to facilitate large-scale image retrieval. Starting from the state-of-the-art Fisher Vector (FV) representation, we replace the computation of sophisticated posterior probabilities with simpler LLC. We then conduct several empirical studies to investigate the effects and benefits of this change and to adapt the other terms in FV for a better trade-off between performance and complexity. The result is a simpler global descriptor that combines the merits of both FV and LLC. Experimental results show that when compared with other similar works, our scheme not only brings performance benefits in mean Average Precision, but also offer complexity advantages.
Yu-Hsing Wu, Wei-Lin Ku, Wen-Hsiao Peng, Hung-Chun Chou
ISCAS3
2013 An Interframe Prediction Technique Combining Template Matching Prediction and Block-Motion Compensation for High-Efficiency Video Coding
abstract
This paper introduces an interframe prediction technique that combines two motion vectors (MVs) derived respectively from template and block matching for overlapped block motion compensation (OBMC). It has a salient feature of not having to signal the template MV, while achieving a prediction performance close to that of bi-prediction. We begin by studying template matching prediction (TMP) from a theoretical perspective. Based on two signal models, the template MV is shown to approximate the pixel true motion around the template centroid, through which we explain why TMP generally outperforms SKIP prediction but is inferior to block-based motion compensation in terms of prediction performance. We then approach the problem of finding another MV to best complement the template MV from both deterministic and statistical viewpoints, the latter leading to the search of its optimal sampling location in the motion field. The result is a search criterion with OBMC window functions forming a geometry-like motion partitioning when the template area is straddled on the top and to the left of a target block. Generalizations to adaptive template design, multihypothesis prediction and motion merging are made to explore the complexity and performance trade-offs. Extensive experiments based on the HM-6.0 software show that the best of them, in terms of compression performance, achieves 1.7-2.0% BD-rate reductions at a cost of 26% and 39% increases in encoding and decoding times, respectively.
Wen-Hsiao Peng, Chun-Chi Chen
IEEE Trans. Circuits Syst. Video Technol.1
2012 Analytical mode-dependent rate and distortion models for H.264/SVC coarse grain scalability
abstract
This paper presents analytical rate and distortion models to estimate the rate-distortion (R-D) behavior of different coding modes for H.264/SVC coarse grain scalability (CGS) by jointly considering the effects of quantization, motion-compensated prediction, and motion partitioning structures. Based on the forward channel model and two statistical signal models, we first derive a pair of mode-dependent rate and distortion models for coding modes in base layer. Then these models are extended to account for the coding modes with inter-layer residual prediction in enhancement layer, especially for CGS. Preliminary results show that the proposed rate and distortion models for both layers have the potentiality to estimate the R-D behavior of real data.
Chung-Hao Wu, Yu-Chen Tseng, Wen-Hsiao Peng
ISCAS3
2012 Parametric OBMC for Pixel-Adaptive Temporal Prediction on Irregular Motion Sampling Grids
abstract
This paper adapts overlapped block motion compensation (OBMC) to suit variable block-size motion partitioning. The motion vectors (MVs) for various partitions are formalized as motion samples taken on an irregular grid. From this viewpoint, determining OBMC weights to associate with these samples becomes an under-determined problem since a distinct solution has to be sought for each prediction pixel. In this paper, we tackle this problem by expressing the optimal weights in closed form based on parametric signal assumptions. In particular, the computation of this solution requires only the geometric relations between the prediction pixel and its nearby block centers, leading to a generic framework capable of reconstructing temporal predictors from any irregularly sampled MVs. A modified implementation is also proposed to address the MV location uncertainty and to reduce computational complexity. Experimental results demonstrate that our scheme performs better than similar previous works, and when compared to the recently proposed Quadtree-based adaptive loop filter and enhanced adaptive interpolation filter, show a comparable gain. Furthermore, the combination of it with either of them gives a combined effect that is almost the sum of their separate improvements.
Yi-Wen Chen, Wen-Hsiao Peng
IEEE Trans. Circuits Syst. Video Technol.2
2011 Bi-prediction combining template and block motion compensations
abstract
This paper introduces a bi-prediction scheme with only a motion overhead as for unidirectional prediction. It combines motion vectors found by template and block matchings with the overlapped block motion compensation (OBMC). An optimal window function is designed based on a model-based framework. Additionally, the concept of adaptive motion merging is incorporated to enable a template-matching-free implementation. Three algorithms featuring different performance and complexity trade-offs are implemented using the TMuC-0.9_HM software and tested with the common test conditions. Relative to the anchor, the best of them achieves an average BD-rate saving of 2.2%, with a minimum of 0.2% and a maximum of 4.7%.
Chung-Lin Lee, Chun-Chi Chen, Yi-Wen Chen, Mu-Hsuan Wu, Chung-Hao Wu, Wen-Hsiao Peng
ICIP6
2011 Fast Bi-Directional Prediction Selection in H.264/MPEG-4 AVC Temporal Scalable Video Coding
abstract
In this paper, we propose a fast algorithm that efficiently selects the temporal prediction type for the dyadic hierarchical-B prediction structure in the H.264/MPEG-4 temporal scalable video coding (SVC). We make use of the strong correlations in prediction type inheritance to eliminate the superfluous computations for the bi-directional (BI) prediction in the finer partitions, 16×8/8×16/8×8 , by referring to the best temporal prediction type of 16 × 16. In addition, we carefully examine the relationship in motion bit-rate costs and distortions between the BI and the uni-directional temporal prediction types. As a result, we construct a set of adaptive thresholds to remove the unnecessary BI calculations. Moreover, for the block partitions smaller than 8 × 8, either the forward prediction (FW) or the backward prediction (BW) is skipped based upon the information of their 8 × 8 partitions. Hence, the proposed schemes can efficiently reduce the extensive computational burden in calculating the BI prediction. As compared to the JSVM 9.11 software, our method saves the encoding time from 48% to 67% for a large variety of test videos over a wide range of coding bit-rates and has only a minor coding performance loss.
Hung-Chih Lin, Hsueh-Ming Hang, Wen-Hsiao Peng
IEEE Trans. Image Process.3
2010 On the analysis and design of motion sampling structure for advanced motion-compensated prediction
abstract
This paper addresses the problem of improving motion sampling efficiency for motion-compensated prediction (MCP). We provide a theoretical framework for analyzing the effect of motion sampling structure on MCP efficiency. It is shown that the sampling grid in-duced by the quadtree partition in H.264/AVC is suboptimal. To improve sampling efficiency, we propose a new pattern, which pro-vides sampling points at both the center and the top-left corner of a macroblcok. When contrasted with conventional NxN/2 block par-tition, the proposed scheme performs consistently and significantly better in subjective and objective quality. Index Terms — Motion Sampling, H.264/AVC, OBMC 1.
Yu-Chen Tseng, Chung-Hao Wu, Yi-Wen Chen, Tse-Wei Wang, Wen-Hsiao Peng
ICIP5
2010 Analysis of template matching prediction and its application to parametric overlapped block motion compensation
abstract
Template matching prediction (TMP), which estimates the motion for a target block by using its surrounding pixels, has been observed to perform efficiently in inter-frame coding. In this paper, we expose, from a more theoretical viewpoint, the factors that determine the prediction efficiency of TMP. It is shown that the motion estimate found by template matching tends to be the motion associated with the template centroid and that TMP consistently outperforms SKIP prediction, but hardly competes with block motion compensation (BMC) unless both the motion and intensity fields are less random or have high spatial correlation. We also demonstrate how template and block motion estimates can jointly be applied in a parametric overlapped block motion compensation (OBMC) framework to further improve temporal prediction. Preliminary results show that combining TMP with OBMC can yield 2-16% reductions in mean-square prediction error, as compared with the single use of OBMC. The gain is even higher (18%) when the performance is compared with that of the standard BMC.
Tse-Wei Wang, Yi-Wen Chen, Wen-Hsiao Peng
ISCAS3
2010 Bit-plane compressive sensing with Bayesian decoding for lossy compression
abstract
This paper addresses the problem of reconstructing a compressively sampled sparse signal from its lossy and possibly insufficient measurements. The process involves estimations of sparsity pattern and sparse representation, for which we derived a vector estimator based on the Maximum a Posteriori Probability (MAP) rule. By making full use of signal prior knowledge, our scheme can use a measurement number close to sparsity to achieve perfect reconstruction. It also shows a much lower error probability of sparsity pattern than prior work, given insufficient measurements. To better recover the most significant part of the sparse representation, we further introduce the notion of bit-plane separation. When applied to image compression, the technique in combination with our MAP estimator shows promising results as compared to JPEG: the difference in compression ratio is seen to be within a factor of two, given the same decoded quality.
Sz-Hsien Wu, Wen-Hsiao Peng, Tihao Chiang
PCS2
2010 Fast Context-Adaptive Mode Decision Algorithm for Scalable Video Coding With Combined Coarse-Grain Quality Scalability (CGS) and Temporal Scalability
abstract
To speed up the H.264/MPEG scalable video coding (SVC) encoder, we propose a layer-adaptive intra/inter mode decision algorithm and a motion search scheme for the hierarchical B-frames in SVC with combined coarse-grain quality scalability (CGS) and temporal scalability. To reduce computation but maintain the same level of coding efficiency, we examine the rate-distortion (R-D) performance contributed by different coding modes at the enhancement layers (EL) and the mode conditional probabilities at different temporal layers. For the intra prediction on inter frames, we can reduce the number of Intra4×4/Intra 8×8 prediction modes by 50% or more, based on the reference/base layer intra prediction directions. For the EL inter prediction, the look-up tables containing inter prediction candidate modes are designed to use the macroblock (MB) coding mode dependence and the reference/base layer quantization parameters (Qp). In addition, to avoid checking all motion estimation (ME) reference frames, the base layer (BL) reference frame index is selectively reused. And according to the EL MB partition, the BL motion vector can be used as the initial search point for the EL ME. Compared with Joint Scalable Video Model 9.11, our proposed algorithm provides a 20× speedup on encoding the EL and an 85% time saving on the entire encoding process with negligible loss in coding efficiency. Moreover, compared with other fast mode decision algorithms, our scheme can demonstrate a 7-41% complexity reduction on the overall encoding process.
Hung-Chih Lin, Wen-Hsiao Peng, Hsueh-Ming Hang
IEEE Trans. Circuits Syst. Video Technol.2
2009 Fast temporal prediction selection for H.264/AVC scalable video coding
abstract
In this paper, we propose a fast algorithm that selects the temporal prediction type for the dyadic hierarchical prediction structure in scalable video coding (SVC). We make use of the strong correlations in the large block partitions to eliminate the unnecessary computations for bi-directional prediction. Moreover, based upon the information of an 8 × 8 partition, either forward or backward prediction is skipped for its smaller block partitions. Comparing to the JSVM 9.11, our method saves the encoding time from 50% to 60% for a number of test videos over a typical range of coding bit-rates and its coding penalty is negligible.
Hung-Chih Lin, Hsueh-Ming Hang, Wen-Hsiao Peng
ICIP3
2009 A comparative study on attention-based rate adaptation for scalable video coding
abstract
We conduct subjective tests to evaluate the performance of scalable video coding with different spatial-domain bit-allocation methods, visual attention models, and motion feature extractors in the literature. For spatial-domain bit allocation, we use the selective enhancement and quality layer assignment methods. For characterizing visual attention, we use the motion attention model and perceptual quality significant map. For motion features, we adopt motion vectors from hierarchical B-picture coding and optical flow. Experimental results show that a more accurate visual attention model leads to better perceptual quality. In cooperation with a visual attention model, the selective enhancement method, compared to the quality layer assignment, achieves better subjective quality when an ROI has enough bit allocation and its texture is not complex. The quality layer assignment method is suitable for region-wise quality enhancement due to its frame-based allocation nature.
Chia-Ming Tsai, Chia-Wen Lin, Weisi Lin, Wen-Hsiao Peng
ICIP4
2009 A Synthesis-Quality-Oriented Depth Refinement Scheme for MPEG Free Viewpoint Television (FTV)
abstract
This paper addresses the problem of refining depth information from the received reference and depth images within the MPEG FTV framework. An analytical model is first developed to approximate the per-pixel synthesis distortion (caused by depth-image compression) as a function of depth-error variances, intensity variations, ground-truth depth and virtual camera locations. We then follow the model to detect unreliable depth pixels by inspecting intensity gradients and to refine their values with a candidate-based block disparity search. Additional side information is transmitted to make both operations robust against compression effects. Experimental results show that our scheme offers an average PSNR improvement of 1.2 dB over MPEG FTV and consistently outperforms the state-of-the-art methods. Moreover, it can remove synthesis artifacts to a great extent, producing a result that is very close in appearance to the ground-truth view image.
Chun-Chi Chen, Yi-Wen Chen, Fu-Yao Yang, Wen-Hsiao Peng
ISM4
2009 A parametric window design for OBMC with variable block size motion estimates
abstract
This paper addresses the problem of adapting overlapped block motion compensation (OBMC) windows for use with variable blocksize motion estimates. We tackle the problem by using a parametric window design, which expresses, based on a statistical motion model, the optimal weights as a function of the distances between the predicted pixel and its nearby block centers. The formula enables both prediction weights and prediction order to be adapted on a pixelbypixel basis. Extensive experiments have been conducted using JM 12.4. Compared with conventional block motion compensation, our scheme shows a bitrate saving of 18% (5% on average) while maintaining the same or even higher PSNR (0.1 dB). It also provides a competitive advantage to variable block size motion compensation. Additionally, a hybrid of the two techniques achieves a further bitrate reduction of 13%. The result, nevertheless, provides only a lower bound on what is achievable since both motion estimation and mode decision were accomplished without considering OBMC. Further improvement is expected by incorporating iterative methods.
Yi-Wen Chen, Tse-Wei Wang, Yu-Chen Tseng, Wen-Hsiao Peng, Suh-Yin Lee
MMSP4
2008 Multidimensional SVC bitstream adaptation and extraction for rate-distortion optimized heterogeneous multicasting and playback
abstract
In this paper, we propose an optimal SVC bitstream extraction scheme that can produce scalable layer representations for different viewing devices scattered over a multicasting network with diverse link bandwidth. Our scheme can determine optimal extraction orders/paths that are implementable at different multicasting nodes by performing successive steps of multiple adaptation. In addition to this basic scheme, we also develop an unambiguous denotation for the optimal extraction paths, and an algorithm for deducing the optimal extraction paths for less capable devices through path truncation. Extensive SVC encoding and adaptation experiments have been performed using JSVM 9 and both objective quality metrics (PSNR and MSE) as well as subjective metric (VQM). The experiment results showed that our scheme works well if convexity of R-D performance is maintained throughout the SVC bitstream.
Wen-Hsiao Peng, John Kar-Kin Zao, Hsueh-Ting Huang, Tse-Wei Wang, Lun-Chia Kuo
ICIP1
2008 Design space exploration of an H.264/AVC-based video embedding transcoder using transaction level modeling
abstract
In this paper, we perform the design space exploration for an H.264/AVC video embedding transcoder. Specifically, the design space is pruned for the sub-modules including inverse transform, inter and intra prediction, and deblocking filter with various microarchitecture designs, processing order, memory hierarchy, and granularity of synchronization. In addition, we propose an efficient deblocking filter suitable for 8x8 block pipeline. Compared to the traditional designs, our proposed deblocking filter reduces memory requirement, processing latency, and access frequency to the local memory. The synthesized logic gate count is only 8K using the 0.18 um technology with the maximum frequency of 162 MHz. For rapid exploration, all the design alternatives are simulated with higher level of abstraction using the transaction level modeling to explore 160 design combinations. Our simulation results provide an extensive tradeoff analysis among processing speed, cost, and utilization. Besides, the cost-normalized hardware utilization where the cost of each sub-module weights its associated utilization assists the system designers to keep a balance across different modules.
Chih-Hung Li, Wen-Hsiao Peng, Tihao Chiang
ICME2
2008 A fast mode decision algorithm with macroblock-adaptive rate-distortion estimation for intra-only scalable video coding
abstract
In this paper, we propose a fast mode decision algorithm with macroblock-adaptive rate-distortion (R-D) estimation for intra-only scalable video coding (SVC). We make use of the log-linear R-D relationship of inter-dependent layers to predict the better performer among the Intra4times4 and Intra8times8 prediction types at the enhancement layers. Based upon the base-layer chosen prediction type, we can further reduce the number of candidate modes. In addition, to ensure the best trade-off between complexity and coding efficiency, the Intra16times16 prediction is retained and enabled only for coding high-resolution videos with smooth image contents. Comparing to the joint scalable video model v.8 (JSVM 8), an encoder time saving from 49% to 64% depending on the encoder configurations, is achieved with negligible penalty in coding efficiency.
Hung-Chih Lin, Wen-Hsiao Peng, Hsueh-Ming Hang
ICME2
2008 A reconfigurable video embedding transcoder based on H.264/AVC: Design tradeoffs and analysis
abstract
In this paper, we propose a system architecture for H.264/AVC video embedding transcoder (VET). In addition, the proposed platform-based design can seamlessly combine the MW-VET and decoder such that it can be dynamically configured to perform video decoding and transcoding alternatively or simultaneously. Furthermore, we perform the pruned design space exploration on the design of inter/intea prediction and the on-chip data bus width. Our proposed architecture provides a better tradeoff among execution cycles, hardware cost, resource utilization, and video quality because of the reconfigurable processing modules and the hybrid pipelining. As compared to the cascaded pixel domain transcoder that has the highest complexity, our hardware efficient VET can significantly reduce the hardware cost while maintaining similar rate-distortion performance. Finally, the proposed architecture is verified at system level using transaction level modeling (TLM) technique. From the simulation results, the proposed architecture with the best tradeoff configuration can achieve a transcoding rate up to 358 frames per second for SD video source while clocking at 162MHz.
Chih-Hung Li, Wen-Hsiao Peng, Tihao Chiang
ISCAS2
2008 A rate-distortion optimization model for SVC inter-layer encoding and bitstream extraction
Wen-Hsiao Peng, John Kar-Kin Zao, Hsueh-Ting Huang, Tse-Wei Wang, Lin-Shung Huang
J. Vis. Commun. Image Represent.1
2007 Layer-Adaptive Mode Decision and Motion Search for Scalable Video Coding with Combined Coarse Granular Scalability (CGS) and Temporal Scalability
abstract
In this paper, we propose a layer-adaptive mode decision algorithm and a motion search scheme for the scalable video coding (SVC) with combined coarse granular scalability (CGS) and temporal scalability. To speed up the encoder while minimizing the loss in coding efficiency, our layer-adaptive mode decision recursively refers to the prediction modes and quantization parameter of the reference/base layer to minimize the number of modes tested at the enhancement layer. Moreover, our motion search scheme adaptively reuses the reference frame indices of the base layer and determines the initial search point using the motion vector at the base layer or the motion vector predictor at the enhancement layer. As compared with JSVM 8, the proposed algorithms provide up to 75% overall time saving and more than 85% time reduction for encoding enhancement layers with negligible loss in coding efficiency.
Hung-Chih Lin, Wen-Hsiao Peng, Hsueh-Ming Hang, Wen-Jen Ho
ICIP (2)2
2006 Adding selective enhancement in scalable video coding for region-of-interest functionality
abstract
In this paper, we propose a lossless, graceful, and arbitrary-shaped region-of-interest (ROI) functionality for the scalable video coding in MPEG-4 Part 10 Amd. 1. Specifically, we propose a prioritized block coding scheme and a layer remapping technique. Within an enhancement layer, the prioritized block coding reshuffles the transform coefficients for ROI. Moreover, the layer remapping technique prioritizes different ROI among enhancement layers. For graceful and arbitrary-shaped properties, additional syntax are coded at the levels of slice, layer, and macroblock. To minimize overhead, an efficient representation for priority information is proposed. Experimental results show that our schemes can offer the ROI functionality over a wide range of bit rates while maintaining coding efficiency.
Wen-Hsiao Peng, Tihao Chiang, Hsueh-Ming Hang
ISCAS1
2006 Trickle: Resilient Real-Time Video Multicasting for Dynamic Peers with Limited or Asymmetric Network Connectivity
abstract
Some of the most challenging scenarios for peer-to-peer multimedia applications arise when the applications require real-time interactions among their users. In those cases, the expectation of sub-second responses prohibits the use of popular P2P IPTV software because those programs invariably use large video buffers to amortize the propagation delays of individual frames and thus cause notable and dispersed viewing latencies among their users. The performance of these programs degrade even further if the users are connected to home networks that offer narrow uplink channels or through wireless links that experience frequent throughput fluctuations. In order to overcome these shortcomings, we develop Trickle, a peer-to-peer real-time media streaming system that can transport H.264 video streams with low link stresses (less than 250Kb/s) and stable sub-second frame delays through the use of erasure correction codes along with the clever construction of multiple multicast trees and the recruitment of many peer helpers. This paper presents the first fruits of our work including the principles and mechanisms of Trickle, its simulated performance based on H.264 video traces and its merit comparisons against SplitStream, the first application layer multicasting protocol for video streaming, and CoolStreaming, a news-making P2P IPTV program that works like BitTorrent
Yu-Hsuang Guo, John Kar-Kin Zao, Wen-Hsiao Peng, Lin-Shung Huang, Fang-Po Kuo, Che-Min Lin
ISM3
2006 A unified systolic architecture for combined inter and intra predictions in H.264/AVC decoder
abstract
This paper presents a unified systolic architecture for inter and intra predictions in H.264/AVC decoder. To increase hardware utilization and minimize cost, we combine inter and intra predictions by a reprogrammable FIR filter, which is further implemented using systolic array. For intra prediction, the boundary pixels are reshuffled before feeding into the systolic array. For inter prediction, the 2-D interpolation is conducted through separable 1-D filtering. As compared with the state-of-the-art approaches, our architecture provides higher performance while maintaining relatively lower cost and input bandwidth. Specifically, up to 4x throughput improvement has been achieved. Moreover, the input bandwidth is significantly reduced. Further, combining inter and intra predictions saves the cost by 22~88%.
Chih-Hung Li, Chih-Chieh Chen, Drew Wei-Chi Su, Ming-Jiun Wang, Wen-Hsiao Peng, Tihao Chiang, Gwo Giun Lee
IWCMC5
2006 A Context Adaptive Bit-Plane Coder With Maximum-Likelihood-Based Stochastic Bit-Reshuffling Technique for Scalable Video Coding
abstract
In this paper, we propose a context adaptive bit-plane coding (CABIC) with a stochastic bit reshuffling (SBR) scheme to deliver higher coding efficiency and better subjective quality for fine granular scalable (FGS) video coding. Traditional bit-plane coding in FGS algorithm suffers from poor coding efficiency and subjective quality. To improve coding efficiency, our CABIC constructs context models based on both the energy distribution in a block and the spatial correlations in the adjacent blocks. Moreover, it exploits the context across bit-planes to save side information. To improve subjective quality, our SBR reorders the coefficient bits by their estimated rate-distortion performance. Particularly, we model transform coefficients with Laplacian distributions and incorporate them into the context probability models for content-aware parameter estimation. Moreover, our SBR is implemented with a dynamic priority management that uses a low-complexity dynamic memory organization. Experimental results show that our CABIC improves the PSNR by 0.5/spl sim/1.0 dB at medium and high bit rates. While maintaining similar or even higher coding efficiency, our SBR improves the subjective quality.
Wen-Hsiao Peng, Tihao Chiang, Hsueh-Ming Hang, Chen-Yi Lee
IEEE Trans. Multim.1
2005 Advances of MPEG Scalable Video Coding Standard
Wen-Hsiao Peng, Chia-Yang Tsai, Tihao Chiang, Hsueh-Ming Hang
KES (4)1
2002 Error drifting reduction in enhanced fine granularity scalability
abstract
We incorporate fading and reset mechanisms in an enhanced fine granularity scalability algorithm to reduce the drifting error at low bit rate while still maintaining 1.5dB PSNR gain at high bit rate over the current MPEG-4 fine granularity scalability. Many previous works use enhancement layers to predict enhancement layers so as to increase the compression efficiency. Drifting error occurs because the enhancement layer, the predictor, is not received as expected. Our fading mechanism linearly combines the current reconstructed base layer and previously reconstructed enhancement layer with fading factors between 0 and 1. Our reset mechanism sets the reference frame for prediction to be the base layer periodically. Our theoretical formulation and experimental results show that drifting error can be distributed more uniformly and maximum accumulated mismatch error is significantly reduced while our mechanisms are turned on. Around 1dB can be improved at low bit rate comparing to the one without any drifting reduction mechanism.
Wen-Hsiao Peng, Yen-Kuang Chen
ICIP (2)1