EDBT 2026 Demo / reviewers in the wild / expert
Martin Benjak
dblp:291/4087
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2025
0009-0009-4303-3623ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploration of Sequence-wise Optimized Parameters for Low Complexity Enhancement Video Coding (LCEVC) on 4K ContentabstractThis paper explores the configuration space of Low Complexity Enhancement Video Coding (LCEVC) for 4K SDR content and the VVC reference software VTM as base layer codec. The configuration space is spanned using a sweep over multiple values for the base layer QP and the LCEVC parameters sublayer 2 step width, scaling mode, number of processed picture planes, transformation type, and sequence-wise optimized upscaling filter coefficients. The best configurations are selected using a convex hull approach. Compared to the default LCEVC configuration, we achieve average BD-rate gains of 10.92% and 2.06% using 1d and 2d upscaling, respectively. However, our best possible LCEVC configuration still yields an average YCbCr BD-rate loss of 15.26% compared to VTM at full resolution. Martin Benjak, Jörn Ostermann |
ICASSP | 1 |
| 2025 | HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video CodingabstractMost frame-based learned video codecs can be interpreted as recurrent neural networks (RNNs) propagating reference information along the temporal dimension. This work revisits the limitations of the current approaches from an RNN perspective. The output-recurrence methods, which propagate decoded frames, are intuitive but impose dual constraints on the output decoded frames, leading to suboptimal rate-distortion performance. In contrast, the hidden-to-hidden connection approaches, which propagate latent features within the RNN, offer greater flexibility but require large buffer sizes. To address these issues, we propose HyTIP, a learned video coding framework that combines both mechanisms. Our hybrid buffering strategy uses explicit decoded frames and a small number of implicit latent features to achieve competitive coding performance. Experimental results show that our HyTIP outperforms the sole use of either output-recurrence or hidden-to-hidden approaches. Furthermore, it achieves comparable performance to state-of-the-art methods but with a much smaller buffer size, and outperforms VTM 17.0 (Low-delay B) in terms of PSNR-RGB and MS-SSIM-RGB. The source code of HyTIP is available at https://github.com/NYCU-MAPL/HyTIP. Yi-Hsin Chen, Yi-Chen Yao, Kuan-Wei Ho, Chun-Hung Wu, Huu-Tai Phung, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
ICCV | 6 |
| 2025 | Learned Hybrid Video Coding for Human Perception and Multiple Machine Vision TasksabstractIn this work, we present a learned multi-task video codec that is optimized for human and machine vision. The codec consists of an encoder that maps images from the pixel domain to a latent representation and multiple decoders that map the latent to either an image for human consumption or multiple task-specific features for different machine vision tasks. This allows a single bitstream to be used for multiple tasks while also reducing the decoder complexity for machine vision tasks. Unlike most learned codecs, our method performs inter-coding at the latent level instead of the pixel domain. Experiments show that the proposed method achieves a compression performance for machine vision tasks comparable to other multi-task codecs designed for machine vision only, while also providing video reconstruction. The code is available at https://github.com/GreenAutoML4FAS/HybridMultiTaskCoding. Martin Benjak, Saifullah Khan, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
ICIP | 1 |
| 2025 | Conditional Residual Coding with Explicit-Implicit Temporal Buffering for Learned Video CompressionabstractThis work proposes a hybrid, explicit-implicit temporal buffering scheme for conditional residual video coding. Recent conditional coding methods propagate implicit temporal information for inter-frame coding, demonstrating superior coding performance to those relying exclusively on previously decoded frames (i.e. the explicit temporal information). However, these methods require substantial memory to store a large number of implicit features. This work presents a hybrid buffering strategy. For inter-frame coding, it buffers one previously decoded frame as the explicit temporal reference and a small number of learned features as implicit temporal reference. Our hybrid buffering scheme for conditional residual coding outperforms the single use of explicit or implicit information. Moreover, it allows the total buffer size to be reduced to the equivalent of two video frames with a negligible performance drop on 2K video sequences. The ablation experiment further sheds light on how these two types of temporal references impact the coding performance. Yi-Hsin Chen, Kuan-Wei Ho, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
ICME | 3 |
| 2025 | Scalable COOL-CHIC: Dual-Resolution Images from a Single Bitstream
Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
PCS | 1 |
| 2025 | A Cross-Framework Study of Temporal Information Buffering Strategies for Learned Video Compression
Kuan-Wei Ho, Yi-Hsin Chen, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
PCS | 3 |
| 2025 | Progressive COOL-CHIC: Efficient Decoding for Dual-Resolution ImagesabstractIn this work, we propose Progressive Cool-Chic (PCC), a scalable overfitted neural image codec that can decode an image at two different resolutions from a single bitstream. Experiments show that our method reduces the necessary bitrate to encode two representations of the same image by up to 31.54% in terms of BD-rate compared to encoding both representations independently using Cool-Chic while also decreasing the necessary decoding time. The bitstream is structured in a way that the low-resolution image can already be decoded, when only a part of the bitstream has been received. The code is available at https://github.com/mbenjak/progressive-CC. Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
VCIP | 1 |
| 2024 | On the Rate-Distortion-Complexity Trade-Offs of Neural Video CodingabstractThis paper aims to delve into the rate-distortion-complexity trade-offs of modern neural video coding. Recent years have witnessed much research effort being focused on exploring the full potential of neural video coding. Conditional auto encoders have emerged as the mainstream approach to efficient neural video coding. The central theme of conditional auto encoders is to leverage both spatial and temporal information for better conditional coding. However, a recent study indicates that conditional coding may suffer from information bottlenecks, potentially performing worse than traditional residual coding. To address this issue, recent conditional coding methods incorporate a large number of high-resolution features as the condition signal, leading to a considerable increase in the number of multiply-accumulate operations, memory footprint, and model size. Taking DCVC as the common code base, we investigate how the newly proposed conditional residual coding, an emerging new school of thought, and its variants may strike a better balance among rate, distortion, and complexity. Yi-Hsin Chen, Kuan-Wei Ho, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
MMSP | 3 |
| 2024 | MaskCRT: Masked Conditional Residual Transformer for Learned Video CompressionabstractConditional coding has lately emerged as the mainstream approach to learned video compression. However, a recent study shows that it may perform worse than residual coding when the information bottleneck arises. Conditional residual coding was thus proposed, creating a new school of thought to improve on conditional coding. Notably, conditional residual coding relies heavily on the assumption that the residual frame has a lower entropy rate than that of the intra frame. Recognizing that this assumption is not always true due to dis-occlusion phenomena or unreliable motion estimates, we propose a masked conditional residual coding scheme. It learns a soft mask to form a hybrid of conditional coding and conditional residual coding in a pixel adaptive manner. We introduce a Transformer-based conditional autoencoder. Several strategies are investigated with regard to how to condition a Transformer-based autoencoder for inter-frame coding, a topic that is largely under-explored. Additionally, we propose a channel transform module (CTM) to decorrelate the image latents along the channel dimension, with the aim of using the simple hyperprior to approach similar compression performance to the channel-wise autoregressive model. Experimental results confirm the superiority of our masked conditional residual transformer (termed MaskCRT) to both conditional coding and conditional residual coding. On commonly used datasets, MaskCRT shows comparable BD-rate results to VTM-17.0 under the low delay P configuration in terms of PSNR-RGB and outperforms VTM-17.0 in terms of MS-SSIM-RGB. It also opens up a new research direction for advancing learned video compression. Yi-Hsin Chen, Hong-Sheng Xie, Cheng-Wei Chen, Zong-Lin Gao, Martin Benjak, Wen-Hsiao Peng, Jörn Ostermann |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Learning-Based Scalable Video Coding with Spatial and Temporal PredictionabstractIn this work, we propose a hybrid learning-based method for layered spatial scalability. Our framework consists of a base layer (BL), which encodes a spatially downsampled representation of the input video using Versatile Video Coding (VVC), and a learning-based enhancement layer (EL), which conditionally encodes the original video signal. The EL is conditioned by two fused prediction signals: a spatial inter-layer prediction signal, that is generated by spatially upsampling the output of the BL using super-resolution, and a temporal inter-frame prediction signal, that is generated by decoder-side motion compensation without signaling any motion vectors. We show that our method outperforms LCEVC and has comparable performance to full-resolution VVC for high-resolution content, while still offering scalability. Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
VCIP | 1 |
| 2023 | Rate Adaptation for Learned Two-layer B-frame Coding without Signaling Motion InformationabstractThis paper explores the potential of a learned two-layer B-frame codec, known as TLZMC. TLZMC is one of the few early attempts that deviate from the hybrid-based coding architecture by skipping motion coding. With TLZMC, a low-resolution base layer is utilized to encode temporally unpredictable information. We address the question of whether adapting the base-layer bitrate can achieve better rate-distortion performance. We apply the feature map modulation technique to enable per-frame bitrate adaptation of the base layer. We then propose and compare three online search strategies for determining the base-layer rate parameter: per-level brute-force search, per-level greedy search, and per-frame greedy search. Experimental results show that our top-performing search strategy achieves 0.6%-15.8% Bjøntegaard-Delta rate savings over TLZMC. Hong-Sheng Xie, Yi-Hsin Chen, Wen-Hsiao Peng, Martin Benjak, Jörn Ostermann |
VCIP | 4 |
| 2022 | Neural Network-based Error Concealment for B-Frames in VVCabstractIn this paper we introduce an error concealment method for VVC that error-conceals B-frames based on the neural frame interpolation network RIFE. The network is trained using the BVI-DVC dataset to infer even full-HD frames. We integrate our proposed model in the VVC reference software VTM for its evaluation. The average error of a whole GOP with a single corrupted frame is decreased by 15% and 24% in terms of PSNR measurement compared to block matching and frame copy, respectively. To our knowledge, our approach is currently the best performing error concealment algorithm for single slice per B-frame settings. Martin Benjak, Niklas Aust, Yasser Samayoa, Jörn Ostermann |
ISCAS | 1 |
| 2021 | Neural Network-Based Error Concealment For VVCabstractIn this paper we introduce an error concealment method for VVC based on deep recurrent neural networks, which employs the PredNet model to estimate missing video frames by using past decoded frames. The network is trained using the BVI-DVC data set to infer even full-HD frames. We integrated our proposed model in the VVC reference software VTM for its evaluation. It performs, in average, 6 dB or up to 5 dB better than the frame copy model in terms of PSNR measurements for a concealed I-frame or P-frame, respectively. Martin Benjak, Yasser Samayoa, Jörn Ostermann |
ICIP | 1 |