EDBT 2026 Demo / reviewers in the wild / expert
Christopher Schroers
dblp:117/5901
· DBLP profile ↗
40ranked-venue papers
2as first author
24since 2021 · last 2025
0000-0003-1473-1878ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 23 · 14 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image CompressionabstractGenerative neural image compression supports data representation at extremely low bitrate, synthesizing details at the client and consistently producing highly realistic images. By leveraging the similarities between quantization error and additive noise, diffusion-based generative image compression codecs can be built using a latent diffusion model to "denoise" the artifacts introduced by quantization. However, we identify three critical gaps in previous approaches following this paradigm (namely, the noise level, noise type, and discretization gaps) that result in the quantized data falling out of the data distribution known by the diffusion model. In this work, we propose a novel quantization-based forward diffusion process with theoretical foundations that tackles all three aforementioned gaps. We achieve this through universal quantization with a carefully tailored quantization schedule and a diffusion model trained with uniform noise. Compared to previous work, our proposal produces consistently realistic and detailed reconstructions, even at very low bitrates. In such a regime, we achieve the best rate-distortion-realism performance, outperforming previous related works. Lucas Relic, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers |
CVPR | 5 |
| 2025 | LDIP: Long Distance Information Propagation for Video Super-ResolutionabstractVideo super-resolution (VSR) methods typically exploit information across multiple frames to achieve high quality upscaling, with recent approaches demonstrating impressive performance. Nevertheless, challenges remain, particularly in effectively leveraging information over long distances. To address this limitation in VSR, we propose a strategy for long distance information propagation with a flexible fusion module that can optionally also assimilate information from additional high resolution reference images. We design our overall approach such that it can leverage existing pre-trained VSR backbones and adapt the feature upscaling module to support arbitrary scaling factors. Our experiments demonstrate that we can achieve state-of-theart results on perceptual metrics and deliver more visually pleasing results compared to existing solutions. Michael Bernasconi, Abdelaziz Djelouah, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers |
ICCV | 5 |
| 2025 | Spatiotemporal Diffusion Priors for Extreme Video CompressionabstractDiffusion models have recently demonstrated impressive results in image compression, where the strong spatial prior enables the synthesis of fine details rather than allocating bits to transmit them. In this work, we propose to extend this paradigm to video compression by utilizing a generative spatiotemporal prior and present the first codec based on a video diffusion model. Our method operates by performing longcontext interpolation guided by sparse inter-frame predictions, thus requiring minimal motion information. To this end, we develop a sparse, bidirectional optical flow which serves as a bitrate-efficient motion conditioning in the diffusion decoding process. The resulting codec can compress videos to extremely low rates (as low as 0.01 bits per pixel) while maintaining realistic textures and motion, and outperforms both neural and traditional baselines on several benchmark datasets. Our method shows state-of-the art performance in perceptually-oriented distortion metrics, and, when considering rate-realism, we achieve an improvement in FID score of up to 73.3 at the same bitrate compared to the leading traditional video codec, VTM. Overall, we present an important first work examining spatiotemporal diffusion priors for video compression. Lucas Relic, André Emmenegger, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers |
PCS | 6 |
| 2025 | CLIP-Fusion: A Spatio-Temporal Quality Metric for Frame Interpolation
Göksel Mert Çökmez, Yang Zhang 0003, Christopher Schroers, Tunç Ozan Aydin |
WACV | 3 |
| 2024 | CoARF: Controllable 3D Artistic Style Transfer for Radiance FieldsabstractCreating artistic 3D scenes can be time-consuming and requires specialized knowledge. To address this, recent works such as ARF [57], use a radiance field-based approach with style constraints to generate 3D scenes that resemble a style image provided by the user. However, these methods lack fine-grained control over the resulting scenes. In this paper, we introduce Controllable Artistic Radiance Fields (CoARF), a novel algorithm for controllable 3D scene stylization. CoARF enables style transfer for specified objects, compositional 3D style transfer and semantic-aware style transfer. We achieve controllability using segmentation masks with different label-dependent loss functions. We also propose a semantic-aware nearest neighbor matching algorithm to improve the style transfer quality. Our extensive experiments demonstrate that CoARF provides user-specified controllability of style transfer and superior style transfer quality with more precise feature matching. Deheng Zhang, Clara Fernandez-Labrador, Christopher Schroers |
3DV | 3 |
| 2024 | Revitalizing Legacy Video Content: Deinterlacing with Bidirectional Information Propagation
Zhaowei Gao, Christopher Schroers, Yang Zhang 0003 |
BMVC | 3 |
| 2024 | DiVAS: Video and Audio Synchronization with Dynamic Frame RatesabstractSynchronization issues between audio and video are one of the most disturbing quality defects in film production and live broadcasting. Even a discrepancy as short as 45 milliseconds can degrade the viewer's experience enough to warrant manual quality checks over entire movies. In this paper, we study the automatic discovery of such issues. Specifically, we focus on the alignment of lip movements with spoken words, targeting realistic production scenarios which can include background noise and music, intricate head poses, excessive makeup, or scenes with multiple individuals where the speaker is unknown. Our model's robustness also extends to various media specifications, including different video frame rates and audio sample rates. To address these challenges, we present a model fully based on Transformers that encodes face crops or full video frames and raw audio using timestamp information, identifies the speaker and provides highly accurate synchronization predictions much faster than previous methods. Clara Fernandez-Labrador, Mertcan Akçay, Eitan Abecassis, Joan Massich Vall, Christopher Schroers |
CVPR | 5 |
| 2024 | QUADify: Extracting Meshes with Pixel-Level Details and Materials from ImagesabstractDespite exciting progress in automatic 3D reconstruction from images, excessive and irregular triangular faces in the resulting meshes still constitute a significant challenge when it comes to adoption in practical artist work-flows. Therefore, we propose a method to extract regular quad-dominant meshes from posed images. More specifically, we generate a high-quality 3D model through de-composition into an easily editable quad-dominant mesh with pixel-level details such as displacement, materials, and lighting. To enable end-to-end learning of shape and quad topology, we QUADify a neural implicit representation using our novel differentiable re-meshing objective. Distinct from previous work, our method exploits artifact-free Catmull-Clark subdivision combined with vertex displacement to extract pixel-level details linked to the base geom-etry. Finally, we apply differentiable rendering techniques for material and lighting decomposition to optimize for image reconstruction. Our experiments show the benefits of end-to-end re-meshing and that our method yields state-of-the-art geometric accuracy while providing lightweight meshes with displacements and textures that are directly compatible with professional renderers and game engines. Maximilian Frühauf, Hayko Riemenschneider, Markus Gross 0001, Christopher Schroers |
CVPR | 4 |
| 2024 | Combining Frame and GOP Embeddings for Neural Video RepresentationabstractImplicit neural representations (INRs) were recently proposed as a new video compression paradigm, with existing approaches performing on par with HEVC. However, such methods only perform well in limited settings, e.g., specific model sizes, fixed aspect ratios, and low-motion videos. We address this issue by proposing T-NeRV, a hybrid video INR that combines framespecific embeddings with GOP-specific features, providing a lever for content-specific fine-tuning. We employ entropy-constrained training to jointly optimize our model for rate and distortion and demonstrate that T-NeRV can thereby automatically adjust this lever during training, effectively fine-tuning itself to the target content. We evaluate T-NeRVon the UVG dataset, where it achieves state-of-the-art results on the video representation task, outperforming previous works by up to 3dB PSNR on challenging high-motion sequences. Further, our method improves on the compression performance of pre vious methods and is the first video INR to outperform HEVC on all UVG sequences. Jens Eirik Saethre, Roberto Azevedo, Christopher Schroers |
CVPR | 3 |
| 2024 | Lossy Image Compression with Foundation Diffusion ModelsabstractAbstract Incorporating diffusion models in the image compression domain has the potential to produce realistic and detailed reconstructions, especially at extremely low bitrates. Previous methods focus on using diffusion models as expressive decoders robust to quantization errors in the conditioning signals. However, achieving competitive results in this manner requires costly training of the diffusion model and long inference times due to the iterative generative process. In this work we formulate the removal of quantization error as a denoising task, using diffusion to recover lost information in the transmitted image latent. Our approach allows us to perform less than 10% of the full diffusion generative process and requires no architectural changes to the diffusion model, enabling the use of foundation models as a strong prior without additional fine tuning of the backbone. Our proposed codec outperforms previous methods in quantitative realism metrics, and we verify that our reconstructions are qualitatively preferred by end users, even when other methods use twice the bitrate. Lucas Relic, Roberto Azevedo, Markus Gross 0001, Christopher Schroers |
ECCV (61) | 4 |
| 2024 | RAST: A Reference-Audio Synchronization Tool for Dubbed Content
David Meyer, Eitan Abecassis, Clara Fernandez-Labrador, Christopher Schroers |
INTERSPEECH | 4 |
| 2024 | Efficient Video Encoder Autotuning via Offline Bayesian Optimization and Supervised LearningabstractModern video encoders are complex software containing dozens of parameters, which allows them to be configured to different scenarios, requirements, or specific titles or scenes. Besides the number of parameters, the inter-dependency between them adds to the complexity of finding a per-title optimized combination of encoding parameters. Even though good practices in the industry have emerged, with the definition of presets per content type (e.g., film vs. cartoon), such practices are suboptimal for specific titles or scenes. Indeed, finding the best encoding parameters for a piece of content is currently a mix of best practices and trial-and-error artwork. We propose an efficient video encoder autotuner based on offline Bayesian optimization and supervised machine learning. Our proposal uses Bayesian optimization to search for a per-title best encoding parameter set offline to generate a dataset. Then, we use the generated dataset to train machine learning models that can map features extracted from the content to the best encoding parameters. Our experiments show that our generated dataset can find a combination of parameters that improves up to approximately −14.49% BD-Rate (0.77 BD-PSNR) and −11.59% BD-Rate (2.12 BD-VMAF) when optimizing for PSNR and VMAF, respectively. In comparison, our prediction models can recover ~80% of such performance while requiring only one fast encoding (compared to hundreds of encodes of a search optimization). Roberto Azevedo, Yuanyi Xue, Xuewei Meng, Scott Labrozzi, Christopher Schroers |
MMSP | 6 |
| 2024 | BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth EstimationabstractBy training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhibit a superior ability to extract details, they struggle in geometrically complex scenes that challenge their geometry prior, trained on less diverse 3D data. To leverage the complementary merits of both worlds, we propose BetterDepth to achieve geometrically correct affine-invariant MDE while capturing fine details. Specifically, BetterDepth is a conditional diffusion-based refiner that takes the prediction from pre-trained MDE models as depth conditioning, in which the global depth layout is well-captured, and iteratively refines details based on the input image. For the training of such a refiner, we propose global pre-alignment and local patch masking methods to ensure BetterDepth remains faithful to the depth conditioning while learning to add fine-grained scene details. With efficient training on small-scale synthetic datasets, BetterDepth achieves state-of-the-art zero-shot MDE performance on diverse public datasets and on in-the-wild scenes. Moreover, BetterDepth can improve the performance of other MDE models in a plug-and-play manner without further re-training. Xiang Zhang 0022, Bingxin Ke, Hayko Riemenschneider, Nando Metzger, Anton Obukhov, Markus Gross 0001, Konrad Schindler, Christopher Schroers |
NeurIPS | 8 |
| 2024 | Stereo Conversion with Disparity-Aware Warping, Compositing and InpaintingabstractDespite of exciting advances in image-based rendering and novel view synthesis, it is still challenging to achieve high-resolution results that can reach production-level quality when applying such methods to the task of stereo conversion. At the same time, only very few dedicated stereo conversion approaches exist, which also fall short in terms of the required quality. Hence, in this paper, we present a novel method for high-resolution 2D-to-3D conversion. It is fully differentiable in all of its stages and performs disparity-informed warping, consistent foreground-background compositing, and background-aware inpainting. To enable temporal consistency in the resulting video, we propose a strategy to integrate information from additional video frames. Extensive ablation studies validate our design choices, leading to a fully automatic model that outperforms existing approaches by a large margin (49-70% LPIPS error reduction). Finally, inspired from current practices in manual stereo conversion, we introduce optional interactive tools into our model, which allow to steer the conversion process and make it significantly more applicable for 3D film production. Lukas Mehl, Andrés Bruhn, Markus Gross 0001, Christopher Schroers |
WACV | 4 |
| 2023 | Kernel Aware ResamplerabstractDeep learning based methods for super-resolution have become state-of-the-art and outperform traditional approaches by a significant margin. From the initial models designed for fixed integer scaling factors (e.g.$\times 2$or$\times 4)$), efforts were made to explore different directions such as modeling blur kernels or addressing non-integer scaling factors. However, existing works do not provide a sound framework to handle them jointly. In this paper we propose a framework for generic image resampling that not only addresses all the above mentioned issues but extends the sets of possible transforms from upscaling to generic transforms. A key aspect to unlock these capabilities is the faithful modeling of image warping and changes of the sampling rate during the training data preparation. This allows a localized representation of the implicit image degradation that takes into account the reconstruction kernel, the local geometric distortion and the anti-aliasing kernel. Using this spatially variant degradation map as conditioning for our resampling model, we can address with the same model both global transformations, such as upscaling or rotation, and locally varying transformations such lens distortion or undistortion. Another important contribution is the automatic estimation of the degradation map in this more complex resampling setting (i.e. blind image resampling). Fi-nally, we show that state-of-the-art results can be achieved by predicting kernels to apply on the input image instead of direct color prediction. This renders our model applicable for different types of data not seen during the training such as normals. Michael Bernasconi, Abdelaziz Djelouah, Farnood Salehi, Markus Gross 0001, Christopher Schroers |
CVPR | 5 |
| 2023 | Video Compression with Entropy-Constrained Neural RepresentationsabstractEncoding videos as neural networks is a recently proposed approach that allows new forms of video processing. However, traditional techniques still outperform such neural video representation (NVR) methods for the task of video compression. This performance gap can be explained by the fact that current NVR methods: i) use architectures that do not efficiently obtain a compact representation of temporal and spatial information; and ii) minimize rate and distortion disjointly (first overfitting a network on a video and then using heuristic techniques such as post-training quantization or weight pruning to compress the model). We propose a novel convolutional architecture for video representation that better represents spatio-temporal information and a training strategy capable of jointly optimizing rate and distortion. All network and quantization parameters are jointly learned end-to-end, and the post-training operations used in previous works are unnecessary. We evaluate our method on the UVG dataset, achieving new state-of-the-art results for video compression with NVRs. Moreover, we deliver the first NVR-based video compression method that improves over the typically adopted HEVC benchmark (x265, disabled b-frames, “medium” preset), closing the gap to autoencoder-based video compression techniques. Carlos Gomes, Roberto Azevedo, Christopher Schroers |
CVPR | 3 |
| 2023 | Frame Interpolation Transformer and Uncertainty GuidanceabstractVideo frame interpolation has seen important progress in recent years, thanks to developments in several directions. Some works leverage better optical flow methods with improved splatting strategies or additional cues from depth, while others have investigated alternative approaches through direct predictions or transformers. Still, the problem remains unsolved in more challenging conditions such as complex lighting or large motion. In this work, we are bridging the gap towards video production with a novel transformer-based interpolation network architecture capable of estimating the expected error together with the interpolated frame. This offers several advantages that are of key importance for frame interpolation usage: First, we obtained improved visual quality over several datasets. The improvement in terms of quality is also clearly demonstrated through a user study. Second, our method estimates error maps for the interpolated frame, which are essential for real-life applications on longer video sequences where problematic frames need to be flagged. Finally, for rendered content a partial rendering pass of the intermediate frame, guided by the predicted error, can be utilized during the interpolation to generate a new frame of superior quality. Through this error estimation, our method can produce even higher-quality intermediate frames using only a fraction of the time compared to a full rendering. Markus Plack, Matthias B. Hullin, Karlis Martins Briedis, Markus Gross 0001, Abdelaziz Djelouah, Christopher Schroers |
CVPR | 6 |
| 2023 | Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersabstractAlthough existing neural video compression~(NVC) methods have achieved significant success, most of them focus on improving either temporal or spatial information separately. They generally use simple operations such as concatenation or subtraction to utilize this information, while such operations only partially exploit spatio-temporal redundancies. This work aims to effectively and jointly leverage robust temporal and spatial information by proposing a new 3D-based transformer module: Spatio-Temporal Cross-Covariance Transformer (ST-XCT). The ST-XCT module combines two individual extracted features into a joint spatio-temporal feature, followed by 3D convolutional operations and a novel spatio-temporal-aware cross-covariance attention mechanism. Unlike conventional transformers, the cross-covariance attention mechanism is applied across the feature channels without breaking down the spatio-temporal features into local tokens. Such design allows for modeling global cross-channel correlations of the spatio-temporal context while lowering the computational requirement. Based on ST-XCT, we introduce a novel transformer-based end-to-end optimized NVC framework. ST-XCT-based modules are integrated into various key coding components of NVC, such as feature extraction, frame reconstruction, and entropy modeling, demonstrating its generalizability. Extensive experiments show that our ST-XCT-based NVC proposal achieves state-of-the-art compression performances on various standard video benchmark datasets. Lucas Relic, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Dong Xu 0001, Luping Zhou, Christopher Schroers |
ACM Multimedia | 8 |
| 2023 | Large-Scale Multi-Site Subjective Assessment on Image Banding ArtifactsabstractBanding largely occurs due to the finite bit depth representation of digital media and displays. It is a unique type of artifact that is both challenging to accurately measure and difficult to properly mitigate. These artifacts can also be observed across a wide spectrum of streaming services, no matter if professional studio productions or user-generated content. Inspired by the need to monitor and measure these banding artifacts, we present a large-scale multi-site subjective study on image banding artifacts. We designed a novel two-question-based subjective evaluation protocol that goes beyond giving ratings at the frame level, but also collects subjective opinions at the sub-frame level. The high-quality subjective data are collected from a cohort of 56 participants across 3 physical sites with calibrated environments, covering diverse backgrounds and experiences. Finally, we demonstrate that none of the common banding metrics are highly correlated with the subjective data. The paper calls for a need of continued efforts in modeling the perceptual effect of banding artifacts. Yuanyi Xue, Roberto Azevedo, Xuchang Huangfu, Yang Zhang 0003, Christopher Schroers, Scott Labrozzi |
QoMEX | 5 |
| 2022 | Contrastive Learning for Controllable Blind Video Restoration
Givi Meishvili, Abdelaziz Djelouah, Shinobu Hattori, Christopher Schroers |
BMVC | 4 |
| 2022 | Learning Dynamic 3D Geometry and Texture for Video Face SwappingabstractAbstract Face swapping is the process of applying a source actor's appearance to a target actor's performance in a video. This is a challenging visual effect that has seen increasing demand in film and television production. Recent work has shown that data‐driven methods based on deep learning can produce compelling effects at production quality in a fraction of the time required for a traditional 3D pipeline. However, the dominant approach operates only on 2D imagery without reference to the underlying facial geometry or texture, resulting in poor generalization under novel viewpoints and little artistic control. Methods that do incorporate geometry rely on pre‐learned facial priors that do not adapt well to particular geometric features of the source and target faces. We approach the problem of face swapping from the perspective of learning simultaneous convolutional facial autoencoders for the source and target identities, using a shared encoder network with identity‐specific decoders. The key novelty in our approach is that each decoder first lifts the latent code into a 3D representation, comprising a dynamic face texture and a deformable 3D face shape, before projecting this 3D face back onto the input image using a differentiable renderer. The coupled autoencoders are trained only on videos of the source and target identities, without requiring 3D supervision. By leveraging the learned 3D geometry and texture, our method achieves face swapping with higher quality than when using off‐the‐shelf monocular 3D face reconstruction, and overall lower FID score than state‐of‐the‐art 2D methods. Furthermore, our 3D representation allows for efficient artistic control over the result, which can be hard to achieve with existing 2D approaches. Christopher Otto, Jacek Naruniec, Leonhard Helminger, Thomas Etterlin, Graziana Mignone, Prashanth Chandran, Gaspard Zoss, Christopher Schroers, Markus Gross 0001, Paulo F. U. Gotardo, Derek Bradley, Romann M. Weber |
Comput. Graph. Forum | 8 |
| 2022 | Deep Adaptive Sampling and Reconstruction Using Analytic DistributionsabstractWe propose an adaptive sampling and reconstruction method for offline Monte Carlo rendering. Our method produces sampling maps constrained by a user-defined budget that minimize the expected future denoising error. Compared to other state-of-the-art methods, which produce the necessary training data on the fly by composing pre-rendered images, our method samples from analytic noise distributions instead. These distributions are compact and closely approximate the pixel value distributions stemming from Monte Carlo rendering. Our method can efficiently sample training data by leveraging only a few per-pixel statistics of the target distribution, which provides several benefits over the current state of the art. Most notably, our analytic distributions' modeling accuracy and sampling efficiency increase with sample count, essential for high-quality offline rendering. Although our distributions are approximate, our method supports joint end-to-end training of the sampling and denoising networks. Finally, we propose the addition of a global summary module to our architecture that accumulates valuable information from image regions outside of the network's receptive field. This information discourages sub-optimal decisions based on local information. Our evaluation against other state-of-the-art neural sampling methods demonstrates denoising quality and data efficiency improvements. Farnood Salehi, Marco Manzi, Gerhard Röthlin, Romann M. Weber, Christopher Schroers, Marios Papas |
ACM Trans. Graph. | 5 |
| 2021 | Microdosing: Knowledge Distillation for GAN Based Compression
Leonhard Helminger, Roberto Azevedo, Abdelaziz Djelouah, Markus Gross 0001, Christopher Schroers |
BMVC | 5 |
| 2021 | Neural frame interpolation for rendered contentabstractThe demand for creating rendered content continues to drastically grow. As it often is extremely computationally expensive and thus costly to render high-quality computer-generated images, there is a high incentive to reduce this computational burden. Recent advances in learning-based frame interpolation methods have shown exciting progress but still have not achieved the production-level quality which would be required to render fewer pixels and achieve savings in rendering times and costs. Therefore, in this paper we propose a method specifically targeted to achieve high-quality frame interpolation for rendered content. In this setting, we assume that we have full input for every n -th frame in addition to auxiliary feature buffers that are cheap to evaluate (e.g. depth, normals, albedo) for every frame. We propose solutions for leveraging such auxiliary features to obtain better motion estimates, more accurate occlusion handling, and to correctly reconstruct non-linear motion between keyframes. With this, our method is able to significantly push the state-of-the-art in frame interpolation for rendered content and we are able to obtain production-level quality results. Karlis Martins Briedis, Abdelaziz Djelouah, Mark Meyer, Ian McGonigal, Markus Gross 0001, Christopher Schroers |
ACM Trans. Graph. | 6 |
| 2020 | High-Resolution Neural Face Swapping for Visual EffectsabstractAbstract In this paper, we propose an algorithm for fully automatic neural face swapping in images and videos. To the best of our knowledge, this is the first method capable of rendering photo‐realistic and temporally coherent results at megapixel resolution. To this end, we introduce a progressively trained multi‐way comb network and a light‐ and contrast‐preserving blending method. We also show that while progressive training enables generation of high‐resolution images, extending the architecture and training data beyond two people allows us to achieve higher fidelity in generated expressions. When compositing the generated expression onto the target face, we show how to adapt the blending strategy to preserve contrast and low‐frequency lighting. Finally, we incorporate a refinement strategy into the face landmark stabilization algorithm to achieve temporal stability, which is crucial for working with high‐resolution videos. We conduct an extensive ablation study to show the influence of our design choices on the quality of the swap and compare our work with popular state‐of‐the‐art methods. Jacek Naruniec, Leonhard Helminger, Christopher Schroers, Romann M. Weber |
Comput. Graph. Forum | 3 |
| 2019 | Neural Inter-Frame Compression for Video CodingabstractWhile there are many deep learning based approaches for single image compression, the field of end-to-end learned video coding has remained much less explored. Therefore, in this work we present an inter-frame compression approach for neural video coding that can seamlessly build up on different existing neural image codecs. Our end-to-end solution performs temporal prediction by optical flow based motion compensation in pixel space. The key insight is that we can increase both decoding efficiency and reconstruction quality by encoding the required information into a latent representation that directly decodes into motion and blending coefficients. In order to account for remaining prediction errors, residual information between the original image and the interpolated frame is needed. We propose to compute residuals directly in latent space instead of in pixel space as this allows to reuse the same image compression network for both key frames and intermediate frames. Our extended evaluation on different datasets and resolutions shows that the rate-distortion performance of our approach is competitive with existing state-of-the-art codecs. Abdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher Schroers |
ICCV | 4 |
| 2019 | Light Field Synthesis Using Inexpensive Surveillance Camera SystemsabstractWe present a light field synthesis technique that achieves accurate reconstruction given a low-cost, wide-baseline camera rig. Our system integrates optical flow with methods for rectification, disparity estimation, and feature extraction, which we then feed to a neural network view synthesis solver with wide-baseline capability. We propose two novel warping methods that improve the accuracy of disparity estimation and view synthesis. The methods enable the use of off-the-shelf surveillance camera hardware in a simplified and expedited capture workflow. A thorough analysis of the process and resulting view synthesis accuracy over state of the art is provided. Frederike Dümbgen, Christopher Schroers, Kenny Mitchell |
ICIP | 2 |
| 2019 | Deep Generative Video CompressionabstractThe usage of deep generative models for image compression has led to impressive performance gains over classical codecs while neural video compression is still in its infancy. Here, we propose an end-to-end, deep generative modeling approach to compress temporal sequences with a focus on video. Our approach builds upon variational autoencoder (VAE) models for sequential data and combines them with recent work on neural image compression. The approach jointly learns to transform the original sequence into a lower-dimensional representation as well as to discretize and entropy code this representation according to predictions of the sequential VAE. Rate-distortion evaluations on small videos from public data sets with varying complexity and diversity show that our model yields competitive results when trained on generic video content. Extreme compression performance is achieved when training the model on specialized content. Salvator Lombardo, Christopher Schroers, Stephan Mandt |
NeurIPS | 3 |
| 2019 | Blind image super-resolution with spatially variant degradationsabstractExisting deep learning approaches to single image super-resolution have achieved impressive results but mostly assume a setting with fixed pairs of high resolution and low resolution images. However, to robustly address realistic upscaling scenarios where the relation between high resolution and low resolution images is unknown, blind image super-resolution is required. To this end, we propose a solution that relies on three components: First, we use a degradation aware SR network to synthesize the HR image given a low resolution image and the corresponding blur kernel. Second, we train a kernel discriminator to analyze the generated high resolution image in order to predict errors present due to providing an incorrect blur kernel to the generator. Finally, we present an optimization procedure that is able to recover both the degradation kernel and the high resolution image by minimizing the error predicted by our kernel discriminator. We also show how to extend our approach to spatially variant degradations that typically arise in visual effects pipelines when compositing content from different sources and how to enable both local and global user interaction in the upscaling process. Victor Cornillère, Abdelaziz Djelouah, Wang Yifan 0001, Olga Sorkine-Hornung, Christopher Schroers |
ACM Trans. Graph. | 5 |
| 2018 | Deep Video Color Propagation
Simone Schaub-Meyer, Victor Cornillère, Abdelaziz Djelouah, Christopher Schroers, Markus Gross 0001 |
BMVC | 4 |
| 2018 | PhaseNet for Video Frame InterpolationabstractMost approaches for video frame interpolation require accurate dense correspondences to synthesize an in-between frame. Therefore, they do not perform well in challenging scenarios with e.g. lighting changes or motion blur. Recent deep learning approaches that rely on kernels to represent motion can only alleviate these problems to some extent. In those cases, methods that use a per-pixel phase-based motion representation have been shown to work well. However, they are only applicable for a limited amount of motion. We propose a new approach, PhaseNet, that is designed to robustly handle challenging scenarios while also coping with larger motion. Our approach consists of a neural network decoder that directly estimates the phase decomposition of the intermediate frame. We show that this is superior to the hand-crafted heuristics previously used in phase-based methods and also compares favorably to recent deep learning based approaches for video frame interpolation on challenging datasets. Simone Schaub-Meyer, Abdelaziz Djelouah, Brian McWilliams, Alexander Sorkine-Hornung, Markus Gross 0001, Christopher Schroers |
CVPR | 6 |
| 2018 | Normalized Cut Loss for Weakly-Supervised CNN SegmentationabstractMost recent semantic segmentation methods train deep convolutional neural networks with fully annotated masks requiring pixel-accuracy for good quality training. Common weakly-supervised approaches generate full masks from partial input (e.g. scribbles or seeds) using standard interactive segmentation methods as preprocessing. But, errors in such masks result in poorer training since standard loss functions (e.g. cross-entropy) do not distinguish seeds from potentially mislabeled other pixels. Inspired by the general ideas in semi-supervised learning, we address these problems via a new principled loss function evaluating network output with criteria standard in "shallow" segmentation, e.g. normalized cut. Unlike prior work, the cross entropy part of our loss evaluates only seeds where labels are known while normalized cut softly evaluates consistency of all pixels. We focus on normalized cut loss where dense Gaussian kernel is efficiently implemented in linear time by fast Bilateral filtering. Our normalized cut loss approach to segmentation brings the quality of weakly-supervised training significantly closer to fully supervised methods. Meng Tang 0001, Abdelaziz Djelouah, Federico Perazzi, Yuri Boykov, Christopher Schroers |
CVPR | 5 |
| 2018 | On Regularized Losses for Weakly-supervised CNN Segmentation
Meng Tang 0001, Federico Perazzi, Abdelaziz Djelouah, Ismail Ben Ayed, Christopher Schroers, Yuri Boykov |
ECCV (16) | 5 |
| 2018 | An Omnistereoscopic Video Pipeline for Capture and Display of Real-World VRabstractIn this article, we describe a complete pipeline for the capture and display of real-world Virtual Reality video content, based on the concept of omnistereoscopic panoramas. We address important practical and theoretical issues that have remained undiscussed in previous works. On the capture side, we show how high-quality omnistereo video can be generated from a sparse set of cameras (16 in our prototype array) instead of the hundreds of input views previously required. Despite the sparse number of input views, our approach allows for high quality, real-time virtual head motion, thereby providing an important additional cue for immersive depth perception compared to static stereoscopic video. We also provide an in-depth analysis of the required camera array geometry in order to meet specific stereoscopic output constraints, which is fundamental for achieving a plausible and fully controlled VR viewing experience. Finally, we describe additional insights on how to integrate omnistereo video panoramas with rendered CG content. We provide qualitative comparisons to alternative solutions, including depth-based view synthesis and the Facebook Surround 360 system. In summary, this article provides a first complete guide and analysis for reimplementing a system for capturing and displaying real-world VR, which we demonstrate on several real-world examples captured with our prototype. Christopher Schroers, Jean-Charles Bazin, Alexander Sorkine-Hornung |
ACM Trans. Graph. | 1 |
| 2017 | Physically inspired depth-from-defocus
Nico Persch, Christopher Schroers, Simon Setzer, Joachim Weickert |
Image Vis. Comput. | 2 |
| 2016 | Point Cloud Noise and Outlier Removal for Image-Based 3D ReconstructionabstractPoint sets generated by image-based 3D reconstruction techniques are often much noisier than those obtained using active techniques like laser scanning. Therefore, they pose greater challenges to the subsequent surface reconstruction (meshing) stage. We present a simple and effective method for removing noise and outliers from such point sets. Our algorithm uses the input images and corresponding depth maps to remove pixels which are geometrically or photometrically inconsistent with the colored surface implied by the input. This allows standard surface reconstruction methods (such as Poisson surface reconstruction) to perform less smoothing and thus achieve higher quality surfaces with more features. Our algorithm is efficient, easy to implement, and robust to varying amounts of noise. We demonstrate the benefits of our algorithm in combination with a variety of state-of-the-art depth and surface reconstruction methods. Katja Wolff, Changil Kim 0001, Henning Zimmer, Christopher Schroers, Mario Botsch, Olga Sorkine-Hornung, Alexander Sorkine-Hornung |
3DV | 4 |
| 2016 | Cyclic Schemes for PDE-Based Image Analysis
Joachim Weickert, Sven Grewenig, Christopher Schroers, Andrés Bruhn |
Int. J. Comput. Vis. | 3 |
| 2014 | A Variational Taxonomy for Surface Reconstruction from Oriented PointsabstractAbstract The problem of reconstructing a watertight surface from a finite set of oriented points has received much attention over the last decades. In this paper, we propose a general higher order framework for surface reconstruction. It is based on the idea that position and normal defined by each oriented point can be used to construct an implicit local description of the unknown surface. On the one hand, this allows us to systematically explain and relate several popular methods, for example implicit moving least squares, smooth signed distance surface reconstruction as well as (screened) Poisson surface reconstruction. On the other hand, it allows to derive and discuss a number of new approaches for reconstructing either the signed distance or the indicator function of the sought object. All of these approaches are able to achieve competitive results but one of them turns out to be especially promising. To improve reconstructions in difficult real world scenarios where point clouds have been estimated from colour images, we introduce a hull constraint that encourages the surface to stay within a given region. Our framework is implemented on the GPU using a recent cyclic scheme called Fast Jacobi, which combines low implementational effort with high efficiency. Christopher Schroers, Simon Setzer, Joachim Weickert |
Comput. Graph. Forum | 1 |
| 2014 | VideoSnapping: interactive synchronization of multiple videosabstractAligning video is a fundamental task in computer graphics and vision, required for a wide range of applications. We present aninteractivemethod for computing optimal nonlinear temporal video alignments of an arbitrary number of videos. We first derive a robust approximation of alignment quality between pairs of clips, computed as a weighted histogram of feature matches. We then find optimal temporal mappings (constituting frame correspondences) using a graph-based approach that allows for very efficient evaluation with artist constraints. This enables an enhancement to the "snapping" interface in video editing tools, where videos in a time-line are now able snap to one another when dragged by an artist based on theircontent, rather than simply start-and-end times. The pairwise snapping is then generalized to multiple clips, achieving a globally optimal temporal synchronization that automatically arranges a series of clips filmed at different times into a single consistent time frame. When followed by a simple spatial registration, we achieve high quality spatiotemporal video alignments at a fraction of the computational complexity compared to previous methods. Assisted temporal alignment is a degree of freedom that has been largely unexplored, but is an important task in video editing. Our approach is simple to implement, highly efficient, and very robust to differences in video content, allowing forinteractiveexploration of the temporal alignment space for multiple real world HD videos. Oliver Wang, Christopher Schroers, Henning Zimmer, Markus Gross 0001, Alexander Sorkine-Hornung |
ACM Trans. Graph. | 2 |
| 2012 | Cross Anisotropic Cost Volume Filtering for Segmentation
Vladislav Kramarev, Oliver Demetz, Christopher Schroers, Joachim Weickert |
ACCV (1) | 3 |