EDBT 2026 Demo / reviewers in the wild / expert
Alexander Sorkine-Hornung
dblp:46/4142 · also Alexander Hornung
· DBLP profile ↗
58ranked-venue papers
7as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
33 papers |
Visual content generation and editing · 22% Image and video processing · 20% Rendering · 16% | |
| Artificial intelligence
15 papers |
3D vision · 57% Video understanding and tracking · 27% Segmentation and scene understanding · 11% |
Topics — the 30 heaviest of 86, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing › style transfer
3d style transfer |
0.8 | 1 | 2024 | Geometry Transfer for Stylizing Radiance Fields · CVPR 2024 |
Geometric modeling and processing › deformation modeling
geometric deformation |
0.8 | 1 | 2024 | Geometry Transfer for Stylizing Radiance Fields · CVPR 2024 |
Visual content generation and editing › style transfer
geometric style transfer |
0.8 | 1 | 2024 | Geometry Transfer for Stylizing Radiance Fields · CVPR 2024 |
Rendering
neural radiance fields |
0.8 | 1 | 2024 | Geometry Transfer for Stylizing Radiance Fields · CVPR 2024 |
Computer vision › Video understanding and tracking
video object segmentation |
0.8 | 3 | 2017 | Learning Video Object Segmentation from Static Images · CVPR 2017 A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation · CVPR 2016 Fully Connected Object Proposals for Video Segmentation · ICCV 2015 |
Visual content generation and editing
video editing |
0.7 | 3 | 2016 | Phase-Based Modification Transfer for Video · ECCV (3) 2016 FaceDirector: Continuous Control of Facial Performance in Video · ICCV 2015 VideoSnapping: interactive synchronization of multiple videos · ACM Trans. Graph. 2014 |
Computer vision › 3D vision
implicit neural representation |
0.6 | 1 | 2022 | PINs: Progressive Implicit Networks for Multi-Scale Neural Representations · ICML 2022 |
Virtual and augmented reality › avatar
avatar acquisition |
0.6 | 2 | 2017 | Demonstration: Rapid one-shot acquisition of dynamic VR avatars · VR 2017 Rapid one-shot acquisition of dynamic VR avatars · VR 2017 |
Rendering › neural rendering
neural scene representation |
0.6 | 1 | 2022 | PINs: Progressive Implicit Networks for Multi-Scale Neural Representations · ICML 2022 |
Image and video processing
video frame interpolation |
0.5 | 2 | 2018 | PhaseNet for Video Frame Interpolation · CVPR 2018 Phase-based frame interpolation for video · CVPR 2015 |
Rendering
novel view synthesis |
0.5 | 2 | 2019 | An integrated 6DoF video camera and system design · ACM Trans. Graph. 2019 Three-Dimensional Video Postproduction and Processing · Proc. IEEE 2011 |
Computer vision › Segmentation and scene understanding
video segmentation |
0.5 | 2 | 2016 | Bilateral Space Video Segmentation · CVPR 2016 Fully Connected Object Proposals for Video Segmentation · ICCV 2015 |
Multimedia systems and quality of experience › multimedia synchronization
video synchronization |
0.4 | 2 | 2016 | ActionSnapping: Motion-Based Video Synchronization · ECCV (5) 2016 VideoSnapping: interactive synchronization of multiple videos · ACM Trans. Graph. 2014 |
Image and video processing
image segmentation |
0.4 | 2 | 2016 | Efficient 3D Object Segmentation from Densely Sampled Light Fields with Applications to 3D Reconstruction · ACM Trans. Graph. 2016 Cache-efficient graph cuts on structured grids · CVPR 2012 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 1 | 2017 | Learning Video Object Segmentation from Static Images · CVPR 2017 |
Computer vision › 3D vision
approximate nearest neighbor |
0.2 | 1 | 2016 | Efficient Large-Scale Approximate Nearest Neighbor Search on the GPU · CVPR 2016 |
Computer vision › 3D vision › 3d reconstruction
image-based 3d reconstruction |
0.2 | 1 | 2016 | Efficient 3D Object Segmentation from Densely Sampled Light Fields with Applications to 3D Reconstruction · ACM Trans. Graph. 2016 |
Computer vision › 3D vision
nearest neighbor search |
0.2 | 1 | 2016 | Efficient Large-Scale Approximate Nearest Neighbor Search on the GPU · CVPR 2016 |
Computational photography and imaging › light field imaging
light field segmentation |
0.2 | 1 | 2016 | Efficient 3D Object Segmentation from Densely Sampled Light Fields with Applications to 3D Reconstruction · ACM Trans. Graph. 2016 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 3 | 2015 | Structure and motion from scene registration · CVPR 2012 Scalable structure from motion for densely sampled videos · CVPR 2015 Image selection for improved Multi-View Stereo · CVPR 2008 |
Virtual and augmented reality › tracking and registration
simultaneous localization and mapping |
0.2 | 1 | 2015 | Scalable structure from motion for densely sampled videos · CVPR 2015 |
Geometric modeling and processing › 3d reconstruction
structure from motion |
0.2 | 1 | 2015 | Scalable structure from motion for densely sampled videos · CVPR 2015 |
Computer vision › 3D vision › 3d reconstruction
volumetric reconstruction |
0.2 | 2 | 2012 | Structure and motion from scene registration · CVPR 2012 Robust and Efficient Photo-Consistency Estimation for Volumetric 3D Reconstruction · ECCV (2) 2006 |
Image and video processing › image sequence processing
temporal alignment |
0.2 | 1 | 2014 | VideoSnapping: interactive synchronization of multiple videos · ACM Trans. Graph. 2014 |
Computer vision › 3D vision › light field
light field reconstruction |
0.2 | 1 | 2013 | Scene reconstruction from high spatio-angular resolution light fields · ACM Trans. Graph. 2013 |
Rendering
image-based rendering |
0.2 | 1 | 2013 | Scene reconstruction from high spatio-angular resolution light fields · ACM Trans. Graph. 2013 |
Computational photography and imaging
image stitching |
0.2 | 1 | 2013 | Megastereo: Constructing High-Resolution Stereo Panoramas · CVPR 2013 |
Geometric modeling and processing › vectorization
line drawing vectorization |
0.2 | 1 | 2013 | Topology-driven vectorization of clean line drawings · ACM Trans. Graph. 2013 |
Computational photography and imaging › multi-perspective imaging
multi-perspective panorama |
0.2 | 1 | 2013 | Megastereo: Constructing High-Resolution Stereo Panoramas · CVPR 2013 |
Computational photography and imaging
panoramic imaging |
0.2 | 1 | 2013 | Megastereo: Constructing High-Resolution Stereo Panoramas · CVPR 2013 |
Methods — techniques the papers use, named apart from their topics
positional encoding · 1.1geometric deformation · 0.8depth map extraction · 0.8graph cuts · 0.6multilayer perceptron · 0.6multi-layer perceptron · 0.6phase-based motion representation · 0.5product quantization · 0.5real-time streaming · 0.4markerless calibration · 0.4neural network decoder · 0.3depth-based view synthesis · 0.3parametric avatar model · 0.3offline and online learning · 0.3convolutional neural network · 0.3vector quantization tree · 0.2reranking · 0.2re-ranking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Geometry Transfer for Stylizing Radiance FieldsabstractShape and geometric patterns are essential in defining stylistic identity. However, current 3D style transfer methods predominantly focus on transferring colors and textures, often overlooking geometric aspects. In this paper, we introduce Geometry Transfer, a novel method that leverages geometric deformation for 3D style transfer. This technique employs depth maps to extract a style guide, subsequently applied to stylize the geometry of radiance fields. Moreover, we propose new techniques that utilize geometric cues from the 3D scene, thereby enhancing aesthetic expressiveness and more accurately reflecting intended styles. Our extensive experiments show that Geometry Transfer enables a broader and more expressive range of stylizations, thereby significantly expanding the scope of 3D style transfer. Hyunyoung Jung 0001, Seonghyeon Nam, Nikolaos Sarafianos, Sungjoo Yoo, Alexander Sorkine-Hornung |
CVPR | 5 |
| 2022 | PINs: Progressive Implicit Networks for Multi-Scale Neural RepresentationsabstractMulti-layer perceptrons (MLP) have proven to be effective scene encoders when combined with higher-dimensional projections of the input, commonly referred to as positional encoding. However, scenes with a wide frequency spectrum remain a challenge: choosing high frequencies for positional encoding introduces noise in low structure areas, while low frequencies results in poor fitting of detailed regions. To address this, we propose a progressive positional encoding, exposing a hierarchical MLP structure to incremental sets of frequency encodings. Our model accurately reconstructs scenes with wide frequency bands and learns a scene representation at progressive level of detail without explicit per-level supervision. The architecture is modular: each level encodes a continuous implicit representation that can be leveraged separately for its respective resolution, meaning a smaller network for coarser reconstructions. Experiments on several 2D and 3D datasets shows improvements in reconstruction accuracy, representational capacity and training speed compared to baselines. Zoe Landgraf, Alexander Sorkine-Hornung, Ricardo Silveira Cabral |
ICML | 2 |
| 2019 | An integrated 6DoF video camera and system designabstractDesigning a fully integrated 360° video camera supporting 6DoF head motion parallax requires overcoming many technical hurdles, including camera placement, optical design, sensor resolution, system calibration, real-time video capture, depth reconstruction, and real-time novel view synthesis. While there is a large body of work describing various system components, such as multi-view depth estimation, our paper is the first to describe a complete, reproducible system that considers the challenges arising when designing, building, and deploying a full end-to-end 6DoF video camera and playback environment. Our system includes a computational imaging software pipeline supporting online markerless calibration, high-quality reconstruction, and real-time streaming and rendering. Most of our exposition is based on a professional 16-camera configuration, which will be commercially available to film producers. However, our software pipeline is generic and can handle a variety of camera geometries and configurations. The entire calibration and reconstruction software pipeline along with example datasets is open sourced to encourage follow-up research in high-quality 6DoF video reconstruction and rendering 1 . Albert Parra Pozo, Michael Toksvig, Terry Filiba Schrager, Joyce Hsu, Uday Mathur, Alexander Sorkine-Hornung, Richard Szeliski, Brian Cabral |
ACM Trans. Graph. | 6 |
| 2018 | PhaseNet for Video Frame InterpolationabstractMost approaches for video frame interpolation require accurate dense correspondences to synthesize an in-between frame. Therefore, they do not perform well in challenging scenarios with e.g. lighting changes or motion blur. Recent deep learning approaches that rely on kernels to represent motion can only alleviate these problems to some extent. In those cases, methods that use a per-pixel phase-based motion representation have been shown to work well. However, they are only applicable for a limited amount of motion. We propose a new approach, PhaseNet, that is designed to robustly handle challenging scenarios while also coping with larger motion. Our approach consists of a neural network decoder that directly estimates the phase decomposition of the intermediate frame. We show that this is superior to the hand-crafted heuristics previously used in phase-based methods and also compares favorably to recent deep learning based approaches for video frame interpolation on challenging datasets. Simone Schaub-Meyer, Abdelaziz Djelouah, Brian McWilliams, Alexander Sorkine-Hornung, Markus Gross 0001, Christopher Schroers |
CVPR | 4 |
| 2018 | An Omnistereoscopic Video Pipeline for Capture and Display of Real-World VRabstractIn this article, we describe a complete pipeline for the capture and display of real-world Virtual Reality video content, based on the concept of omnistereoscopic panoramas. We address important practical and theoretical issues that have remained undiscussed in previous works. On the capture side, we show how high-quality omnistereo video can be generated from a sparse set of cameras (16 in our prototype array) instead of the hundreds of input views previously required. Despite the sparse number of input views, our approach allows for high quality, real-time virtual head motion, thereby providing an important additional cue for immersive depth perception compared to static stereoscopic video. We also provide an in-depth analysis of the required camera array geometry in order to meet specific stereoscopic output constraints, which is fundamental for achieving a plausible and fully controlled VR viewing experience. Finally, we describe additional insights on how to integrate omnistereo video panoramas with rendered CG content. We provide qualitative comparisons to alternative solutions, including depth-based view synthesis and the Facebook Surround 360 system. In summary, this article provides a first complete guide and analysis for reimplementing a system for capturing and displaying real-world VR, which we demonstrate on several real-world examples captured with our prototype. Christopher Schroers, Jean-Charles Bazin, Alexander Sorkine-Hornung |
ACM Trans. Graph. | 3 |
| 2017 | Learning Video Object Segmentation from Static ImagesabstractInspired by recent advances of deep learning in instance segmentation and object tracking, we introduce the concept of convnet-based guidance applied to video object segmentation. Our model proceeds on a per-frame basis, guided by the output of the previous frame towards the object of interest in the next frame. We demonstrate that highly accurate object segmentation in videos can be enabled by using a convolutional neural network (convnet) trained with static images only. The key component of our approach is a combination of offline and online learning strategies, where the former produces a refined mask from the previous frame estimate and the latter allows to capture the appearance of the specific object instance. Our method can handle different types of input annotations such as bounding boxes and segments while leveraging an arbitrary amount of annotated frames. Therefore our system is suitable for diverse applications with different requirements in terms of accuracy and efficiency. In our extensive evaluation, we obtain competitive results on three different datasets, independently from the type of input annotation. Federico Perazzi, Anna Khoreva, Rodrigo Benenson, Bernt Schiele, Alexander Sorkine-Hornung |
CVPR | 5 |
| 2017 | Rapid one-shot acquisition of dynamic VR avatarsabstractWe present a system for rapid acquisition of bespoke, animatable, full-body avatars including face texture and shape. A blendshape rig with a skeleton is used as a template for customization. Identity blendshapes are used to customize the body and face shape at the fitting stage, while animation blendshapes allow the face to be animated. The subject assumes a T-pose and a single snapshot is captured using a stereo RGB plus depth sensor rig. Our system automatically aligns a photo texture and fits the 3D shape of the face. The body shape is stylized according to body dimensions estimated from segmented depth. The face identity blendweights are optimised according to image-based facial landmarks, while a custom texture map for the face is generated by warping the input images to a reference texture according to the facial landmarks. The total capture and processing time is under 10 seconds and the output is a light-weight, game-engine-ready avatar which is recognizable as the subject. We demonstrate our system in a VR environment in which each user sees the other users' animated avatars through a VR headset with real-time audio-based facial animation and live body motion tracking, affording an enhanced level of presence and social engagement compared to generic avatars. Charles Malleson, Maggie Kosek, Martin Klaudiny, Ivan Huerta Casado, Jean-Charles Bazin, Alexander Sorkine-Hornung, Mark Mine, Kenny Mitchell |
VR | 6 |
| 2017 | Demonstration: Rapid one-shot acquisition of dynamic VR avatarsabstractIn this demonstration, we showcase a system for rapid acquisition of bespoke avatars for each participant (subject) in a social VR environment is presented. For each subject, the system automatically customizes a parametric avatar model to match the captured subject by adjusting its overall height, body and face shape parameters and generating a custom face texture. Charles Malleson, Maggie Kosek, Martin Klaudiny, Ivan Huerta Casado, Jean-Charles Bazin, Alexander Sorkine-Hornung, Mark Mine, Kenny Mitchell |
VR | 6 |
| 2016 | Point Cloud Noise and Outlier Removal for Image-Based 3D ReconstructionabstractPoint sets generated by image-based 3D reconstruction techniques are often much noisier than those obtained using active techniques like laser scanning. Therefore, they pose greater challenges to the subsequent surface reconstruction (meshing) stage. We present a simple and effective method for removing noise and outliers from such point sets. Our algorithm uses the input images and corresponding depth maps to remove pixels which are geometrically or photometrically inconsistent with the colored surface implied by the input. This allows standard surface reconstruction methods (such as Poisson surface reconstruction) to perform less smoothing and thus achieve higher quality surfaces with more features. Our algorithm is efficient, easy to implement, and robust to varying amounts of noise. We demonstrate the benefits of our algorithm in combination with a variety of state-of-the-art depth and surface reconstruction methods. Katja Wolff, Changil Kim 0001, Henning Zimmer, Christopher Schroers, Mario Botsch, Olga Sorkine-Hornung, Alexander Sorkine-Hornung |
3DV | 7 |
| 2016 | Depth from Gradients in Dense Light Fields for Object ReconstructionabstractObjects with thin features and fine details are challenging for most multi-view stereo techniques, since such features occupy small volumes and are usually only visible in a small portion of the available views. In this paper, we present an efficient algorithm to reconstruct intricate objects using densely sampled light fields. At the heart of our technique lies a novel approach to compute per-pixel depth values by exploiting local gradient information in densely sampled light fields. This approach can generate accurate depth values for very thin features, and can be run for each pixel in parallel. We assess the reliability of our depth estimates using a novel two-sided photoconsistency measure, which can capture whether the pixel lies on a texture or a silhouette edge. This information is then used to propagate the depth estimates at high gradient regions to smooth parts of the views efficiently and reliably using edge-aware filtering. In the last step, the per-image depth values and color information are aggregated in 3D space using a voting scheme, allowing the reconstruction of a globally consistent mesh for the object. Our approach can process large video datasets very efficiently and at the same time generates high quality object reconstructions that compare favorably to the results of state-of-the-art multi-view stereo methods. Kaan Yücer, Changil Kim 0001, Alexander Sorkine-Hornung, Olga Sorkine-Hornung |
3DV | 3 |
| 2016 | Bilateral Space Video SegmentationabstractIn this work, we propose a novel approach to video segmentation that operates in bilateral space. We design a new energy on the vertices of a regularly sampled spatiotemporal bilateral grid, which can be solved efficiently using a standard graph cut label assignment. Using a bilateral formulation, the energy that we minimize implicitly approximates long-range, spatio-temporal connections between pixels while still containing only a small number of variables and only local graph edges. We compare to a number of recent methods, and show that our approach achieves state-of-the-art results on multiple benchmarks in a fraction of the runtime. Furthermore, our method scales linearly with image size, allowing for interactive feedback on real-world high resolution video. Nicolas Marki, Federico Perazzi, Oliver Wang, Alexander Sorkine-Hornung |
CVPR | 4 |
| 2016 | A Benchmark Dataset and Evaluation Methodology for Video Object SegmentationabstractOver the years, datasets and benchmarks have proven their fundamental importance in computer vision research, enabling targeted progress and objective comparisons in many fields. At the same time, legacy datasets may impend the evolution of a field due to saturated algorithm performance and the lack of contemporary, high quality data. In this work we present a new benchmark dataset and evaluation methodology for the area of video object segmentation. The dataset, named DAVIS (Densely Annotated VIdeo Segmentation), consists of fifty high quality, Full HD video sequences, spanning multiple occurrences of common video object segmentation challenges such as occlusions, motionblur and appearance changes. Each video is accompanied by densely annotated, pixel-accurate and per-frame ground truth segmentation. In addition, we provide a comprehensive analysis of several state-of-the-art segmentation approaches using three complementary metrics that measure the spatial extent of the segmentation, the accuracy of the silhouette contours and the temporal coherence. The results uncover strengths and weaknesses of current approaches, opening up promising directions for future works. Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross 0001, Alexander Sorkine-Hornung |
CVPR | 6 |
| 2016 | Efficient Large-Scale Approximate Nearest Neighbor Search on the GPUabstractWe present a new approach for efficient approximate nearest neighbor (ANN) search in high dimensional spaces, extending the idea of Product Quantization. We propose a two level product and vector quantization tree that reduces the number of vector comparisons required during tree traversal. Our approach also includes a novel highly parallelizable re-ranking method for candidate vectors by efficiently reusing already computed intermediate values. Due to its small memory footprint during traversal the method lends itself to an efficient, parallel GPU implementation. This Product Quantization Tree (PQT) approach significantly outperforms recent state of the art methods for high dimensional nearest neighbor queries on standard reference datasets. Ours is the first work that demonstrates GPU performance superior to CPU performance on high dimensional, large scale ANN problems in time-critical real-world applications, like loop-closing in videos. Patrick Wieschollek, Oliver Wang, Alexander Sorkine-Hornung, Hendrik P. A. Lensch |
CVPR | 3 |
| 2016 | ActionSnapping: Motion-Based Video Synchronization
Jean-Charles Bazin, Alexander Sorkine-Hornung |
ECCV (5) | 2 |
| 2016 | Phase-Based Modification Transfer for Video
Simone Schaub-Meyer, Alexander Sorkine-Hornung, Markus Gross 0001 |
ECCV (3) | 2 |
| 2016 | Efficient 3D Object Segmentation from Densely Sampled Light Fields with Applications to 3D ReconstructionabstractPrecise object segmentation in image data is a fundamental problem with various applications, including 3D object reconstruction. We present an efficient algorithm to automatically segment a static foreground object from highly cluttered background in light fields. A key insight and contribution of our article is that a significant increase of the available input data can enable the design of novel, highly efficient approaches. In particular, the central idea of our method is to exploit high spatio-angular sampling on the order of thousands of input frames, for example, captured as a hand-held video, such that new structures are revealed due to the increased coherence in the data. We first show how purely local gradient information contained in slices of such a dense light field can be combined with information about the camera trajectory to make efficient estimates of the foreground and background. These estimates are then propagated to textureless regions using edge-aware filtering in the epipolar volume. Finally, we enforce global consistency in a gathering step to derive a precise object segmentation in both 2D and 3D space, which captures fine geometric details even in very cluttered scenes. The design of each of these steps is motivated by efficiency and scalability, allowing us to handle large, real-world video datasets on a standard desktop computer. We demonstrate how the results of our method can be used for considerably improving the speed and quality of image-based 3D reconstruction algorithms, and we compare our results to state-of-the-art segmentation and multiview stereo methods. Kaan Yücer, Alexander Sorkine-Hornung, Oliver Wang, Olga Sorkine-Hornung |
ACM Trans. Graph. | 2 |
| 2015 | Phase-based frame interpolation for videoabstractStandard approaches to computing interpolated (in-between) frames in a video sequence require accurate pixel correspondences between images e.g. using optical flow. We present an efficient alternative by leveraging recent developments in phase-based methods that represent motion in the phase shift of individual pixels. This concept allows in-between images to be generated by simple per-pixel phase modification, without the need for any form of explicit correspondence estimation. Up until now, such methods have been limited in the range of motion that can be interpolated, which fundamentally restricts their usefulness. In order to reduce these limitations, we introduce a novel, bounded phase shift correction method that combines phase information across the levels of a multi-scale pyramid. Additionally, we propose extensions for phase-based image synthesis that yield smoother transitions between the interpolated images. Our approach avoids expensive global optimization typical of optical flow methods, and is both simple to implement and easy to parallelize. This allows us to interpolate frames at a fraction of the computational cost of traditional optical flow-based solutions, while achieving similar quality and in some cases even superior results. Our method fails gracefully in difficult interpolation settings, e.g., significant appearance changes, where flow-based methods often introduce serious visual artifacts. Due to its efficiency, our method is especially well suited for frame interpolation and retiming of high resolution, high frame rate video. Simone Schaub-Meyer, Oliver Wang, Henning Zimmer, Max Grosse, Alexander Sorkine-Hornung |
CVPR | 5 |
| 2015 | Scalable structure from motion for densely sampled videosabstractVideos consisting of thousands of high resolution frames are challenging for existing structure from motion (SfM) and simultaneous-localization and mapping (SLAM) techniques. We present a new approach for simultaneously computing extrinsic camera poses and 3D scene structure that is capable of handling such large volumes of image data. The key insight behind this paper is to effectively exploit coherence in densely sampled video input. Our technical contributions include robust tracking and selection of confident video frames, a novel window bundle adjustment, frame-to-structure verification for globally consistent reconstructions with multi-loop closing, and utilizing efficient global linear camera pose estimation in order to link both consecutive and distant bundle adjustment windows. To our knowledge we describe the first system that is capable of handling high resolution, high frame-rate video data with close to real-time performance. In addition, our approach can robustly integrate data from different video sequences, allowing multiple video streams to be simultaneously calibrated in an efficient and globally optimal way. We demonstrate high quality alignment on large scale challenging datasets, e.g., 2-20 megapixel resolution at frame rates of 25-120 Hz with thousands of frames. Benjamin Resch, Hendrik P. A. Lensch, Oliver Wang, Marc Pollefeys, Alexander Sorkine-Hornung |
CVPR | 5 |
| 2015 | FaceDirector: Continuous Control of Facial Performance in VideoabstractWe present a method to continuously blend between multiple facial performances of an actor, which can contain different facial expressions or emotional states. As an example, given sad and angry video takes of a scene, our method empowers the movie director to specify arbitrary weighted combinations and smooth transitions between the two takes in post-production. Our contributions include (1) a robust nonlinear audio-visual synchronization technique that exploits complementary properties of audio and visual cues to automatically determine robust, dense spatiotemporal correspondences between takes, and (2) a seamless facial blending approach that provides the director full control to interpolate timing, facial expression, and local appearance, in order to generate novel performances after filming. In contrast to most previous works, our approach operates entirely in image space, avoiding the need of 3D facial reconstruction. We demonstrate that our method can synthesize visually believable performances with applications in emotion transition, performance correction, and timing control. Charles Malleson, Jean-Charles Bazin, Oliver Wang, Derek Bradley, Thabo Beeler, Adrian Hilton 0001, Alexander Sorkine-Hornung |
ICCV | 7 |
| 2015 | Fully Connected Object Proposals for Video SegmentationabstractWe present a novel approach to video segmentation using multiple object proposals. The problem is formulated as a minimization of a novel energy function defined over a fully connected graph of object proposals. Our model combines appearance with long-range point tracks, which is key to ensure robustness with respect to fast motion and occlusions over longer video sequences. As opposed to previous approaches based on object proposals, we do not seek the best per-frame object hypotheses to perform the segmentation. Instead, we combine multiple, potentially imperfect proposals to improve overall segmentation accuracy and ensure robustness to outliers. Overall, the basic algorithm consists of three steps. First, we generate a very large number of object proposals for each video frame using existing techniques. Next, we perform an SVM-based pruning step to retain only high quality proposals with sufficiently discriminative power. Finally, we determine the fore-and background classification by solving for the maximum a posteriori of a fully connected conditional random field, defined using our novel energy function. Experimental results on a well established dataset demonstrate that our method compares favorably to several recent state-of-the-art approaches. Federico Perazzi, Oliver Wang, Markus Gross 0001, Alexander Sorkine-Hornung |
ICCV | 4 |
| 2015 | Online view sampling for estimating depth from light fieldsabstractGeometric information such as depth obtained from light fields finds more applications recently. Where and how to sample images to populate a light field is an important problem to maximize the usability of information gathered for depth reconstruction. We propose a simple analysis model for view sampling and an adaptive, online sampling algorithm tailored to light field depth reconstruction. Our model is based on the trade-off between visibility and depth resolvability for varying sampling locations, and seeks the optimal locations that best balance the two conflicting criteria. Changil Kim 0001, Kartic Subr, Kenny Mitchell, Alexander Sorkine-Hornung, Markus Gross 0001 |
ICIP | 4 |
| 2015 | Panoramic Video from Unstructured Camera ArraysabstractAbstract We describe an algorithm for generating panoramic video from unstructured camera arrays. Artifact‐free panorama stitching is impeded by parallax between input views. Common strategies such as multi‐level blending or minimum energy seams produce seamless results on quasi‐static input. However, on video input these approaches introduce noticeable visual artifacts due to lack of global temporal and spatial coherence. In this paper we extend the basic concept of local warping for parallax removal. Firstly, we introduce an error measure with increased sensitivity to stitching artifacts in regions with pronounced structure. Using this measure, our method efficiently finds an optimal ordering of pair‐wise warps for robust stitching with minimal parallax artifacts. Weighted extrapolation of warps in non‐overlap regions ensures temporal stability, while at the same time avoiding visual discontinuities around transitions between views. Remaining global deformation introduced by the warps is spread over the entire panorama domain using constrained relaxation, while staying as close as possible to the original input views. In combination, these contributions form the first system for spatiotemporally stable panoramic video stitching from unstructured camera array input. Federico Perazzi, Alexander Sorkine-Hornung, Henning Zimmer, Peter Kaufmann 0001, Oliver Wang, Scott Watson, Markus Gross 0001 |
Comput. Graph. Forum | 2 |
| 2015 | Path-space Motion Estimation and Decomposition for Robust Animation FilteringabstractAbstract Renderings of animation sequences with physics‐based Monte Carlo light transport simulations are exceedingly costly to generate frame‐by‐frame, yet much of this computation is highly redundant due to the strong coherence in space, time and among samples. A promising approach pursued in prior work entails subsampling the sequence in space, time, and number of samples, followed by image‐based spatio‐temporal upsampling and denoising. These methods can provide significant performance gains, though major issues remain: firstly, in a multiple scattering simulation, the final pixel color is the composite of many different light transport phenomena, and this conflicting information causes artifacts in image‐based methods. Secondly, motion vectors are needed to establish correspondence between the pixels in different frames, but it is unclear how to obtain them for most kinds of light paths (e.g. an object seen through a curved glass panel). To reduce these ambiguities, we propose a general decomposition framework, where the final pixel color is separated into components corresponding to disjoint subsets of the space of light paths. Each component is accompanied by motion vectors and other auxiliary features such as reflectance and surface normals. The motion vectors of specular paths are computed using a temporal extension of manifold exploration and the remaining components use a specialized variant of optical flow. Our experiments show that this decomposition leads to significant improvements in three image‐based applications: denoising, spatial upsampling, and temporal interpolation. Henning Zimmer, Fabrice Rousselle, Wenzel Jakob, Oliver Wang, David Adler, Wojciech Jarosz, Olga Sorkine-Hornung, Alexander Sorkine-Hornung |
Comput. Graph. Forum | 8 |
| 2015 | Sampling based scene-space video processingabstractMany compelling video processing effects can be achieved if per-pixel depth information and 3D camera calibrations are known. However, the success of such methods is highly dependent on the accuracy of this "scene-space" information. We present a novel, sampling-based framework for processing video that enables high-quality scene-space video effects in the presence of inevitable errors in depth and camera pose estimation. Instead of trying to improve the explicit 3D scene representation, the key idea of our method is to exploit the high redundancy of approximate scene information that arises due to most scene points being visible multiple times across many frames of video. Based on this observation, we propose a novel pixel gathering and filtering approach. The gathering step is general and collects pixel samples in scene-space, while the filtering step is application-specific and computes a desired output video from the gathered sample sets. Our approach is easily parallelizable and has been implemented on GPU, allowing us to take full advantage of large volumes of video data and facilitating practical runtimes on HD video using a standard desktop computer. Our generic scene-space formulation is able to comprehensively describe a multitude of video processing applications such as denoising, deblurring, super resolution, object removal, computational shutter functions, and other scene-space camera effects. We present results for various casually captured, hand-held, moving, compressed, monocular videos depicting challenging scenes recorded in uncontrolled environments. Felix Klose, Oliver Wang, Jean-Charles Bazin, Marcus A. Magnor, Alexander Sorkine-Hornung |
ACM Trans. Graph. | 5 |
| 2014 | Memory Efficient Stereoscopy from Light FieldsabstractWe address the problem of stereoscopic content generation from light fields using multi-perspective imaging. Our proposed method takes as input a light field and a target disparity map, and synthesizes a stereoscopic image pair by selecting light rays that fulfill the given target disparity constraints. We formulate this as a variational convex optimization problem. Compared to previous work, our method makes use of multi-view input to composite the new view with occlusions and disocclusions properly handled, does not require any correspondence information such as scene depth, is free from undesirable artifacts such as grid bias or image distortion, and is more efficiently solvable. In particular, our method is about ten times more memory efficient than the previous art, and is capable of processing higher resolution input. This is essential to make the proposed method practically applicable to realistic scenarios where HD content is standard. We demonstrate the effectiveness of our method experimentally. Changil Kim 0001, Ulrich Muller, Henning Zimmer, Yael Pritch, Alexander Sorkine-Hornung, Markus Gross 0001 |
3DV | 5 |
| 2014 | VideoSnapping: interactive synchronization of multiple videosabstractAligning video is a fundamental task in computer graphics and vision, required for a wide range of applications. We present aninteractivemethod for computing optimal nonlinear temporal video alignments of an arbitrary number of videos. We first derive a robust approximation of alignment quality between pairs of clips, computed as a weighted histogram of feature matches. We then find optimal temporal mappings (constituting frame correspondences) using a graph-based approach that allows for very efficient evaluation with artist constraints. This enables an enhancement to the "snapping" interface in video editing tools, where videos in a time-line are now able snap to one another when dragged by an artist based on theircontent, rather than simply start-and-end times. The pairwise snapping is then generalized to multiple clips, achieving a globally optimal temporal synchronization that automatically arranges a series of clips filmed at different times into a single consistent time frame. When followed by a simple spatial registration, we achieve high quality spatiotemporal video alignments at a fraction of the computational complexity compared to previous methods. Assisted temporal alignment is a degree of freedom that has been largely unexplored, but is an important task in video editing. Our approach is simple to implement, highly efficient, and very robust to differences in video content, allowing forinteractiveexploration of the temporal alignment space for multiple real world HD videos. Oliver Wang, Christopher Schroers, Henning Zimmer, Markus Gross 0001, Alexander Sorkine-Hornung |
ACM Trans. Graph. | 5 |
| 2013 | Megastereo: Constructing High-Resolution Stereo PanoramasabstractWe present a solution for generating high-quality stereo panoramas at mega pixel resolutions. While previous approaches introduced the basic principles, we show that those techniques do not generalise well to today's high image resolutions and lead to disturbing visual artefacts. As our first contribution, we describe the necessary correction steps and a compact representation for the input images in order to achieve a highly accurate approximation to the required ray space. Our second contribution is a flow-based up sampling of the available input rays which effectively resolves known aliasing issues like stitching artefacts. The required rays are generated on the fly to perfectly match the desired output resolution, even for small numbers of input images. In addition, the up sampling is real-time and enables direct interactive control over the desired stereoscopic depth effect. In combination, our contributions allow the generation of stereoscopic panoramas at high output resolutions that are virtually free of artefacts such as seams, stereo discontinuities, vertical parallax and other mono-/stereoscopic shape distortions. Our process is robust, and other types of multiperspective panoramas, such as linear panoramas, can also benefit from our contributions. We show various comparisons and high-resolution results. Christian Richardt, Yael Pritch, Henning Zimmer, Alexander Sorkine-Hornung |
CVPR | 4 |
| 2013 | Content-aware compression using saliency-driven image retargetingabstractIn this paper we propose a novel method to compress video content based on image retargeting. First, a saliency map is extracted from the video frames either automatically or according to user input. Next, nonlinear image scaling is performed which assigns a higher pixel count to salient image regions and fewer pixels to non-salient regions. The non-linearly downscaled images can then be compressed using existing compression techniques and decoded and upscaled at the receiver. To this end we introduce a non-uniform antialiasing technique that significantly improves the image resampling quality. The overall process is complementary to existing compression methods and can be seamlessly incorporated into existing pipelines. We compare our method to JPEG 2000 and H.264/AVC-10 and show that, at the cost of visual quality in non-salient image regions, our method achieves a significant improvement of the visual quality of salient image regions in terms of Structural Similarity (SSIM) and Peak Signal-to-Noise-Ratio (PSNR) quality measures, in particular for scenarios with high compression ratios. Fabio Zünd, Yael Pritch, Alexander Sorkine-Hornung, Stefan Mangold, Thomas R. Gross |
ICIP | 3 |
| 2013 | Finite Element Image WarpingabstractAbstract We introduce a single unifying framework for a wide range of content‐aware image warping tasks using a finite element method (FEM). Existing approaches commonly define error terms over vertex finite differences and can be expressed as a special case of our general FEM model. In this work, we exploit the full generality of FEMs, gaining important advantages over prior methods. These advantages include arbitrary mesh connectivity allowing for adaptive meshing and efficient large‐scale solutions, a well‐defined continuous problem formulation that enables clear analysis of existing warping error functions and allows us to propose improved ones, and higher order basis functions that allow for smoother warps with fewer degrees of freedom. To support per‐element basis functions of varying degree and complex mesh connectivity with hanging nodes, we also introduce a novel use of discontinuous Galerkin FEM. We demonstrate the utility of our method by showing examples in video retargeting and camera stabilization applications, and compare our results with previous state of the art methods. Peter Kaufmann 0001, Oliver Wang, Alexander Sorkine-Hornung, Olga Sorkine-Hornung, Aljoscha Smolic, Markus Gross 0001 |
Comput. Graph. Forum | 3 |
| 2013 | Scalable Music: Automatic Music Retargeting and SynthesisabstractAbstract In this paper we propose a method for dynamic rescaling of music, inspired by recent works on image retargeting, video reshuffling and character animation in the computer graphics community. Given the desired target length of a piece of music and optional additional constraints such as position and importance of certain parts, we build on concepts from seam carving, video textures and motion graphs and extend them to allow for a global optimization of jumps in an audio signal. Based on an automatic feature extraction and spectral clustering for segmentation, we employ length‐constrained least‐costly path search via dynamic programming to synthesize a novel piece of music that best fulfills all desired constraints, with imperceptible transitions between reshuffled parts. We show various applications of music retargeting such as part removal, decreasing or increasing music duration, and in particular consistent joint video and audio editing. Simon Wenner, Jean-Charles Bazin, Alexander Sorkine-Hornung, Changil Kim 0001, Markus Gross 0001 |
Comput. Graph. Forum | 3 |
| 2013 | Scene reconstruction from high spatio-angular resolution light fieldsabstractThis paper describes a method for scene reconstruction of complex, detailed environments from 3D light fields. Densely sampled light fields in the order of 10 9 light rays allow us to capture the real world in unparalleled detail, but efficiently processing this amount of data to generate an equally detailed reconstruction represents a significant challenge to existing algorithms. We propose an algorithm that leverages coherence in massive light fields by breaking with a number of established practices in image-based reconstruction. Our algorithm first computes reliable depth estimates specifically around object boundaries instead of interior regions, by operating on individual light rays instead of image patches. More homogeneous interior regions are then processed in a fine-to-coarse procedure rather than the standard coarse-to-fine approaches. At no point in our method is any form of global optimization performed. This allows our algorithm to retain precise object contours while still ensuring smooth reconstructions in less detailed areas. While the core reconstruction method handles general unstructured input, we also introduce a sparse representation and a propagation scheme for reliable depth estimates which make our algorithm particularly effective for 3D input, enabling fast and memory efficient processing of "Gigaray light fields" on a standard GPU. We show dense 3D reconstructions of highly detailed scenes, enabling applications such as automatic segmentation and image-based rendering, and provide an extensive evaluation and comparison to existing image-based reconstruction techniques. Changil Kim 0001, Henning Zimmer, Yael Pritch, Alexander Sorkine-Hornung, Markus Gross 0001 |
ACM Trans. Graph. | 4 |
| 2013 | Painting by feature: texture boundaries for example-based image creationabstractIn this paper we propose a reinterpretation of the brush and the fill tools for digital image painting. The core idea is to provide an intuitive approach that allows users to paint in the visual style of arbitrary example images. Rather than a static library of colors, brushes, or fill patterns, we offer users entire images as their palette, from which they can select arbitrary contours or textures as their brush or fill tool in their own creations. Compared to previous example-based techniques related to the painting-by-numbers paradigm we propose a new strategy where users can generate salient texture boundaries by our randomized graph-traversal algorithm and apply a content-aware fill to transfer textures into the delimited regions. This workflow allows users of our system to intuitively create visually appealing images that better preserve the visual richness and fluidity of arbitrary example images. We demonstrate the potential of our approach in various applications including interactive image creation, editing and vector image stylization. Michal Lukác, Jakub Fiser, Jean-Charles Bazin, Ondrej Jamriska, Alexander Sorkine-Hornung, Daniel Sýkora |
ACM Trans. Graph. | 5 |
| 2013 | Topology-driven vectorization of clean line drawingsabstractVectorization provides a link between raster scans of pencil-and-paper drawings and modern digital processing algorithms that require accurate vector representations. Even when input drawings are comprised of clean, crisp lines, inherent ambiguities near junctions make vectorization deceptively difficult. As a consequence, current vectorization approaches often fail to faithfully capture the junctions of drawn strokes. We propose a vectorization algorithm specialized for clean line drawings that analyzes the drawing's topology in order to overcome junction ambiguities. A gradient-based pixel clustering technique facilitates topology computation. This topological information is exploited during centerline extraction by a new “reverse drawing” procedure that reconstructs all possible drawing states prior to the creation of a junction and then selects the most likely stroke configuration. For cases where the automatic result does not match the artist's interpretation, our drawing analysis enables an efficient user interface to easily adjust the junction location. We demonstrate results on professional examples and evaluate the vectorization quality with quantitative comparison to hand-traced centerlines as well as the results of leading commercial algorithms. Gioacchino Noris, Alexander Sorkine-Hornung, Robert W. Sumner, Maryann Simmons, Markus Gross 0001 |
ACM Trans. Graph. | 2 |
| 2013 | Sketch-based generation and editing of quad meshesabstractCoarse quad meshes are the preferred representation for animating characters in movies and video games. In these scenarios, artists want explicit control over the edge flows and the singularities of the quad mesh. Despite the significant advances in recent years, existing automatic quad remeshing algorithms are not yet able to achieve the quality of manually created remeshings. We present an interactive system for manual quad remeshing that provides the user with a high degree of control while avoiding the tediousness involved in existing manual tools. With our sketch-based interface the user constructs a quad mesh by defining patches consisting of individual quads. The desired edge flow is intuitively specified by the sketched patch boundaries, and the mesh topology can be adjusted by varying the number of edge subdivisions at patch boundaries. Our system automatically inserts singularities inside patches if necessary, while providing the user with direct control of their topological and geometrical locations. We developed a set of novel user interfaces that assist the user in constructing a curve network representing such patch boundaries. The effectiveness of our system is demonstrated through a user evaluation with professional artists. Our system is also useful for editing automatically generated quad meshes. Kenshi Takayama, Daniele Panozzo, Alexander Sorkine-Hornung, Olga Sorkine-Hornung |
ACM Trans. Graph. | 3 |
| 2012 | Structure and motion from scene registrationabstractWe propose a method for estimating the 3D structure and the dense 3D motion (scene flow) of a dynamic nonrigid 3D scene, using a camera array. The core idea is to use a dense multi-camera array to construct a novel, dense 3D volumetric representation of the 3D space where each voxel holds an estimated intensity value and a confidence measure of this value. The problem of 3D structure and 3D motion estimation of a scene is thus reduced to a nonrigid registration of two volumes - hence the term ”Scene Registration”. Registering two dense 3D scalar volumes does not require recovering the 3D structure of the scene as a preprocessing step, nor does it require explicit reasoning about occlusions. From this nonrigid registration we accurately extract the 3D scene flow and the 3D structure of the scene, and successfully recover the sharp discontinuities in both time and space. We demonstrate the advantages of our method on a number of challenging synthetic and real data sets. Tali Dekel, Shai Avidan, Alexander Sorkine-Hornung, Wojciech Matusik |
CVPR | 3 |
| 2012 | Cache-efficient graph cuts on structured gridsabstractFinding minimal cuts on graphs with a grid-like structure has become a core task for solving many computer vision and graphics related problems. However, computation speed and memory consumption oftentimes limit the effective use in applications requiring high resolution grids or interactive response. In particular, memory bandwidth represents one of the major bottlenecks even in today's most efficient implementations. We propose a compact data structure with cache-efficient memory layout for the representation of graph instances that are based on regular N-D grids with topologically identical neighborhood systems. For this common class of graphs our data structure allows for 3 to 12 times higher grid resolutions and a 3- to 9-fold speedup compared to existing approaches. Our design is agnostic to the underlying algorithm, and hence orthogonal to other optimizations such as parallel and hierarchical processing. We evaluate the performance gain on a variety of typical problems including 2D/3D segmentation, colorization, and stereo. All experiments show an unconditional improvement in terms of speed and memory consumption, with graceful performance degradation for graphs with increasing topological irregularities. Ondrej Jamriska, Daniel Sýkora, Alexander Sorkine-Hornung |
CVPR | 3 |
| 2012 | Saliency filters: Contrast based filtering for salient region detectionabstractSaliency estimation has become a valuable tool in image processing. Yet, existing approaches exhibit considerable variation in methodology, and it is often difficult to attribute improvements in result quality to specific algorithm properties. In this paper we reconsider some of the design choices of previous methods and propose a conceptually clear and intuitive algorithm for contrast-based saliency estimation. Our algorithm consists of four basic steps. First, our method decomposes a given image into compact, perceptually homogeneous elements that abstract unnecessary detail. Based on this abstraction we compute two measures of contrast that rate the uniqueness and the spatial distribution of these elements. From the element contrast we then derive a saliency measure that produces a pixel-accurate saliency map which uniformly covers the objects of interest and consistently separates fore- and background. We show that the complete contrast and saliency estimation can be formulated in a unified way using high-dimensional Gaussian filters. This contributes to the conceptual simplicity of our method and lends itself to a highly efficient implementation with linear complexity. In a detailed experimental evaluation we analyze the contribution of each individual feature and show that our method outperforms all state-of-the-art approaches. Federico Perazzi, Philipp Krähenbühl, Yael Pritch, Alexander Sorkine-Hornung |
CVPR | 4 |
| 2012 | Smart Scribbles for Sketch SegmentationabstractAbstract We present ‘Smart Scribbles’—a new scribble‐based interface for user‐guided segmentation of digital sketchy drawings. In contrast to previous approaches based on simple selection strategies, Smart Scribbles exploits richer geometric and temporal information, resulting in a more intuitive segmentation interface. We introduce a novel energy minimization formulation in which both geometric and temporal information from digital input devices is used to define stroke‐to‐stroke and scribble‐to‐stroke relationships. Although the minimization of this energy is, in general, an NP‐hard problem, we use a simple heuristic that leads to a good approximation and permits an interactive system able to produce accurate labellings even for cluttered sketchy drawings. We demonstrate the power of our technique in several practical scenarios such as sketch editing, as‐rigid‐as‐possible deformation and registration, and on‐the‐fly labelling based on pre‐classified guidelines. Gioacchino Noris, Daniel Sýkora, Arik Shamir, Stelian Coros, Brian Whited, Maryann Simmons, Alexander Sorkine-Hornung, Markus Gross 0001, Robert W. Sumner |
Comput. Graph. Forum | 7 |
| 2012 | Transfusive image manipulationabstractWe present a method for consistent automatic transfer of edits applied to one image to many other images of the same object or scene. By introducing novel, content-adaptive weight functions we enhance the non-rigid alignment framework of Lucas-Kanade to robustly handle changes of view point, illumination and non-rigid deformations of the subjects. Our weight functions are content-aware and possess high-order smoothness, enabling to define high-quality image warping with a low number of parameters using spatially-varying weighted combinations of affine deformations. Optimizing the warp parameters leads to subpixel-accurate alignment while maintaining computation efficiency. Our method allows users to perform precise, localized edits such as simultaneous painting on multiple images in real-time, relieving them from tedious and repetitive manual reapplication to each individual image. Kaan Yücer, Alec Jacobson, Alexander Sorkine-Hornung, Olga Sorkine-Hornung |
ACM Trans. Graph. | 3 |
| 2011 | Extending SVC by Content-adaptive Spatial ScalabilityabstractThis paper provides details on a complete integration of Content- adaptive Spatial Scalability (CASS) into the scalable video coding extension of H.264/AVC (SVC). CASS enables the efficient encoding of a high-quality bit stream that contains several versions of an original image sequence. Thereby, each such image sequence has been created by content-adaptive and art directed retargeting to different display aspect-ratios and/or resolutions. Non-linear dependencies between spatial layers, which have been introduced through content-adaptive retargeting, are exploited by a generalization of the three inter-layer prediction tools of SVC, i.e. by content-adaptive inter-layer texture, motion and residual prediction. The CASS extended SVC enables the transmission of video content which has been specifically adapted in an art- directed way to multiple display configurations (e.g. to SD and HD displays with 4:3 and 16:9 aspect-ratios, respectively) using a single compressed bit stream. With our extension, video content of higher semantic quality can be transmitted in a scalable way by introducing an average overhead in bit rate of 9.3%. Yongzhe Wang, Nikolce Stefanoski, Manuel Lang, Alexander Sorkine-Hornung, Aljoscha Smolic, Markus Gross 0001 |
ICIP | 4 |
| 2011 | Evaluation of backward mapping DIBR for FVV applicationsabstractIn this paper, we explore the challenges posed by wide baseline camera configurations for depth-image-based rendering, which should provide greater freedom for choosing the virtual viewpoint in a Free Viewpoint Video context, compared with the usual camera configurations intended for use in 3DTV settings. We implement a backward mapping approach with a custom filtering scheme based on median filters. Whilst the results back our initial assumption that this camera configuration provides good mobility, we show that the usual encoding for depth information referred to a global reference system is wrong and reference systems local to each camera should be used instead. Daniel Berjón, Alexander Sorkine-Hornung, Francisco Morán, Aljoscha Smolic |
ICME | 2 |
| 2011 | Automatic content creation for multiview autostereoscopic displays using image domain warpingabstractContent creation for autostereoscopic displays is a widely unresolved task. Typical methods rely on view synthesis based on depth image based rendering. Our method applies purely image domain warping instead. Input video is analyzed and information about sparse disparity, vertical edges and saliency is extracted. A constrained energy minimization problem is formulated and efficiently solved. The resulting image warping functions are used to synthesize novel views. Our approach is fully automatic, accurate, and reliable. Disocclusions and related artifacts are avoided due to smooth, saliency-driven warping functions. Our method also works well for extrapolation of views in a limited range, thus supporting multiview creation from stereo input, which is the most relevant use case scenario. Miquel A. Farre, Oliver Wang, Manuel Lang, Nikolce Stefanoski, Alexander Sorkine-Hornung, Aljoscha Smolic |
ICME | 5 |
| 2011 | Three-Dimensional Video Postproduction and ProcessingabstractThis paper gives an overview of the state-of-the-art in 3-D video postproduction and processing as well as an outlook to remaining challenges and opportunities. First, fundamentals of stereography are outlined that set the rules for proper 3-D content creation. Manipulation of the depth composition of a given stereo pair via view synthesis is identified as the key functionality in this context. Basic algorithms are described to adapt and correct fundamental stereo properties such as geometric distortions, color alignment, and stereo geometry. Then, depth image-based rendering is explained as the widely applied solution for view synthesis in 3-D content creation today. Recent improvements of depth estimation already provide very good results. However, in most cases, still interactive workflows dominate. Warping-based methods may become an alternative for some applications in the future, which do not rely on dense and accurate depth estimation. Finally, 2-D to 3-D conversion is covered, which is an important special area for reuse of existing legacy 2-D content in 3-D. Here various advanced algorithms are combined in interactive workflows. Aljoscha Smolic, Peter Kauff, Sebastian Knorr, Alexander Sorkine-Hornung, Matthias Kunter, Marcus Müller 0001, Manuel Lang |
Proc. IEEE | 4 |
| 2011 | Multi-perspective stereoscopy from light fieldsabstractThis paper addresses stereoscopic view generation from a light field. We present a framework that allows for the generation of stereoscopic image pairs with per-pixel control over disparity, based on multi-perspective imaging from light fields. The proposed framework is novel and useful for stereoscopic image processing and post-production. The stereoscopic images are computed as piecewise continuous cuts through a light field, minimizing an energy reflecting prescribed parameters such as depth budget, maximum disparity gradient, desired stereoscopic baseline, and so on. As demonstrated in our results, this technique can be used for efficient and flexible stereoscopic post-processing, such as reducing excessive disparity while preserving perceived depth, or retargeting of already captured scenes to various view settings. Moreover, we generalize our method to multiple cuts, which is highly useful for content creation in the context of multi-view autostereoscopic displays. We present several results on computer-generated content as well as live-action content. Changil Kim 0001, Alexander Sorkine-Hornung, Simon Heinzle, Wojciech Matusik, Markus Gross 0001 |
ACM Trans. Graph. | 2 |
| 2011 | OSCAM - optimized stereoscopic camera control for interactive 3DabstractThis paper presents a controller for camera convergence and interaxial separation that specifically addresses challenges ininteractivestereoscopic applications like games. In such applications, unpredictable viewer- or object-motion often compromises stereopsis due to excessive binocular disparities. We derive constraints on the camera separation and convergence that enable our controller to automatically adapt to any given viewing situation and 3D scene, providing an exact mapping of the virtual content into a comfortable depth range around the display. Moreover, we introduce an interpolation function that linearizes the transformation of stereoscopic depth over time, minimizing nonlinear visual distortions. We describe how to implement the complete control mechanism on the GPU to achieve running times below 0.2ms for full HD. This provides a practical solution even for demanding real-time applications. Results of a user study show a significant increase of stereoscopic comfort, without compromising perceived realism. Our controller enables 'fail-safe' stereopsis, provides intuitive control to accommodate to personal preferences, and allows to properly display stereoscopic content on differently sized output devices. Thomas Oskam, Alexander Sorkine-Hornung, Huw Bowles, Kenny Mitchell, Markus Gross 0001 |
ACM Trans. Graph. | 2 |
| 2010 | Non-linear warping and warp coding for content-adaptive prediction in advanced video coding applicationsabstractThis paper presents a new concept for scalable video coding, which is content adaptive and art-directable. Video retargeting is applied to scale video between different resolutions and aspect ratios without introducing inacceptable distortions or cutting off content. The non-linear warping operations are integrated into a spatial scalability framework, which includes two new building blocks, i.e. non-linear warping prediction and warp coding. Efficient algorithms for both processes are presented, tested and optimized. The presented results indicate that our non-linear scaling and warp coding algorithms provide efficient performance compared to standard linear scaling methods. Further, our advanced scaling algorithms, i.e. EWA splatting in combination with backward mapping, may be very useful for linear scaling as well. Aljoscha Smolic, Yongzhe Wang, Nikolce Stefanoski, Manuel Lang, Alexander Sorkine-Hornung, Markus Gross 0001 |
ICIP | 5 |
| 2010 | Articulated Billboards for Video-based RenderingabstractAbstract We present a novel representation and rendering method for free‐viewpoint video of human characters based on multiple input video streams. The basic idea is to approximate the articulated 3D shape of the human body using a subdivision into textured billboards along the skeleton structure. Billboards are clustered to fans such that each skeleton bone contains one billboard per source camera. We call this representationarticulated billboards. In the paper we describe a semi‐automatic, data‐driven algorithm to construct and render this representation, which robustly handles even challenging acquisition scenarios characterized by sparse camera positioning, inaccurate camera calibration, low video resolution, or occlusions in the scene. First, for each input view, a 2D pose estimation based on image silhouettes, motion capture data, and temporal video coherence is used to create a segmentation mask for each body part. Then, from the 2D poses and the segmentation, the actual articulated billboard model is constructed by a 3D joint optimization and compensation for camera calibration errors. The rendering method includes a novel way of blending the textural contributions of each billboard and features an adaptive seam correction to eliminate visible discontinuities between adjacent billboards textures. Our articulated billboards do not only minimize ghosting artifacts known from conventional billboard rendering, but also alleviate restrictions to the setup and sensitivities to errors of more complex 3D representations and multiview reconstruction techniques. Our results demonstrate the flexibility and the robustness of our approach with high quality free‐viewpoint video generated from broadcast footage of challenging, uncontrolled environments. Marcel Germann, Alexander Sorkine-Hornung, Richard Keiser, Remo Ziegler, Stephan Würmlin, Markus Gross 0001 |
Comput. Graph. Forum | 2 |
| 2010 | Nonlinear disparity mapping for stereoscopic 3DabstractThis paper addresses the problem of remapping the disparity range of stereoscopic images and video. Such operations are highly important for a variety of issues arising from the production, live broadcast, and consumption of 3D content. Our work is motivated by the observation that the displayed depth and the resulting 3D viewing experience are dictated by a complex combination of perceptual, technological, and artistic constraints. We first discuss the most important perceptual aspects of stereo vision and their implications for stereoscopic content creation. We then formalize these insights into a set of basic disparity mapping operators. These operators enable us to control and retarget the depth of a stereoscopic scene in a nonlinear and locally adaptive fashion. To implement our operators, we propose a new strategy based on stereoscopic warping of the input video streams. From a sparse set of stereo correspondences, our algorithm computes disparity and image-based saliency estimates, and uses them to compute a deformation of the input views so as to meet the target disparities. Our approach represents a practical solution for actual stereo production and display that does not require camera calibration, accurate dense depth maps, occlusion handling, or inpainting. We demonstrate the performance and versatility of our method using examples from live action post-production, 3D display size adaptation, and live broadcast. An additional user study and ground truth comparison further provide evidence for the quality and practical relevance of the presented work. Manuel Lang, Alexander Sorkine-Hornung, Oliver Wang, Steven Poulakos, Aljoscha Smolic, Markus Gross 0001 |
ACM Trans. Graph. | 2 |
| 2009 | Interactive Pixel-Accurate Free Viewpoint Rendering from Images with Silhouette Aware SamplingabstractAbstract We present an integrated, fully GPU‐based processing pipeline to interactively render new views of arbitrary scenes from calibrated but otherwise unstructured input views. In a two‐step procedure, our method first generates for each input view a dense proxy of the scene using a new multi‐view stereo formulation. Each scene proxy consists of a structured cloud of feature aware particles which automatically have their image space footprints aligned to depth discontinuities of the scene geometry and hence effectively handle sharp object boundaries and occlusions. We propose a particle optimization routine combined with a special parameterization of the view space that enables an efficient proxy generation as well as robust and intuitive filter operators for noise and outlier removal. Moreover, our generic proxy generation allows us to flexibly handle scene complexities ranging from small objects up to complete outdoor scenes. The second phase of the algorithm combines these particle clouds in real‐time into a view‐dependent proxy for the desired output view and performs a pixel‐accurate accumulation of the colour contributions from each available input view. This makes it possible to reconstruct even fine‐scale view‐dependent illumination effects. We demonstrate how all these processing stages of the pipeline can be implemented entirely on the GPU with memory efficient, scalable data structures for maximum performance. This allows us to generate new output renderings of high visual quality from input images in real‐time. Alexander Sorkine-Hornung, Leif Kobbelt |
Comput. Graph. Forum | 1 |
| 2009 | A system for retargeting of streaming videoabstractWe present a novel, integrated system for content-aware video retargeting. A simple and interactive framework combines key frame based constraint editing with numerous automatic algorithms for video analysis. This combination gives content producers high level control of the retargeting process. The central component of our framework is a non-uniform, pixel-accurate warp to the target resolution which considers automatic as well as interactively defined features. Automatic features comprise video saliency, edge preservation at the pixel resolution, and scene cut detection to enforce bilateral temporal coherence. Additional high level constraints can be added by the producer to guarantee a consistent scene composition across arbitrary output formats. For high quality video display we adopted a 2D version of EWA splatting eliminating aliasing artifacts known from previous work. Our method seamlessly integrates into postproduction and computes the reformatting in real-time. This allows us to retarget annotated video streams at a high quality to arbitary aspect ratios while retaining the intended cinematographic scene composition. For evaluation we conducted a user study which revealed a strong viewer preference for our method. Philipp Krähenbühl, Manuel Lang, Alexander Sorkine-Hornung, Markus Gross 0001 |
ACM Trans. Graph. | 3 |
| 2008 | Image selection for improved Multi-View StereoabstractThe Middlebury multi-view stereo evaluation clearly shows that the quality and speed of most multi-view stereo algorithms depends significantly on the number and selection of input images. In general, not all input images contribute equally to the quality of the output model, since several images may often contain similar and hence overly redundant visual information. This leads to unnecessarily increased processing times. On the other hand, a certain degree of redundancy can help to improve the reconstruction in more ldquodifficultrdquo regions of a model. In this paper we propose an image selection scheme for multi-view stereo which results in improved reconstruction quality compared to uniformly distributed views. Our method is tuned towards the typical requirements of current multi-view stereo algorithms, and is based on the idea of incrementally selecting images so that the overall coverage of a simultaneously generated proxy is guaranteed without adding too much redundant information. Critical regions such as cavities are detected by an estimate of the local photo-consistency and are improved by adding additional views. Our method is highly efficient, since most computations can be out-sourced to the GPU. We evaluate our method with four different methods participating in the Middlebury benchmark and show that in each case reconstructions based on our selected images yield an improved output quality while at the same time reducing the processing time considerably. Alexander Sorkine-Hornung, Boyi Zeng, Leif Kobbelt |
CVPR | 1 |
| 2007 | Character animation from 2D pictures and 3D motion dataabstractThis article presents a new method to animate photos of 2D characters using 3D motion capture data. Given a single image of a person or essentially human-like subject, our method transfers the motion of a 3D skeleton onto the subject's 2D shape in image space, generating the impression of a realistic movement. We present robust solutions to reconstruct a projective camera model and a 3D model pose which matches best to the given 2D image. Depending on the reconstructed view, a 2D shape template is selected which enables the proper handling of occlusions. After fitting the template to the character in the input image, it is deformed as-rigid-as-possible by taking the projected 3D motion data into account. Unlike previous work, our method thereby correctly handles projective shape distortion. It works for images from arbitrary views and requires only a small amount of user interaction. We present animations of a diverse set of human (and nonhuman) characters with different types of motions, such as walking, jumping, or dancing. Alexander Sorkine-Hornung, Ellen Dekkers, Leif Kobbelt |
ACM Trans. Graph. | 1 |
| 2006 | Hierarchical Volumetric Multi-view Stereo Reconstruction of Manifold Surfaces based on Dual Graph EmbeddingabstractThis paper presents a new volumetric stereo algorithm to reconstruct the 3D shape of an arbitrary object. Our method is based on finding the minimum cut in an octahedral graph structure embedded into the volumetric grid, which establishes a well defined relationship between the integrated photo-consistency function of a region in space and the corresponding edge weights of the embedded graph. This new graph structure allows for a highly efficient hierarchical implementation supporting high volumetric resolutions and large numbers of input images. Furthermore we will show how the resulting cut surface can be directly converted into a consistent, closed and manifold mesh. Hence this work provides a complete multi-view stereo reconstruction pipeline. We demonstrate the robustness and efficiency of our technique by a number of high quality reconstructions of real objects. Alexander Sorkine-Hornung, Leif Kobbelt |
CVPR (1) | 1 |
| 2006 | Robust and Efficient Photo-Consistency Estimation for Volumetric 3D Reconstruction
Alexander Sorkine-Hornung, Leif Kobbelt |
ECCV (2) | 1 |
| 2006 | Robust reconstruction of watertight 3D models from non-uniformly sampled point clouds without normal information
Alexander Sorkine-Hornung, Leif Kobbelt |
Symposium on Geometry Processing | 1 |
| 2005 | Self-Calibrating Optical Motion Tracking for Articulated BodiesabstractBuilding intuitive user-interfaces for virtual reality applications is a difficult task, as one of the main purposes is to provide a natural, yet efficient input device to interact with the virtual environment. One particularly interesting approach is to track and retarget the complete motion of a subject. Established techniques for full body motion capture like optical motion tracking exist. However, due to their computational complexity and their reliance on pre-specified models, they fail to meet the demanding requirements of virtual reality environments such as real-time response, immersion, and ad hoc configurability. Our goal is to support the use of motion capture as a general input device for virtual reality applications. In this paper we present a self-calibrating framework for optical motion capture, enabling the reconstruction and tracking of arbitrary articulated objects in real-time. Our method automatically estimates all relevant model parameters on-the-fly without any information on the initial tracking setup or the marker distribution, and computes the geometry and topology of multiple tracked skeletons. Moreover, we show how the model can make the motion capture phase robust against marker occlusions by exploiting the redundancy in the skeleton model and by reconstructing missing inner limbs and joints of the subject from partial information. Meeting the above requirements our system is well applicable to a wide range of virtual reality based applications, where unconstrained tracking and flexible retargeting of motion data is desirable. Alexander Sorkine-Hornung, Sandip Sar-Dessai, Leif Kobbelt |
VR | 1 |
| 2001 | Visualization of eclipses and planetary conjunction events. The interplay between model coherence, scaling and animation
Walter Oberschelp, Alexander Sorkine-Hornung, Horst Samulowitz |
Vis. Comput. | 2 |
| 2000 | Visualization of Eclipses and Planetary Conjunction Events: The Interplay between Model Coherence, Scaling and AnimationabstractThe problem of an instructive and realistic animation and visualization of the shadow and color-conditions during conjunctions of actively and passively illuminated cosmic objects has found only partially satisfying solutions so far. As an example we study a total solar eclipse. There are didactic shortcomings of specialized astronomical software, even though solutions have been given, which are very impressive for experts. Using the possibilities of commercial 3D animation software we present an object oriented partial solution. In order to get correct astronomical representations, we model (for different tasks) the object space under cinematic aspects with parameters for spatial and temporal scaling, for illumination and coloring under couplings of varying strength. The adaptation of the parameters to optimal acceptance of the spectator must be done a posteriori. Walter Oberschelp, Alexander Sorkine-Hornung, Horst Samulowitz |
Computer Graphics International | 2 |