VLDB 2026 Research / reviewers in the wild / expert
Remo Ziegler
dblp:30/3117
· DBLP profile ↗
18ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-authorArtificial intelligence and machine learning · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 71% Face, body and person analysis · 17% Planning, search and constraint satisfaction · 12% | |
| Computer graphics and multimedia
3 papers |
Virtual and augmented reality · 51% Rendering · 32% Computational photography and imaging · 8% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d shape representation |
0.3 | 1 | 2017 | Human Shape from Silhouettes Using Generative HKS Descriptors and Cross-Modal Neural Networks · CVPR 2017 |
Computer vision › 3D vision › 3d human reconstruction
human shape reconstruction |
0.3 | 1 | 2017 | Human Shape from Silhouettes Using Generative HKS Descriptors and Cross-Modal Neural Networks · CVPR 2017 |
Computer vision › 3D vision › 3d reconstruction
shape from silhouette |
0.3 | 1 | 2017 | Human Shape from Silhouettes Using Generative HKS Descriptors and Cross-Modal Neural Networks · CVPR 2017 |
Computer vision › 3D vision › human mesh recovery
human body shape estimation |
0.2 | 1 | 2016 | Shape from Selfies: Human Body Shape Estimation Using CCA Regression Forests · ECCV (4) 2016 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search
branch-and-bound search |
0.2 | 1 | 2014 | Foreground Consistent Human Pose Estimation Using Branch and Bound · ECCV (5) 2014 |
Computer vision › Face, body and person analysis
human pose estimation |
0.2 | 1 | 2014 | Foreground Consistent Human Pose Estimation Using Branch and Bound · ECCV (5) 2014 |
Computer vision › Face, body and person analysis › human body analysis
human body shape analysis |
0.1 | 1 | 2016 | Shape from Selfies: Human Body Shape Estimation Using CCA Regression Forests · ECCV (4) 2016 |
Virtual and augmented reality › immersive display
CAVE |
0.1 | 1 | 2007 | Low-Cost Telepresence for Collaborative Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2007 |
Rendering › physically based rendering › wave optics rendering › computer-generated holography
holographic rendering |
0.1 | 1 | 2007 | A Framework for Holographic Scene Representation and Image Synthesis · IEEE Trans. Vis. Comput. Graph. 2007 |
Virtual and augmented reality
immersive display |
0.1 | 1 | 2007 | Low-Cost Telepresence for Collaborative Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2007 |
Virtual and augmented reality
telepresence |
0.1 | 1 | 2007 | Low-Cost Telepresence for Collaborative Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2007 |
Virtual and augmented reality › avatar
video avatars |
0.1 | 1 | 2007 | Low-Cost Telepresence for Collaborative Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2007 |
Rendering
image-based rendering |
0.0 | 1 | 2002 | Image-based 3D photography using opacity hulls · ACM Trans. Graph. 2002 |
Geometric modeling and processing
shape representation |
0.0 | 1 | 2002 | Image-based 3D photography using opacity hulls · ACM Trans. Graph. 2002 |
Computational photography and imaging
depth estimation |
0.0 | 1 | 2007 | Low-Cost Telepresence for Collaborative Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2007 |
Image and video processing
stereo vision |
0.0 | 1 | 2007 | Low-Cost Telepresence for Collaborative Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2007 |
Methods — techniques the papers use, named apart from their topics
multi-view correlation · 0.3cross-modal mapping · 0.3convolutional neural network · 0.3random forest · 0.2canonical correlation analysis · 0.2branch-and-bound · 0.2view-dependent depth image computation · 0.1stereo compression · 0.1physical camera modeling · 0.1infrared-based image segmentation · 0.1image warping · 0.1digital holography · 0.1relighting · 0.0multi-background matting · 0.0alpha matte · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Convolutional Autoencoders for Human Motion InfillingabstractIn this paper we propose a convolutional autoencoder to address the problem of motion infilling for 3D human motion data. Given a start and end sequence, motion infilling aims to complete the missing gap in between, such that the filled in poses plausibly forecast the start sequence and naturally transition into the end sequence. To this end, we propose a single, end-to-end trainable convolutional autoencoder. We show that a single model can be used to create natural transitions between different types of activities. Furthermore, our method is not only able to fill in entire missing frames, but it can also be used to complete gaps where partial poses are available (e.g. from end effectors), or to clean up other forms of noise (e.g. Gaussian). Also, the model can fill in an arbitrary number of gaps that potentially vary in length. In addition, no further post-processing on the model's outputs is necessary such as smoothing or closing discontinuities at the end of the gap. At the heart of our approach lies the idea to cast motion infilling as an inpainting problem and to train a convolutional de-noising autoencoder on image-like representations of motion sequences. At training time, blocks of columns are removed from such images and we ask the model to fill in the gaps. We demonstrate the versatility of the approach via a number of complex motion sequences and report on thorough evaluations performed to better understand the capabilities and limitations of the proposed approach. Manuel Kaufmann, Emre Aksan, Jie Song 0006, Fabrizio Pece, Remo Ziegler, Otmar Hilliges |
3DV | 5 |
| 2017 | Human Shape from Silhouettes Using Generative HKS Descriptors and Cross-Modal Neural NetworksabstractIn this work, we present a novel method for capturing human body shape from a single scaled silhouette. We combine deep correlated features capturing different 2D views, and embedding spaces based on 3D cues in a novel convolutional neural network (CNN) based architecture. We first train a CNN to find a richer body shape representation space from pose invariant 3D human shape descriptors. Then, we learn a mapping from silhouettes to this representation space, with the help of a novel architecture that exploits correlation of multi-view data during training time, to improve prediction at test time. We extensively validate our results on synthetic and real data, demonstrating significant improvements in accuracy as compared to the state-of-the-art, and providing a practical system for detailed human body measurements from a single image. Endri Dibra, Himanshu Jain, A. Cengiz Öztireli, Remo Ziegler, Markus Gross 0001 |
CVPR | 4 |
| 2017 | DeepGarment : 3D Garment Shape Estimation from a Single Imageabstract3D garment capture is an important component for various applications such as free-view point video, virtual avatars, online shopping, and virtual cloth fitting. Due to the complexity of the deformations, capturing 3D garment shapes requires controlled and specialized setups. A viable alternative is image-based garment capture. Capturing 3D garment shapes from a single image, however, is a challenging problem and the current solutions come with assumptions on the lighting, camera calibration, complexity of human or mannequin poses considered, and more importantly a stable physical state for the garment and the underlying human body. In addition, most of the works require manual interaction and exhibit high run-times. We propose a new technique that overcomes these limitations, making garment shape estimation from an image a practical approach for dynamic garment capture. Starting from synthetic garment shape data generated through physically based simulations from various human bodies in complex poses obtained through Mocap sequences, and rendered under varying camera positions and lighting conditions, our novel method learns a mapping from rendered garment images to the underlying 3D garment model. This is achieved by training Convolutional Neural Networks (CNN-s) to estimate 3D vertex displacements from a template mesh with a specialized loss function. We illustrate that this technique is able to recover the global shape of dynamic 3D garments from a single image under varying factors such as challenging human poses, self occlusions, various camera poses and lighting conditions, at interactive rates. Improvement is shown if more than one view is integrated. Additionally, we show applications of our method to videos. R. Danerek, Endri Dibra, A. Cengiz Öztireli, Remo Ziegler, Markus Gross 0001 |
Comput. Graph. Forum | 4 |
| 2016 | HS-Nets: Estimating Human Body Shape from Silhouettes with Convolutional Neural NetworksabstractWe represent human body shape estimation from binary silhouettes or shaded images as a regression problem, and describe a novel method to tackle it using CNNs. Utilizing a parametric body model, we train CNNs to learn a global mapping from the input to shape parameters used to reconstruct the shapes of people, in neutral poses, with the application of garment fitting in mind. This results in an accurate, robust and automatic system, orders of magnitude faster than methods we compare to, enabling interactive applications. In addition, we show how to combine silhouettes from two views to improve prediction over a single view. The method is extensively evaluated on thousands of synthetic shapes and real data and compared to state of-art approaches, clearly outperforming methods based on global fitting and strongly competing with more expensive local fitting based ones. Endri Dibra, Himanshu Jain, A. Cengiz Öztireli, Remo Ziegler, Markus Gross 0001 |
3DV | 4 |
| 2016 | Shape from Selfies: Human Body Shape Estimation Using CCA Regression Forests
Endri Dibra, A. Cengiz Öztireli, Remo Ziegler, Markus Gross 0001 |
ECCV (4) | 3 |
| 2014 | Joint Camera Pose Estimation and 3D Human Pose Estimation in a Multi-camera Setup
Jens Puwein, Luca Ballan, Remo Ziegler, Marc Pollefeys |
ACCV (2) | 3 |
| 2014 | Foreground Consistent Human Pose Estimation Using Branch and Bound
Jens Puwein, Luca Ballan, Remo Ziegler, Marc Pollefeys |
ECCV (5) | 3 |
| 2012 | PTZ camera network calibration from moving people in sports broadcastsabstractIn sports broadcasts, networks consisting of pan-tilt-zoom (PTZ) cameras usually exhibit very wide baselines, making standard matching techniques for camera calibration very hard to apply. If, additionally, there is a lack of texture, finding corresponding image regions becomes almost impossible. However, such networks are often set up to observe dynamic scenes on a ground plane. Corresponding image trajectories produced by moving objects need to fulfill specific geometric constraints, which can be leveraged for camera calibration. We present a method which combines image trajectory matching with the self-calibration of rotating and zooming cameras, effectively reducing the remaining degrees of freedom in the matching stage to a 2D similarity transformation. Additionally, lines on the ground plane are used to improve the calibration. In the end, all extrinsic and intrinsic camera parameters are refined in a final bundle adjustment. The proposed algorithm was evaluated both qualitatively and quantitatively on four different soccer sequences. Jens Puwein, Remo Ziegler, Luca Ballan, Marc Pollefeys |
WACV | 2 |
| 2012 | Novel-View Synthesis of Outdoor Sport Events Using an Adaptive View-Dependent GeometryabstractAbstract We propose a novel fully automatic method for novel‐viewpoint synthesis. Our method robustly handles multi‐camera setups featuring wide‐baselines in an uncontrolled environment. In a first step, robust and sparse point correspondences are found based on an extension of the Daisy features [ TLF10 ]. These correspondences together with back‐projection errors are used to drive a novel adaptive coarse to fine reconstruction method, allowing to approximate detailed geometry while avoiding an extreme triangle count. To render the scene from arbitrary viewpoints we use a view‐dependent blending of color information in combination with a view‐dependent geometry morph. The view‐dependent geometry compensates for misalignments caused by calibration errors. We demonstrate that our method works well under arbitrary lighting conditions with as little as two cameras featuring wide‐baselines. The footage taken from real sports broadcast events contains fine geometric structures, which result in nice novel‐viewpoint renderings despite of the low resolution in the images. Marcel Germann, Tiberiu Popa, Richard Keiser, Remo Ziegler, Markus Gross 0001 |
Comput. Graph. Forum | 4 |
| 2011 | Robust multi-view camera calibration for wide-baseline camera networksabstractReal-world camera networks are often characterized by very wide baselines covering a wide range of viewpoints. We describe a method not only calibrating each camera sequence added to the system automatically, but also taking advantage of multi-view correspondences to make the entire calibration framework more robust. Novel camera sequences can be seamlessly integrated into the system at any time, adding to the robustness of future computations. One of the challenges consists in establishing correspondences between cameras. Initializing a bag of features from a calibrated frame, correspondences between cameras are established in a two-step procedure. First, affine invariant features of camera sequences are warped into a common coordinate frame and a coarse matching is obtained between the collected features and the incrementally built and updated bag of features. This allows us to warp images to a common view. Second, scale invariant features are extracted from the warped images. This leads to both more numerous and more accurate correspondences. Finally, the parameters are optimized in a bundle adjustment. Adding the feature descriptors and the optimized 3D positions to the bag of features, we obtain a feature-based scene abstraction, allowing for the calibration of novel sequences and the correction of drift in single-view calibration tracking. We demonstrate that our approach can deal with wide baselines. Novel sequences can seamlessly be integrated in the calibration framework. Jens Puwein, Remo Ziegler, Julia Vogel, Marc Pollefeys |
WACV | 2 |
| 2010 | Articulated Billboards for Video-based RenderingabstractAbstract We present a novel representation and rendering method for free‐viewpoint video of human characters based on multiple input video streams. The basic idea is to approximate the articulated 3D shape of the human body using a subdivision into textured billboards along the skeleton structure. Billboards are clustered to fans such that each skeleton bone contains one billboard per source camera. We call this representationarticulated billboards. In the paper we describe a semi‐automatic, data‐driven algorithm to construct and render this representation, which robustly handles even challenging acquisition scenarios characterized by sparse camera positioning, inaccurate camera calibration, low video resolution, or occlusions in the scene. First, for each input view, a 2D pose estimation based on image silhouettes, motion capture data, and temporal video coherence is used to create a segmentation mask for each body part. Then, from the 2D poses and the segmentation, the actual articulated billboard model is constructed by a 3D joint optimization and compensation for camera calibration errors. The rendering method includes a novel way of blending the textural contributions of each billboard and features an adaptive seam correction to eliminate visible discontinuities between adjacent billboards textures. Our articulated billboards do not only minimize ghosting artifacts known from conventional billboard rendering, but also alleviate restrictions to the setup and sensitivities to errors of more complex 3D representations and multiview reconstruction techniques. Our results demonstrate the flexibility and the robustness of our approach with high quality free‐viewpoint video generated from broadcast footage of challenging, uncontrolled environments. Marcel Germann, Alexander Sorkine-Hornung, Richard Keiser, Remo Ziegler, Stephan Würmlin, Markus Gross 0001 |
Comput. Graph. Forum | 4 |
| 2008 | Lighting and Occlusion in a Wave-Based FrameworkabstractAbstract We present novel methods to enhance Computer Generated Holography (CGH) by introducing a complex‐valued wave‐based occlusion handling method. This offers a very intuitive and efficient interface to introduce optical elements featuring physically‐based light interaction exhibiting depth‐of‐field, diffraction, and glare effects. Fur‐thermore, an efficient and flexible evaluation of lit objects on a full‐parallax hologram leads to more convincing images. Previous illumination methods for CGH are not able to change the illumination settings of rendered holo‐grams. In this paper we propose a novel method for real‐time lighting of rendered holograms in order to change the appearance of a previously captured holographic scene. These functionalities are features of a bigger wave‐based rendering framework which can be combined with 2D framebuffer graphics. We present an algorithm which uses graphics hardware to accelerate the rendering. Remo Ziegler, Simone Croci, Markus Gross 0001 |
Comput. Graph. Forum | 1 |
| 2007 | A Bidirectional Light Field - Hologram TransformabstractAbstract In this paper, we propose a novel framework to represent visual information. Extending the notion of conventional image‐based rendering, our framework makes joint use of both light fields and holograms as complementary representations. We demonstrate how light fields can be transformed into holograms, and vice versa. By exploiting the advantages of either representation, our proposed dual representation and processing pipeline is able to overcome the limitations inherent to light fields and holograms alone. We show various examples from synthetic and real light fields to digital holograms demonstrating advantages of either representation, such as speckle‐free images, ghosting‐free images, aliasing‐free recording, natural light recording, aperture‐dependent effects and real‐time rendering which can all be achieved using the same framework. Capturing holograms under white light illumination is one promising application for future work. Remo Ziegler, Simon Bucheli, Lukas Ahrenberg, Marcus A. Magnor, Markus Gross 0001 |
Comput. Graph. Forum | 1 |
| 2007 | Low-Cost Telepresence for Collaborative Virtual EnvironmentsabstractWe present a novel low-cost method for visual communication and telepresence in a CAVE -like environment, relying on 2D stereo-based video avatars. The system combines a selection of proven efficient algorithms and approximations in a unique way, resulting in a convincing stereoscopic real-time representation of a remote user acquired in a spatially immersive display. The system was designed to extend existing projection systems with acquisition capabilities requiring minimal hardware modifications and cost. The system uses infrared-based image segmentation to enable concurrent acquisition and projection in an immersive environment without a static background. The system consists of two color cameras and two additional b/w cameras used for segmentation in the near-IR spectrum. There is no need for special optics as the mask and color image are merged using image-warping based on a depth estimation. The resulting stereo image stream is compressed, streamed across a network, and displayed as a frame-sequential stereo texture on a billboard in the remote virtual environment. Seon-Min Rhee, Remo Ziegler, Jiyoung Park 0002, Martin Näf, Markus Gross 0001, Myoung-Hee Kim |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2007 | A Framework for Holographic Scene Representation and Image SynthesisabstractWe present a framework for the holographic representation and display of graphics objects. As opposed to traditional graphics representations, our approach reconstructs the light wave reflected or emitted by the original object directly from the underlying digital hologram. Our novel holographic graphics pipeline consists of several stages including the digital recording of a full-parallax hologram, the reconstruction and propagation of its wavefront, and rendering of the final image onto conventional, framebuffer-based displays. The required view-dependent depth image is computed from the phase information inherently represented in the complex-valued wavefront. Our model also comprises a correct physical modeling of the camera taking into account optical elements, such as lens and aperture. It thus allows for a variety of effects including depth of field, diffraction, interference, and features built-in anti-aliasing. A central feature of our framework is its seamless integration into conventional rendering and display technology which enables us to elegantly combine traditional 3D object or scene representations with holograms. The presented work includes the theoretical foundations and allows for high quality rendering of objects consisting of large numbers of elementary waves while keeping the hologram at a reasonable size. Remo Ziegler, Peter Kaufmann 0001, Markus Gross 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2005 | Adaptive Instant Displays: Continuously Calibrated Projections Using Per-Pixel Light ControlabstractWe present a framework for achieving user-defined on-demand displays in setups containing bricks of movable cameras and DLP-projectors. A dynamic calibration procedure is introduced, which handles cameras and projectors in a unified way and allows continuous flexible setup changes, while seamless projection alignment and blending is performed simultaneously. For interaction, an intuitive laser pointer based technique is developed, which can be combined with real-time 3D information acquired from the scene. All these tasks can be performed concurrently with the display of a user-chosen application in a non-disturbing way. This is achieved by using an imperceptible structured light approach enabling pixel-based surface light control suited for a wide range of computer graphics and vision algorithms. To ensure scalability of light control in the same working space, multiple projectors are multiplexed. Daniel Cotting, Henry Fuchs, Remo Ziegler, Markus Gross 0001 |
Comput. Graph. Forum | 3 |
| 2003 | 3D Reconstruction Using Labeled Image Regions
Remo Ziegler, Wojciech Matusik, Hanspeter Pfister, Leonard McMillan |
Symposium on Geometry Processing | 1 |
| 2002 | Image-based 3D photography using opacity hullsabstractWe have built a system for acquiring and displaying high quality graphical models of objects that are impossible to scan with traditional scanners. Our system can acquire highly specular and fuzzy materials, such as fur and feathers. The hardware set-up consists of a turntable, two plasma displays, an array of cameras, and a rotating array of directional lights. We use multi-background matting techniques to acquire alpha mattes of the object from multiple viewpoints. The alpha mattes are used to construct an opacity hull. The opacity hull is a new shape representation, defined as the visual hull of the object with view-dependent opacity. It enables visualization of complex object silhouettes and seamless blending of objects into new environments. Our system also supports relighting of objects with arbitrary appearance using surface reflectance fields, a purely image-based appearance representation. Our system is the first to acquire and render surface reflectance fields under varying illumination from arbitrary viewpoints. We have built three generations of digitizers with increasing sophistication. In this paper, we present our results from digitizing hundreds of models. Wojciech Matusik, Hanspeter Pfister, Addy Ngan, Paul A. Beardsley, Remo Ziegler, Leonard McMillan |
ACM Trans. Graph. | 5 |