VLDB 2026 Research / reviewers in the wild / expert
Vladislav Golyanik
dblp:180/6438
· DBLP profile ↗
102ranked-venue papers
15as first author
71since 2021 · last 2026
0000-0003-1630-2006ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 15 first-author · 64 since 2021Artificial intelligence and machine learning · 54 · 5 first-author · 42 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Layered Quantum Architecture Search for 3D Point Cloud ClassificationabstractWe introduce layered Quantum Architecture Search (layered-QAS), a strategy inspired by classical network morphism that designs Parametrised Quantum Circuit (PQC) architectures by progressively growing and adapting them. PQCs offer strong expressiveness with relatively few parameters, yet they lack standard architectural layers (e.g., convolution, attention) that encode inductive biases for a given learning task. To assess the effectiveness of our method, we focus on 3D point cloud classification as a challenging yet highly structured problem. Whereas prior work on this task has used PQCs only as feature extractors for classical classifiers, our approach uses the PQC as the main building block of the classification model. Simulations show that our layered-QAS mitigates barren plateau, outperforms quantum-adapted local and evolutionary QAS baselines, and achieves state-of-the-art results among PQC-based methods on the ModelNet dataset11project page: https://4dqv.mpi-inf.mpg.de/LQAS/. Natacha Kuete Meli, Jovita Lukasik, Vladislav Golyanik, Michael Möller 0001 |
3DV | 3 |
| 2026 | GRMM: Real-Time High-Fidelity Gaussian Morphable Head Model with Learned Residualsabstract3D Morphable Models (3DMMs) enable controllable editing of facial geometry and expressions for reconstruction, animation, and AR/VR, but traditional PCA-based mesh models are limited in resolution, detail, and photorealism. Neural volumetric methods improve realism but remain too slow for interactive use. Recent Gaussian Splatting (3DGS)- based facial models achieve fast, high-quality rendering but still rely solely on a meshbased 3DMM prior for expression control, thereby limiting their ability to capture fine-grained geometry, expressions, and full-head coverage. We introduce GRMM, the first full-head Gaussian 3D morphable model that augments a base 3DMM with residual geometry and appearance components, additive refinements that recover high-frequency details such as wrinkles, fine skin texture, and hairline variations. GRMM provides disentangled control through low-dimensional, interpretable parameters (e.g., identity shape, facial expressions) while separately modelling residuals that capture subject- and expression-specific detail beyond the base model's capacity. Coarse decoders produce vertex-level mesh deformations, fine decoders represent per-Gaussian appearance,and a lightweight CNN refines rasterised images for enhanced realism, all while maintaining 75 FPS real-time rendering. To learn consistent, high-fidelity residuals, we present EXPRESS-50, the first dataset with 60 aligned expressions across 50 identities, enabling robust disentanglement of identity and expression in Gaussian-based 3DMMs. Across monocular 3D face reconstruction, novel-view synthesis, and expression transfer, GRMM surpasses state-of-the-art methods in fidelity and expression accuracy while delivering interactive real-time performance. Project page: https://mohitm1994.github.io/GRMM/ Mohit Mendiratta, Mayur Deshmukh, Kartik Teotia, Vladislav Golyanik, Adam Kortylewski, Christian Theobalt |
3DV | 4 |
| 2026 | EventNeuS: 3D Mesh Reconstruction From a Single Event CameraabstractEvent cameras offer a considerable alternative to RGB cameras in many scenarios. While there are recent works on event-based novel-view synthesis, dense 3D mesh reconstruction remains scarcely explored and existing event-based techniques are severely limited in their 3D reconstruction accuracy. To address this limitation, we present EventNeuS, a self-supervised neural model for learning 3D representations from monocular colour event streams. Our approach, for the first time, combines 3D signed distance function and density field learning with event-based supervision. Furthermore, we introduce spherical harmonics encodings into our model for enhanced handling of view-dependent effects. EventNeuS outperforms existing approaches by a significant margin, achieving 34% lower Chamfer distance and 31% lower mean absolute error on average compared to the best previous method. Shreyas Sachan, Viktor Rudnev, Mohamed A. Elgharib, Christian Theobalt, Vladislav Golyanik |
3DV | 5 |
| 2026 | Quantum Multiple Rotation Averaging
Shuteng Wang, Natacha Kuete Meli, Michael Möller 0001, Vladislav Golyanik |
3DV | 4 |
| 2025 | Betsu-Betsu: Multi-View Separable 3D Reconstruction of Two Interacting ObjectsabstractSeparable 3D reconstruction of multiple objects from multi-view RGB images—resulting in two different 3D shapes for the two objects with a clear separation between them—remains a sparsely researched problem. It is challenging due to severe mutual occlusions and ambiguities along the objects' interaction boundaries. This paper investigates the setting and introduces a new neuro-implicit method that can reconstruct the geometry and appearance of two objects undergoing close interactions while disjoining both in 3D, avoiding surface inter-penetrations and enabling novel-view synthesis of the observed scene. The framework is end-to-end trainable and supervised using a novel alpha-blending regularisation that ensures that the two geometries are well separated even under extreme occlusions. Our reconstruction method is markerless and can be applied to rigid as well as articulated objects. We introduce a new dataset consisting of close interactions between a human and an object and also evaluate on two scenes of humans performing martial arts. The experiments confirm the effectiveness of our framework and substantial improvements using 3D and novel view synthesis metrics compared to several existing approaches applicable in our setting11 Project page: https://vcai.mpi-inf.mpg.de/projects/separable-recon/. Suhas Gopal, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt |
3DV | 3 |
| 2025 | E-3DGS: Event-Based Novel View Rendering of Large-Scale Scenes Using 3D Gaussian SplattingabstractNovel view synthesis techniques predominantly utilize RGB cameras, inheriting their limitations such as the need for sufficient lighting, susceptibility to motion blur, and restricted dynamic range. In contrast, event cameras are significantly more resilient to these limitations but have been less explored in this domain, particularly in large-scale settings. Current methodologies primarily focus on front-facing or object-oriented (360-degree view) scenarios. For the first time, we introduce 3D Gaussians for event-based novel view synthesis. Our method reconstructs large and unbounded scenes with high visual quality. We contribute the first real and synthetic event datasets tailored for this setting. Our method demonstrates superior novel view synthesis and consistently outperforms the baseline EventNeRF by a margin of 11–25% in PSNR (dB) while being orders of magnitude faster in reconstruction and rendering. Sohaib Zahid, Viktor Rudnev, Eddy Ilg, Vladislav Golyanik |
3DV | 4 |
| 2025 | Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation Fieldsabstract3D reconstruction of highly deformable surfaces (e.g. cloths) from monocular RGB videos is a challenging problem, and no solution provides a consistent and accurate recovery of fine-grained surface details. To account for the ill-posed nature of the setting, existing methods use deformation models with statistical, neural, or physical priors. They also predominantly rely on nonadaptive discrete surface representations (e.g. polygonal meshes), perform frame-by-frame optimisation leading to error propagation, and suffer from poor gradients of the mesh-based differentiable renderers. Consequently, fine surface details such as cloth wrinkles are often not recovered with the desired accuracy. In response to these limitations, we propose Thin-Shell-SfT, a new method for non-rigid 3D tracking that represents a surface as an implicit and continuous spatiotemporal neural field. We incorporate continuous thin shell physics prior based on the Kirchhoff-Love model for spatial regularisation, which starkly contrasts the discretised alternatives of earlier works. Lastly, we leverage 3D Gaussian splatting to differentiably render the surface into image space and optimise the deformations based on analysis-by-synthesis principles. Our Thin-Shell-SfT outperforms prior works qualitatively and quantitatively thanks to our continuous surface formulation in conjunction with a specially tailored simulation prior and surface-induced 3D Gaussians. See our project page at https://4dqv.mpi-inf.mpg.de/ThinShellSfT. Navami Kairanda, Marc Habermann, Shanthika Naik, Christian Theobalt, Vladislav Golyanik |
CVPR | 5 |
| 2025 | QuCOOP: A Versatile Framework for Solving Composite and Binary-Parametrised Problems on Quantum AnnealersabstractThere is growing interest in solving computer vision problems such as mesh or point set alignment using Adiabatic Quantum Computing (AQC). Unfortunately, modern experimental AQC devices such as D-Wave only support Quadratic Unconstrained Binary Optimisation (QUBO) problems, which severely limits their applicability. This paper proposes a new way to overcome this limitation and introduces QuCOOP, an optimisation framework extending the scope of AQC to composite and binary-parametrised, possibly non-quadratic problems. The key idea of QuCOOP is to iteratively approximate the original objective function by a sequel of local (intermediate) QUBO forms, whose binary parameters can be sampled on AQC devices. We experiment with quadratic assignment problems, shape matching and point set registration without knowing the correspondences in advance. Our approach achieves state-of-the-art results across multiple instances of tested problems1. Natacha Kuete Meli, Vladislav Golyanik, Marcel Seelbach Benkner, Michael Möller 0001 |
CVPR | 2 |
| 2025 | BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated ObjectsabstractWe present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-based contact maps conditioned on the object trajectory with an articulation-aware feature representation, revealing rich bimanual patterns for manipulation. The learned contact prior is then used to guide our hand motion generator, producing diverse and realistic bimanual motions for object movement and articulation. Our work offers key insights into feature representation and contact prior for articulated objects, demonstrating their effectiveness in taming the complex, high-dimensional space of bimanual hand-object interactions. Through comprehensive quantitative experiments, we demonstrate a clear step towards simplified and high-quality hand-object animations that surpass the state of the art in motion quality and diversity. Project page: https://vcai.mpi-inf.mpg.de/projects/bimart/. Wanyue Zhang, Rishabh Dabral, Vladislav Golyanik, Vasileios Choutas, Eduardo Alvarado, Thabo Beeler, Marc Habermann, Christian Theobalt |
CVPR | 3 |
| 2025 | Denoising Functional Maps: Diffusion Models for Shape CorrespondenceabstractEstimating correspondences between pairs of deformable shapes remains a challenging problem. Despite substantial progress, existing methods lack broad generalization capabilities and require category-specific training data. To address these limitations, we propose a fundamentally new approach to shape correspondence based on denoising diffusion models. In our method, a diffusion model learns to directly predict the functional map, a low-dimensional representation of a point-wise map between shapes. We use a large dataset of synthetic human meshes for training and employ two steps to reduce the number of functional maps that need to be learned. First, the maps refer to a template rather than shape pairs. Second, the functional map is defined in a basis of eigenvectors of the Laplacian, which is not unique due to sign ambiguity. Therefore, we introduce an unsupervised approach to select a specific basis by correcting the signs of eigenvectors based on surface features. Our model achieves competitive performance on standard human datasets, meshes with anisotropic connectivity, non-isometric humanoid shapes, as well as animals compared to existing descriptor-based and large-scale shape deformation methods. See our project page1for the source code2and the datasets. Aleksei Zhuravlev, Zorah Lähner, Vladislav Golyanik |
CVPR | 3 |
| 2025 | Bring Your Rear Cameras for Egocentric 3D Human Pose EstimationabstractEgocentric 3D human pose estimation has been actively studied using cameras installed in front of a head-mounted device (HMD). While frontal placement is the optimal and the only option for some tasks, such as hand tracking, it remains unclear if the same holds for full-body tracking due to self-occlusion and limited field-of-view coverage. Notably, even the state-of-the-art methods often fail to estimate accurate 3D poses in many scenarios, such as when HMD users tilt their heads upward -- a common motion in human activities. A key limitation of existing HMD designs is their neglect of the back of the body, despite its potential to provide crucial 3D reconstruction cues. Hence, this paper investigates the usefulness of rear cameras for full-body tracking. We also show that simply adding rear views to the frontal inputs is not optimal for existing methods due to their dependence on individual 2D joint detectors without effective multi-view integration. To address this issue, we propose a new transformer-based method that refines 2D joint heatmap estimation with multi-view information and heatmap uncertainty, thereby improving 3D pose tracking. Also, we introduce two new large-scale datasets, Ego4View-Syn and Ego4View-RW, for a rear-view evaluation. Our experiments show that the new camera configurations with back views provide superior support for 3D pose tracking compared to only frontal placements. The proposed method achieves significant improvement over the current state of the art (>10% on MPJPE). The source code, trained models, and datasets are available on our project page at https://4dqv.mpi-inf.mpg.de/EgoRear/. Hiroyasu Akada, Vladislav Golyanik, Christian Theobalt |
ICCV | 3 |
| 2025 | HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View Synthesis
Timo Teufel, Pulkit Gera, Xilong Zhou 0001, Umar Iqbal 0001, Pramod Rao, Jan Kautz, Vladislav Golyanik, Christian Theobalt |
ICCV | 7 |
| 2025 | DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single ImageabstractReconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, complex deformations, and the ambiguity of the single-view setting. The previous state-of-the-art, Decaf, employs a global fitting optimization guided by contact and deformation estimation networks trained on studio-collected data with 3D annotations. However, Decaf suffers from a time-consuming optimization process and limited generalization capability due to its reliance on 3D annotations of hand-face interaction data. To address these issues, we present DICE, the first end-to-end method for Deformation-aware hand-face Interaction reCovEry from a single image. DICE estimates the poses of hands and faces, contacts, and deformations simultaneously using a Transformer-based architecture. It features disentangling the regression of local deformation fields and global mesh vertex locations into two network branches, enhancing deformation and contact estimation for precise and robust hand-face mesh recovery. To improve generalizability, we propose a weakly-supervised training approach that augments the training set using in-the-wild images without 3D ground-truth annotations, employing the depths of 2D keypoints estimated by off-the-shelf models and adversarial priors of poses for supervision. Our experiments demonstrate that DICE achieves state-of-the-art performance on a standard benchmark and in-the- wild data in terms of accuracy and physical plausibility. Additionally, our method operates at an interactive rate (20 fps) on an Nvidia 4090 GPU, whereas Decaf requires more than 15 seconds for a single image. The code will be available at: https://github.com/Qingxuan-Wu/DICE. Qingxuan Wu, Zhiyang Dou, Sirui Xu 0002, Soshi Shimada, Chen Wang 0049, Zhengming Yu, Yuan Liu 0025, Cheng Lin 0001, Zeyu Cao, Taku Komura, Vladislav Golyanik, Christian Theobalt, Wenping Wang 0001, Lingjie Liu |
ICLR | 11 |
| 2025 | Attention (as Discrete-Time Markov) ChainsabstractWe introduce a new interpretation of the attention matrix as a discrete-time Markov chain. Our interpretation sheds light on common operations involving attention scores such as selection, summation, and averaging in a unified framework. It further extends them by considering indirect attention, propagated through the Markov chain, as opposed to previous studies that only model immediate effects. Our key observation is that tokens linked to semantically similar regions form metastable states, i.e., regions where attention tends to concentrate, while noisy attention scores dissipate. Metastable states and their prevalence can be easily computed through simple matrix multiplication and eigenanalysis, respectively. Using these lightweight tools, we demonstrate state-of-the-art zero-shot segmentation. Lastly, we define TokenRank---the steady state vector of the Markov chain, which measures global token importance. We show that TokenRank enhances unconditional image generation, improving both quality (IS) and diversity (FID), and can also be incorporated into existing segmentation techniques to improve their performance over existing benchmarks. We believe our framework offers a fresh view of how tokens are being attended in modern visual transformers. Yotam Erel, Olaf Dünkel, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, Amit Bermano |
NeurIPS | 4 |
| 2025 | Quantum Visual Fields with Neural Amplitude EncodingabstractQuantum Implicit Neural Representations (QINRs) have emerged as a promising paradigm that leverages parametrised quantum circuits to encode and process classical information. However, significant challenges remain in areas such as ansatz architecture design, the effective utility of quantum-mechanical properties, training efficiency, and the integration with classical modules. This paper advances the field by introducing a novel QINR architecture for 2D image and 3D geometric field learning, which we collectively refer to as Quantum Visual Field (QVF). QVF encodes classical data into quantum statevectors using neural amplitude encoding grounded in a learnable energy manifold, ensuring meaningful Hilbert space embeddings. Our ansatz follows a fully entangled design of learnable parametrised quantum circuits, with quantum (unitary) operations performed in the real Hilbert space, resulting in numerically stable training with fast convergence. QVF does not rely on classical post-processing---in contrast to the previous QINR learning approach---and directly employs measurements to extract learned signals encoded in the ansatz. Experiments on a quantum hardware simulator demonstrate that QVF outperforms existing quantum approach and competes widely used classical foundational baselines in terms of visual representation accuracy across various metrics and model characteristics. We also show applications of QVF in 2D and 3D field completion and 3D shape interpolation, highlighting its practical potential. Project page: \url{https://4dqv.mpi-inf.mpg.de/QVF/}. Shuteng Wang, Christian Theobalt, Vladislav Golyanik |
NeurIPS | 3 |
| 2025 | D-NPC: Dynamic Neural Point Clouds for Non-Rigid View Synthesis from Monocular VideoabstractAbstract Dynamic reconstruction and spatiotemporal novel‐view synthesis of non‐rigidly deforming scenes recently gained increased attention. While existing work achieves impressive quality and performance on multi‐view or teleporting camera setups, most methods fail to efficiently and faithfully recover motion and appearance from casual monocular captures. This paper contributes to the field by introducing a new method for dynamic novel view synthesis from monocular video, such as casual smartphone captures. Our approach represents the scene as a dynamic neural point cloud, an implicit time‐conditioned point distribution that encodes local geometry and appearance in separate hash‐encoded neural feature grids for static and dynamic regions. By sampling a discrete point cloud from our model, we can efficiently render high‐quality novel views using a fast differentiable rasterizer and neural rendering network. Similar to recent work, we leverage advances in neural scene analysis by incorporating data‐driven priors like monocular depth estimation and object segmentation to resolve motion and depth ambiguities originating from the monocular captures. In addition to guiding the optimization process, we show that these priors can be exploited to explicitly initialize our scene representation to drastically improve optimization speed and final image quality. As evidenced by our experimental evaluation, our dynamic point cloud model not only enables fast optimization and real‐time frame rates for interactive applications, but also achieves competitive image quality on monocular benchmark sequences. Our code and data are available online https://moritzkappel.github.io/projects/dnpc/ . Moritz Kappel, Florian Hahlbohm, Timon Scholz, Susana Castillo 0001, Christian Theobalt, Martin Eisemann, Vladislav Golyanik, Marcus A. Magnor |
Comput. Graph. Forum | 7 |
| 2025 | EventEgo3D++: 3D Human Motion Capture from A Head-Mounted Event CameraabstractMonocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, which are common in head-mounted device applications. Existing methods that rely on RGB cameras often fail under these conditions. To address these limitations, we introduce EventEgo3D++, the first approach that leverages a monocular event camera with a fisheye lens for 3D human motion capture. Event cameras excel in high-speed scenarios and varying illumination due to their high temporal resolution, providing reliable cues for accurate 3D human motion capture. EventEgo3D++ leverages the LNES representation of event streams to enable precise 3D reconstructions. We have also developed a mobile head-mounted device (HMD) prototype equipped with an event camera, capturing a comprehensive dataset that includes real event observations from both controlled studio environments and in-the-wild settings, in addition to a synthetic dataset. Additionally, to provide a more holistic dataset, we include allocentric RGB streams that offer different perspectives of the HMD wearer, along with their corresponding SMPL body model. Our experiments demonstrate that EventEgo3D++ achieves superior 3D accuracy and robustness compared to existing solutions, even in challenging conditions. Moreover, our method supports real-time 3D pose updates at a rate of 140Hz. This work is an extension of the EventEgo3D approach (CVPR 2024) and further advances the state of the art in egocentric 3D human motion capture. For more details, visit the project page at https://eventego3d.mpi-inf.mpg.de. Christen Millerdurai, Hiroyasu Akada, Jian Wang 0042, Diogo C. Luvizon, Alain Pagani, Didier Stricker, Christian Theobalt, Vladislav Golyanik |
Int. J. Comput. Vis. | 8 |
| 2025 | PractiLight: Practical Light Control Using Foundational Diffusion ModelsabstractLight control in generated images is a difficult task, posing specific challenges, spanning over the entire image and frequency spectrum. Most approaches tackle this problem by training on extensive yet domain-specific datasets, limiting the inherent generalization and applicability of the foundational backbones used. Instead, PractiLight is a practical approach, effectively leveraging foundational understanding of recent generative models for the task. Our key insight is that lighting relationships in an image are similar in nature to token interaction in self-attention layers, and hence are best represented there. Based on this and other analyses regarding the importance of early diffusion iterations, PractiLight trains a lightweight LoRA regressor to produce the direct-irradiance map for a given image, using a small set of training images. We then employ this regressor to incorporate the desired lighting into the generation process of another image using Classifier Guidance. This careful design generalizes well to diverse conditions and image domains. We demonstrate state-of-the-art performance in terms of quality and control with proven parameter and data efficiency compared to leading works over a wide variety of scene types. We hope this work affirms that image lighting can feasibly be controlled by tapping into foundational knowledge, enabling practical and general relighting. Yotam Erel, Rishabh Dabral, Vladislav Golyanik, Amit Bermano, Christian Theobalt |
ACM Trans. Graph. | 3 |
| 2025 | Fast Non-Rigid Radiance Fields From Monocularized DataabstractThe reconstruction and novel view synthesis of dynamic scenes recently gained increased attention. As reconstruction from large-scale multi-view data involves immense memory and computational requirements, recent benchmark datasets provide collections of single monocular views per timestamp sampled from multiple (virtual) cameras. We refer to this form of inputs asmonocularizeddata. Existing work shows impressive results for synthetic setups and forward-facing real-world data, but is often limited in the training speed and angular range for generating novel views. This paper addresses these limitations and proposes a new method for full$360^{\circ}$inward-facing novel view synthesis of non-rigidly deforming scenes. At the core of our method are: 1) An efficient deformation module that decouples the processing of spatial and temporal information for accelerated training and inference; and 2) A static module representing the canonical scene as a fast hash-encoded neural radiance field. In addition to existing synthetic monocularized data, we systematically analyze the performance on real-world inward-facing scenes using a newly recorded challenging dataset sampled from a synchronized large-scale multi-view rig. In both cases, our method is significantly faster than previous methods, converging in less than 7 minutes and achieving real-time framerates at 1K resolution, while obtaining a higher visual accuracy for generated novel views. Our code and dataset are available online:https://github.com/MoritzKappel/MoNeRF. Moritz Kappel, Vladislav Golyanik, Susana Castillo 0001, Christian Theobalt, Marcus A. Magnor |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Quantum-Hybrid Stereo Matching With Nonlinear Regularization and Spatial PyramidsabstractQuantum visual computing is advancing rapidly. This paper presents a new formulation for stereo matching with nonlinear regularizers and spatial pyramids on quantum annealers as a maximum a posteriori inference problem that minimizes the energy of a Markov Random Field. Our approach is hybrid (i.e., quantum-classical) and is compatible with modern D-Wave quantum annealers, i.e., it includes a quadratic unconstrained binary optimization (QUBO) objective. Previous quantum annealing techniques for stereo matching are limited to using linear regularizers, and thus, they do not exploit the fundamental advantages of the quantum computing paradigm in solving combinatorial optimization problems. In contrast, our method utilizes the full potential of quantum annealing for stereo matching, as nonlinear regularizers create optimization problems which are $\mathcal{N}\mathcal{P}$-hard. On the Middlebury benchmark, we achieve an improved root mean squared accuracy over the previous state of the art in quantum stereo matching of $2 \%$ and $22.5 \%$ when using different solvers. Cameron Braunstein, Eddy Ilg, Vladislav Golyanik |
3DV | 3 |
| 2024 | 3D Pose Estimation of Two Interacting Hands from a Monocular Event Cameraabstract3D hand tracking from a monocular video is a very challenging problem due to hand interactions, occlusions, left-right hand ambiguity, and fast motion. Most existing methods rely on RGB inputs, which have severe limitations under low-light conditions and suffer from motion blur. In contrast, event cameras capture local brightness changes instead of full image frames and do not suffer from the described effects. Unfortunately, existing image-based techniques cannot be directly applied to events due to significant differences in the data modalities. In response to these challenges, this paper introduces the first framework for 3D tracking of two fast-moving and interacting hands from a single monocular event camera. Our approach tackles the left-right hand ambiguity with a novel semi-supervised feature-wise attention mechanism and integrates an intersection loss to fix hand collisions. To facilitate advances in this research domain, we release a new synthetic large-scale dataset of two interacting hands, Ev2Hands-S, and a new real benchmark with real event streams and ground-truth 3D annotations, Ev2Hands-R. Our approach outperforms existing methods in terms of the 3D reconstruction accuracy and generalises to real data under severe light conditions1.1https://4dqv.mpi-inf.mpg.de/Ev2Hands/ Christen Millerdurai, Diogo C. Luvizon, Viktor Rudnev, André Jonas, Jiayi Wang 0001, Christian Theobalt, Vladislav Golyanik |
3DV | 7 |
| 2024 | MACS: Mass Conditioned 3D Hand and Object Motion SynthesisabstractThe physical properties of an object, such as mass, significantly affect how we manipulate it with our hands. Surprisingly, this aspect has so far been neglected in prior work on 3D motion synthesis. To improve the naturalness of the synthesized 3D hand-object motions, this work proposes MACS–the first MAss Conditioned 3D hand and object motion Synthesis approach. Our approach is based on cascaded diffusion models and generates interactions that plausibly adjust based on the object’s mass and interaction type. MACS also accepts a manually drawn 3D object trajectory as input and synthesizes the natural 3D hand motions conditioned by the object’s mass. This flexibility enables MACS to be used for various downstream applications, such as generating synthetic training data for ML tasks, fast animation of hands for graphics workflows, and generating character interactions for computer games. We show experimentally that a small-scale dataset is sufficient for MACS to reasonably generalize across interpolated and extrapolated object masses unseen during the training. Furthermore, MACS shows moderate generalization to unseen objects, thanks to the mass-conditioned contact labels generated by our surface contact synthesis model ConNet. Our comprehensive user study confirms that the synthesized 3D hand-object interactions are highly plausible and realistic. Soshi Shimada, Franziska Mueller 0001, Jan Bednarík, Bardia Doosti, Bernd Bickel, Danhang Tang, Vladislav Golyanik, Jonathan Taylor 0001, Christian Theobalt, Thabo Beeler |
3DV | 7 |
| 2024 | SceNeRFlow: Time-Consistent Reconstruction of General Dynamic ScenesabstractExisting methods for the 4D reconstruction of general, non-rigidly deforming objects focus on novel-view synthesis and neglect correspondences. However, time consistency enables advanced downstream tasks like 3D editing, motion analysis, or virtual-asset creation. We propose SceNeRFlow to reconstruct a general, non-rigid scene in a time-consistent manner. Our dynamic-NeRF method takes multi-view RGB videos and background images from static cameras with known camera parameters as input. It then reconstructs the deformations of an estimated canonical model of the geometry and appearance in an online fashion. Since this canonical model is time-invariant, we obtain correspondences even for long-term, long-range motions. We employ neural scene representations to parametrize the components of our method. Like prior dynamic-NeRF methods, we use a backwards deformation model. We find nontrivial adaptations of this model necessary to handle larger motions: We decompose the deformations into a strongly regularized coarse component and a weakly regularized fine component, where the coarse component also extends the deformation field into the space surrounding the object, which enables tracking over time. We show experimentally that, unlike prior work that only handles small motion, our method enables the reconstruction of studio-scale motions. Edith Tretschk, Vladislav Golyanik, Michael Zollhöfer, Aljaz Bozic, Christoph Lassner, Christian Theobalt |
3DV | 2 |
| 2024 | ROAM: Robust and Object-Aware Motion Generation Using Neural Pose DescriptorsabstractExisting automatic approaches for 3D virtual character motion synthesis supporting scene interactions do not gen-eralise well to new objects outside training distributions, even when trained on extensive motion capture datasets with diverse objects and annotated interactions. This paper addresses this limitation and shows that robustness and generalisation to novel scene objects in 3D object-aware character synthesis can be achieved by training a motion model with as few as one reference object. We leverage an implicit feature representation trained on object-only datasets, which encodes an SE(3)-equivariant descriptor field around the object. Given an unseen object and a reference pose-object pair, we optimise for the object-aware pose that is closest in the feature space to the reference pose. Finally, we use l-NSM, i.e., our motion generation model that is trained to seamlessly transition from locomotion to object interaction with the proposed bidirectional pose blending scheme. Through comprehensive numerical comparisons to state-of-the-art methods and in a user study, we demonstrate substantial improvements in 3D virtual character motion and interaction quality and robustness to scenarios with unseen objects. Our project page is available at https://vcai.mpi-inf.mpg.de/projects/ROAM/. Wanyue Zhang, Rishabh Dabral, Thomas Leimkühler, Vladislav Golyanik, Marc Habermann, Christian Theobalt |
3DV | 4 |
| 2024 | Projected Stochastic Gradient Descent with Quantum Annealed Binary Gradients
Maximilian Krahn, Michele Sasdelli, Frances Fengyi Yang, Vladislav Golyanik, Juho Kannala, Tat-Jun Chin, Tolga Birdal |
BMVC | 4 |
| 2024 | 3D Human Pose Perception from Egocentric Stereo VideosabstractWhile head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from egocentric views. In this work, we propose a new transformer-based framework to improve egocentric stereo 3D human pose estimation, which leverages the scene information and temporal context of egocentric stereo videos. Specifically, we utilize 1) depth features from our 3D scene reconstruction module with uniformly sampled windows of egocentric stereo frames, and 2) human joint queries enhanced by temporal features of the video inputs. Our method is able to accurately estimate human poses even in challenging scenarios, such as crouching and sitting. Furthermore, we introduce two new benchmark datasets, i.e., UnrealEgo2 and UnrealEgo-RW (RealWorld). The proposed datasets offer a much larger number of egocentric stereo views with a wider variety of human motions than the existing datasets, allowing comprehensive evaluation of existing and upcoming methods. Our extensive experiments show that the proposed approach significantly outperforms previous methods. UnrealEgo2, UnrealEgo-RW, and trained models are available on our project page11https://4dqv.mpi-inf.mpg.de/UnrealEgo2/ and Benchmark Challenge22https://unrealego.mpi-inf.mpg.de/. Hiroyasu Akada, Jian Wang 0042, Vladislav Golyanik, Christian Theobalt |
CVPR | 3 |
| 2024 | VINECS: Video-based Neural Character SkinningabstractRigging and skinning clothed human avatars is a challenging task and traditionally requires a lot of manual work and expertise. Recent methods addressing it either generalize across different characters or focus on capturing the dynamics of a single character observed under different pose configurations. However, the former methods typically predict solely static skinning weights, which perform poorly for highly articulated poses, and the latter ones either require dense 3D character scans in different poses or cannot generate an explicit mesh with vertex correspondence over time. To address these challenges, we propose a fully automated approach for creating a fully rigged character with pose-dependent skinning weights, which can be solely learned from multi-view video. Therefore, we first acquire a rigged template, which is then statically skinned. Next, a coordinate-based MLP learns a skinning weights field parameterized over the position in a canonical pose space and the respective pose. Moreover, we introduce our pose- and view-dependent appearance field allowing us to differentiably render and supervise the posed mesh using multi-view imagery. We show that our approach outperforms state-of-the-art while not relying on dense 4D scans. More details can be found on our project page11https://people.mpi-inf.mpg.de/~mhaberma/projects/2023-Vinecs. Zhouyingcheng Liao, Vladislav Golyanik, Marc Habermann, Christian Theobalt |
CVPR | 2 |
| 2024 | EventEgo3D: 3D Human Motion Capture from Egocentric Event StreamsabstractMonocular egocentric 3D human motion capture is a challenging and actively researched problem. Existing methods use synchronously operating visual sensors (e.g. RGB cameras) and often fail under low lighting and fast motions, which can be restricting in many applications involving head-mounted devices. In response to the existing limitations, this paper 1) introduces a new problem, i.e. 3D human motion capture from an egocentric monoc-ular event camera with a fish eye lens, and 2) proposes the first approach to it called EventEgo 3D (EE3D). Event streams have high temporal resolution and provide reliable cues for 3D human motion capture under high-speed hu-man motions and rapidly changing illumination. The pro-posed EE3D framework is specifically tailored for learning with event streams in the LNES representation, enabling high 3D reconstruction accuracy. We also design a proto-type of a mobile head-mounted device with an event cam-era and record a real dataset with event observations and the ground-truth 3D human poses (in addition to the syn-thetic dataset). Our EE3D demonstrates robustness and su-perior 3D accuracy compared to existing solutions across various challenging experiments while supporting real-time 3D pose update rates of 140Hz.11https://4dqv.mpi-inf.mpg.de/EventEgo3D/ Christen Millerdurai, Hiroyasu Akada, Jian Wang 0042, Diogo C. Luvizon, Christian Theobalt, Vladislav Golyanik |
CVPR | 6 |
| 2024 | Holoported Characters: Real-Time Free-Viewpoint Rendering of Humans from Sparse RGB CamerasabstractWe present the first approach to render highly realistic free-viewpoint videos of a human actor in general apparel, from sparse multi-view recording to display, in real-time at an unprecedented 4K resolution. At inference, our method only requires four camera views of the moving actor and the respective 3D skeletal pose. It handles actors in wide clothing, and reproduces even fine-scale dynamic detail, e.g. clothing wrinkles, face expressions, and hand gestures. At training time, our learning-based approach expects dense multi- view video and a rigged static surface scan of the actor. Our method comprises three main stages. Stage 1 is a skeleton-driven neural approach for high-quality capture of the detailed dynamic mesh geometry. Stage 2 is a novel solution to create a view-dependent texture using four test-time camera views as input. Finally, stage 3 comprises a new image-based refinement network rendering the final 4K image given the output from the previous stages. Our approach establishes a new benchmark for real-time rendering resolution and quality using sparse input camera views, unlocking possibilities for immersive telepresence. Code and data is available on our project page. Ashwath Shetty, Marc Habermann, Guoxing Sun 0001, Diogo C. Luvizon, Vladislav Golyanik, Christian Theobalt |
CVPR | 5 |
| 2024 | REMOS: 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions
Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, Philipp Slusallek |
ECCV (36) | 3 |
| 2024 | Relightable Neural Actor with Intrinsic Decomposition and Pose Control
Diogo C. Luvizon, Vladislav Golyanik, Adam Kortylewski, Marc Habermann, Christian Theobalt |
ECCV (59) | 2 |
| 2024 | NeuralClothSim: Neural Deformation Fields Meet the Thin Shell TheoryabstractDespite existing 3D cloth simulators producing realistic results, they predominantly operate on discrete surface representations (e.g. points and meshes) with a fixed spatial resolution, which often leads to large memory consumption and resolution-dependent simulations. Moreover, back-propagating gradients through the existing solvers is difficult and they hence cannot be easily integrated into modern neural architectures. In response, this paper re-thinks physically plausible cloth simulation: We propose NeuralClothSim, i.e., a new quasistatic cloth simulator using thin shells, in which surface deformation is encoded in neural network weights in form of a neural field. Our memory-efficient solver operates on a new continuous coordinate-based surface representation called neural deformation fields (NDFs); it supervises NDF equilibria with the laws of the non-linear Kirchhoff-Love shell theory with a non-linear anisotropic material model. NDFs are adaptive: They 1) allocate their capacity to the deformation details and 2) allow surface state queries at arbitrary spatial resolutions without re-training. We show how to train NeuralClothSim while imposing hard boundary conditions and demonstrate multiple applications, such as material interpolation and simulation editing. The experimental results highlight the effectiveness of our continuous neural formulation. Navami Kairanda, Marc Habermann, Christian Theobalt, Vladislav Golyanik |
NeurIPS | 4 |
| 2024 | State of the Art on Diffusion Models for Visual ComputingabstractAbstract The field of visual computing is rapidly advancing due to the emergence of generative artificial intelligence (AI), which unlocks unprecedented capabilities for the generation, editing, and reconstruction of images, videos, and 3D scenes. In these domains, diffusion models are the generative AI architecture of choice. Within the last year alone, the literature on diffusion‐based tools and applications has seen exponential growth and relevant papers are published across the computer graphics, computer vision, and AI communities with new works appearing daily on arXiv. This rapid growth of the field makes it difficult to keep up with all recent developments. The goal of this state‐of‐the‐art report (STAR) is to introduce the basic mathematical concepts of diffusion models, implementation details and design choices of the popular Stable Diffusion model, as well as overview important aspects of these generative AI tools, including personalization, conditioning, inversion, among others. Moreover, we give a comprehensive overview of the rapidly growing literature on diffusion‐based generation and editing, categorized by the type of generated medium, including 2D images, videos, 3D objects, locomotion, and 4D scenes. Finally, we discuss available datasets, metrics, open challenges, and social implications. This STAR provides an intuitive starting point to explore this exciting topic for researchers, artists, and practitioners alike. Ryan Po, Wang Yifan 0001, Vladislav Golyanik, Kfir Aberman, Jonathan T. Barron, Amit Bermano, Eric R. Chan, Tali Dekel, Aleksander Holynski, Angjoo Kanazawa, C. Karen Liu, Lingjie Liu, Ben Mildenhall, Matthias Nießner, Björn Ommer, Christian Theobalt, Peter Wonka, Gordon Wetzstein |
Comput. Graph. Forum | 3 |
| 2024 | Recent Trends in 3D Reconstruction of General Non-Rigid ScenesabstractAbstract Reconstructing models of the real world, including 3D geometry, appearance, and motion of real scenes, is essential for computer graphics and computer vision. It enables the synthesizing of photorealistic novel views, useful for the movie industry and AR/VR applications. It also facilitates the content creation necessary in computer games and AR/VR by avoiding laborious manual design processes. Further, such models are fundamental for intelligent computing systems that need to interpret real‐world scenes and actions to act and interact safely with the human world. Notably, the world surrounding us is dynamic, and reconstructing models of dynamic, non‐rigidly moving scenes is a severely underconstrained and challenging problem. This state‐of‐the‐art report (STAR) offers the reader a comprehensive summary of state‐of‐the‐art techniques with monocular and multi‐view inputs such as data from RGB and RGB‐D sensors, among others, conveying an understanding of different approaches, their potential applications, and promising further research directions. The report covers 3D reconstruction of general non‐rigid scenes and further addresses the techniques for scene decomposition, editing and controlling, and generalizable and generative modeling. More specifically, we first review the common and fundamental concepts necessary to understand and navigate the field and then discuss the state‐of‐the‐art techniques by reviewing recent approaches that use traditional and machine‐learning‐based neural representations, including a discussion on the newly enabled applications. The STAR is concluded with a discussion of the remaining limitations and open challenges. Raza Yunus, Jan Eric Lenssen, Michael Niemeyer, Yiyi Liao, Christian Rupprecht 0001, Christian Theobalt, Gerard Pons-Moll, Jia-Bin Huang 0001, Vladislav Golyanik, Eddy Ilg |
Comput. Graph. Forum | 9 |
| 2024 | Promptable Game Models: Text-guided Game Simulation via Masked Diffusion ModelsabstractNeural video game simulators emerged as powerful tools to generate and edit videos. Their idea is to represent games as the evolution of an environment’s state driven by the actions of its agents. While such a paradigm enables users toplaya game action-by-action, its rigidity precludes more semantic forms of control. To overcome this limitation, we augment game models withpromptsspecified as a set ofnatural languageactions anddesired states. The result—a Promptable Game Model (PGM)—makes it possible for a user toplaythe game by prompting it with high- and low-level action sequences. Most captivatingly, our PGM unlocks thedirector’s mode, where the game is played by specifying goals for the agents in the form of a prompt. This requires learning “game AI,” encapsulated by our animation model, to navigate the scene using high-level constraints, play against an adversary, and devise a strategy to win a point. To render the resulting state, we use a compositional NeRF representation encapsulated in our synthesis model. To foster future research, we present newly collected, annotated and calibrated Tennis and Minecraft datasets. Our method significantly outperforms existing neural video game simulators in terms of rendering quality and unlocks applications beyond the capabilities of the current state-of-the-art. Our framework, data, and models are available at snap-research.github.io/promptable-game-models. Willi Menapace, Aliaksandr Siarohin, Stéphane Lathuilière, Panos Achlioptas, Vladislav Golyanik, Sergey Tulyakov, Elisa Ricci 0001 |
ACM Trans. Graph. | 5 |
| 2023 | Fully Quantum Auto-Encoding of 3D Shapes
Lakshika Rathi, Edith Tretschk, Christian Theobalt, Rishabh Dabral, Vladislav Golyanik |
BMVC | 5 |
| 2023 | CCuantuMM: Cycle-Consistent Quantum-Hybrid Matching of Multiple ShapesabstractJointly matching multiple, non-rigidly deformed 3D shapes is a challenging,$\mathcal{NP}$-hard problem. A perfect matching is necessarily cycle-consistent: Following the pairwise point correspondences along several shapes must end up at the starting vertex of the original shape. Unfortunately, existing quantum shape-matching methods do not support multiple shapes and even less cycle consistency. This paper addresses the open challenges and introduces the first quantum-hybrid approach for 3D shape multi-matching; in addition, it is also cycle-consistent. Its iterative formulation is admissible to modern adiabatic quantum hardware and scales linearly with the total number of input shapes. Both these characteristics are achieved by reducing the N-shape case to a sequence of three-shape matchings, the derivation of which is our main technical contribution. Thanks to quantum annealing, high-quality solutions with low energy are retrieved for the intermediate$\mathcal{NP}$- hard objectives. On benchmark datasets, the proposed approach significantly outperforms extensions to multi-shape matching of a previous quantum-hybrid two-shape matching method and is on-par with classical multi-matching methods. Our source code is available at 4dqv.mpiinf.mpg.de/CCuantuMM/. Harshil Bhatia, Edith Tretschk, Zorah Lähner, Marcel Seelbach Benkner, Michael Möller 0001, Christian Theobalt, Vladislav Golyanik |
CVPR | 7 |
| 2023 | MoFusion: A Framework for Denoising-Diffusion-Based Motion SynthesisabstractConventional methods for human motion synthesis have either been deterministic or have had to struggle with the trade-off between motion diversity vs motion quality. In response to these limitations, we introduce MoFusion, i.e., a new denoising-diffusion-based framework for high-quality conditional human motion synthesis that can synthesise long, temporally plausible, and semantically accurate motions based on a range of conditioning contexts (such as music and text). We also present ways to introduce well-known kinematic losses for motion plausibility within the motion-diffusion framework through our scheduled weighting strategy. The learned latent space can be used for several interactive motion-editing applications like in-betweening, seed-conditioning, and text-based editing, thus, providing crucial abilities for virtual-character animation and robotics. Through comprehensive quantitative evaluations and a perceptual user study, we demonstrate the effectiveness of MoFusion compared to the state of the art on established benchmarks in the literature. We urge the reader to watch our supplementary video at https://vcai.mpi-inf.mpg.de/projects/MoFusion/. Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, Christian Theobalt |
CVPR | 3 |
| 2023 | Quantum Multi-Model FittingabstractGeometric model fitting is a challenging but fundamental computer vision problem. Recently, quantum optimization has been shown to enhance robust fitting for the case of a single model, while leaving the question of multi-model fitting open. In response to this challenge, this paper shows that the latter case can significantly benefit from quantum hardware and proposes the first quantum approach to multimodel fitting (MMF). We formulate MMF as a problem that can be efficiently sampled by modern adiabatic quantum computers without the relaxation of the objective function. We also propose an iterative and decomposed version of our method, which supports real-world-sized problems. The experimental evaluation demonstrates promising results on a variety of datasets. The source code is available at: https://github.com/FarinaMatteo/qmmf. Matteo Farina, Luca Magri 0002, Willi Menapace, Elisa Ricci 0001, Vladislav Golyanik, Federica Arrigoni |
CVPR | 5 |
| 2023 | Self-Supervised Pre-Training with Masked Shape Prediction for 3D Scene UnderstandingabstractMasked signal modeling has greatly advanced self-supervised pre-training for language and 2D images. However, it is still not fully explored in 3D scene understanding. Thus, this paper introduces Masked Shape Prediction (MSP), a new framework to conduct masked signal modeling in 3D scenes. MSP uses the essential 3D semantic cue, i.e., geometric shape, as the prediction target for masked points. The context-enhanced shape target consisting of explicit shape context and implicit deep shape feature is proposed to facilitate exploiting contextual cues in shape prediction. Meanwhile, the pre-training architecture in MSP is carefully designed to alleviate the masked shape leakage from point coordinates. Experiments on multiple 3D understanding tasks on both indoor and outdoor datasets demon-strate the effectiveness of MSP in learning good feature representations to consistently boost downstream performance. Li Jiang 0009, Zetong Yang, Shaoshuai Shi, Vladislav Golyanik, Dengxin Dai, Bernt Schiele |
CVPR | 4 |
| 2023 | EventNeRF: Neural Radiance Fields from a Single Colour Event CameraabstractAsynchronously operating event cameras find many applications due to their high dynamic range, vanishingly low motion blur, low latency and low data bandwidth. The field saw remarkable progress during the last few years, and existing event-based 3D reconstruction approaches recover sparse point clouds of the scene. However, such sparsity is a limiting factor in many cases, especially in computer vision and graphics, that has not been addressed satisfactorily so far. Accordingly, this paper proposes the first approach for 3D-consistent, dense and photorealistic novel view synthesis using just a single colour event stream as input. At its core is a neural radiance field trained entirely in a self-supervised manner from events while preserving the original resolution of the colour event channels. Next, our ray sampling strategy is tailored to events and allows for data-efficient training. At test, our method produces results in the RGB space at unprecedented quality. We evaluate our method qualitatively and numerically on several challenging synthetic and real scenes and show that it produces significantly denser and more visually appealing renderings than the existing methods. We also demonstrate robustness in challenging scenarios with fast motion and under low lighting conditions. We release the newly recorded dataset and our source code to facilitate the research field, see https://4dqv.mpi-inf.mpg.de/EventNeRF. Viktor Rudnev, Mohamed A. Elgharib, Christian Theobalt, Vladislav Golyanik |
CVPR | 4 |
| 2023 | QuAnt: Quantum Annealing with Learnt Couplings
Marcel Seelbach Benkner, Maximilian Krahn, Edith Tretschk, Zorah Lähner, Michael Möller 0001, Vladislav Golyanik |
ICLR | 6 |
| 2023 | Discovering Fatigued Movements for Virtual Character AnimationabstractVirtual character animation and movement synthesis have advanced rapidly during recent years, especially through a combination of extensive motion capture datasets and machine learning. A remaining challenge is interactively simulating characters that fatigue when performing extended motions, which is indispensable for the realism of generated animations. However, capturing such movements is problematic, as performing movements like backflips with fatigued variations up to exhaustion raises capture cost and risk of injury. Surprisingly, little research has been done on faithful fatigue modeling. To address this, we propose a deep reinforcement learning-based approach, which—for the first time in literature—generates control policies for full-body physically simulated agents aware of cumulative fatigue. For this, we first leverage Generative Adversarial Imitation Learning (GAIL) to learn an expert policy for the skill; Second, we learn a fatigue policy by limiting the generated constant torque bounds based on endurance time to non-linear, state- and time-dependent limits in the joint-actuation space using a Three-Compartment Controller (3CC) model. Our results demonstrate that agents can adapt to different fatigue and rest rates interactively, and discover realistic recovery strategies without the need for any captured data of fatigued movement. Noshaba Cheema, Nam Hee Kim, Perttu Hämäläinen, Vladislav Golyanik, Marc Habermann, Christian Theobalt, Philipp Slusallek |
SIGGRAPH Asia | 5 |
| 2023 | IMoS: Intent-Driven Full-Body Motion Synthesis for Human-Object InteractionsabstractAbstract Can we make virtual characters in a scene interact with their surrounding objects through simple instructions? Is it possible to synthesize such motion plausibly with a diverse set of objects and instructions? Inspired by these questions, we present the first framework to synthesize the full‐body motion of virtual human characters performing specified actions with 3D objects placed within their reach. Our system takes textual instructions specifying the objects and the associated ‘intentions’ of the virtual characters as input and outputs diverse sequences of full‐body motions. This contrasts existing works, where full‐body action synthesis methods generally do not consider object interactions, and human‐object interaction methods focus mainly on synthesizing hand or finger movements for grasping objects. We accomplish our objective by designing an intent‐driven full‐body motion generator, which uses a pair of decoupled conditional variational auto‐regressors to learn the motion of the body parts in an autoregressive manner. We also optimize the 6‐DoF pose of the objects such that they plausibly fit within the hands of the synthesized characters. We compare our proposed method with the existing methods of motion synthesis and establish a new and stronger state‐of‐the‐art for the task of intent‐driven motion synthesis. Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, Philipp Slusallek |
Comput. Graph. Forum | 3 |
| 2023 | Scene-Aware 3D Multi-Human Motion Capture from a Single CameraabstractAbstract In this work, we consider the problem of estimating the 3D position of multiple humans in a scene as well as their body shape and articulation from a single RGB video recorded with a static camera. In contrast to expensive marker‐based or multi‐view systems, our lightweight setup is ideal for private users as it enables an affordable 3D motion capture that is easy to install and does not require expert knowledge. To deal with this challenging setting, we leverage recent advances in computer vision using large‐scale pre‐trained models for a variety of modalities, including 2D body joints, joint angles, normalized disparity maps, and human segmentation masks. Thus, we introduce the first non‐linear optimization‐based approach that jointly solves for the 3D position of each human, their articulated pose, their individual shapes as well as the scale of the scene. In particular, we estimate the scene depth and person scale from normalized disparity predictions using the 2D body joints and joint angles. Given the per‐frame scene depth, we reconstruct a point‐cloud of the static scene in 3D space. Finally, given the per‐frame 3D estimates of the humans and scene point‐cloud, we perform a space‐time coherent optimization over the video to ensure temporal, spatial and physical plausibility. We evaluate our method on established multi‐person 3D human pose benchmarks where we consistently outperform previous methods and we qualitatively demonstrate that our method is robust to in‐the‐wild conditions including challenging scenes with people of different sizes. Code: https://github.com/dluvizon/scene‐aware‐3d‐multi‐human Diogo C. Luvizon, Marc Habermann, Vladislav Golyanik, Adam Kortylewski, Christian Theobalt |
Comput. Graph. Forum | 3 |
| 2023 | State of the Art in Dense Monocular Non-Rigid 3D ReconstructionabstractAbstract 3D reconstruction of deformable (ornon‐rigid) scenes from a set of monocular 2D image observations is a long‐standing and actively researched area of computer vision and graphics. It is an ill‐posed inverse problem, since—without additional prior assumptions—it permits infinitely many solutions leading to accurate projection to the input 2D images. Non‐rigid reconstruction is a foundational building block for downstream applications like robotics, AR/VR, or visual content creation. The key advantage of using monocular cameras is their omnipresence and availability to the end users as well as their ease of use compared to more sophisticated camera set‐ups such as stereo or multi‐view systems. This survey focuses on state‐of‐the‐art methods for dense non‐rigid 3D reconstruction of various deformable objects and composite scenes from monocular videos or sets of monocular views. It reviews the fundamentals of 3D reconstruction and deformation modeling from 2D image observations. We then start from general methods—that handle arbitrary scenes and make only a few prior assumptions—and proceed towards techniques making stronger assumptions about the observed objects and types of deformations (e.g. human faces, bodies, hands, and animals). A significant part of this STAR is also devoted to classification and a high‐level comparison of the methods, as well as an overview of the datasets for training and evaluation of the discussed techniques. We conclude by discussing open challenges in the field and the social aspects associated with the usage of the reviewed methods. Edith Tretschk, Navami Kairanda, Mallikarjun B. R. 0001, Rishabh Dabral, Adam Kortylewski, Bernhard Egger 0001, Marc Habermann, Pascal Fua, Christian Theobalt, Vladislav Golyanik |
Comput. Graph. Forum | 10 |
| 2023 | AvatarStudio: Text-Driven Editing of 3D Dynamic Human Head AvatarsabstractCapturing and editing full-head performances enables the creation of virtual characters with various applications such as extended reality and media production. The past few years witnessed a steep rise in the photorealism of human head avatars. Such avatars can be controlled through different input data modalities, including RGB, audio, depth, IMUs, and others. While these data modalities provide effective means of control, they mostly focus on editing the head movements such as the facial expressions, head pose, and/or camera viewpoint. In this paper, we propose AvatarStudio, a text-based method for editing the appearance of a dynamic full head avatar. Our approach builds on existing work to capture dynamic performances of human heads using Neural Radiance Field (NeRF) and edits this representation with a text-to-image diffusion model. Specifically, we introduce an optimization strategy for incorporating multiple keyframes representing different camera viewpoints and time stamps of a video performance into a single diffusion model. Using this personalized diffusion model, we edit the dynamic NeRF by introducing view-and-time-aware Score Distillation Sampling (VT-SDS) following a model-based guidance approach. Our method edits the full head in a canonical space and then propagates these edits to the remaining time steps via a pre-trained deformation network. We evaluate our method visually and numerically via a user study, and results show that our method outperforms existing approaches. Our experiments validate the design choices of our method and highlight that our edits are genuine, personalized, as well as 3D- and time-consistent. Mohit Mendiratta, Xingang Pan, Mohamed A. Elgharib, Kartik Teotia, Mallikarjun B. R. 0001, Ayush Tewari, Vladislav Golyanik, Adam Kortylewski, Christian Theobalt |
ACM Trans. Graph. | 7 |
| 2023 | Decaf: Monocular Deformation Capture for Face and Hand InteractionsabstractExisting methods for 3D tracking from monocular RGB videos predominantly consider articulated and rigid objects ( e.g. , two hands or humans interacting with rigid environments). Modelling dense non-rigid object deformations in this setting ( e.g. when hands are interacting with a face), remained largely unaddressed so far, although such effects can improve the realism of the downstream applications such as AR/VR, 3D virtual avatar communications, and character animations. This is due to the severe ill-posedness of the monocular view setting and the associated challenges ( e.g. , in acquiring a dataset for training and evaluation or obtaining the reasonable non-uniform stiffness of the deformable object). While it is possible to naïvely track multiple non-rigid objects independently using 3D templates or parametric 3D models, such an approach would suffer from multiple artefacts in the resulting 3D estimates such as depth ambiguity, unnatural intra-object collisions and missing or implausible deformations. Hence, this paper introduces the first method that addresses the fundamental challenges depicted above and that allows tracking human hands interacting with human faces in 3D from single monocular RGB videos. We model hands as articulated objects inducing non-rigid face deformations during an active interaction. Our method relies on a new hand-face motion and interaction capture dataset with realistic face deformations acquired with a markerless multi-view camera system. As a pivotal step in its creation, we process the reconstructed raw 3D shapes with position-based dynamics and an approach for non-uniform stiffness estimation of the head tissues, which results in plausible annotations of the surface deformations, hand-face contact regions and head-hand positions. At the core of our neural approach are a variational auto-encoder supplying the hand-face depth prior and modules that guide the 3D tracking by estimating the contacts and the deformations. Our final 3D hand and face reconstructions are realistic and more plausible compared to several baselines applicable in our setting, both quantitatively and qualitatively. https://vcai.mpi-inf.mpg.de/projects/Decaf Soshi Shimada, Vladislav Golyanik, Patrick Pérez, Christian Theobalt |
ACM Trans. Graph. | 2 |
| 2023 | EgoLocate: Real-time Motion Capture, Localization, and Mapping with Sparse Body-mounted SensorsabstractHuman and environment sensing are two important topics in Computer Vision and Graphics. Human motion is often captured by inertial sensors, while the environment is mostly reconstructed using cameras. We integrate the two techniques together in EgoLocate, a system that simultaneously performs human motion capture (mocap), localization, and mapping in real time from sparse body-mounted sensors, including 6 inertial measurement units (IMUs) and a monocular phone camera. On one hand, inertial mocap suffers from large translation drift due to the lack of the global positioning signal. EgoLo-cate leverages image-based simultaneous localization and mapping (SLAM) techniquesto locate the human in the reconstructed scene. Onthe other hand, SLAM often fails when the visual feature is poor. EgoLocate involves inertial mocap to provide a strong prior for the camera motion. Experiments show that localization, a key challenge for both two fields, is largely improved by our technique, compared with the state of the art of the two fields. Our codes are available for research at https://xinyu-yi.github.io/EgoLocate/. Xinyu Yi, Yuxiao Zhou 0001, Marc Habermann, Vladislav Golyanik, Shaohua Pan 0002, Christian Theobalt, Feng Xu 0005 |
ACM Trans. Graph. | 4 |
| 2022 | MoCapDeform: Monocular 3D Human Motion Capture in Deformable Scenesabstract3D human motion capture from monocular RGB images respecting interactions of a subject with complex and possibly deformable environments is a very challenging, illposed and under-explored problem. Existing methods address it only weakly and do not model possible surface deformations often occurring when humans interact with scene surfaces. In contrast, this paper proposes MoCapDe-form, i.e., a new framework for monocular 3D human motion capture that is the first to explicitly model non-rigid deformations of a 3D scene for improved 3D human pose estimation and deformable environment reconstruction. Mo-CapDeform accepts a monocular RGB video and a 3D scene mesh aligned in the camera space. It first localises a subject in the input monocular video along with dense contact labels using a new raycasting based strategy. Next, our human-environment interaction constraints are leveraged to jointly optimise global 3D human poses and non-rigid surface deformations. MoCapDeform achieves superior accuracy than competing methods on several datasets, including our newly recorded one with deforming background scenes. Zhi Li 0055, Soshi Shimada, Bernt Schiele, Christian Theobalt, Vladislav Golyanik |
3DV | 5 |
| 2022 | HiFECap: Monocular High-Fidelity and Expressive Capture of Human Performances
Marc Habermann, Vladislav Golyanik, Christian Theobalt |
BMVC | 3 |
| 2022 | $\phi$-SfT: Shape-from-Template with a Physics-Based Deformation ModelabstractShape-from-Template (SfT) methods estimate 3D surface deformations from a single monocular RGB camera while assuming a 3D state known in advance (a template). This is an important yet challenging problem due to the under-constrained nature of the monocular setting. Existing SfT techniques predominantly use geometric and simplified deformation models, which often limits their reconstruction abilities. In contrast to previous works, this paper proposes a new SfT approach explaining 2D observations through physical simulations accounting for forces and material properties. Our differentiable physics simulator regularises the surface evolution and optimises the material elastic properties such as bending coefficients, stretching stiffness and density. We use a differentiable renderer to minimise the dense reprojection error between the estimated 3D states and the input images and recover the deformation parameters using an adaptive gradient-based optimisation. For the evaluation, we record with an RGB-D camera challenging real surfaces exposed to physical forces with various material properties and textures. Our approach significantly reduces the 3D reconstruction error compared to multiple competing methods. For the source code and data, see https://4dqv.mpi-inf.mpg.de/phi-SfT/. Navami Kairanda, Edith Tretschk, Mohamed A. Elgharib, Christian Theobalt, Vladislav Golyanik |
CVPR | 5 |
| 2022 | Playable Environments: Video Manipulation in Space and TimeabstractWe present Playable Environments-a new representation for interactive video generation and manipulation in space and time. With a single image at inference time, our novel framework allows the user to move objects in 3D while generating a video by providing a sequence of desired actions. The actions are learnt in an unsupervised manner. The camera can be controlled to get the desired viewpoint. Our method builds an environment state for each frame, which can be manipulated by our proposed action mod-ule and decoded back to the image space with volumetric rendering. To support diverse appearances of objects, we extend neural radiance fields with style-based modulation. Our method trains on a collection of various monocular videos requiring only the estimated camera parameters and 2D object locations. To set a challenging benchmark, we in-troduce two large scale video datasets with significant cam-era movements. As evidenced by our experiments, playable environments enable several creative applications not at-tainable by prior video synthesis works, including playable 3D video generation, stylization and manipulation11willi-menapace.github.io/playable-environments-website. Willi Menapace, Stéphane Lathuilière, Aliaksandr Siarohin, Christian Theobalt, Sergey Tulyakov, Vladislav Golyanik, Elisa Ricci 0001 |
CVPR | 6 |
| 2022 | Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsabstractMotion capture from sparse inertial sensors has shown great potential compared to image-based approaches since occlusions do not lead to a reduced tracking quality and the recording space is not restricted to be within the viewing frustum of the camera. However, capturing the motion and global position only from a sparse set of inertial sensors is inherently ambiguous and challenging. In consequence, recent state-of-the-art methods can barely handle very long period motions, and unrealistic artifacts are common due to the unawareness of physical constraints. To this end, we present the first method which combines a neural kinematics estimator and a physics-aware motion optimizer to track body motions with only 6 inertial sensors. The kinematics module first regresses the motion status as a reference, and then the physics module refines the motion to satisfy the physical constraints. Experiments demonstrate a clear improvement over the state of the art in terms of capture accuracy, temporal stability, and physical correctness. Xinyu Yi, Yuxiao Zhou 0001, Marc Habermann, Soshi Shimada, Vladislav Golyanik, Christian Theobalt, Feng Xu 0005 |
CVPR | 5 |
| 2022 | UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture
Hiroyasu Akada, Jian Wang 0042, Soshi Shimada, Masaki Takahashi 0001, Christian Theobalt, Vladislav Golyanik |
ECCV (6) | 6 |
| 2022 | Quantum Motion Segmentation
Federica Arrigoni, Willi Menapace, Marcel Seelbach Benkner, Elisa Ricci 0001, Vladislav Golyanik |
ECCV (29) | 5 |
| 2022 | NeRF for Outdoor Scene Relighting
Viktor Rudnev, Mohamed A. Elgharib, William A. P. Smith, Lingjie Liu, Vladislav Golyanik, Christian Theobalt |
ECCV (16) | 5 |
| 2022 | HULC: 3D HUman Motion Capture with Pose Manifold SampLing and Dense Contact Guidance
Soshi Shimada, Vladislav Golyanik, Zhi Li 0055, Patrick Pérez, Weipeng Xu, Christian Theobalt |
ECCV (22) | 2 |
| 2022 | Q-FW: A Hybrid Classical-Quantum Frank-Wolfe for Quadratic Binary Optimization
Alp Yurtsever, Tolga Birdal, Vladislav Golyanik |
ECCV (23) | 3 |
| 2022 | Advances in Neural RenderingabstractAbstract Synthesizing photo‐realistic images and videos is at the heart of computer graphics and has been the focus of decades of research. Traditionally, synthetic images of a scene are generated using rendering algorithms such as rasterization or ray tracing, which take specifically defined representations of geometry and material properties as input. Collectively, these inputs define the actual scene and what is rendered, and are referred to as the scene representation (where a scene consists of one or more objects). Example scene representations are triangle meshes with accompanied textures (e.g., created by an artist), point clouds (e.g., from a depth sensor), volumetric grids (e.g., from a CT scan), or implicit surface functions (e.g., truncated signed distance fields). The reconstruction of such a scene representation from observations using differentiable rendering losses is known as inverse graphics or inverse rendering. Neural rendering is closely related, and combines ideas from classical computer graphics and machine learning to create algorithms for synthesizing images from real‐world observations. Neural rendering is a leap forward towards the goal of synthesizing photo‐realistic image and video content. In recent years, we have seen immense progress in this field through hundreds of publications that show different ways to inject learnable components into the rendering pipeline. This state‐of‐the‐art report on advances in neural rendering focuses on methods that combine classical rendering principles with learned 3D scene representations, often now referred to as neural scene representations. A key advantage of these methods is that they are 3D‐consistent by design, enabling applications such as novel viewpoint synthesis of a captured scene. In addition to methods that handle static scenes, we cover neural scene representations for modeling non‐rigidly deforming objects and scene editing and composition. While most of these approaches are scene‐specific, we also discuss techniques that generalize across object classes and can be used for generative tasks. In addition to reviewing these state‐of‐the‐art methods, we provide an overview of fundamental concepts and definitions used in the current literature. We conclude with a discussion on open challenges and social implications. Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul P. Srinivasan, Edith Tretschk, Wang Yifan 0001, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, Tomas Simon, Christian Theobalt, Matthias Nießner, Jonathan T. Barron, Gordon Wetzstein, Michael Zollhöfer, Vladislav Golyanik |
Comput. Graph. Forum | 17 |
| 2022 | HandVoxNet++: 3D Hand Shape and Pose Estimation Using Voxel-Based Neural Networksabstract3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. Existing methods addressing it directly regress hand meshes via 2D convolutional neural networks, which leads to artifacts due to perspective distortions in the images. To address the limitations of the existing methods, we develop HandVoxNet++, i.e., a voxel-based deep network with 3D and graph convolutions trained in a fully supervised manner. The input to our network is a 3D voxelized-depth-map-based on the truncated signed distance function (TSDF). HandVoxNet++ relies on two hand shape representations. The first one is the 3D voxelized grid of hand shape, which does not preserve the mesh topology and which is the most accurate representation. The second representation is the hand surface that preserves the mesh topology. We combine the advantages of both representations by aligning the hand surface to the voxelized hand shape either with a new neural Graph-Convolutions-based Mesh Registration (GCN-MeshReg) or classical segment-wise Non-Rigid Gravitational Approach (NRGA++) which does not rely on training data. In extensive evaluations on three public benchmarks, i.e., SynHand5M, depth-based HANDS19 challenge and HO-3D, the proposed HandVoxNet++ achieves the state-of-the-art performance. In this journal extension of our previous approach presented at CVPR 2020, we gain 41.09% and 13.7% higher shape alignment accuracy on SynHand5M and HANDS19 datasets, respectively. Our method is ranked first on the HANDS19 challenge dataset (Task 1: Depth-Based 3D Hand Pose Estimation) at the moment of the submission of our results to the portal in August 2020. Jameel Malik, Soshi Shimada, Ahmed Elhayek, Sk Aziz Ali, Christian Theobalt, Vladislav Golyanik, Didier Stricker |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Convex Joint Graph Matching and Clustering via Semidefinite RelaxationsabstractThis paper proposes a new algorithm for simultaneous graph matching and clustering. For the first time in the literature, these two problems are solved jointly and synergetically without relying on any training data, which brings advantages for identifying similar arbitrary objects in compound 3D scenes and matching them. For joint reasoning, we first rephrase graph matching as a rigid point set registration problem operating on spectral graph embeddings. Consequently, we utilise efficient convex semidefinite program relaxations for aligning points in Hilbert spaces and add coupling constraints to model the mutual dependency and exploit synergies between both tasks. We outperform state of the art in challenging cases with non-perfectly matching and noisy graphs, and we show successful applications on real compound scenes with multiple 3D elements. Our source code and data are publicly available. Maximilian Krahn, Florian Bernard 0001, Vladislav Golyanik |
3DV | 3 |
| 2021 | HumanGAN: A Generative Model of Human ImagesabstractGenerative adversarial networks achieve great performance in photorealistic image synthesis in various domains, including human images. However, they usually employ latent vectors that encode the sampled outputs globally. This does not allow convenient control of semantically relevant individual parts of the image, and cannot draw samples that only differ in partial aspects, such as clothing style. We address these limitations and present a generative model for images of dressed humans offering control over pose, local body part appearance and garment style. This is the first method to solve various aspects of human image generation, such as global appearance sampling, pose transfer, parts and garment transfer, and part sampling jointly in a unified framework. As our model encodes part-based latent appearance vectors in a normalized pose-independent space and warps them to different poses, it preserves body and clothing appearance under varying posture. Experiments show that our flexible and general generative method outperforms task-specific baselines for pose-conditioned image generation, pose transfer and part sampling in terms of realism and output resolution. Kripasindhu Sarkar, Lingjie Liu, Vladislav Golyanik, Christian Theobalt |
3DV | 3 |
| 2021 | Quantum Permutation SynchronizationabstractWe present QuantumSync, the first quantum algorithm for solving a synchronization problem in the context of computer vision. In particular, we focus on permutation synchronization which involves solving a non-convex optimization problem in discrete variables. We start by formulating synchronization into a quadratic unconstrained binary optimization problem (QUBO). While such formulation respects the binary nature of the problem, ensuring that the result is a set of permutations requires extra care. Hence, we: (i) show how to insert permutation constraints into a QUBO problem and (ii) solve the constrained QUBO problem on the current generation of the adiabatic quantum computers D-Wave. Thanks to the quantum annealing, we guarantee global optimality with high probability while sampling the energy landscape to yield confidence estimates. Our proof-of-concepts realization on the adiabatic D-Wave computer demonstrates that quantum machines offer a promising way to solve the prevalent yet difficult synchronization problems. Tolga Birdal, Vladislav Golyanik, Christian Theobalt, Leonidas J. Guibas |
CVPR | 2 |
| 2021 | High-Fidelity Neural Human Motion Transfer From Monocular VideoabstractVideo-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing dynamics, including fine and high-frequency details, significantly limits the attainable visual quality. We address these limitations for the first time in the literature and present a new framework which performs high-fidelity and temporally-consistent human motion transfer with natural pose-dependent non-rigid deformations, for several types of loose garments. In contrast to the previous techniques, we perform image generation in three subsequent stages: synthesizing human shape, structure, and appearance. Given a monocular RGB video of an actor, we train a stack of recurrent deep neural networks that generate these intermediate representations from 2D poses and their temporal derivatives. Splitting the difficult motion transfer problem into subtasks that are aware of the temporal motion context helps us to synthesize results with plausible dynamics and pose-dependent detail. It also allows artistic control of results by manipulation of individual framework stages. In the experimental results, we significantly outperform the state-of-the-art in terms of video realism. The source code is available at https://graphics.tu-bs.de/publications/kappel2020high-fidelity. Moritz Kappel, Vladislav Golyanik, Mohamed A. Elgharib, Jann-Ole Henningson, Hans-Peter Seidel, Susana Castillo 0001, Christian Theobalt, Marcus A. Magnor |
CVPR | 2 |
| 2021 | Pose-Guided Human Animation From a Single Image in the WildabstractWe present a new pose transfer method for synthesizing a human animation from a single image of a person controlled by a sequence of body poses. Existing pose transfer methods exhibit significant visual artifacts when applying to a novel scene, resulting in temporal inconsistency and failures in preserving the identity and textures of the person. To address these limitations, we design a compositional neural network that predicts the silhouette, garment labels, and textures. Each modular network is explicitly dedicated to a subtask that can be learned from the synthetic data. At the inference time, we utilize the trained network to produce a unified representation of appearance and its labels in UV coordinates, which remains constant across poses. The unified representation provides an incomplete yet strong guidance to generating the appearance in response to the pose change. We use the trained network to complete the appearance and render it with the background. With these strategies, we are able to synthesize human animations that can preserve the identity and appearance of the person in a temporally coherent way without any fine-tuning of the network on the testing scene. Experiments show that our method outperforms the state-of-the-arts in terms of synthesis quality, temporal coherence, and generalization ability. Jae Shin Yoon, Lingjie Liu, Vladislav Golyanik, Kripasindhu Sarkar, Hyun Soo Park, Christian Theobalt |
CVPR | 3 |
| 2021 | Q-Match: Iterative Shape Matching via Quantum AnnealingabstractFinding shape correspondences can be formulated as an ${\mathcal{N}}{\mathcal{P}} - hard$ quadratic assignment problem (QAP) that becomes infeasible for shapes with high sampling density. A promising research direction is to tackle such quadratic optimization problems over binary variables with quantum annealing, which allows for some problems a more efficient search in the solution space. Unfortunately, enforcing the linear equality constraints in QAPs via a penalty significantly limits the success probability of such methods on currently available quantum hardware. To address this limitation, this paper proposes Q-Match, i.e., a new iterative quantum method for QAPs inspired by the α-expansion algorithm, which allows solving problems of an order of magnitude larger than current quantum methods. It implicitly enforces the QAP constraints by updating the current estimates in a cyclic fashion. Further, Q-Match can be applied iteratively, on a subset of well-chosen correspondences, al-lowing us to scale to real-world problems. Using the latest quantum annealer, the D-Wave Advantage, we evaluate the proposed method on a subset of QAPLIB as well as on isometric shape matching problems from the FAUST dataset. Marcel Seelbach Benkner, Zorah Lähner, Vladislav Golyanik, Christof Wunderlich, Christian Theobalt, Michael Möller 0001 |
ICCV | 3 |
| 2021 | Gravity-Aware Monocular 3D Human-Object ReconstructionabstractThis paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free flight. In contrast to existing monocular methods, we can recover scale, object trajectories as well as human bone lengths in meters and the ground plane's orientation, thanks to the awareness of the gravity constraining object motions. Our objective function is parametrised by the object's initial velocity and position, gravity direction and focal length, and jointly optimised for one or several free flight episodes. The proposed human-object interaction constraints ensure geometric consistency of the 3D reconstructions and improved physical plausibility of human poses compared to the unconstrained case. We evaluate GraviCap on a new dataset with ground-truth annotations for persons and different objects undergoing free flights. In the experiments, our approach achieves state-of-the-art accuracy in 3D human motion capture on various metrics. We urge the reader to watch our supplementary video. Both the source code and the dataset are released; see http://4dqv.mpi-inf.mpg.de/GraviCap/. Rishabh Dabral, Soshi Shimada, Arjun Jain, Christian Theobalt, Vladislav Golyanik |
ICCV | 5 |
| 2021 | EventHands: Real-Time Neural 3D Hand Pose Estimation from an Event Streamabstract3D hand pose estimation from monocular videos is a long-standing and challenging problem, which is now seeing a strong upturn. In this work, we address it for the first time using a single event camera, i.e., an asynchronous vision sensor reacting on brightness changes. Our EventHands approach has characteristics previously not demonstrated with a single RGB or depth camera such as high temporal resolution at low data throughputs and real-time performance at 1000 Hz. Due to the different data modality of event cameras compared to classical cameras, existing methods cannot be directly applied to and re-trained for event streams. We thus design a new neural approach which accepts a new event stream representation suitable for learning, which is trained on newly-generated synthetic event streams and can generalise to real data. Experiments show that EventHands outperforms recent monocular methods using a colour (or depth) camera in terms of accuracy and its ability to capture hand motions of unprecedented speed. Our method, the event stream simulator and the dataset are publicly available (see https://4dqv.mpi-inf.mpg.de/EventHands/). Viktor Rudnev, Vladislav Golyanik, Jiayi Wang 0001, Hans-Peter Seidel, Franziska Mueller 0001, Mohamed A. Elgharib, Christian Theobalt |
ICCV | 2 |
| 2021 | Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular VideoabstractWe present Non-Rigid Neural Radiance Fields (NR-NeRF), a reconstruction and novel view synthesis approach for general non-rigid dynamic scenes. Our approach takes RGB images of a dynamic scene as input (e.g., from a monocular video recording), and creates a high-quality space-time geometry and appearance representation. We show that a single handheld consumer-grade camera is sufficient to synthesize sophisticated renderings of a dynamic scene from novel virtual camera views, e.g. a ‘bullet-time’ video effect. NR-NeRF disentangles the dynamic scene into a canonical volume and its deformation. Scene deformation is implemented as ray bending, where straight rays are deformed non-rigidly. We also propose a novel rigidity network to better constrain rigid regions of the scene, leading to more stable results. The ray bending and rigidity network are trained without explicit supervision. Our formulation enables dense correspondence estimation across views and time, and compelling video editing applications such as motion exaggeration. Our code will be open sourced. Edith Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer, Christoph Lassner, Christian Theobalt |
ICCV | 3 |
| 2021 | Neural monocular 3D human motion capture with physical awarenessabstractWe present a new trainable system for physically plausible markerless 3D human motion capture, which achieves state-of-the-art results in a broad range of challenging scenarios. Unlike most neural methods for human motion capture, our approach, which we dub "physionical", is aware of physical and environmental constraints. It combines in a fully-differentiable way several key innovations, i.e. , 1) a proportional-derivative controller, with gains predicted by a neural network, that reduces delays even in the presence of fast motions, 2) an explicit rigid body dynamics model and 3) a novel optimisation layer that prevents physically implausible foot-floor penetration as a hard constraint. The inputs to our system are 2D joint keypoints, which are canonicalised in a novel way so as to reduce the dependency on intrinsic camera parameters---both at train and test time. This enables more accurate global translation estimation without generalisability loss. Our model can be finetuned only with 2D annotations when the 3D annotations are not available. It produces smooth and physically-principled 3D motions in an interactive frame rate in a wide variety of challenging scenes, including newly recorded ones. Its advantages are especially noticeable on in-the-wild sequences that significantly differ from common 3D pose estimation benchmarks such as Human 3.6M and MPI-INF-3DHP. Qualitative results are provided in the supplementary video. Soshi Shimada, Vladislav Golyanik, Weipeng Xu, Patrick Pérez, Christian Theobalt |
ACM Trans. Graph. | 2 |
| 2020 | Adiabatic Quantum Graph Matching with Permutation Matrix ConstraintsabstractMatching problems on 3D shapes and images are challenging as they are frequently formulated as combinatorial quadratic assignment problems (QAPs) with permutation matrix constraints, which are NP-hard. In this work, we address such problems with emerging quantum computing technology and propose several reformulations of QAPs as unconstrained problems suitable for efficient execution on quantum hardware. We investigate several ways to inject permutation matrix constraints in a quadratic unconstrained binary optimization problem which can be mapped to quantum hardware. We focus on obtaining a sufficient spectral gap, which further increases the probability to measure optimal solutions and valid permutation matrices in a single run. We perform our experiments on the quantum computer D-Wave 2000Q(211qubits, adiabatic). Despite the observed discrepancy between simulated adiabatic quantum computing and execution on real quantum hardware, our reformulation of permutation matrix constraints increases the robustness of the numerical computations over other penalty approaches in our experiments. The proposed algorithm has the potential to scale to higher dimensions on future quantum computing architectures, which opens up multiple new directions for solving matching problems in 3D computer vision and graphics. Marcel Seelbach Benkner, Vladislav Golyanik, Christian Theobalt, Michael Möller 0001 |
3DV | 2 |
| 2020 | Intrinsic Dynamic Shape Prior for Dense Non-Rigid Structure from MotionabstractWhile dense non-rigid structure from motion (NRSfM) has been extensively studied from the perspective of the reconstructability problem over the recent years, almost no attempts have been undertaken to bring it into the practical realm. The reasons for the slow dissemination are the severe ill-posedness, high sensitivity to motion and deformation cues, and the difficulty to obtain reliable point tracks in the vast majority of practical scenarios.To fill this gap, we propose a new framework that first extracts prior knowledge from an input image sequence with NRSfM. Our Dynamic Shape Prior Reconstruction (DSPR) approach then uses the obtained 3D reconstructions as a dynamic shape prior for sequential surface recovery in scenarios with recurrence. DSPR can be combined with existing dense NRSfM techniques while its energy functional is optimised with multi-start gradient descent at real-time frame rates for new incoming point tracks. The proposed versatile framework with a new core NRSfM approach outperforms several other methods in the ability to handle inaccurate and noisy point tracks, provided we have access to a representative (in terms of the deformation variety) image sequence. Comprehensive experiments highlight convergence properties and the accuracy of DSPR under different disturbing effects. We also perform a joint study of tracking and reconstruction and show applications to shape compression and heart reconstruction under occlusions. We achieve state-of-the-art metrics (accuracy and compression ratios) in different scenarios. Vladislav Golyanik, André Jonas, Didier Stricker, Christian Theobalt |
3DV | 1 |
| 2020 | Fast Simultaneous Gravitational Alignment of Multiple Point SetsabstractThe problem of simultaneous rigid alignment of multiple unordered point sets which is unbiased towards any of the inputs has recently attracted increasing interest, and several reliable methods have been newly proposed. While being remarkably robust towards noise and clustered outliers, current approaches require sophisticated initialisation schemes and do not scale well to large point sets. This paper proposes a new resilient technique for simultaneous registration of multiple point sets by interpreting the latter as particle swarms rigidly moving in the mutually induced force fields. Thanks to the improved simulation with altered physical laws and acceleration of globally multiply-linked point interactions with a 2D-tree (D is the space dimensionality), our Multi-Body Gravitational Approach (MBGA) is robust to noise and missing data while supporting more massive point sets than previous methods (with 105 points and more). In various experimental settings, MBGA is shown to outperform several baseline point set alignment approaches in terms of accuracy and runtime. We make our source code available for the community to facilitate the reproducibility of the results1.1http://gvv.mpi-inf.mpg.de/projects/MBGA/. Vladislav Golyanik, Soshi Shimada, Christian Theobalt |
3DV | 1 |
| 2020 | A Quantum Computational Approach to Correspondence Problems on Point SetsabstractModern adiabatic quantum computers (AQC) are already used to solve difficult combinatorial optimisation problems in various domains of science. Currently, only a few applications of AQC in computer vision have been demonstrated. We review AQC and derive a new algorithm for correspondence problems on point sets suitable for execution on AQC. Our algorithm has a subquadratic computational complexity of the state preparation. Examples of successful transformation estimation and point set alignment by simulated sampling are shown in the systematic experimental evaluation. Finally, we analyse the differences in the solutions and the corresponding energy values. Vladislav Golyanik, Christian Theobalt |
CVPR | 1 |
| 2020 | HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth Mapabstract3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. The state-of-the-art methods directly regress 3D hand meshes from 2D depth images via 2D convolutional neural networks, which leads to artefacts in the estimations due to perspective distortions in the images. In contrast, we propose a novel architecture with 3D convolutions trained in a weakly-supervised manner. The input to our method is a 3D voxelized depth map, and we rely on two hand shape representations. The first one is the 3D voxelized grid of the shape which is accurate but does not preserve the mesh topology and the number of mesh vertices. The second representation is the 3D hand surface which is less accurate but does not suffer from the limitations of the first representation. We combine the advantages of these two representations by registering the hand surface to the voxelized hand shape. In the extensive experiments, the proposed approach improves over the state of the art by47.8% on the SynHand5M dataset. Moreover, our augmentation policy for voxelized depth maps further enhances the accuracy of 3D hand pose estimation on real data. Our method produces visually more reasonable and realistic hand shapes on NYU and BigHand2.2M datasets compared to the existing approaches. Jameel Malik, Ibrahim Abdelaziz, Ahmed Elhayek, Soshi Shimada, Sk Aziz Ali, Vladislav Golyanik, Christian Theobalt, Didier Stricker |
CVPR | 6 |
| 2020 | EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event CameraabstractThe high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap - the first approach for 3D capturing of high-speed human motions using a single event camera. Our method combines model-based optimization and CNN-based human pose detection to capture high frequency motion details and to reduce the drifting in the tracking. As a result, we can capture fast motions at millisecond resolution with significantly higher data efficiency than using high frame rate videos. Experiments on our new event-based fast human motion dataset demonstrate the effectiveness and accuracy of our method, as well as its robustness to challenging lighting conditions. Lan Xu 0003, Weipeng Xu, Vladislav Golyanik, Marc Habermann, Lu Fang 0001, Christian Theobalt |
CVPR | 3 |
| 2020 | HTML: A Parametric Hand Texture Model for 3D Hand Reconstruction and Personalization
Neng Qian, Jiayi Wang 0001, Franziska Mueller 0001, Florian Bernard 0001, Vladislav Golyanik, Christian Theobalt |
ECCV (11) | 5 |
| 2020 | Neural Re-rendering of Humans from a Single Image
Kripasindhu Sarkar, Dushyant Mehta, Weipeng Xu, Vladislav Golyanik, Christian Theobalt |
ECCV (11) | 4 |
| 2020 | Neural Dense Non-Rigid Structure from Motion with Latent Space Constraints
Vikramjit Sidhu, Edith Tretschk, Vladislav Golyanik, Antonio Agudo, Christian Theobalt |
ECCV (16) | 3 |
| 2020 | PatchNets: Patch-Based Generalizable Deep Implicit 3D Shape Representations
Edith Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer, Carsten Stoll, Christian Theobalt |
ECCV (16) | 3 |
| 2020 | DEMEA: Deep Mesh Autoencoders for Non-rigidly Deforming Objects
Edith Tretschk, Ayush Tewari, Michael Zollhöfer, Vladislav Golyanik, Christian Theobalt |
ECCV (4) | 4 |
| 2020 | Egocentric videoconferencingabstractWe introduce a method for egocentric videoconferencing that enables hands-free video calls, for instance by people wearing smart glasses or other mixed-reality devices. Videoconferencing portrays valuable non-verbal communication and face expression cues, but usually requires a front-facing camera. Using a frontal camera in a hands-free setting when a person is on the move is impractical. Even holding a mobile phone camera in the front of the face while sitting for a long duration is not convenient. To overcome these issues, we propose a low-cost wearable egocentric camera setup that can be integrated into smart glasses. Our goal is to mimic a classical video call, and therefore, we transform the egocentric perspective of this camera into a front facing video. To this end, we employ a conditional generative adversarial neural network that learns a transition from the highly distorted egocentric views to frontal views common in videoconferencing. Our approach learns to transfer expression details directly from the egocentric view without using a complex intermediate parametric expressions model, as it is used by related face reenactment methods. We successfully handle subtle expressions, not easily captured by parametric blendshape-based solutions, e.g., tongue movement, eye movements, eye blinking, strong expressions and depth varying movements. To get control over the rigid head movements in the target view, we condition the generator on synthetic renderings of a moving neutral face. This allows us to synthesis results at different head poses. Our technique produces temporally smooth video-realistic renderings in real-time using a video-to-video translation network in conjunction with a temporal discriminator. We demonstrate the improved capabilities of our technique by comparing against related state-of-the art approaches. Mohamed A. Elgharib, Mohit Mendiratta, Justus Thies, Matthias Nießner, Hans-Peter Seidel, Ayush Tewari, Vladislav Golyanik, Christian Theobalt |
ACM Trans. Graph. | 7 |
| 2020 | PhysCap: physically plausible monocular 3D motion capture in real timeabstractMarker-less 3D human motion capture from a single colour camera has seen significant progress. However, it is a very challenging and severely ill-posed problem. In consequence, even the most accurate state-of-the-art approaches have significant limitations. Purely kinematic formulations on the basis of individual joints or skeletons, and the frequent frame-wise reconstruction in state-of-the-art methods greatly limit 3D accuracy and temporal stability compared to multi-view or marker-based motion capture. Further, captured 3D poses are often physically incorrect and biomechanically implausible, or exhibit implausible environment interactions (floor penetration, foot skating, unnatural body leaning and strong shifting in depth), which is problematic for any use case in computer graphics. We, therefore, present PhysCap , the first algorithm for physically plausible, real-time and marker-less human 3D motion capture with a single colour camera at 25 fps. Our algorithm first captures 3D human poses purely kinematically. To this end, a CNN infers 2D and 3D joint positions, and subsequently, an inverse kinematics step finds space-time coherent joint angles and global 3D pose. Next, these kinematic reconstructions are used as constraints in a real-time physics-based pose optimiser that accounts for environment constraints ( e.g. , collision handling and floor placement), gravity, and biophysical plausibility of human postures. Our approach employs a combination of ground reaction force and residual force for plausible root control, and uses a trained neural network to detect foot contact events in images. Our method captures physically plausible and temporally stable global 3D human motion, without physically implausible postures, floor penetrations or foot skating, from video in real time and in general scenes. PhysCap achieves state-of-the-art accuracy on established pose benchmarks, and we propose new metrics to demonstrate the improved physical plausibility and temporal stability. Soshi Shimada, Vladislav Golyanik, Weipeng Xu, Christian Theobalt |
ACM Trans. Graph. | 2 |
| 2019 | Optimising for Scale in Globally Multiply-Linked Gravitational Point Set Registration Leads to SingularitiesabstractThe recently introduced gravitational point set registration methods can robustly align point sets with high noise ratios. While many results with rotation and translation resolution have been presented in the literature so far, the uncertainty about the scale resolving capability of globally multiply-linked gravitational point set registration methods is remaining. To address the uncertainty, we analyse the gravitational potential energy (GPE) functional of the system with two point sets with multivariate calculus and come to the conclusion that GPE of a singularity, i.e., the state when the template collapses to a single point, always has a lower energy than the state of the optimal alignment. Moreover, the GPE as a function of scale monotonically increases and does not have equienergetic states, which we prove for several hollow and volumetric geometric primitives in R 2 and R 3 including a unit circle, a unit sphere, a unit disk and a unit ball. We perform a series of experiments with various regular shapes and different combinations of references and templates to validate our findings. The consequence is that in practice, it is highly unlikely for globally multiply-linked gravitational alignment to resolve the scale accurately. We propose ways to overcome the limitation in scale resolution, which can be considered in the next generation of gravitational approaches. Vladislav Golyanik, Christian Theobalt |
3DV | 1 |
| 2019 | DispVoxNets: Non-Rigid Point Set Alignment with Supervised Learning ProxiesabstractWe introduce a supervised-learning framework for nonrigid point set alignment of a new kind - Displacements on Voxels Networks (DispVoxNets) - which abstracts away from the point set representation and regresses 3D displacement fields on regularly sampled proxy 3D voxel grids. Thanks to recently released collections of deformable objects with known intra-state correspondences, DispVoxNets learn a deformation model and further priors (e.g., weak point topology preservation) for different object categories such as cloths, human bodies and faces. DispVoxNets cope with large deformations, noise and clustered outliers more robustly than the state-of-the-art. At test time, our approach runs orders of magnitude faster than previous techniques. All properties of DispVoxNets are ascertained numerically and qualitatively in extensive experiments and comparisons to several previous methods. Soshi Shimada, Vladislav Golyanik, Edith Tretschk, Didier Stricker, Christian Theobalt |
3DV | 2 |
| 2019 | Accelerated Gravitational Point Set Alignment With Altered Physical LawsabstractThis work describes Barnes-Hut Rigid Gravitational Approach (BH-RGA) - a new rigid point set registration method relying on principles of particle dynamics. Interpreting the inputs as two interacting particle swarms, we directly minimise the gravitational potential energy of the system using non-linear least squares. Compared to solutions obtained by solving systems of second-order ordinary differential equations, our approach is more robust and less dependent on the parameter choice. We accelerate otherwise exhaustive particle interactions with a Barnes-Hut tree and efficiently handle massive point sets in quasilinear time while preserving the globally multiply-linked character of interactions. Among the advantages of BH-RGA is the possibility to define boundary conditions or additional alignment cues through varying point masses. Systematic experiments demonstrate that BH-RGA surpasses performances of baseline methods in terms of the convergence basin and accuracy when handling incomplete, noisy and perturbed data. The proposed approach also positively compares to the competing method for the alignment with prior matches. Vladislav Golyanik, Christian Theobalt, Didier Stricker |
ICCV | 1 |
| 2019 | Face It!: A Pipeline for Real-Time Performance-Driven Facial AnimationabstractThis paper presents a new lightweight approach for real-time performance-driven facial animation from monocular videos. We transfer facial expressions from 2D images to a 3D virtual character, by estimating the rigid head pose and non-rigid face deformation from detected and tracked 2D facial landmarks. We map the input face into the facial expression space of the 3D head model using blendshape models and formulate a lightweight energy-based optimization problem, which is solved by non-linear least squares at 18 FPS on a single CPU. Our method robustly handles varying head poses and different facial expressions, including moderately asymmetric ones. Compared to related methods, our approach does not require training data, specialised camera setups or graphics cards, and is suitable for embedded systems. We support our claims with several experiments. Jilliam María Díaz Barros, Vladislav Golyanik, Kiran Varanasi, Didier Stricker |
ICIP | 2 |
| 2018 | NRGA: Gravitational Approach for Non-rigid Point Set RegistrationabstractRecovery of correspondences between point sets which differ by some non-rigid transformation is an ill-posed problem. Many existing methods underperform on noisy or corrupted input data. In this study, a novel physics-based approach -- Non-Rigid Gravitational Approach (NRGA) -- for non-rigid point set registration is introduced which is robust to the mentioned artifacts. Thereafter, a distributed N-body simulation and iterative Procrustes alignment non-rigidly transform and register the template point set. Furthermore, in the force field evolution, per-point Gaussian curvature serves as a shape matching descriptor whereas the displacement fields are regularized by coherent collective motion. The optimal alignment is referred to as the state of minimum gravitational potential energy between the point sets. A thorough experimental evaluation and comparison are provided with widely used state-of-the-art methods on 2D and 3D data sets. Experiments show NRGA's robustness against uniform outliers and missing data. Sk Aziz Ali, Vladislav Golyanik, Didier Stricker |
3DV | 2 |
| 2018 | Improving Time-of-Flight Sensor for Specular Surfaces with Shape from PolarizationabstractTime-of-Flight (ToF) sensors can obtain depth values for diffuse objects. However, the essential problem is that the sensor can not receive active light from specular surfaces due to specular reflections. In this paper, we propose a new depth reconstruction framework for specular objects that combines ToF cues and Shape from Polarization (SfP). To overcome the ill-posedness of SfP with a single view, we integrate superpixel segmentation with planarity constraints for every superpixel. Experimental results demonstrate the effectiveness of the depth reconstruction algorithm for both controlled environment data and real vehicle data in a parking area. Tomonari Yoshida, Vladislav Golyanik, Oliver Wasenmüller, Didier Stricker |
ICIP | 2 |
| 2017 | Scalable Dense Monocular Surface ReconstructionabstractThis paper reports on a novel template-free monocular non-rigid surface reconstruction approach. Existing techniques using motion and deformation cues rely on multiple prior assumptions, are often computationally expensive and do not perform equally well across the variety of data sets. In contrast, the proposed Scalable Monocular Surface Reconstruction (SMSR) combines strengths of several algorithms, i.e., it is scalable with the number of points, can handle sparse and dense settings as well as different types of motions and deformations. We estimate camera pose by singular value thresholding and proximal gradient. Our formulation adopts alternating direction method of multipliers which converges in linear time for large point track matrices. In the proposed SMSR, trajectory space constraints are integrated by smoothing of the measurement matrix. In the extensive experiments, SMSR is demonstrated to consistently achieve state-of-the-art accuracy on a wide variety of data sets. Mohammad Dawud Ansari, Vladislav Golyanik, Didier Stricker |
3DV | 2 |
| 2017 | Multiframe Scene Flow with Piecewise Rigid MotionabstractWe introduce a novel multiframe scene flow approach that jointly optimizes the consistency of the patch appearances and their local rigid motions from RGB-D image sequences. In contrast to the competing methods, we take advantage of an oversegmentation of the reference frame and robust optimization techniques. We formulate scene flow recovery as a global non-linear least squares problem which is iteratively solved by a damped Gauss-Newton approach. As a result, we obtain a qualitatively new level of accuracy in RGB-D based scene flow estimation which can potentially run in real-time. Our method can handle challenging cases with rigid, piecewise rigid, articulated and moderate non-rigid motion, and does not rely on prior knowledge about the types of motions and deformations. Extensive experiments on synthetic and real data show that our method outperforms state-of-the-art. Vladislav Golyanik, Robert Maier 0001, Matthias Nießner, Didier Stricker, Jan Kautz |
3DV | 1 |
| 2017 | High Dimensional Space Model for Dense Monocular Surface RecoveryabstractDense surface reconstruction from monocular image sequences - known as Non-Rigid Structure from Motion (NRSfM) - is a highly ill-posed inverse problem. The objective of NRSfM is to learn 3D shapes from 2D point tracks in an unsupervised manner. While existing methods rely on low-rank models, we propose the concept of High Dimensional Space Model (HDSM). In HDSM, time-varying geometry is encoded by a high-dimensional static structure projected into different metric subspaces. To express non-rigid deformations, instead of directly modelling in the 3D space, we gradually increase space dimensionality as the complexity of the scene increases. HDSM allows for a compact representation with deformation localisation and can be interpreted as a generalisation of the previously proposed models for NRSfM. Relying on HDSM, we develop an algorithm for dense monocular surface recovery. Experiments show that the proposed method achieves high accuracy while allowing for the fine-grained control. Vladislav Golyanik, Didier Stricker |
3DV | 1 |
| 2017 | Introduction to Coherent Depth Fields for Dense Monocular Surface Recovery
Vladislav Golyanik, Torben Fetzer, Didier Stricker |
BMVC | 1 |
| 2017 | Towards scheduling hard real-time image processing tasks on a single GPUabstractGraphics Processing Units (GPU) are becoming the key hardware accelerators in the emerging image processing applications such as self-driving cars and mobile augmented reality systems. As GPUs execute launched workloads non-preemptively, their usage in safety-critical systems with hard real-time constraints is impeded. The existing solutions for scheduling real-time tasks on a single GPU focus on soft real-time systems. In this paper, we consider real-time systems with a single dedicated GPU handling sporadic tasks with hard deadlines and propose a scheduling approach based on time division multiplexing called the GPU-TDMh - a lightweight middleware framework located between the application and the GPU driver layers. We evaluate the proposed approach on a matrix multiplication benchmark on a heterogeneous platform. The experiments demonstrate the effectiveness of our method as well as superiority over the non-preemptive online scheduling policies. Vladislav Golyanik, Mitra Nasri, Didier Stricker |
ICIP | 1 |
| 2017 | Accurate 3D Reconstruction of Dynamic Scenes from Monocular Image Sequences with Severe OcclusionsabstractThe paper introduces an accurate solution to dense orthographic Non-Rigid Structure from Motion (NRSfM) in scenarios with severe occlusions or, likewise, inaccurate correspondences. We integrate a shape prior term into variational optimisation framework. It allows to penalize irregularities of the time-varying structure on the per-pixel level if correspondence quality indicator such as an occlusion tensor is available. We make a realistic assumption that several non-occluded views of the scene are sufficient to estimate an initial shape prior, though the entire observed scene may exhibit non-rigid deformations. Experiments on synthetic and real image data show that the proposed framework significantly outperforms state of the art methods for correspondence establishment in combination with the state of the art NRSfM methods. Together with the profound insights into optimisation methods, implementation details for heterogeneous platforms are provided. Vladislav Golyanik, Torben Fetzer, Didier Stricker |
WACV | 1 |
| 2017 | Dense Batch Non-Rigid Structure from Motion in a SecondabstractIn this paper, we show how to minimise a quadratic function on a set of orthonormal matrices using an efficient semidefinite programming solver with application to dense non-rigid structure from motion. Thanks to the proposed technique, a new form of the convex relaxation for the Metric Projections (MP) algorithm is obtained. The modification results in an efficient single-core CPU implementation enabling dense factorisations of long image sequences with tens of thousands of points into camera pose and non-rigid shape in seconds, i.e., at least two orders of magnitude faster than the runtimes reported in the literature so far. The proposed implementation can be useful for interactive or real-time robotic and other applications, where monocular non-rigid reconstruction is required. In a narrow sense, our paper complements research on MP, though the proposed convex relaxation methodology can also be useful in other computer vision tasks. The experimental part providing runtime evaluation and qualitative analysis concludes the paper. Vladislav Golyanik, Didier Stricker |
WACV | 1 |
| 2016 | NRSfM-Flow: Recovering Non-Rigid Scene Flow from Monocular Image Sequences
Vladislav Golyanik, Aman S. Mathur, Didier Stricker |
BMVC | 1 |
| 2016 | Gravitational Approach for Point Set RegistrationabstractIn this paper a new astrodynamics inspired rigid point set registration algorithm is introduced-the Gravitational Approach (GA). We formulate point set registration as a modified N-body problem with additional constraints and obtain an algorithm with unique properties which is fully scalable with the number of processing cores. In GA, a template point set moves in a viscous medium under gravitational forces induced by a reference point set. Pose updates are completed by numerically solving the differential equations of Newtonian mechanics. We discuss techniques for efficient implementation of the new algorithm and evaluate it on several synthetic and real-world scenarios. GA is compared with the widely used Iterative Closest Point and the state of the art rigid Coherent Point Drift algorithms. Experiments evidence that the new approach is robust against noise and can handle challenging scenarios with structured outliers. Vladislav Golyanik, Sk Aziz Ali, Didier Stricker |
CVPR | 1 |
| 2016 | Joint pre-alignment and robust rigid point set registrationabstractWe present an elegant solution to joint pre-alignment and rigid point set registration, given prior matches. Instead of performing pre-alignment and the actual registration in the separate steps, prior matches explicitly influence the registration procedure in our approach. This results in several advantages. Firstly, our approach solves the pre-alignment task - an approximate resolving of rotation and translation - with an insufficient number of prior correspondences, when other methods fail. Secondly, it produces more accurate rigid registrations of noisy point sets than the state of the art Coherent Point Drift method. Combined with application specific methods for correspondence establishment, we demonstrate superiority of our approach in several synthetic and real-world scenarios. Vladislav Golyanik, Bertram Taetz, Didier Stricker |
ICIP | 1 |
| 2016 | Extended coherent point drift algorithm with correspondence priors and optimal subsamplingabstractThe problem of dense point set registration, given a sparse set of prior correspondences, often arises in computer vision tasks. Unlike in the rigid case, integrating prior knowledge into a registration algorithm is especially demanding in the non-rigid case due to the high variability of motion and deformation. In this paper we present the Extended Coherent Point Drift registration algorithm. It enables, on the one hand, to couple correspondence priors into the dense registration procedure in a closed form and, on the other hand, to process large point sets in reasonable time through adopting an optimal coarse-to-fine strategy. Combined with a suitable keypoint extractor during the preprocessing step, our method allows for non-rigid registrations with increased accuracy for point sets with structured outliers. We demonstrate advantages of our approach against other non-rigid point set registration methods in synthetic and real-world scenarios. Vladislav Golyanik, Bertram Taetz, Gerd Reis, Didier Stricker |
WACV | 1 |
| 2016 | Occlusion-aware video registration for highly non-rigid objectsabstractThis paper addresses the problem of video registration for dense non-rigid structure from motion under suboptimal conditions, such as noise, self-occlusions, considerable external occlusions or specularities, i.e. the computation of optical flow between the reference image and each of the subsequent images in a video sequence when the camera observes a highly deformable object. We tackle this challenging task by improving previously proposed variational optimization techniques for multi-frame optical flow (MFOF) through detection, tracking and handling of uncertain flow field estimates. This is based on a novel Bayesian inference approach incorporated into the MFOF. At the same time, computational costs are significantly reduced through iterative pre-computation of the flow fields. As shown through experiments, the resulting method performs superior to other state-of-the-art (MF)OF methods on video sequences showing a highly non-rigidly deforming object with considerable occlusions. Bertram Taetz, Gabriele Bleser-Taetz, Vladislav Golyanik, Didier Stricker |
WACV | 3 |