EDBT 2026 Demo / reviewers in the wild / expert
Peter Eisert
dblp:27/1029
· DBLP profile ↗
104ranked-venue papers
17as first author
26since 2021 · last 2025
0000-0001-8378-4805ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 17 first-author · 21 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint ExtractionabstractWe introduce RIPE, an innovative reinforcement learning-based framework for weakly-supervised training of a keypoint extractor that excels in both detection and description tasks. In contrast to conventional training regimes that depend heavily on artificial transformations, pre-generated models, or 3D data, RIPE requires only a binary label indicating whether paired images represent the same scene. This minimal supervision significantly expands the pool of training data, enabling the creation of a highly generalized and robust keypoint extractor. RIPE utilizes the encoder's intermediate layers for the description of the keypoints with a hyper-column approach to integrate information from different scales. Additionally, we propose an auxiliary loss to enhance the discriminative capability of the learned descriptors. Comprehensive evaluations on standard benchmarks demonstrate that RIPE simplifies data preparation while achieving competitive performance compared to state-of-the-art techniques, marking a significant advancement in robust keypoint extraction and description. To support further research, we have made our code publicly available at https://github.com/fraunhoferhhi/RIPE. Johannes Künzel, Anna Hilsmann, Peter Eisert |
ICCV | 3 |
| 2025 | Batch-Aware Active Learning for Object DetectionabstractWe propose a Batch-Aware Active Learning (BAAL) framework to optimize the training of object detection models, reducing annotation costs while maintaining strong model performance. The framework adapts different uncertainty sampling strategies to the specific challenges of object detection, including multi-class labelling and spatial localization. By combining uncertainty with diversity, leveraging feature representations and clustering, our method ensures diverse and informative batch selection. The non-invasive, plug-and-play design supports seamless integration with any object detection model without architectural modifications. Evaluations on COCO and Pascal VOC datasets with SSD, Faster R-CNN, YOLOv8, and RetinaNet demonstrate that our approach is not only efficient and robust but also comparable to, and in some cases exceeds, current state-of-the-art solutions. Mykyta Kovalenko, Peter Eisert, Anna Hilsmann, Sebastian Bosse |
ICIP | 2 |
| 2025 | CGS-GAN: 3D Consistent Gaussian Splatting GANs for High Resolution Human Head SynthesisabstractRecently, 3D GANs based on 3D Gaussian splatting have been proposed for high quality synthesis of human heads. However, existing methods stabilize training and enhance rendering quality from steep viewpoints by conditioning the random latent vector on the current camera position. This compromises 3D consistency, as we observe significant identity changes when re-synthesizing the 3D head with each camera shift. Conversely, fixing the camera to a single viewpoint yields high-quality renderings for that perspective but results in poor performance for novel views. Removing view-conditioning typically destabilizes GAN training, often causing the training to collapse. In response to these challenges, we introduce CGS-GAN, a novel 3D Gaussian Splatting GAN framework that enables stable training and high-quality 3D-consistent synthesis of human heads without relying on view-conditioning. To ensure training stability, we introduce a multi-view regularization technique that enhances generator convergence with minimal computational overhead. Additionally, we adapt the conditional loss used in existing 3D Gaussian splatting GANs and propose a generator architecture designed to not only stabilize training but also facilitate efficient rendering and straightforward scaling, enabling output resolutions up to $2048^2$. To evaluate the capabilities of CGS-GAN, we curate a new dataset derived from FFHQ. This dataset enables very high resolutions, focuses on larger portions of the human head, reduces view-dependent artifacts for improved 3D consistency, and excludes images where subjects are obscured by hands or other objects. As a result, our approach achieves very high rendering quality, supported by competitive FID scores, while ensuring consistent 3D scene generation. Florian Barthel, Wieland Morgenstern, Paul Hinzer, Anna Hilsmann, Peter Eisert |
NeurIPS | 5 |
| 2025 | AI-based Denoising and Interpolation of Magnetic UXO DataabstractMagnetic surveys are a key tool in detecting buried objects such as unexploded ordnance (UXO), where dense magnetic maps must be reconstructed from sparsely sampled gradiometer data. We present a deep learning-based approach that outperforms classical interpolation methods in both accuracy and speed. Trained on synthetic magnetic fields simulating realistic UXO signatures and measurement noise, our modified U-Net with ResNet-34 encoding reconstructs high-resolution magnetic maps from sparse inputs. Compared to state-of-the-art gridding methods, our model achieves 3–5% higher reconstruction accuracy on average while operating up to 80× faster than SOTA algorithms, enabling more efficient and interpretable UXO detection in real-world survey conditions. Mykyta Kovalenko, David Przewozny, Paul Chojecki, Anna Hilsmann, Peter Eisert, Sebastian Bosse |
SMC | 5 |
| 2025 | Adaptive and Temporally Consistent Gaussian Surfels for Multi-View Dynamic Reconstructionabstract3D Gaussian Splatting has recently achieved notable success in novel view synthesis for dynamic scenes and ge-ometry reconstruction in static scenes. Building on these advancements, early methods have been developed for dy-namic surface reconstruction by globally optimizing entire sequences. However, reconstructing dynamic scenes with significant topology changes, emerging or disappearing ob-jects, and rapid movements remains a substantial chal-lenge, particularly for long sequences. To address these issues, we propose AT-GS, a novel method for reconstructing high-quality dynamic surfaces from multi-view videos through per-frame incremental optimization. To avoid local minima across frames, we introduce a unified and adaptive gradient-aware densification strategy that integrates the strengths of conventional cloning and splitting techniques. Additionally, we reduce temporal jittering in dy-namic surfaces by ensuring consistency in curvature maps across consecutive frames. Our method achieves superior accuracy and temporal coherence in dynamic surface re-construction, delivering high-fidelity space-time novel view synthesis, even in complex and challenging scenes. Extensive experiments on diverse multi-view video datasets demonstrate the effectiveness of our approach, showing clear advantages over baseline methods. Project page: https://fraunhoferhhi.github.io/AT-GS Decai Chen, Brianne Oberson, Ingo Feldmann, Oliver Schreer, Anna Hilsmann, Peter Eisert |
WACV | 6 |
| 2025 | 3DGS.zip: A survey on 3D Gaussian Splatting Compression MethodsabstractAbstract 3D Gaussian Splatting (3DGS) has emerged as a cutting‐edge technique for real‐time radiance field rendering, offering state‐of‐the‐art performance in terms of both quality and speed. 3DGS models a scene as a collection of three‐dimensional Gaussians, with additional attributes optimized to conform to the scene's geometric and visual properties. Despite its advantages in rendering speed and image fidelity, 3DGS is limited by its significant storage and memory demands. These high demands make 3DGS impractical for mobile devices or headsets, reducing its applicability in important areas of computer graphics. To address these challenges and advance the practicality of 3DGS, this state‐of‐the‐art report (STAR) provides a comprehensive and detailed examination of two complementary yet fundamentally distinct strategies: compression and compaction. Compression techniques focus on reducing the file size by encoding Gaussian attributes more efficiently. In contrast, compaction methods directly optimize the scene's structure by optimizing the number of Gaussian primitives. Notably, while methods in both categories aim to maintain or improve quality, each while minimizing its respective attributes—file size for compression and the number of Gaussians for compaction—compaction does not necessarily lead to smaller file sizes; it specifically targets improved efficiency during rendering, making it distinct from compression. We introduce the basic mathematical concepts underlying the analyzed methods, as well as key implementation details and design choices. Our report thoroughly discusses similarities and differences among the methods, as well as their respective advantages and disadvantages. We establish a consistent framework for comparing the surveyed methods based on key performance metrics and datasets. Specifically, since these methods have been developed in parallel and over a short period of time, currently, no comprehensive comparison exists. This survey, for the first time, presents a unified framework to evaluate 3DGS compression techniques. To facilitate the continuous monitoring of emerging methodologies, we maintain a dedicated website that will be regularly updated with new techniques and revisions of existing findings. Overall, this STAR provides an intuitive starting point for researchers interested in exploring the rapidly growing field of 3DGS compression. By comprehensively categorizing and evaluating existing compression and compaction strategies, our work advances the understanding and practical application of 3DGS in computationally constrained environments. Milena T. Bagdasarian, Paul Knoll, Yi-Hsin Li, Florian Barthel, Anna Hilsmann, Peter Eisert, Wieland Morgenstern |
Comput. Graph. Forum | 6 |
| 2025 | Real-time fusion of stereo vision and hyperspectral imaging for objective decision support during surgeryabstractWe present a real-time stereo hyperspectral imaging (stereo-HSI) system for intraoperative tissue and organ analysis that integrates multispectral snapshot imaging with stereo vision to support clinical decision-making. The system visualize both RGB and high-dimensional spectral data while simultaneously reconstructing 3D surfaces, offering a compact, non-contact solution for seamless integration into surgical workflows. A modular processing pipeline enables robust demosaicing, spectral and spatial fusion, and pixel-wise medical assessment, including perfusion and tissue classification. Our spectral warping algorithm leverages a custom learned mapping, our white-balance network method is the first for snapshot MSI cameras, and our fusion CNN employs spectral-attention modules to exploit the rich hyperspectral domain. Clinical feasibility was demonstrated in 57 surgical procedures, including kidney transplantation, parotidectomy, and neck dissection, achieving high spatial and spectral resolution under standard surgical lighting conditions. The system enables visualization of oxygenation and tissue composition in real-time, offering surgeons a novel tool for image-guided interventions. This study establishes the stereo-HSI platform as a clinically viable and effective method for enhancing intraoperative insight and surgical precision. Eric L. Wisotzky, Jost Triller, Michael Knoke, Brigitta Globke, Anna Hilsmann, Peter Eisert |
Comput. Vis. Image Underst. | 6 |
| 2025 | EdgeRegNet: Edge Feature-Based Multimodal Registration Network Between Images and LiDAR Point CloudsabstractCross-modal data registration has long been a critical task in computer vision, with extensive applications in autonomous driving and robotics. Accurate and robust registration methods are essential for aligning data from different modalities, forming the foundation for multimodal sensor data fusion and enhancing perception systems' accuracy and reliability. The registration task between 2D images captured by cameras and 3D point clouds captured by Light Detection and Ranging (LiDAR) sensors is usually treated as a visual pose estimation problem. High-dimensional feature similarities from different modalities are leveraged to identify pixel-point correspondences, followed by pose estimation techniques using least squares methods. However, existing approaches often resort to downsampling the original point cloud and image data due to computational constraints, inevitably leading to a loss in precision. Additionally, high-dimensional features extracted using different feature extractors from various modalities require specific techniques to mitigate cross-modal differences for effective matching. To address these challenges, we propose a method that uses edge information from the original point clouds and images for cross-modal registration. We retain crucial information from the original data by extracting edge points and pixels, enhancing registration accuracy while maintaining computational efficiency. The use of edge points and edge pixels allows us to introduce an attention-based feature exchange block to eliminate cross-modal disparities. Furthermore, we incorporate an optimal matching layer to improve correspondence identification. We validate the accuracy of our method on the KITTI and nuScenes datasets, demonstrating its state-of-the-art performance. Our code is publicly available on GitHub athttps://github.com/ESRSchao/EdgeRegNet. Yuanchao Yue, Hui Yuan 0001, Qinglong Miao, Xiaolong Mao, Raouf Hamzaoui, Peter Eisert |
IEEE Trans. Multim. | 6 |
| 2024 | SPVLoc: Semantic Panoramic Viewport Matching for 6D Camera Localization in Unseen Environments
Niklas Gard, Anna Hilsmann, Peter Eisert |
ECCV (73) | 3 |
| 2024 | Compact 3D Scene Representation via Self-Organizing Gaussian Grids
Wieland Morgenstern, Florian Barthel, Anna Hilsmann, Peter Eisert |
ECCV (85) | 4 |
| 2024 | Multi-View Gesture Recognition in Conflict SituationsabstractNon-verbal cues play a crucial role in social interactions and can influence conflict dynamics. For law enforcement officers, recognizing these cues is essential for effective deescalation, yet traditional training may not fully address their complexity. This paper focuses on body gestures and presents an automated system for recognizing specific body gestures relevant to social conflict situations, aiming to foster awareness for unconsciously performed body gestures and thereby enabling the training of de-escalation strategies. Karam Tomotaki-Dawoud, Birgit Nierula, Farelle Toumaleu Siewe, Daniel Johannes Meyer, Andreas Bock, Marianne Heinze, Daniela Knuth, Denis Martin, Julia Schander, Anna Hilsmann, Peter Eisert, Sebastian Bosse |
ISM | 12 |
| 2024 | Creating Sorted Grid Layouts with Gradient-based Optimizationabstract1199 Kai Uwe Barthel, Florian Barthel, Peter Eisert, Nico Hezel, Konstantin Schall |
ICMR | 3 |
| 2024 | Animatable Virtual Humans: Learning Pose-Dependent Human Representations in UV Space for Interactive Performance SynthesisabstractWe propose a novel representation of virtual humans for highly realistic real-time animation and rendering in 3D applications. We learn pose dependent appearance and geometry from highly accurate dynamic mesh sequences obtained from state-of-the-art multiview-video reconstruction. Learning pose-dependent appearance and geometry from mesh sequences poses significant challenges, as it requires the network to learn the intricate shape and articulated motion of a human body. However, statistical body models like SMPL provide valuable a-priori knowledge which we leverage in order to constrain the dimension of the search space, enabling more efficient and targeted learning and to define pose-dependency. Instead of directly learning absolute pose-dependent geometry, we learn the difference between the observed geometry and the fitted SMPL model. This allows us to encode both pose-dependent appearance and geometry in the consistent UV space of the SMPL model. This approach not only ensures a high level of realism but also facilitates streamlined processing and rendering of virtual humans in real-time scenarios. Wieland Morgenstern, Milena T. Bagdasarian, Anna Hilsmann, Peter Eisert |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | But That's Not Why: Inference Adjustment by Interactive Prototype Revision
Michael Gerstenberger, Thomas Wiegand 0001, Peter Eisert, Sebastian Bosse |
CIARP | 3 |
| 2023 | Dynamic Multi-View Scene Reconstruction Using Neural Implicit SurfaceabstractReconstructing general dynamic scenes is important for many computer vision and graphics applications. Recent works represent the dynamic scene with neural radiance fields for photorealistic view synthesis, while their surface geometry is under-constrained and noisy. Other works introduce surface constraints to the implicit neural representation to disentangle the ambiguity of geometry and appearance field for static scene reconstruction. To bridge the gap between rendering dynamic scenes and recovering static surface geometry, we propose a template-free method to reconstruct surface geometry and appearance using neural implicit representations from multi-view videos. We leverage topology-aware deformation and the signed distance field to learn complex dynamic surfaces via differentiable volume rendering without scene-specific prior knowledge like template models. Furthermore, we propose a novel mask-based ray selection strategy to significantly boost the optimization on challenging time-varying regions. Experiments on different multi-view video datasets demonstrate that our method achieves high-fidelity surface reconstruction as well as photorealistic novel view synthesis. Decai Chen, Haofei Lu, Ingo Feldmann, Oliver Schreer, Peter Eisert |
ICASSP | 5 |
| 2023 | A Differentiable Gaussian Prototype Layer for Explainable Fruit SegmentationabstractWe introduce a Gaussian Prototype Layer for gradient-based prototype learning and demonstrate two novel network architectures for explainable segmentation one of which relies on region proposals. Both models are evaluated on agricultural datasets. While Gaussian Mixture Models (GMMs) have been used to model latent distributions of neural networks before, they are typically fitted using the EM algorithm. Instead, the proposed prototype layer relies on gradient-based optimization and hence allows for end-to-end training. This facilitates development and allows to use the full potential of a trainable deep feature extractor. We show that it can be used as a novel building block for explainable neural networks. We employ our Gaussian Prototype Layer in (1) a model where prototypes are detected in the latent grid and (2) a model inspired by Fast-RCNN with SLIC superpixels as region proposals. The earlier achieves a similar performance as compared to the state-of-the art while the latter has the benefit of a more precise prototype localization that comes at the cost of slightly lower accuracies. By introducing a gradient-based GMM layer we combine the benefits of end-to-end training with the simplicity and theoretical foundation of GMMs which will allow to adapt existing semi-supervised learning strategies for prototypical part models in future. Michael Gerstenberger, Steffen Maaß, Peter Eisert, Sebastian Bosse |
ICIP | 3 |
| 2023 | Pre-Training with Fractal Images Facilitates Learned Image Quality EstimationabstractToday’s image quality estimation is widely dominated by learning-based approaches. The availability of annotated, i.e. rated, images is often a bottleneck in training data-driven visual quality models and hinders their generalization power. This paper proposed a novel pre-training scheme for learning-based quality estimation that does not rely on human-annotated datasets, but leverages synthetic fractal images. These images can be synthesized inexhaustibly and are inherently labeled during generation. We evaluate the pre-training strategy on a popular neural network-based quality model and show that the training effort can be reduced significantly, resulting in better final accuracy and faster convergence speed. Malte Silbernagel, Thomas Wiegand 0001, Peter Eisert, Sebastian Bosse |
ICIP | 3 |
| 2023 | Fooling State-of-the-art Deepfake Detection with High-quality DeepfakesabstractDue to the rising threat of deepfakes to security and privacy, it is most important to develop robust and reliable detectors. In this paper, we examine the need for high-quality samples in the training datasets of such detectors. Accordingly, we show that deepfake detectors proven to generalize well on multiple research datasets still struggle in real-world scenarios with well-crafted fakes. First, we propose a novel autoencoder for face swapping alongside an advanced face blending technique, which we utilize to generate 90 high-quality deepfakes. Second, we feed those fakes to a state-of-the-art detector, causing its performance to decrease drastically. Moreover, we fine-tune the detector on our fakes and demonstrate that they contain useful clues for the detection of manipulations. Overall, our results provide insights into the generalization of deepfake detectors and suggest that their training datasets should be complemented by high-quality fakes since training on mere research data is insufficient. Arian Beckmann, Anna Hilsmann, Peter Eisert |
IH&MMSec | 3 |
| 2023 | Recovering Fine Details for Neural Implicit Surface ReconstructionabstractRecent works on implicit neural representations have made significant strides. Learning implicit neural surfaces using volume rendering has gained popularity in multi-view reconstruction without 3D supervision. However, accurately recovering fine details is still challenging, due to the underlying ambiguity of geometry and appearance representation. In this paper, we present D-NeuS, a volume rendering-base neural implicit surface reconstruction method capable to recover fine geometry details, which extends NeuS by two additional loss functions targeting enhanced reconstruction quality. First, we encourage the rendered surface points from alpha compositing to have zero signed distance values, alleviating the geometry bias arising from transforming SDF to density for volume rendering. Second, we impose multi-view feature consistency on the surface points, derived by interpolating SDF zerocrossings from sampled points along rays. Extensive quantitative and qualitative results demonstrate that our method reconstructs high-accuracy surfaces with details, and outperforms the state of the art.1 Decai Chen, Ingo Feldmann, Oliver Schreer, Peter Eisert |
WACV | 5 |
| 2023 | Unsupervised learning of style-aware facial animation from real acting performancesabstractThis paper presents a novel approach for text/speech-driven animation of a photo-realistic head model based on blend-shape geometry, dynamic textures, and neural rendering. Training a VAE for geometry and texture yields a parametric model for accurate capturing and realistic synthesis of facial expressions from a latent feature vector. Our animation method is based on a conditional CNN that transforms text or speech into a sequence of animation parameters. In contrast to previous approaches, our animation model learns disentangling/synthesizing different acting-styles in an unsupervised manner, requiring only phonetic labels that describe the content of training sequences. For realistic real-time rendering, we train a U-Net that refines rasterization-based renderings by computing improved pixel colors and a foreground matte. We compare our framework qualitatively/quantitatively against recent methods for head modeling as well as facial animation and evaluate the perceived rendering/animation quality in a user-study, which indicates large improvements compared to state-of-the-art approaches. Wolfgang Paier, Anna Hilsmann, Peter Eisert |
Graph. Model. | 3 |
| 2023 | Imposing temporal consistency on deep monocular body shape and pose estimationabstractAccurate and temporally consistent modeling of human bodies is essential for a wide range of applications, including character animation, understanding human social behavior, and AR/VR interfaces. Capturing human motion accurately from a monocular image sequence remains challenging; modeling quality is strongly influenced by temporal consistency of the captured body motion. Our work presents an elegant solution to integrating temporal constraints during fitting. This increases both temporal consistency and robustness during optimization. In detail, we derive parameters of a sequence of body models, representing shape and motion of a person. We optimize these parameters over the complete image sequence, fitting a single consistent body shape while imposing temporal consistency on the body motion, assuming body joint trajectories to be linear over short time. Our approach enables the derivation of realistic 3D body models from image sequences, including jaw pose, facial expression, and articulated hands. Our experiments show that our approach accurately estimates body shape and motion, even for challenging movements and poses. Further, we apply it to the particular application of sign language analysis, where accurate and temporally consistent motion modelling is essential, and show that the approach is well-suited to this kind of application. Alexandra Zimmer, Anna Hilsmann, Wieland Morgenstern, Peter Eisert |
Comput. Vis. Media | 4 |
| 2022 | CASAPose: Class-Adaptive and Semantic-Aware Multi-Object Pose Estimation
Niklas Gard, Anna Hilsmann, Peter Eisert |
BMVC | 3 |
| 2022 | Multi-View Mesh Reconstruction with Neural Deferred ShadingabstractWe propose an analysis-by-synthesis method for fast multi-view 3D reconstruction of opaque objects with arbitrary materials and illumination. State-of-the-art methods use both neural surface representations and neural rendering. While flexible, neural surface representations are a significant bottleneck in optimization runtime. Instead, we represent surfaces as triangle meshes and build a differentiable rendering pipeline around triangle rasterization and neural shading. The renderer is used in a gradient descent optimization where both a triangle mesh and a neural shader are jointly optimized to reproduce the multi-view images. We evaluate our method on a public 3D reconstruction dataset and show that it can match the reconstruction accuracy of traditional baselines and neural approaches while surpassing them in optimization runtime. Additionally, we investigate the shader and find that it learns an interpretable representation of appearance, enabling applications such as 3D material editing. Markus Worchel, Weiwen Hu, Oliver Schreer, Ingo Feldmann, Peter Eisert |
CVPR | 6 |
| 2022 | Latency Compensation Through Image Warping For Remote Rendering-Based Volumetric Video StreamingabstractRendering multiple high-quality volumetric videos is still a challenge for today’s mobile devices. Remote rendering offloads complex rendering operations to a powerful server and provides the final result to the end device as a 2D video stream. A drawback of remote rendering is the significant increase of interaction latency that can degrade the user experience. We present a planar homography-based approach that compensates minor changes of the user’s head pose due to the interaction latency by warping the transmitted image on the client-side, just before it is sent to the display. In detail, we use the homography between the initial head pose when the image is rendered at the server and the latest available head pose of the user at the client. We perform controlled experiments using artificial camera traces to evaluate our approach. The results show that the proposed approach reduces the rendering errors significantly in terms of the mean-squared error between the rendered and reference images, especially combined with initial head motion prediction. Serhan Gul, Cornelius Hellge, Peter Eisert |
ICIP | 3 |
| 2021 | Zero in on Shape: A Generic 2D-3D Instance Similarity Metric Learned from Synthetic DataabstractWe present a network architecture which compares RGB images and untextured 3D models by the similarity of the represented shape. Our system is optimised for zero-shot retrieval, meaning it can recognise shapes never shown in training. We use a view-based shape descriptor and a siamese network to learn object geometry from pairs of 3D models and 2D images.Due to scarcity of datasets with exact photograph-mesh correspondences, we train our network with only synthetic data.Our experiments investigate the effect of different qualities and quantities of training data on retrieval accuracy and present insights from bridging the domain gap. We show that increasing the variety of synthetic data improves retrieval accuracy and that our system’s performance in zero-shot mode can match that of the instance-aware mode, as far as narrowing down the search to the top 10% of objects. Maciej Janik, Niklas Gard, Anna Hilsmann, Peter Eisert |
ICIP | 4 |
| 2021 | Can You Do Real-Time Gesture Recognition with 5 Watts?abstractAccurate and reliable gesture recognition is a central problem in human-computer interaction (HCI). Many applications that make use of gesture recognition call for mobile devices with reduced power consumption, weight and form factors. Recent advances in computer vision were particularly brought by deep neural networks and come at the cost of high computational complexity that hinders the employment on mobile devices. In this study, we evaluate the usability of a low-cost Raspberry Pi 4B amended by a Coral USB Accelerator, or a Neural Compute Stick 2, respectively, for low power real-time gesture recognition. To this end we evaluate the accuracy, inference time and power consumption for two different deep neural network-based recognition models and compare the results to other computer systems available as standard. Our experiments show that a combination of a Raspberry Pi 4B and Coral USB Accelerator allows for hand gesture recognition at frame rates of up to 30 frames per second at a power consumption of less than 5 Watts. Azrin Rahman, Mykyta Kovalenko, David Przewozny, Karam Tomotaki-Dawoud, Paul Chojecki, Peter Eisert, Sebastian Bosse |
SMC | 6 |
| 2020 | EEG-Based Assessment of Perceived Realness in Stylized Face ImagesabstractIn this paper, we investigate the perception of realness in rendered face images experimentally using electroencephalography. To this end, we presented ten subjects with 36 character images based on six different faces (varying in gender and emotional expression) rendered at six different levels of realness ranging from abstract, cartoon-like renderings to real photographs. In the first psychophysical part of our study, we asked participants to rate perceived realness, appeal, familiarity, reassurance, and attractiveness for the presented characters. In the second part, we recorded the electroencephalogram when presenting the character images at a stimulation frequency of fstim= 5 Hz. We show that the amplitudes of the odd harmonics of the elicited steady-state visual evoked potential correlate with the psychophysical responses (|ρ| = 0.83, p < 0.05). Milena T. Bagdasarian, Anna Hilsmann, Peter Eisert, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001, Sebastian Bosse |
QoMEX | 3 |
| 2020 | Going beyond free viewpoint: creating animatable volumetric video of human performancesabstractAn end‐to‐end pipeline for the creation of high‐quality animatable volumetric video of human performances is presented. Going beyond the application of free‐viewpoint video, the authors allow re‐animation and alteration of an actor's performance through the enrichment of the captured data with semantics and animation properties. Hybrid geometry‐ and video‐based animation methods are applied that allow a direct animation of the high‐quality data itself instead of creating a CG model that resembles the captured data. Semantic enrichment and animation are achieved by establishing temporal consistency followed by automatic rigging of each 3D frame using a parametric human body model. The hybrid approach combines the flexibility of classical CG animation with the realism of real captured data. For the face, coarse movements are modelled in the geometry only, while very fine and subtle details, often lacking in purely geometric methods, are captured in video textures, which can interactively be combined to form new facial expressions. On top of that, regions that are challenging to synthesise, such as the teeth or the eyes, are learned and filled in realistically in an autoencoder‐based approach. This study covers the full pipeline from capturing, volumetric video production, and enrichment with semantics for the final hybrid animation. Anna Hilsmann, Philipp Fechteler, Wieland Morgenstern, Wolfgang Paier, Ingo Feldmann, Oliver Schreer, Peter Eisert |
IET Comput. Vis. | 7 |
| 2020 | Interactive facial animation with deep neural networksabstractCreating realistic animations of human faces is still a challenging task in computer graphics. While computer graphics (CG) models capture much variability in a small parameter vector, they usually do not meet the necessary visual quality. This is due to the fact, that geometry‐based animation often does not allow fine‐grained deformations and fails in difficult areas (mouth, eyes) to produce realistic renderings. Image‐based animation techniques avoid these problems by using dynamic textures that capture details and small movements that are not explained by geometry. This comes at the cost of high‐memory requirements and limited flexibility in terms of animation because dynamic texture sequences need to be concatenated seamlessly, which is not always possible and prone to visual artefacts. In this study, the authors present a new hybrid animation framework that exploits recent advances in deep learning to provide an interactive animation engine that can be used via a simple and intuitive visualisation for facial expression editing. The authors describe an automatic pipeline to generate training sequences that consist of dynamic textures plus sequences of consistent three‐dimensional face models. Based on this data, they train a variational autoencoder to learn a low‐dimensional latent space of facial expressions that is used for interactive facial animation. Wolfgang Paier, Anna Hilsmann, Peter Eisert |
IET Comput. Vis. | 3 |
| 2020 | Accurate and robust neural networks for face morphing attack detectionabstractArtificial neural networks tend to use only what they need for a task. For example, to recognize a rooster, a network might only considers the rooster’s red comb and wattle and ignores the rest of the animal. This makes them vulnerable to attacks on their decision making process and can worsen their generality. Thus, this phenomenon has to be considered during the training of networks, especially in safety and security related applications. In this paper, we propose neural network training schemes, which are based on different alternations of the training data, to increase robustness and generality. Precisely, we limit the amount and position of information available to the neural network for the decision making process and study their effects on the accuracy, generality, and robustness against semantic and black box attacks for the particular example of face morphing attacks. In addition, we exploit layer-wise relevance propagation (LRP) to analyze the differences in the decision making process of the differently trained neural networks. A face morphing attack is an attack on a biometric facial recognition system, where the system is fooled to match two different individuals with the same synthetic face image. Such a synthetic image can be created by aligning and blending images of the two individuals that should be matched with this image. We train neural networks for face morphing attack detection using our proposed training schemes and show that they lead to an improvement of robustness against attacks on neural networks. Using LRP, we show that the improved training forces the networks to develop and use reliable models for all regions of the analyzed image. This redundancy in representation is of crucial importance to security related applications. Clemens Seibold, Wojciech Samek, Anna Hilsmann, Peter Eisert |
J. Inf. Secur. Appl. | 4 |
| 2019 | Capture and 3D Video Processing of Volumetric VideoabstractVolumetric video is regarded worldwide as the next important development step in the field of media production. Especially in the context of the extremely rapid development of the Virtual Reality (VR) and Augmented Reality (AR) markets, volumetric video is becoming a key technology. In this paper, a new capture and processing system for volumetric video is presented, called 3D Human Body Reconstruction (3DHBR). The system is based on 16 stereo pairs of high-resolution cameras capturing a moving person in 360 degree. A novel stereo approach provides depth information from all perspectives, which is then fused to a single consistent 3D point cloud. A meshing and mesh reduction algorithm finally produces a sequence of meshes that can be integrated into common render engines. Given that, an integration of realistic dynamic 3D reconstructions of moving persons in VR and AR applications is possible. Oliver Schreer, Ingo Feldmann, Sylvain Renault, Marcus Zepp, Markus Worchel, Peter Eisert, Peter Kauff |
ICIP | 6 |
| 2019 | Interactive and Multimodal-based Augmented Reality for Remote Assistance using a Digital Surgical MicroscopeabstractWe present an interactive and multimodal-based augmented reality system for computer-assisted surgery in the context of ear, nose and throat (ENT) treatment. The proposed processing pipeline uses fully digital stereoscopic imaging devices, which support multispectral and white light imaging to generate high resolution image data, and consists of five modules. Input/output data handling, a hybrid multimodal image analysis and a bi-directional interactive augmented reality (AR) and mixed reality (MR) interface for local and remote surgical assistance are of high relevance for the complete framework. The hybrid multimodal 3D scene analysis module uses different wavelengths to classify tissue structures and combines this spectral data with metric 3D information. Additionally, we propose a zoom-independent intraoperative tool for virtual ossicular prosthesis insertion (e.g. stapedectomy) guaranteeing very high metric accuracy in sub-millimeter range (1/10 mm). A bi-directional interactive AR/MR communication module guarantees low latency, while consisting surgical information and avoiding informational overload. Display agnostic AR/MR visualization can show our analyzed data synchronized inside the digital binocular, the 3D display or any connected head-mounted-display (HMD). In addition, the analyzed data can be enriched with annotations by involving external clinical experts using AR/MR and furthermore an accurate registration of preoperative data. The benefits of such a collaborative surgical system are manifold and will lead to a highly improved patient outcome through an easier tissue classification and reduced surgery risk. Eric L. Wisotzky, Jean-Claude Rosenthal, Peter Eisert, Anna Hilsmann, Falko Schmid, Armin Schneider, Florian C. Uecker |
VR | 3 |
| 2019 | Markerless Multiview Motion Capture with 3D Shape Model AdaptationabstractAbstract In this paper, we address simultaneous markerless motion and shape capture from 3D input meshes of partial views onto a moving subject. We exploit a computer graphics model based on kinematic skinning as template tracking model. This template model consists of vertices, joints and skinning weights learned a priori from registered full‐body scans, representing true human shape and kinematics‐based shape deformations. Two data‐driven priors are used together with a set of constraints and cues for setting up sufficient correspondences. A Gaussian mixture model‐based pose prior of successive joint configurations is learned to soft‐constrain the attainable pose space to plausible human poses. To make the shape adaptation robust to outliers and non‐visible surface regions and to guide the shape adaptation towards realistically appearing human shapes, we use a mesh‐Laplacian‐based shape prior. Both priors are learned/extracted from the training set of the template model learning phase. The output is a model adapted to the captured subject with respect to shape and kinematic skeleton as well as the animation parameters to resemble the observed movements. With example applications, we demonstrate the benefit of such footage. Experimental evaluations on publicly available datasets show the achieved natural appearance and accuracy. Philipp Fechteler, Anna Hilsmann, Peter Eisert |
Comput. Graph. Forum | 3 |
| 2019 | Projection Distortion-based Object Tracking in Shader Lamp ScenariosabstractShader lamp systems augment the real environment by projecting new textures on known target geometries. In dynamic scenes, object tracking maintains the illusion if the physical and virtual objects are well aligned. However, traditional trackers based on texture or contour information are often distracted by the projected content and tend to fail. In this paper, we present a model-based tracking strategy, which directly takes advantage from the projected content for pose estimation in a projector-camera system. An iterative pose estimation algorithm captures and exploits visible distortions caused by object movements. In a closed-loop, the corrected pose allows the update of the projection for the subsequent frame. Synthetic frames simulating the projection on the model are rendered and an optical flow-based method minimizes the difference between edges of the rendered and the camera image. Since the thresholds automatically adapt to the synthetic image, a complicated radiometric calibration can be avoided. The pixel-wise linear optimization is designed to be easily implemented on the GPU. Our approach can be combined with a regular contour-based tracker and is transferable to other problems, like the estimation of the extrinsic pose between projector and camera. We evaluate our procedure with real and synthetic images and obtain very precise registration results. Niklas Gard, Anna Hilsmann, Peter Eisert |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Animatable 3D Model Generation from 2D Monocular Visual DataabstractIn this paper, we present an approach for creating animatable 3D models from temporal monocular image acquisitions of non-rigid objects. During deformation, the object of interest is captured with only a single camera under full perspective projection. The aim of the presented framework is to obtain a shape deformation model in terms of joints and skinning weights that can finally be used for animating the model vertices. First, the monocular rigid shape estimation problem is solved by computing a template model of the object in rest pose from an image sequence. Next, the unknown external camera parameters and the deformation for each vertex are estimated alternately in a sequential approach. The resulting consistent non-rigid shape geometries are used to compute a kinematic skeleton control structure including skinning weights and optimized shape. For that, a completely data-driven optimization scheme is used, which iterates over three steps: (a) optimization of pose for each frame as well as joint parameters consistent over the entire sequence, (b) optimization of rest pose vertices to enhance the shape and (c) optimization of skinning weights for improved deformation characteristics. With experimental results on publicly available synthetic as well as real-world datasets, we demonstrate the quality of the proposed approach. The resulting models with fixed topology and rigged with skeleton and skinning weights can be animated in existing render engines. Philipp Fechteler, Lisa Kausch, Anna Hilsmann, Peter Eisert |
ICIP | 4 |
| 2018 | Markerless Closed-Loop Projection Plane Tracking for Mobile Projector-Camera SystemsabstractThe recent trend towards miniaturization of mobile projectors is allowing new forms of information presentation and interaction. Projectors can easily be moved freely in space either by humans or by mobile robots. This paper presents a technique to dynamically track the orientation and position of the projection plane only by analyzing the distortion of the projection by itself, independent of the presented content. It allows distortion-free projection with a fixed metric size for moving projector-camera systems. To do this, an optical flow-based model is extended to the geometry of a projector-camera unit. Solving an overdetermined system of equations on pixel level leads to the pose offset between two images. In order to reach a high invariance to illumination changes we use adaptive edge images. Image pyramids allow a fast pose estimation. Because of the global optimization, there is no dependence of the availability of local feature points. Niklas Gard, Peter Eisert |
ICIP | 2 |
| 2018 | Automatic Analysis of Sewer Pipes Based on Unrolled Monocular Fisheye ImagesabstractThe task of detecting and classifying damages in sewer pipes offers an important application area for computer vision algorithms. This paper describes a system, which is capable of accomplishing this task solely based on low quality and severely compressed fisheye images from a pipe inspection robot. Relying on robust image features, we estimate camera poses, model the image lighting, and exploit this information to generate high quality cylindrical unwraps of the pipes' surfaces. Based on the generated images, we apply semantic labeling based on deep convolutional neural networks to detect and classify defects as well as structural elements. Johannes Künzel, Thomas Werner, Peter Eisert, Jan Waschnewski |
WACV | 3 |
| 2018 | Surface tracking assessment and interaction in texture spaceabstractIn this paper, we present a novel approach for assessing and interacting with surface tracking algorithms targeting video manipulation in post-production. As tracking inaccuracies are unavoidable, we enable the user to provide small hints to the algorithms instead of correcting erroneous results afterwards. Based on 2D mesh warp-based optical flow estimation, we visualize results and provide tools for user feedback in a consistent reference system, texture space. In this space, accurate tracking results are reflected by static appearance, and errors can easily be spotted as apparent change. A variety of established tools can be utilized to visualize and assess the change between frames. User interaction to improve tracking results becomes more intuitive in texture space, as it can focus on a small region rather than a moving object. We show how established tools can be implemented for interaction in texture space to provide a more intuitive interface allowing more effective and accurate user feedback. Johannes Furch, Anna Hilsmann, Peter Eisert |
Comput. Vis. Media | 3 |
| 2017 | Detection of Face Morphing Attacks by Deep Learning
Clemens Seibold, Wojciech Samek, Anna Hilsmann, Peter Eisert |
IWDW | 4 |
| 2017 | Model-based motion blur estimation for the improvement of motion trackingabstractVideo tracking is an important task in many automated or semi-automated applications, like cinematic post production, surveillance or traffic monitoring. Most established video tracking methods fail or lead to an inaccurate estimate when motion blur occurs in the video, as they assume, that the object appears constantly sharp in the video. In this paper, we present a novel motion tracking method with explicit modeling of motion blur, estimating the continuous motion of a rigid 3-D object with known geometry in a monocular video as well as the sharp object texture. Instead of treating motion blur as a potential source of errors, we take advantage of it and consider motion blur as an additional information source, providing information about the motion of the tracked object during the exposure. In an analysis-by-synthesis approach we explicitly model the effects of motion blur reconstructing the captured frames, in order to accomplish a more accurate estimation. We design our algorithm to be capable to run in parallel on the GPU using the common rendering pipeline and considering each frame individually to handle also long videos. We tested our approach on both synthetic and real videos. In both cases, we achieve significant improvements of accuracy and reductions of frame reconstruction error compared to the estimated motion of a rigid body tracker, without motion blur handling. Clemens Seibold, Anna Hilsmann, Peter Eisert |
Comput. Vis. Image Underst. | 3 |
| 2017 | Introduction to the Special Section on Augmented VideoabstractMerging computer-generated content with real-world visual data is one of the main challenges in fields like augmented reality or visual effects and is increasingly important in broadcasting, gaming, medical, automotive, maintenance, and learning applications. Although augmented reality (AR) has been investigated for a long time, it has recently emerged as a hot topic, with significant commercial interest. One reason for that is the availability of new camera-equipped devices like smart phones or tablets that enable see-through capabilities. Combined with powerful graphics capabilities, sensors, and tracking methods, AR now becomes available to everyone. In addition, glasses-based systems, like Microsoft’s HoloLens, allow for hands-free visualization, enabling many new applications. Peter Eisert, Yebin Liu, Kyuong Mu Lee, Didier Stricker, Graham A. Thomas |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | A Hybrid Approach for Facial Performance Analysis and EditingabstractAs of today, fine-grained editing of facial performances in movie and video production requires either retouching every single frame or creating a highly detailed CGI model of the actor, both of which is restricted to high-budget productions. In this paper, we present an example-based approach for facial performance editing that achieves realistic results with standard equipment and very little manual intervention. Based on a model-free surface tracking approach, temporally consistent dynamic texture sequences are extracted from multiple video streams. Using such geometry-plus-texture sequences allows transferring facial expressions/performances between videos, which enables the editor, for example, to change the facial expression while leaving the rest of the video untouched. Moreover, concatenating and/or looping these sequences in a motion-graph-like manner offers a convenient way of composing novel facial performances from multiple source videos. Finally, we present a blending method to seamlessly concatenate texture/mesh sequences and to insert a composed facial performance in a target video. Wolfgang Paier, Markus Kettern, Anna Hilsmann, Peter Eisert |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Real-time avatar animation with dynamic face texturingabstractIn this paper, we present a system to capture and animate a highly realistic avatar model of a user in real-time. The animated human model consists of a rigged 3D mesh and a texture map. The system is based on KinectV2 input which captures the skeleton of the current pose of the subject in order to animate the human shape model. An additional high-resolution RGB camera is used to capture the face for updating the texture map on each frame. With this combination of image based rendering with computer graphics we achieve photo-realistic animations in real-time. Additionally, this approach is well suited for networked scenarios, because of the low per frame amount of data to animate the model, which consists of motion capture parameters and a video frame. With experimental results, we demonstrate the high degree of realism of the presented approach. Philipp Fechteler, Wolfgang Paier, Anna Hilsmann, Peter Eisert |
ICIP | 4 |
| 2016 | Real-time 3D body reconstruction for immersive TVabstractIn this work, a novel and fast algorithm for real-time 3D body reconstruction from stereo sequences is proposed. The main contributions of this work consist of a novel approach for a statistically guided stereo processing and a data parallel iteration scheme for 3D estimation that includes temporal predecessors from a local spatial neighborhood. A purely GPU based implementation is provided that exhibits a nearly linear scaling of the runtime with respect to the number of GPUs. This leads to an inherent sub-pixel processing due to the availability of hardware supported texture lookups. Our implementation is able to process 4K (UHD) stereo streams on a 4×4 grid with 30 fps on a single state-of-the-art consumer graphics card. The algorithmic performance of our approach is demonstrated in the context of an immersive TV application. Wolfgang Waizenegger, Ingo Feldmann, Oliver Schreer, Peter Kauff, Peter Eisert |
ICIP | 5 |
| 2015 | A framework for image-based asset generation and animationabstractCreating digital animatable models of real-world objects and characters is important for many applications, ranging from highly expensive movie productions to low-cost real-time applications like computer games and augmented reality. However, achieving real photorealism with convincing appearance and deformation behavior requires sophisticated capturing, elaborate manual modeling and time-consuming simulation. This can only be achieved in well funded film productions, while in low-cost applications, animated objects usually lack visual quality. In this paper, we present a new framework for image-based animatable asset generation which avoids these time-consuming processes both in the modeling and the simulation stage. Real-time photo-realistic animation is enabled by the use of captured images and shifting computational complexity to an a-priori training phase. Our paper covers the complete pipeline of content creation, asset generation and representation, and a real-time animation and rendering implementation. Johannes Furch, Anna Hilsmann, Peter Eisert |
ICIP | 3 |
| 2015 | Interactive Scene Flow Editing for Improved Image-based Rendering and Virtual Spacetime NavigationabstractHigh-quality stereo and optical flow maps are essential for a multitude of tasks in visual media production, e.g. virtual camera navigation, disparity adaptation or scene editing. Rather than estimating stereo and optical flow separately, scene flow is a valid alternative since it combines both spatial and temporal information and recently surpassed the former two in terms of accuracy. However, since automated scene flow estimation is non-accurate in a number of situations, resulting rendering artifacts have to be corrected manually in each output frame, an elaborate and time-consuming task. We propose a novel workflow to edit the scene flow itself, catching the problem at its source and yielding a more flexible instrument for further processing. By integrating user edits in early stages of the optimization, we allow the use of approximate scribbles instead of accurate editing, thereby reducing interaction times. Our results show that editing the scene flow improves the quality of visual results considerably while requiring vastly less editing effort. Kai Ruhl, Martin Eisemann, Anna Hilsmann, Peter Eisert, Marcus A. Magnor |
ACM Multimedia | 4 |
| 2014 | Articulated 3D model tracking with on-the-fly texturingabstractIn this paper, we present a framework for capturing and tracking humans based on RGBD input data. The two contributions of our approach are: (a) a method for robustly and accurately fitting an articulated computer graphics model to captured depth-images and (b) on-the-fly texturing of the geometry based on the sensed RGB data. Such a representation is especially useful in the context of 3D telepresence applications since model-parameter and texture updates require only low bandwidth. Additionally, this rigged model can be controlled through interpretable parameters and allows automatic generation of naturally appearing animations. Our experimental results demonstrate the high quality of this model-based rendering. Philipp Fechteler, Wolfgang Paier, Peter Eisert |
ICIP | 3 |
| 2014 | High-resolution depth for binocular image-based modeling
David Blumenthal-Barby, Peter Eisert |
Comput. Graph. | 2 |
| 2014 | Guest editorial: Advances in 3D video processing
Shang-Hong Lai, Gene Cheung, Dinei A. F. Florêncio, Peter Eisert, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Real-time generation of multi-view video plus depth content using mixed narrow and wide baseline
Frederik Zilly, Christian Riechert, Marcus Müller 0001, Peter Eisert, Thomas Sikora, Peter Kauff |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | An Iterative Method for Improving Feature MatchesabstractFinding reliable and well distributed keypoint correspondences between images of non-static scenes is an important task in Computer Vision. We present an iterative algorithm that improves a descriptor Based matching result by enforcing local smoothness. During the optimization process, a Delaunay triangulation of the current set of matches is dynamically maintained. This 2D mesh provides natural neighborhoods and local affine transformations that are used to remove outliers and to resolve ambiguities. The optimization results in a decrease of incorrect correspondences and a significant increase in the total number of matches. The runtime of the overall algorithm is by far dictated by the descriptor Based matching. Johannes Furch, Peter Eisert |
3DV | 2 |
| 2013 | Scene flow constrained multi-prior patch-sweeping for real-time upper body 3D reconstructionabstractWe present a high performance real-time approach for robust 3D structure estimation of human bodies. Our idea extends state of the art high precision patch-sweep techniques by exploiting temporal information. By parameterizing 3D patches in the spatio-temporal domain we gain increased quality and robustness, while the computational complexity decreases drastically. The key to these improvements is the utilization of a global scene flow like prior called spatio-temporal object (STO). The target application of our method reaches from video communication and virtual eye contact scenarios till future digital cinema 3D post-production and real-time person modelling and modification. The result of this work extends a novel 3D scene representation developed in the EC research project SCENE. Wolfgang Waizenegger, Ingo Feldmann, Oliver Schreer, Peter Eisert |
ICIP | 4 |
| 2013 | Pose Space Image Based RenderingabstractAbstract This paper introduces a new image‐based rendering approach for articulated objects with complex pose‐dependent appearance, such as clothes. Our approach combines body‐pose‐dependent appearance and geometry to synthesize images of new poses from a database of examples. A geometric model allows animation and view interpolation, while small details as well as complex shading and reflection properties are modeled by pose‐dependent appearance examples in a database. Correspondences between the images are represented as mesh‐based warps, both in the spatial and intensity domain. For rendering, these warps are interpolated inpose space, i.e. the space of body poses, using scattered data interpolation methods. Warp estimation as well as geometry reconstruction is performed in an offline procedure, thus shifting computational complexity to an a‐priori training phase. Anna Hilsmann, Philipp Fechteler, Peter Eisert |
Comput. Graph. Forum | 3 |
| 2012 | Model based 3D gaze estimation for provision of virtual eye contactabstractIn recent years, video communication has received a rapidly increasing interest on the market. Still unsolved is the problem of eye contact. The conferee still needs to decide whether to look into the camera or directly to the screen. Recently, a solution to this problem was presented which is based on a real-time 3D modeling of the conferees [1]. In order to achieve direct eye contact the authors defined a virtual camera directly on the screen in the eyes of the remote conferee. This paper discusses the problem of adequately positioning this virtual camera. A new approach will be presented which performs an eye and gaze tracking directly on the real-time 3D model rather than on the 2D image. Our methods not only provides robust and highly accurate results but is also able to additionally measure the distance between the conferees eye and the display with high precision. Wolfgang Waizenegger, Nicole Atzpadin, Oliver Schreer, Ingo Feldmann, Peter Eisert |
ICIP | 5 |
| 2012 | The Ultimate Immersive Experience: Panoramic 3D Video Acquisition
Christian Weissig, Oliver Schreer, Peter Eisert, Peter Kauff |
MMM | 3 |
| 2011 | Template-free Shape From Texture with Perspective Camerasabstract10 S. Anna Hilsmann, David C. Schneider, Peter Eisert |
BMVC | 3 |
| 2011 | System for the automated segmentation of heads from arbitrary backgroundabstractWe propose a system for the fully automated segmentation of frontal human head portraits from arbitrary unknown background. No user interaction is required at all, as the system is initialized using a standard eye detector. Using this semantic information, the head region is projected into a normalized polar reference frame. Regional and boundary models are learned from the image data to setup an energy function for segmentation. A robust non-local boundary detection scheme is proposed, which minimizes the similarity of fore - and background regions. Additionally, a shape model learned from a large set of manually segmented images is employed as prior information to encourage the segmentation of plausible head shapes. Segmentation is performed as an iterative optimization process, using two different graph-based algorithms. Benjamin Prestele, David C. Schneider, Peter Eisert |
ICIP | 3 |
| 2010 | Virtual jewel rendering for augmented reality environmentsabstractVirtual Mirrors are augmented reality applications that capture a viewer's image and add 2D or 3D overlaid virtual elements, such as cloth patterns, or replace existing pieces of garments, such as shoes, by virtual ones. The purpose of our study is to focus on the addition of jewels that present complex rendering issues due to reflection and translucency, especially for real-time applications. We propose two complementary approaches which are both suited for particular objects. The image-based rendering technique relies on an image data-set of photos taken on a semi-circle of camera positions and a 3D reconstruction of the jewel. This approach is exploited to demonstrate real jewels with sophisticated refractions and highlights. The analytical 3D rendering technique can also visualize configurable virtual jewels and is based on the decomposition of the jewel into parts without non-local reflection or transparency. A graph-based rendering chain takes in consideration some of the self-reflection and lighting effects. Peter Eisert, Christian Jacquemin, Anna Hilsmann |
ICIP | 1 |
| 2010 | Accelerated video encoding using render context informationabstractIn this paper, we present a method to speed up video encoding of GPU rendered 3D scenes, which is particularly suited for the efficient and low-delay encoding of 3D game output as a video stream. The main idea of our approach is to calculate motion vectors directly from the 3D scene information used during rendering of the scene. This allows the omission of the computationally expensive motion estimation search algorithms found in most of today's video encoders. The presented method intercepts the graphics commands during runtime of 3D computer games to capture the required projection information without requiring any modification of the game executable. We demonstrate that this approach is applicable to games based on Linux/OpenGL as well as Windows/DirectX. In experimental results we show an acceleration of video encoding performance of approximately 25% with almost no degradation in image quality. Philipp Fechteler, Peter Eisert |
ICIP | 2 |
| 2010 | Patch-based reconstruction and rendering of human headsabstractReconstructing the 3D shape of human faces is an intensively researched topic. Most approaches aim at generating a closed surface representation of geometry, i.e. a mesh, which is texture-mapped for rendering. However, if free viewpoint rendering is the primary purpose of the reconstruction, representations other than meshes are possible. In this paper a coarse patch-based approach to both reconstruction and rendering is explored and applied not only to the face but the whole human head. The approach has advantages on parts of the scene that are traditionally difficult to reconstruct and render, which is the case for hair when it comes to human heads. In the paper, reconstruction of a patch is posed as a parameter estimation problem which is solved in a generic image-based optimization framework using the Levenberg-Marquard algorithm. In order to improve robustness, the Huber error metric is used and a geometric regularization strategy is introduced. Initial values for the optimization, which are crucial for the method's success, are obtained by triangulation of SIFT feature points and a recursive expansion scheme. David C. Schneider, Anna Hilsmann, Peter Eisert |
ICIP | 3 |
| 2010 | The Stereoscopic Analyzer - An image-based assistance tool for stereo shooting and 3D productionabstractThe paper discusses an assistance system for stereo shooting and 3D production, called Stereoscopic Analyzer (STAN). A feature-based scene analysis estimates in real-time the relative pose of the two cameras in order to allow optimal camera alignment and lens settings directly at the set. It automatically eliminates undesired vertical disparities and geometrical distortions through image rectification. In addition, it detects the position of near- and far objects in the scene to derive the optimal inter-axial distance (stereo baseline), and gives a framing alert in case of stereoscopic window violation. Against this background the paper describes the system architecture, explains the theoretical background and discusses future developments. Frederik Zilly, Marcus Müller 0001, Peter Eisert, Peter Kauff |
ICIP | 3 |
| 2010 | Realistic cloth augmentation in single view video under occlusions
Anna Hilsmann, David C. Schneider, Peter Eisert |
Comput. Graph. | 3 |
| 2009 | Depth map enhanced macroblock partitioning for H.264 video coding of computer graphics contentabstractIn this paper, we present a method to speed up video encoding of GPU rendered scenes. Modern video codecs, like H.264/AVC, are based on motion compensation and support partitioning of macroblocks, e.g. 16×16, 16×8, 8×8, 8×4 etc. In general, encoders use expensive search methods to determine suitable motion vectors and compare the rate-distortion score for possible macroblock partitionings, which results in high computational encoder load. We present a method to accelerate this process for the case of streaming graphical output of unmodified commercially available 3D games which use a Skybox or Skydome rendering technique. For rendered images, usually additional information from the render context of OpenGL resp. DirectX is available which helps in the encoding process. By incorporating the depth map from the graphics board, such regions can be uniquely identified. By adapting the macroblock partitioning accordingly, the computationally expensive search methods can often be avoided. Further reduction of encoding load is achieved by additionally capturing the projection matrices during the Skybox rending and using them to directly calculate a motion vector which is usually the result of expensive search methods. In experiential results, we demonstrate the reduced computational encoder load. Philipp Fechteler, Peter Eisert |
ICIP | 2 |
| 2009 | Precise head segmentation on arbitrary backgroundsabstractWe propose a method for segmentation of frontal human portraits from arbitrary unknown backgrounds. Semantic information is used to project the face into a normalized reference frame. A shape model learned from a set of manually segmented faces is used to compute a rough initial segmentation using a fast iterative algorithm. The rough initial cutout is refined with a boundary based algorithm called ¿cluster cutting¿. Cluster cutting uses a cost function derived from clustering pixels along the normal of the initial segmentation path with a tree-building algorithm. The result can be refined by the user with an interactive variant of the same algorithm. David C. Schneider, Benjamin Prestele, Peter Eisert |
ICIP | 3 |
| 2009 | Parallel high resolution real-time Visual Hull on GPUabstractIn this paper we present an efficient high resolution image based visual hull (IBVH) algorithm that entirely runs in real-time on a single consumer graphics card. The target application is a real-time 3D video conferencing system. One major contribution of this paper is a novel caching strategy for the reduction of line segment intersection tests. In contrast to existing approaches, it additionally allows us to pre-select a close estimation of the set of relevant pixels in the desired view. Based on this, we obtain a significant computational speedup, especially for high resolutions. Further, we propose an efficient way to use IBVH for the generation of voxel equivalent 3D models. We compare our techniques in terms of resolution and runtime with state of the art real-time multi GPU voxel based approaches. Our experiments show that we achieve a speed-up by a factor of five and more for high resolutions. Wolfgang Waizenegger, Ingo Feldmann, Peter Eisert, Peter Kauff |
ICIP | 3 |
| 2008 | 3-D Tracking of shoes for Virtual Mirror applicationsabstractIn this paper, augmented reality techniques are used in order to create a virtual mirror for the real-time visualization of customized sports shoes. Similar to looking into a mirror when trying on new shoes in a shop, we create the same impression but for virtual shoes that the customer can design individually. For that purpose, we replace the real mirror by a large display that shows the mirrored input of a camera capturing the legs and shoes of a person. 3-D tracking of both feet and exchanging the real shoes by computer graphics models gives the impression of actually wearing the virtual shoes. The 3-D motion tracker presented in this paper, exploits mainly silhouette information to achieve robust estimates for both shoes from a single camera view. The use of a hierarchical approach in an image pyramid enables real-time estimation at frame rates of more than 30 frames per second. Peter Eisert, Philipp Fechteler, Jürgen Rurainsky |
CVPR | 1 |
| 2008 | Low delay streaming of computer graphicsabstractIn this paper, we present a graphics streaming system for remote gaming in a local area network. The framework aims at creating a networked game platform for home and hotel environments. A local PC based server executes a computer game and streams the graphical output to local devices in the rooms, such that the users can play everywhere in the network. Since delay is extremely crucial in interactive gaming, efficient encoding and caching of the commands is necessary. In our system we also address the round trip time problem of commands requiring feedback from the graphics board by simulating the graphics state at the server. This results in a system that enables interactive game play over the network. Peter Eisert, Philipp Fechteler |
ICIP | 1 |
| 2008 | Optical flow based tracking and retexturing of garmentsabstractIn this paper, we present a method for tracking and retexturing of garments that exploits the entire image information using the optical flow constraint instead of working with distinct features. In a hierarchical framework we refine the motion model with every level. The motion model is used to regularize the optical flow field such that finding the best transformation amounts in minimizing an error function that can be solved in a least squares sense. Knowledge about the position and deformation of the garment in 2D allows us to erase the old texture and replace it by a new one with correct deformation and shading properties without 3D reconstruction. Additionally, it provides an estimation of the irradiance such that the new texture can be illuminated realistically. Anna Hilsmann, Peter Eisert |
ICIP | 2 |
| 2008 | Real-Time Vision and Speech Driven Avatars for Multimedia ApplicationsabstractRecent progress in advanced video communication services and multimedia applications is grounded on novel human machine interfaces, improved usability, and user friendliness driven by user centric research and development. In this paper, we describe a complete system concept and algorithmic details of an example application within this area. The key features of the system are vision and speech based interfaces, which are used to animate an avatar for an audio-visual representation of a communication partner. The system is applied in two application scenarios, namely video chat and customer care services. Both applications are mass-market oriented and therefore careful design and development of robust and supporting user interfaces are required. The presented approach is integrated into a complete real-time prototype system, which is permanently demonstrated in the showcase at the head quarter of Deutsche Telekom, Bonn, Germany. Oliver Schreer, Roman Englert, Peter Eisert, Ralf Tanger |
IEEE Trans. Multim. | 3 |
| 2007 | Virtual Mirror: Real-Time Tracking of Shoes in Augmented Reality EnvironmentsabstractIn this paper, we present a system that enhances the visualization of customized sports shoes using augmented reality techniques. Instead of viewing yourself in a real mirror, sophisticated 3D image processing techniques are used to verify the appearance of new shoe models. A single camera captures the person and outputs the mirrored images onto a large display which replaces the real mirror. The 3D motion of both feet are tracked in real-time with a new motion tracking algorithm. Computer graphics models of the shoes are augmented into the video such that the person seems to wear the virtual shoes. Peter Eisert, Jürgen Rurainsky, Philipp Fechteler |
ICIP (2) | 1 |
| 2007 | Fast and High Resolution 3D Face ScanningabstractIn this work, we present a framework to capture 3D models of faces in high resolutions with low computational load. The system captures only two pictures of the face, one illuminated with a colored stripe pattern and one with regular white light. The former is needed for the depth calculation, the latter is used as texture. Having these two images a combination of specialized algorithms is applied to generate a 3D model. The results are shown in different views: simple surface, wire grid respective polygon mesh or textured 3D surface. Philipp Fechteler, Peter Eisert, Jürgen Rurainsky |
ICIP (3) | 2 |
| 2007 | Detection Strategies for Image Cube Trajectory AnalysisabstractImage cube trajectory (ICT) analysis is a new and robust method to estimate the 3D structure of a scene from a set of 2D images. For a moving camera each 3D point is represented by a trajectory in a so called image cube. In our previous work we have shown that it is possible to reconstruct the 3D scene from the parameters of these trajectories. A key component for this process is the trajectory detection within the cube. It is based on the image cube parameterization as well as the robust estimation of the trajectory color and trajectory color variation. In this paper we will focus on the second problem in more detail. We propose an algorithm which estimates the trajectory parameters in sub-pixel resolution with high accuracy. The corresponding 3D scene structure can be reconstructed with high level of detail even for complex scenes, multiple occlusions and very fine structures. Ingo Feldmann, Peter Kauff, Peter Eisert |
ICIP (2) | 3 |
| 2007 | Mirror-Based Multi-View Analysis of Facial MotionsabstractWe present our system for the capturing and analysis of 3D facial motion. A high speed camera is used as capture unit in combination with two surface mirrors. The mirrors provide two additional virtual views of the face without the need of multiple cameras and to avoid synchronization problems. We use this system to capture the motion of a person's face while speaking. Investigations of these facial motions are presented and rigid and non-rigid motion are analyzed. In order to extract only facial deformation independent from head pose, we use a new and simple approach for separating rigid and non-rigid motion named weight-compensated motion estimation (WCME). This approach weights the data points according to their influence to the desired motion model. We also present first results of our model-based facial deformation analysis. Such results can be used for facial animations in order to achieve a higher degree of quality. Jürgen Rurainsky, Peter Eisert |
ICIP (3) | 2 |
| 2006 | Towards Robust Intuitive Vision-Based User InterfacesabstractIn future video communication services, the user's communication device, such as PC, laptop, PDA or mobile phone is equipped with new interaction modalities. These can be cameras and microphones on the capturing side and speech synthesis and video/3D graphics on the rendering side. Haptic and tactile interfaces become also available. These modalities help the user to interact more intuitive with complex devices and tools and provide new services. Hence, a key challenge of new modalities is robustness and stability under general conditions in arbitrary environments. Furthermore, inexperienced users should be able to use these new capabilities without dedicated knowledge of device settings or algorithms. In this paper, we will present some key components for a robust vision-based user interface, which are integrated in an advanced future video communication service Oliver Schreer, Peter Eisert, Peter Kauff, Ralf Tanger, Roman Englert |
ICME | 2 |
| 2006 | 3D Video and Free Viewpoint Video - Technologies, Applications and MPEG StandardsabstractAn overview of 3D and free viewpoint video is given in this paper with special focus on related standardization activities in MPEG. Free viewpoint video allows the user to freely navigate within real world visual scenes, as known from virtual worlds in computer graphics. Examples are shown, highlighting standards conform realization using MPEG-4. Then the principles of 3D video are introduced providing the user with a 3D depth impression of the observed scene. Example systems are described again focusing on their realization based on MPEG-4. Finally multi-view video coding is described as a key component for 3D and free viewpoint video systems. The conclusion is that the necessary technology including standard media formats for 3D and free viewpoint is available or will be available in the near future, and that there is a clear demand from industry and user side for such applications. 3D TV at home and free viewpoint video on DVD will be available soon, and will create huge new markets Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Christoph Fehn, Peter Kauff, Peter Eisert, Thomas Wiegand 0001 |
ICME | 6 |
| 2006 | Creation of High-Resolution Video Panoramas of Sport EventsabstractThis paper describes an approach for creating high-resolution video panoramas of large-scale sport events. In an exemplary football scenario, we used two ARRIFLEX D-20 "film-style" digital cameras to capture both sides of the playing field as well as large parts of the stands. From the recorded left and right view images, we created a joint panoramic view with a resolution of 501 6x1400 pel (5k) using sophisticated image processing algorithms. These highly immersive football panoramas were screened with our modular, high-resolution multiprojection system during the FIFA World Cup 2006 in a 600 seat CinemaxX movie theater in Berlin providing the viewers with the feeling of actually watching the game from a good seat on the stands. Christoph Fehn, Christian Weissig, Ingo Feldmann, Marcus Müller 0001, Peter Eisert, Peter Kauff, Hans Bloß |
ISM | 5 |
| 2006 | Geometry-assisted image-based rendering for facial analysis and synthesis
Peter Eisert, Jürgen Rurainsky |
Signal Process. Image Commun. | 1 |
| 2006 | Rate-distortion-optimized predictive compression of dynamic 3D mesh sequences
Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 4 |
| 2005 | Image-based rendering and tracking of facesabstractIn this paper, we present an image-based method for the tracking and rendering of faces. We use the algorithm in an immersive video conferencing system where multiple participants are placed at a virtual table. This requires viewpoint modification of dynamic objects. Since hair and uncovered areas are difficult to model by pure 3-D geometry-based warping, we add image-based rendering techniques to the system. By interpolating novel views from a 3-D image volume, natural looking results can be achieved. The image-based component is embedded into a geometry-based approach that models temporally changing facial features. Both geometry and image cube information are jointly exploited in facial expression analysis and synthesis. Peter Eisert, Jürgen Rurainsky |
ICIP (1) | 1 |
| 2005 | Towards arbitrary camera movements for image cube trajectory analysisabstractImage cube trajectory (ICT) Analysis is a new and robust method to estimate the 3D structure of a scene from a set of 2D images. For a moving camera each 3D point is represented by a trajectory in an image cube. In previous publications we have shown that it is possible to reconstruct the 3D scene from the parameters of these trajectories. Nevertheless, the algorithm was restricted to simple parameterized camera movements. In this paper we discuss the problem of arbitrary camera motion. We benefit from the fact that in many cases, in particular for hand-held cameras, the image deviations caused by rotational variation of the camera parameters is much higher than for translational variation. We show that it is possible to compensate such rotational and translational deviations by transformation and resampling of the image cube. We obtain more uniform and smooth trajectory structures which can be analyzed by standard ICT analysis algorithms. Ingo Feldmann, Peter Eisert, Peter Kauff |
ICIP (3) | 2 |
| 2005 | Predictive compression of dynamic 3D meshesabstractAn efficient algorithm for compression of dynamic time-consistent 3D meshes is presented. Such a sequence of meshes contains a large degree of temporal statistical dependencies that can be exploited for compression using DPCM. The vertex positions are predicted at the encoder from a previously decoded mesh. The difference vectors are further clustered in an octree approach. Only a representative for a cluster of difference vectors is further processed providing a significant reduction of data rate. The representatives are scaled and quantized and finally entropy coded using CABAC, the arithmetic coding technique used in H.264/MPEG4-AVC. The mesh is then reconstructed at the encoder for prediction of the next mesh. In our experiments we compare the efficiency of the proposed algorithm in terms of bit-rate and quality compared to static mesh coding and interpolator compression indicating a significant improvement in compression efficiency. Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
ICIP (1) | 4 |
| 2004 | 3-d geometry enhancement by contour optimization in turntable sequencesabstractA method for the enhancement of geometry accuracy in shape-from-shading frameworks is presented. For the particular case of turntable scenarios, an optimization scheme is presented that minimizes silhouette deviations which correspond to shape errors. Only three unknown parameters have to be optimized leading to a robust and relatively fast framework. In spite of the small number of parameters, experiments have shown that the silhouette error can be reduced by a factor of more than 10 even after an already quite accurate camera calibration step. The quality of an additional texture mapping can also be drastically improved making the proposed scheme applicable as a preprocessing step in many different 3-D multimedia applications. Peter Eisert |
ICIP | 1 |
| 2004 | Optimized space sampling for circular image cube trajectory analysis
Ingo Feldmann, Peter Kauff, Peter Eisert |
ICIP | 3 |
| 2004 | Free viewpoint video extraction, representation, coding, and renderingabstractFree viewpoint video provides the possibility to freely navigate within dynamic real world video scenes by choosing arbitrary viewpoints and view directions. So far, related work only considered free viewpoint video extraction, representation, and rendering methods. Compression and transmission has not yet been studied in detail and combined with the other components into one complete system. In this paper, we present such a complete system for efficient free viewpoint video extraction, representation, coding, and interactive rendering. Data representation is based on 3D mesh models and view-dependent texture mapping using video textures. The geometry extraction is based on a shape-from-silhouette algorithm. The resulting voxel models are converted into 3D meshes that are coded using MPEG-4 SNHC tools. The corresponding video textures are coded using an H.264/AVC codec. Our algorithms for view-dependent texture mapping have been adopted as an extension of MPEG-4 AFX. The presented results illustrate that based on the proposed methods a complete transmission system for efficient free viewpoint video can be built. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Tobias Rein, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
ICIP | 6 |
| 2004 | Representation, coding, and rendering of 3D video objects with MPEG-4 and H.264/AVCabstract3D video objects provide the same functionalities as virtual computer graphics objects but depict the motion and appearance of real world moving objects. They can be viewed interactively from any direction and integrated in complete 3D scenes with other virtual and real world elements. So far, related work only considered extraction, representation, and rendering methods. Compression and transmission has not yet been studied in detail and combined with the other components into one complete system. In this paper, we present such a complete system for efficient 3D video object extraction, representation, coding, and interactive rendering. Data representation is based on 3D mesh models and view-dependent texture mapping using video textures. The geometry extraction is based on a shape-from-silhouette algorithm. The resulting voxel models are converted into 3D meshes that are coded using MPEG-4 SNHC tools. The corresponding video textures are preprocessed taking the object's shape into account and coded using an H.264/AVC codec. The presented results illustrate that based on the proposed methods a complete transmission system for 3D video objects can be built. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Tobias Rein, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
MMSP | 6 |
| 2003 | Extension of epipolar image analysis to circular camera movementsabstractEpipolar image analysis is a robust method for 3D scene depth reconstruction that uses all available views of an image sequence simultaneously. It is restricted to horizontal, linear, and equidistant camera movements. In this paper, we present a concept for an extension of epipolar image analysis to other camera configurations like, e.g., circular movements. Instead of searching for straight lines in the epipolar image, we explicitly compute the trajectories of particular points through the image cube. Variation of the unknown depth leads to different curves. From those, the one corresponding to the true depth is selected by evaluating color constancy along the curve. In order to handle occlusions correctly an explicit occlusion compatible ordering scheme is derived for the case of circular movements. To compensate the influence of perspective projection we introduce a depth corrected epipolar image analysis algorithm which we call image cube trajectory analysis (ICT). Ingo Feldmann, Peter Kauff, Peter Eisert |
ICIP (3) | 3 |
| 2003 | Immersive 3D video conferencing: challenges, concepts, and implementation
Peter Eisert |
VCIP | 1 |
| 2002 | Geometry refinement for light field compressionabstractIn geometry-aided light field compression, a geometry model is used for disparity-compensated prediction of light field images from already encoded light field images. This geometry model, however, may have limited accuracy. We present an algorithm that refines a geometry model to improve the overall light field compression efficiency. This algorithm uses an optical-flow technique to explicitly minimize the disparity-compensated prediction error. Results from experiments performed on both real and synthetic data sets show bit-rate reductions of approximately 10% using the improved geometry model over a silhouette-reconstructed geometry model. Peter Eisert, Prashant Ramanathan, Eckehard G. Steinbach, Bernd Girod |
ICIP (2) | 1 |
| 2002 | A real-time Internet streaming media testbedabstractWe describe a real-time LAN-based testbed that allows us to investigate the behavior of streaming media applications under various network conditions. For commercially available streaming media applications, we are interested to see how they perform over next-generation wireline and wireless networks. For future streaming media applications, the testbed is an invaluable tool for the development and verification of new algorithms. Our testbed implementation is based on Linux Divert Sockets and supports a straightforward integration of various packet erasure and delay models. Individual IP-packets are diverted to a user process where they are delayed or deleted according to the desired channel model. The testbed has been used to investigate the flow-control behavior of existing streaming media systems over wireless networks. Our experiments confirm that flow-control algorithms that consider lost packets to be the result of network congestion, as employed today in wireline streaming, are not suited for wireless networks, where loss is mainly due to link impairments. Wolfgang Kellerer, Eckehard G. Steinbach, Peter Eisert, Bernd Girod |
ICME (2) | 3 |
| 2002 | Model-based enhancement of lighting conditions in image sequences
Peter Eisert, Bernd Girod |
VCIP | 1 |
| 2001 | Multiview image coding with depth maps and 3D geometry for prediction
Marcus A. Magnor, Peter Eisert, Bernd Girod |
VCIP | 2 |
| 2000 | Model-Aided Coding: Using 3-D Scene Models in Motion-Compensated Video CodingabstractWe show that traditional waveform coding and 3-D model-based coding are not competing alternatives but should be combined to support and complement each other. Both approaches are combined such that the generality of waveform coding and the efficiency of 3-D model-based coding are available where needed. The combination is achieved by providing the block-based video coder with a second reference frame for prediction which is synthesized by the model-based coder. Since the coding gain of this approach is directly related to the quality of the synthetic frame, we have extended the model-aided coder to cope with illumination changes and multiple objects. Remaining model failures and objects that are not known at the decoder are handled by standard block-based motion-compensated prediction. Experimental results show that bit-rate savings of up to 45% are achieved at equal average PSNR when comparing the model-aided codec to TMN-10, the test model of the H.263 standard. Peter Eisert, Thomas Wiegand 0001, Bernd Girod |
ICIP | 1 |
| 2000 | Model-Aided Coding of Multi-Viewpoint Image DataabstractS.919-922 Marcus A. Magnor, Peter Eisert, Bernd Girod |
ICIP | 2 |
| 2000 | 3-D Reconstruction of Real-World Objects Using Extended VoxelsabstractIn this paper we present a voxel-based 3-D reconstruction technique that computes a set of non-transparent object surface voxels from a given set of calibrated camera views. We show that the quality of the reconstruction strongly depends on the accuracy of the computed voxel projection in the image plane and discuss different approximations of the exact projection. The most simple and computationally least demanding approximation is obtained when projecting point voxels, i.e., voxels without spatial extent. However, correct occlusion handling is not possible for point voxel volumes leading to reconstruction artifacts. The most accurate projection is obtained by computing the exact outline of the projected voxels. This projection is computationally most demanding but allows correct occlusion handling during reconstruction. Experimental results that compare the reconstruction quality for point and exact voxel projection show that it is worthwhile computing the tract image plane footprint of the projected voxels. Eckehard G. Steinbach, Bernd Girod, Peter Eisert, Arnulf Betz |
ICIP | 3 |
| 2000 | 3-D Object Reconstruction Using Spatially Extended Voxels and Multi-Hypothesis Voxel ColoringabstractWe describe a voxel-based 3-D reconstruction technique from multiple calibrated camera views that makes explicit use of the finite size footprint of a voxel when projected into the image plane. We derive a class of computationally efficient axis-aligned volume traversal orders that ensure that a processed voxel cannot occlude previously processed voxels. For each view, one out of 79 different cases of volume traversal is identified depending on the relative position between camera and voxel volume. Views belonging to the same visibility class can be processed simultaneously. Our voxel coloring strategy is based on a color hypothesis test that ensures the consistency of the projected reconstruction with the original images. A surface voxel list is constantly updated during reconstruction ensuring that only a minimum number of voxels has to be processed. Experimental results that compare the reconstruction quality for voxels with and without spatial extent underscore the conclusion that it is worthwhile taking into account the exact footprint of the projected voxels. Eckehard G. Steinbach, Bernd Girod, Peter Eisert, Arnulf Betz |
ICPR | 3 |
| 2000 | Model-aided coding: a new approach to incorporate facial animation into motion-compensated video codingabstractWe show that traditional waveform coding and 3-D model-based coding are not competing alternatives, but should be combined to support and complement each other. Both approaches are combined such that the generality of waveform coding and the efficiency of 3-D model-based coding are available where needed. The combination is achieved by providing the block-based video coder with a second reference frame for prediction, which is synthesized by the model-based coder. The model-based coder uses a parameterized 3-D head model, specifying the shape and color of a person. We therefore restrict our investigations to typical videotelephony scenarios that show head-and-shoulder scenes. Motion and deformation of the 3-D head model constitute facial expressions which are represented by facial animation parameters (FAPs) based on the MPEG-4 standard. An intensity gradient-based approach that exploits the 3-D model information is used to estimate the FAPs, as well as illumination parameters, that describe changes of the brightness in the scene. Model failures and objects that are not known at the decoder are handled by standard block-based motion-compensated prediction, which is not restricted to a special scene content, but results in lower coding efficiency. A Lagrangian approach is employed to determine the most efficient prediction for each block from either the synthesized model frame or the previous decoded frame. Experiments on five video sequences show that bit rate savings of about 35% are achieved at equal average peak signal-to-noise ratio (PSNR) when comparing the model-aided codec to TMN-10, the state-of-the-art test model of the M.263 standard. This corresponds to a gain of 2-3 dB in PSNR when encoding at the same average bit rate. Peter Eisert, Bernd Girod, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2000 | Automatic reconstruction of stationary 3-D objects from multiple uncalibrated camera viewsabstractA system for the automatic reconstruction of real-world objects from multiple uncalibrated camera views is presented. The camera position and orientation for all views, the 3-D shape of the rigid object, as well as the associated color information, are recovered from the image sequence. The system proceeds in four steps. First, the internal camera parameters describing the imaging geometry are calibrated using a reference object. Second, an initial 3-D description of the object is computed from two views. This model information is then used in a third step to estimate the camera positions for all available views using a novel linear 3-D motion and shape estimation algorithm. The main feature of this third step is the simultaneous estimation of 3-D camera-motion parameters and object shape refinement with respect to the initial 3-D model. The initial 3-D shape model exhibits only a few degrees of freedom and the object shape refinement is defined as flexible deformation of the initial shape model. Our formulation of the shape deformation allows the object texture to slide on the surface, which differs from traditional flexible body modeling. This novel combined shape and motion estimation using sliding texture considerably improves the calibration data of the individual views in comparison to fixed-shape model based camera-motion estimation. Since the shape model used for model based camera-motion estimation is only approximate, a volumetric 3-D reconstruction process is initiated in the fourth step that combines the information from ail views simultaneously. The recovered object consists of a set of voxels with associated color information that describes even fine structures and details of the object. New views of the object can be rendered from the recovered 3-D model, which has potential applications in virtual reality or multimedia systems and the emerging field of video coding using 3-D scene models. Peter Eisert, Eckehard G. Steinbach, Bernd Girod |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Multi-hypothesis, volumetric reconstruction of 3-D objects from multiple calibrated camera viewsabstractIn this paper we present a volumetric method for the 3-D reconstruction of real world objects from multiple calibrated camera views. The representation of the objects is fully volume-based and no explicit surface description is needed. The approach is based on multi-hypothesis tests of the voxel model back-projected into the image planes. All camera views are incorporated in the reconstruction process simultaneously and no explicit data fusion is needed. In a first step each voxel of the viewing volume is filled with several color hypotheses originating from different camera views. This leads to an overcomplete representation of the 3-D object and each voxel typically contains multiple hypotheses. In a second step only those hypotheses remain in the voxels which are consistent with all camera views where the voxel is visible. Voxels without a valid hypothesis are considered to be transparent. The methodology of our approach combines the advantages of silhouette-based and image feature-based methods. Experimental results on real and synthetic image data show the excellent visual quality of the voxel-based 3-D reconstruction. Peter Eisert, Eckehard G. Steinbach, Bernd Girod |
ICASSP | 1 |
| 1999 | Rate-Distortion-Efficient Video Compression Using a 3-D Head ModelabstractIn this paper we combine model-based video synthesis with block-based motion-compensated prediction (MCP). Two frames utilized for prediction where one frame is the previous decoded one and the other frame is provided by a model-based coder. The approach is integrated into an H.263-based video codec. Rate-distortion optimization is employed for the coding control. Hence, the coding efficiency does not decrease below H.263 even if the model based coder cannot describe the current scene. On the other hand, if the objects in the scene correspond to the model-based coder, significant gains in coding efficiency can be obtained compared to TMN-10, the test model of the H.263 standard. This is verified by experiments with natural head-and-shoulder sequences. Bit-rate savings of about 35% are achieved at equal average PSNR. When encoding at equal bit-rate, significant improvements in terms of subjective quality are visible. Peter Eisert, Thomas Wiegand 0001, Bernd Girod |
ICIP (4) | 1 |
| 1999 | 3-D Image Models and Compression: Synthetic Hybrid or Natural Fit?abstractThis paper highlights recent advances in image compression aided by 3-D geometry information. As two examples, we present a model-aided video coder for efficient compression of head-and-shoulder scenes and a geometry-aided coder for 4-D light fields for image-based rendering. Both examples illustrate that an explicit representation of 3-D geometry is advantageous if many views of the same 3-D object or scene have to be encoded. Waveform-coding and 3-D model-based coding can be combined in a rate-distortion framework, such that the generality of waveform coding and the efficiency of 3-D models are available where needed. Bernd Girod, Peter Eisert, Marcus A. Magnor, Eckehard G. Steinbach, Thomas Wiegand 0001 |
ICIP (2) | 2 |
| 1998 | Digital watermarking of MPEG-4 facial animation parameters
Frank Hartung, Peter Eisert, Bernd Girod |
Comput. Graph. | 2 |
| 1998 | Motion-based analysis and segmentation of image sequences using 3-D scene models
Eckehard G. Steinbach, Peter Eisert, Bernd Girod |
Signal Process. | 2 |
| 1997 | Model-Based Estimation of Facial Expression Parameters from Image SequencesabstractWe present a model-based algorithm for the estimation of 3D motion and the analysis of facial expressions of a speaking person. A set of facial animation parameters based on the MPEG-4 standard is determined from two successive video frames using a hierarchical optical flow based method. The motion in the image plane is constrained by a 3D triangular B-spline model that defines shape, texture and facial expressions of an individual person. The computational requirement for this solution is low due to the linear structure of the algorithm. Peter Eisert, Bernd Girod |
ICIP (2) | 1 |
| 1997 | Model-Based Synthetic View Generation from a Monocular Video SequenceabstractIn this paper a model-based multi-view image generation system for video conferencing is presented. The system assumes that a 3-D model of the person in front of the camera is available. It extracts texture from speaking person sequence images and maps it to the static 3-D model during the videoconference session. Since only the incrementally updated texture information is transmitted during the whole session, the bandwidth requirement is very small. Based on the experimental results one can conclude that the proposed system is very promising for practical applications. Chun-Jen Tsai, Aggelos K. Katsaggelos, Peter Eisert, Bernd Girod |
ICIP (1) | 3 |